Academy/VCF 9.0 Support (2V0-15.25)/vSAN Disk Failure Recovery
This lab targets VCF 9.0

vSAN Disk Failure Recovery

VCF 9.0Intermediatevcp-foundationvcp-support⏱ 120 min

Comprehensive vSAN ESA disk failure diagnosis, recovery, and capacity impact analysis.

Objectives

  • Interpret vSAN ESA disk state transitions (Healthy, Degraded, Absent, Decommissioned)
  • Diagnose disk failures using esxcli vsan, vSAN health checks, and smart.log analysis
  • Execute disk removal and replacement procedures in vSAN ESA storage pools
  • Calculate resync impact on cluster capacity and VM performance
  • Configure vSAN proactive rebalancing and rebuild rate tuning
  • Document disk failure recovery using RCAR methodology

Prerequisites

VCF lab environment with 4+ host vSAN ESA cluster

Required skills:

  • SSH access to ESXi hosts
  • Basic vSAN architecture knowledge (ESA storage pools vs OSA disk groups)
  • esxcli command familiarity

Lab Environment

VCF workload domain with 4-host vSAN ESA cluster, FTT=1 RAID-1 default policy

Tasks

Task 1 vSAN ESA Disk Architecture & Health Assessment

Foundation task. Panelists probe: What changed from OSA to ESA? Why does ESA eliminate the cache/capacity tier distinction? How does single-tier NVMe affect failure domains?

Understand vSAN ESA storage pool architecture vs. legacy OSA disk groups, map disk states, and perform baseline health assessment before simulating failures.

Step 1

SSH to an ESXi host in the vSAN cluster. Run 'esxcli vsan storage list' to enumerate all vSAN storage devices. Document: device name, device type (SSD/NVMe), capacity, status, and storage pool membership.

vSAN ESA shows NVMe devices in a single-tier storage pool — no separate cache and capacity tiers. Each host contributes one or more NVMe devices to the pool. Device status should be 'Healthy' for all devices. Minimum configuration: 1 NVMe TLC device per host.
ESA key difference from OSA: no disk groups, no cache tier. All NVMe devices contribute to a unified storage pool. This simplifies disk management but changes failure behavior — a single disk failure reduces both capacity and performance simultaneously.
Step 2

Run 'esxcli vsan health cluster list' to check overall cluster health. Examine each health check category: Network, Physical Disk, Data, Limits, Stretch Cluster. Document any warnings or errors. Pay special attention to 'Disk balance' and 'Component metadata health'.

Healthy cluster shows green across all categories. Key health checks: vSAN object health (all objects accessible), disk balance (data distributed evenly across hosts), network health (multicast/unicast functioning), and hardware compatibility (all devices on HCL).
Run health checks BEFORE and AFTER any disk operation. Comparing pre/post health reports is essential for validating recovery completeness.
Step 3

Check vSAN capacity status. In vSphere Client navigate to Cluster > Monitor > vSAN > Capacity. Document: total raw capacity, used capacity, free capacity, dedup/compression savings ratio, and slack space reservation. Also run: esxcli vsan debug disk list.

Capacity view shows raw vs. usable space. FTT=1 RAID-1 uses 2x raw capacity. vSAN reserves 25-30% slack space for rebuilds. Actual usable capacity = (raw - metadata 2% - slack 25-30% - filesystem 2%) / FTT multiplier.
Capacity planning is critical for disk failure resilience. If the cluster is above 70% used, a single disk failure may not have enough slack space for rebuild — creating a data-at-risk scenario.
Step 4

Document vSAN disk state machine: Healthy (normal operation) -> Degraded (disk errors detected, data being rebuilt to other disks) -> Absent (disk not responding, 60-minute timer starts) -> Decommissioned (explicitly removed, data fully evacuated). Understand the 60-minute absent timer: after 60 minutes, vSAN begins full rebuild of components from the absent disk. Before 60 minutes, vSAN assumes the disk might return.

State transition diagram with timers: Absent triggers 60-minute wait (configurable via CLOM repair delay). After timer expires, vSAN initiates full component rebuild. During rebuild, cluster must have sufficient slack space.
The 60-minute timer is a key operational parameter. In production, reduce it for SSDs (they rarely recover) and increase it for maintenance windows to prevent unnecessary rebuilds.

Validation Gate

Check: Complete vSAN ESA architecture documentation and baseline health assessment

Expected: All storage devices enumerated, health checks passed, capacity breakdown documented, disk state machine understood.

Common Errors

Confusing vSAN ESA storage pools with OSA disk groups — ESA has no cache tier
Not accounting for slack space when calculating usable capacity — missing 25-30% reservation
Ignoring the 60-minute absent timer — leads to unexpected rebuild operations during maintenance

Task 2 Disk Failure Simulation & Diagnosis

Hands-on troubleshooting task. Panelists ask: How do you distinguish a transient disk error from permanent failure? What is the impact on VMs with FTT=1 when one disk fails? When does data become at risk?

Simulate disk failure scenarios in the lab, diagnose using multiple tools, and understand the cascade effects on vSAN objects and VM availability.

Step 1

Before simulating failure, document the current state of vSAN objects. Run: esxcli vsan debug object health summary. Note the count of healthy, degraded, and absent objects. Also run: esxcli vsan debug object list | head -50 to see individual object placement.

All objects should be healthy with correct number of components per FTT policy. FTT=1 RAID-1 objects have 2 data components + 1 witness component distributed across different hosts. Document the object count as baseline.
Object health is the true measure of vSAN resilience — even if disks are healthy, misplaced or degraded objects indicate underlying issues.
Step 2

Simulate a disk failure. In the Holodeck environment, identify a vSAN data device on one host. Remove the virtual disk from the ESXi VM configuration (simulates physical disk failure). Immediately check: esxcli vsan storage list (disk now shows Absent), esxcli vsan debug object health summary (objects on that disk now degraded).

Removed disk transitions to Absent state. Objects that had components on that disk are now degraded (reduced below FTT policy). vSAN health check shows warning for 'Component metadata health' and 'Data' categories. 60-minute repair timer starts.
In a real production scenario, disk failure is detected by SMART monitoring, ESXi kernel error messages, and vSAN health checks. The simulation skips detection but the recovery process is identical.
Step 3

Analyze the impact on VMs. Check which VMs have components on the failed disk: esxcli vsan debug object list --filter-property=objectStatus --filter-value=degraded. For each affected VM, document: VM name, object type (vmdk, vmx, vswp), current compliance status, and whether the VM is still accessible.

VMs with FTT=1 RAID-1 remain accessible after single disk failure — the mirror copy on another host serves reads. However, they are now unprotected (FTT=0 effective) until rebuild completes. If a second disk fails before rebuild, data loss occurs.
Critical VCDX point: FTT=1 protects against ONE failure. After first failure, data is at risk until rebuild completes. This is why vSAN monitors and alerting are essential — immediate awareness of degraded state enables proactive response.
Step 4

Monitor the repair timer countdown. Check repair progress: esxcli vsan debug resync summary. Note: before the 60-minute timer expires, no repair activity occurs (vSAN waits for the disk to potentially return). After timer expiry or manual trigger, rebuild begins. Monitor rebuild progress: component count, data movement rate (MB/s), estimated time remaining.

Resync summary shows: objects to sync, bytes to sync, completion percentage, and estimated time. Rebuild rate depends on: cluster load, number of hosts, network bandwidth, and configured rebuild rate limit. Typical rebuild for 500 GB per disk: 2-8 hours.
Rebuild time formula: data_on_failed_disk / (rebuild_rate * num_hosts_participating). Default rebuild rate is 100 Mbps per host. Increasing it speeds recovery but impacts VM I/O performance.
Step 5

Calculate the capacity impact of disk failure. Document: (a) raw capacity lost (failed disk size), (b) usable capacity lost (raw / FTT multiplier), (c) remaining cluster free space, (d) whether remaining capacity is sufficient for rebuild (need free space >= data on failed disk). If cluster is above 80% used, calculate whether rebuild can complete.

Single disk failure in a 4-host cluster with 4 TB per disk: raw capacity drops from 16 TB to 12 TB. Usable capacity (FTT=1 RAID-1): drops from ~4.8 TB to ~3.6 TB. If current usage is 3 TB, rebuild has sufficient space. If usage is 4 TB, cluster enters data-at-risk state.
This capacity analysis is exactly what VCDX panelists expect. Show the math: raw capacity - metadata - slack - FTT overhead = usable. Demonstrate you can calculate failure impact quickly.

Validation Gate

Check: Successfully simulate disk failure, diagnose impact, and calculate capacity implications

Expected: Disk failure simulated, affected objects identified, VM impact assessed, capacity calculations documented, repair timer behavior understood.

Common Errors

Not documenting pre-failure baseline — makes it impossible to compare post-recovery state
Confusing Absent (disk might return) with Degraded (data being rebuilt) states
Forgetting that FTT=1 objects are unprotected during rebuild — a second failure causes data loss
Not calculating whether cluster has sufficient free space for rebuild before triggering it

Task 3 Disk Replacement & Rebuild Operations

Operational execution task. Panelists expect precise command sequences and understanding of maintenance mode implications: ensure data availability vs. full data migration.

Execute the complete disk replacement workflow in vSAN ESA — from removing the failed disk to adding the replacement to verifying full cluster recovery.

Step 1

Remove the failed disk from the vSAN storage pool. Use esxcli: 'esxcli vsan storage remove -d <device_naa_id>'. Verify removal: 'esxcli vsan storage list' should no longer show the device. Check that objects previously on this disk are being rebuilt to remaining disks.

Disk removal is immediate. vSAN begins rebuilding components that were on the removed disk to other healthy disks in the cluster. Rebuild progress visible in esxcli vsan debug resync summary.
In ESA, disk removal from the storage pool is simpler than OSA disk group removal — no cache tier concerns. But verify that the cluster has sufficient capacity before removing.
Step 2

While waiting for rebuild, understand vSAN maintenance mode options. Three modes: (1) Ensure accessibility — fastest, moves minimum data to maintain FTT=0 access; (2) Full data migration — slowest, moves ALL data off the host (required for host decommission); (3) No data migration — fastest, used for quick reboots when host returns quickly. Document when each mode is appropriate.

Ensure accessibility: use for disk replacement (host stays, just one disk gone). Full data migration: use for host removal from cluster (permanent). No data migration: use for firmware update reboot (<60 minutes). Pre-check available: cluster validates whether the requested mode can complete.
VCDX design point: maintenance mode choice directly impacts maintenance window duration. Full data migration for a 4 TB host takes 4-8 hours. Ensure accessibility takes minutes. Choose wisely based on the operation.
Step 3

Simulate adding a replacement disk. In Holodeck, add a new virtual disk to the ESXi VM. Claim it for vSAN: 'esxcli vsan storage add -d <new_device_naa_id> -s <storage_pool_uuid>'. Verify: 'esxcli vsan storage list' shows the new device in Healthy state.

New disk joins the storage pool immediately. vSAN begins proactive rebalancing to distribute data across all disks including the new one. Rebalancing runs at lower priority than rebuild to minimize VM impact.
After adding a new disk, rebalancing may take hours depending on data volume. Monitor with: esxcli vsan debug resync summary. Do not add another disk until rebalancing completes.
Step 4

Monitor the complete recovery. Track: (a) resync progress (esxcli vsan debug resync summary), (b) object health returning to fully compliant (esxcli vsan debug object health summary), (c) cluster capacity returning to expected level (vSphere Client capacity view), (d) VM I/O performance during rebuild (esxtop disk view). Document the total recovery timeline.

Full recovery sequence: disk removal (immediate) -> rebuild to remaining disks (2-8 hours) -> replacement disk added (immediate) -> rebalancing to new disk (2-8 hours) -> all objects compliant (final state). Total recovery window: 4-16 hours depending on data volume.
Rebuild rate tuning: esxcli vsan policy setdefault -c resyncThrottleRate -P <value>. Higher rate = faster rebuild but more VM I/O impact. Default balances both. During off-peak hours, temporarily increase for faster recovery.
Step 5

Perform post-recovery validation. Run complete health check: esxcli vsan health cluster list. Verify: all objects healthy, disk balance within 5% variance, no stale components, capacity matches expected values. Compare with pre-failure baseline documented in Task 1.

Post-recovery cluster should match pre-failure baseline: same object count (healthy), same capacity (with new disk), all health checks green. Disk balance may take additional time to equalize across all hosts.
Document the complete recovery timeline as a reference for future incidents and capacity planning. This data is invaluable for SLA calculations and VCDX design justification.

Validation Gate

Check: Complete disk replacement workflow with verified cluster recovery

Expected: Failed disk removed, replacement added, all objects rebuilt to compliance, cluster health restored to baseline.

Common Errors

Removing disk without verifying cluster has capacity for rebuild — triggers data-at-risk state
Using 'Full data migration' maintenance mode for disk replacement — unnecessary, wastes hours
Adding replacement disk before rebuild completes — causes competing I/O and extends total recovery time
Not running post-recovery health check — may miss partially rebuilt objects or balance issues

Task 4 Proactive Monitoring Design & VCDX Defense

Synthesis task connecting operational procedures to architectural design. Panelists expect monitoring strategy, SLA impact analysis, and justification for FTT policy choices.

Design a comprehensive vSAN disk health monitoring strategy and prepare to defend disk failure recovery design decisions in VCDX context.

Step 1

Design a vSAN disk health monitoring framework. Define alerts for: (a) SMART pre-failure warnings (disk showing errors before failure), (b) disk state transition alerts (Healthy to Degraded/Absent), (c) capacity threshold alerts (70% warning, 80% critical, 85% emergency), (d) rebuild progress alerts (rebuild started, 50% complete, completed), (e) performance impact alerts (VM latency exceeding baseline during rebuild). Map each alert to Aria Operations or vCenter alarm configuration.

Monitoring framework with 5 alert categories, each mapped to specific metrics, thresholds, and notification targets. SMART monitoring catches 60-80% of disk failures before they occur. Capacity alerts prevent rebuild failures due to insufficient space.
Proactive monitoring is the key differentiator between reactive and mature operations. VCDX panelists specifically ask: How do you prevent disk failure incidents? SMART monitoring and capacity alerts are the answers.
Step 2

Calculate SLA impact for different failure scenarios. Build a table: Scenario | Detection Time | Repair Time | Total Impact | SLA Effect. Scenarios: (a) single disk failure with FTT=1 RAID-1 (data accessible, rebuild 4h), (b) single disk failure with FTT=1 RAID-5 (data accessible, rebuild 6h due to parity calculation), (c) dual disk failure with FTT=1 (data loss if on same object), (d) host failure (all disks on host lost, rebuild from remaining 3 hosts). Correlate with SLA targets: 99.9% = 8.76h/year, 99.99% = 52.6 min/year.

SLA impact table showing that FTT=1 handles single-failure scenarios within SLA. Dual failure before rebuild completes is the risk scenario. Host failure with FTT=1 is equivalent to disk failure (1 of 2 copies lost). FTT=2 provides dual-failure protection but costs 3x capacity (RAID-1) or 1.5x (RAID-6).
VCDX defense: always connect storage policy (FTT level) to SLA requirement. FTT=1 for 99.9%, FTT=2 for 99.99%+ with financial workloads. Show the cost-capacity tradeoff.
Step 3

Document the complete disk failure recovery runbook for VCF operations team. Include: (a) incident detection and classification, (b) impact assessment checklist, (c) step-by-step recovery procedure with esxcli commands, (d) communication template for stakeholders, (e) post-incident review template. This runbook should be usable by L2 support without escalation.

Runbook covers: detection (alert fires) to resolution (all objects compliant) with specific commands, expected outputs, and decision points. Communication template includes: what happened, what is the impact, what is being done, when will it be resolved.
A well-documented runbook reduces MTTR from hours to minutes. VCDX panelists value operational maturity — show that your design includes operational procedures, not just architecture diagrams.
Step 4

Prepare four VCDX defense responses: (1) 'Why FTT=1 instead of FTT=2 for this workload domain?' — cost analysis shows FTT=2 RAID-1 requires 3x capacity (vs 2x for FTT=1), adding $X per TB. Risk analysis shows single-failure probability is 10x higher than dual-failure, making FTT=1 appropriate for non-critical workloads. (2) 'What happens if a disk fails during a VCF upgrade?' — disk failure during ESXi upgrade triggers vSAN rebuild on remaining hosts; upgrade pauses if cluster health drops below threshold; resume after rebuild stabilizes. (3) 'How do you handle disk failure in a 4-host minimum cluster?' — 4-host cluster can tolerate 1 host/disk failure with FTT=1; during rebuild, cluster operates at reduced capacity; if cluster drops below 3 healthy hosts, new VM provisioning is blocked. (4) 'Why ESA over OSA for this design?' — ESA eliminates cache/capacity tier complexity, supports larger NVMe devices (up to 32 TB), delivers higher IOPS per device, and simplifies disk replacement procedures.

Four defense responses with quantitative backing. Each demonstrates storage design understanding connected to business requirements.
Always lead with the business requirement, then the technical justification, then the risk mitigation. This RCAR flow is natural for VCDX defense.

Validation Gate

Check: Complete monitoring framework, SLA analysis, operational runbook, and VCDX defense preparation

Expected: Monitoring framework with 5 alert categories documented. SLA impact table created. Recovery runbook ready for L2 operations. Four VCDX defense responses prepared.

Common Errors

Not connecting FTT policy to SLA requirements — storage policy must trace to business availability targets
Overlooking SMART monitoring — proactive detection prevents 60-80% of disk failure incidents
Writing runbooks that require L3/architect escalation for routine operations — defeats the purpose of runbook automation
Ignoring rebuild performance impact on VMs — rebuild generates significant I/O that competes with production workloads

Final Validation

Complete vSAN disk failure recovery with proactive monitoring and VCDX defense readiness

✓ vSAN ESA architecture understood → Storage pool architecture, disk states, and ESA vs OSA differences documented

✓ Disk failure diagnosed → Failure simulated, objects identified, capacity impact calculated

✓ Recovery completed → Disk replaced, objects rebuilt, cluster restored to baseline health

✓ Monitoring framework designed → SMART alerts, capacity thresholds, and rebuild monitoring configured

✓ VCDX defense prepared → Four defense responses with quantitative analysis documented

Cleanup / Restore

• Verify all vSAN objects are healthy: esxcli vsan debug object health summary

• Confirm cluster capacity matches expected values

• Check disk balance across all hosts

• Revert to snapshot if needed for next lab

Design Reflection (VCDX)

vSAN disk failure recovery demonstrates operational maturity in VCF storage management. The design connects storage policy (FTT level) to SLA requirements, implements proactive SMART monitoring, and provides graduated recovery procedures. ESA storage pool architecture simplifies disk management compared to OSA disk groups.

Requirements

  • R-001: vSAN cluster must tolerate single disk failure without VM downtime
  • R-002: Disk failure recovery must complete within 8 hours (SLA target)
  • R-003: Proactive monitoring must detect 80%+ of disk failures before data loss

Constraints

  • Minimum 4 hosts per vSAN cluster (VCF requirement)
  • vSAN ESA requires NVMe devices (no SAS/SATA support)
  • Rebuild rate must be balanced against production VM performance

Assumptions

  • Spare NVMe disks available within 4-hour delivery window
  • Operations team trained on vSAN disk replacement procedures
  • Aria Operations or equivalent monitoring deployed and configured

Risks

  • Dual disk failure before rebuild completes causes data loss with FTT=1 — mitigate with FTT=2 for critical workloads
  • Cluster at >80% capacity may not have sufficient space for rebuild — proactive capacity monitoring required
  • Firmware bugs on NVMe devices can cause simultaneous multi-disk failure — use mixed firmware versions

Self-Assessment Discussion Prompts

  1. At what cluster utilization percentage would you block new VM provisioning to protect rebuild capacity?
  2. How does vSAN ESA change your disk failure response compared to OSA?
  3. What is the cost-benefit breakpoint for FTT=1 RAID-5 vs FTT=1 RAID-1 in a capacity-sensitive environment?
  4. How would you design disk failure resilience for a stretched vSAN cluster?

Extensions

vSAN RAID-5/6 Erasure Coding Recovery

Configure a storage policy with FTT=1 RAID-5 (erasure coding). Simulate disk failure and compare rebuild time and I/O impact with RAID-1 mirroring. Document the rebuild overhead difference: RAID-5 requires parity recalculation, making rebuild ~50% slower than RAID-1.

Multi-Disk Failure Scenario

Simulate simultaneous failure of 2 disks on different hosts in an FTT=2 cluster. Verify data remains accessible. Document the degraded capacity and rebuild timeline for dual-failure recovery.

Automated Disk Replacement with vSAN Proactive Operations

Configure Aria Operations to detect SMART warnings, automatically create a change ticket, and pre-stage the disk replacement workflow. Design the end-to-end automation from detection to replacement verification.

⚠ Known Pitfalls (from Community KB)

Removing a failed disk without verifying cluster has sufficient slack space for rebuild — if cluster is above 80% used, rebuild may fail and trigger data-at-risk state
Using 'No data migration' maintenance mode for permanent disk removal — data on that disk is lost; use only for temporary host reboots
Not monitoring rebuild progress — rebuilds can stall due to network congestion or capacity constraints, requiring manual intervention
Ignoring SMART pre-failure warnings — proactive replacement during maintenance window is 10x less risky than reactive replacement after failure

References

  • vSAN 9.0 Administration Guide — Disk Management and Maintenance chapter
  • vSAN 9.0 Troubleshooting Guide — Object Health and Repair
  • VMware KB 2107713 — vSAN disk failure troubleshooting and replacement procedures
  • vSAN ESA Architecture and Operations Guide
  • VMware VCF 9.0 Storage Management — vSAN Lifecycle Operations
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.