vSAN Disk Failure Recovery
Objectives
- Interpret vSAN ESA disk state transitions (Healthy, Degraded, Absent, Decommissioned)
- Diagnose disk failures using esxcli vsan, vSAN health checks, and smart.log analysis
- Execute disk removal and replacement procedures in vSAN ESA storage pools
- Calculate resync impact on cluster capacity and VM performance
- Configure vSAN proactive rebalancing and rebuild rate tuning
- Document disk failure recovery using RCAR methodology
Prerequisites
VCF lab environment with 4+ host vSAN ESA cluster
Required skills:
- SSH access to ESXi hosts
- Basic vSAN architecture knowledge (ESA storage pools vs OSA disk groups)
- esxcli command familiarity
Lab Environment
VCF workload domain with 4-host vSAN ESA cluster, FTT=1 RAID-1 default policy
Tasks
Task 1 vSAN ESA Disk Architecture & Health Assessment
Understand vSAN ESA storage pool architecture vs. legacy OSA disk groups, map disk states, and perform baseline health assessment before simulating failures.
SSH to an ESXi host in the vSAN cluster. Run 'esxcli vsan storage list' to enumerate all vSAN storage devices. Document: device name, device type (SSD/NVMe), capacity, status, and storage pool membership.
Run 'esxcli vsan health cluster list' to check overall cluster health. Examine each health check category: Network, Physical Disk, Data, Limits, Stretch Cluster. Document any warnings or errors. Pay special attention to 'Disk balance' and 'Component metadata health'.
Check vSAN capacity status. In vSphere Client navigate to Cluster > Monitor > vSAN > Capacity. Document: total raw capacity, used capacity, free capacity, dedup/compression savings ratio, and slack space reservation. Also run: esxcli vsan debug disk list.
Document vSAN disk state machine: Healthy (normal operation) -> Degraded (disk errors detected, data being rebuilt to other disks) -> Absent (disk not responding, 60-minute timer starts) -> Decommissioned (explicitly removed, data fully evacuated). Understand the 60-minute absent timer: after 60 minutes, vSAN begins full rebuild of components from the absent disk. Before 60 minutes, vSAN assumes the disk might return.
Validation Gate
Check: Complete vSAN ESA architecture documentation and baseline health assessment
Expected: All storage devices enumerated, health checks passed, capacity breakdown documented, disk state machine understood.
Common Errors
Task 2 Disk Failure Simulation & Diagnosis
Simulate disk failure scenarios in the lab, diagnose using multiple tools, and understand the cascade effects on vSAN objects and VM availability.
Before simulating failure, document the current state of vSAN objects. Run: esxcli vsan debug object health summary. Note the count of healthy, degraded, and absent objects. Also run: esxcli vsan debug object list | head -50 to see individual object placement.
Simulate a disk failure. In the Holodeck environment, identify a vSAN data device on one host. Remove the virtual disk from the ESXi VM configuration (simulates physical disk failure). Immediately check: esxcli vsan storage list (disk now shows Absent), esxcli vsan debug object health summary (objects on that disk now degraded).
Analyze the impact on VMs. Check which VMs have components on the failed disk: esxcli vsan debug object list --filter-property=objectStatus --filter-value=degraded. For each affected VM, document: VM name, object type (vmdk, vmx, vswp), current compliance status, and whether the VM is still accessible.
Monitor the repair timer countdown. Check repair progress: esxcli vsan debug resync summary. Note: before the 60-minute timer expires, no repair activity occurs (vSAN waits for the disk to potentially return). After timer expiry or manual trigger, rebuild begins. Monitor rebuild progress: component count, data movement rate (MB/s), estimated time remaining.
Calculate the capacity impact of disk failure. Document: (a) raw capacity lost (failed disk size), (b) usable capacity lost (raw / FTT multiplier), (c) remaining cluster free space, (d) whether remaining capacity is sufficient for rebuild (need free space >= data on failed disk). If cluster is above 80% used, calculate whether rebuild can complete.
Validation Gate
Check: Successfully simulate disk failure, diagnose impact, and calculate capacity implications
Expected: Disk failure simulated, affected objects identified, VM impact assessed, capacity calculations documented, repair timer behavior understood.
Common Errors
Task 3 Disk Replacement & Rebuild Operations
Execute the complete disk replacement workflow in vSAN ESA — from removing the failed disk to adding the replacement to verifying full cluster recovery.
Remove the failed disk from the vSAN storage pool. Use esxcli: 'esxcli vsan storage remove -d <device_naa_id>'. Verify removal: 'esxcli vsan storage list' should no longer show the device. Check that objects previously on this disk are being rebuilt to remaining disks.
While waiting for rebuild, understand vSAN maintenance mode options. Three modes: (1) Ensure accessibility — fastest, moves minimum data to maintain FTT=0 access; (2) Full data migration — slowest, moves ALL data off the host (required for host decommission); (3) No data migration — fastest, used for quick reboots when host returns quickly. Document when each mode is appropriate.
Simulate adding a replacement disk. In Holodeck, add a new virtual disk to the ESXi VM. Claim it for vSAN: 'esxcli vsan storage add -d <new_device_naa_id> -s <storage_pool_uuid>'. Verify: 'esxcli vsan storage list' shows the new device in Healthy state.
Monitor the complete recovery. Track: (a) resync progress (esxcli vsan debug resync summary), (b) object health returning to fully compliant (esxcli vsan debug object health summary), (c) cluster capacity returning to expected level (vSphere Client capacity view), (d) VM I/O performance during rebuild (esxtop disk view). Document the total recovery timeline.
Perform post-recovery validation. Run complete health check: esxcli vsan health cluster list. Verify: all objects healthy, disk balance within 5% variance, no stale components, capacity matches expected values. Compare with pre-failure baseline documented in Task 1.
Validation Gate
Check: Complete disk replacement workflow with verified cluster recovery
Expected: Failed disk removed, replacement added, all objects rebuilt to compliance, cluster health restored to baseline.
Common Errors
Task 4 Proactive Monitoring Design & VCDX Defense
Design a comprehensive vSAN disk health monitoring strategy and prepare to defend disk failure recovery design decisions in VCDX context.
Design a vSAN disk health monitoring framework. Define alerts for: (a) SMART pre-failure warnings (disk showing errors before failure), (b) disk state transition alerts (Healthy to Degraded/Absent), (c) capacity threshold alerts (70% warning, 80% critical, 85% emergency), (d) rebuild progress alerts (rebuild started, 50% complete, completed), (e) performance impact alerts (VM latency exceeding baseline during rebuild). Map each alert to Aria Operations or vCenter alarm configuration.
Calculate SLA impact for different failure scenarios. Build a table: Scenario | Detection Time | Repair Time | Total Impact | SLA Effect. Scenarios: (a) single disk failure with FTT=1 RAID-1 (data accessible, rebuild 4h), (b) single disk failure with FTT=1 RAID-5 (data accessible, rebuild 6h due to parity calculation), (c) dual disk failure with FTT=1 (data loss if on same object), (d) host failure (all disks on host lost, rebuild from remaining 3 hosts). Correlate with SLA targets: 99.9% = 8.76h/year, 99.99% = 52.6 min/year.
Document the complete disk failure recovery runbook for VCF operations team. Include: (a) incident detection and classification, (b) impact assessment checklist, (c) step-by-step recovery procedure with esxcli commands, (d) communication template for stakeholders, (e) post-incident review template. This runbook should be usable by L2 support without escalation.
Prepare four VCDX defense responses: (1) 'Why FTT=1 instead of FTT=2 for this workload domain?' — cost analysis shows FTT=2 RAID-1 requires 3x capacity (vs 2x for FTT=1), adding $X per TB. Risk analysis shows single-failure probability is 10x higher than dual-failure, making FTT=1 appropriate for non-critical workloads. (2) 'What happens if a disk fails during a VCF upgrade?' — disk failure during ESXi upgrade triggers vSAN rebuild on remaining hosts; upgrade pauses if cluster health drops below threshold; resume after rebuild stabilizes. (3) 'How do you handle disk failure in a 4-host minimum cluster?' — 4-host cluster can tolerate 1 host/disk failure with FTT=1; during rebuild, cluster operates at reduced capacity; if cluster drops below 3 healthy hosts, new VM provisioning is blocked. (4) 'Why ESA over OSA for this design?' — ESA eliminates cache/capacity tier complexity, supports larger NVMe devices (up to 32 TB), delivers higher IOPS per device, and simplifies disk replacement procedures.
Validation Gate
Check: Complete monitoring framework, SLA analysis, operational runbook, and VCDX defense preparation
Expected: Monitoring framework with 5 alert categories documented. SLA impact table created. Recovery runbook ready for L2 operations. Four VCDX defense responses prepared.
Common Errors
Final Validation
Complete vSAN disk failure recovery with proactive monitoring and VCDX defense readiness
✓ vSAN ESA architecture understood → Storage pool architecture, disk states, and ESA vs OSA differences documented
✓ Disk failure diagnosed → Failure simulated, objects identified, capacity impact calculated
✓ Recovery completed → Disk replaced, objects rebuilt, cluster restored to baseline health
✓ Monitoring framework designed → SMART alerts, capacity thresholds, and rebuild monitoring configured
✓ VCDX defense prepared → Four defense responses with quantitative analysis documented
Cleanup / Restore
• Verify all vSAN objects are healthy: esxcli vsan debug object health summary
• Confirm cluster capacity matches expected values
• Check disk balance across all hosts
• Revert to snapshot if needed for next lab
Design Reflection (VCDX)
vSAN disk failure recovery demonstrates operational maturity in VCF storage management. The design connects storage policy (FTT level) to SLA requirements, implements proactive SMART monitoring, and provides graduated recovery procedures. ESA storage pool architecture simplifies disk management compared to OSA disk groups.
Requirements
- R-001: vSAN cluster must tolerate single disk failure without VM downtime
- R-002: Disk failure recovery must complete within 8 hours (SLA target)
- R-003: Proactive monitoring must detect 80%+ of disk failures before data loss
Constraints
- Minimum 4 hosts per vSAN cluster (VCF requirement)
- vSAN ESA requires NVMe devices (no SAS/SATA support)
- Rebuild rate must be balanced against production VM performance
Assumptions
- Spare NVMe disks available within 4-hour delivery window
- Operations team trained on vSAN disk replacement procedures
- Aria Operations or equivalent monitoring deployed and configured
Risks
- Dual disk failure before rebuild completes causes data loss with FTT=1 — mitigate with FTT=2 for critical workloads
- Cluster at >80% capacity may not have sufficient space for rebuild — proactive capacity monitoring required
- Firmware bugs on NVMe devices can cause simultaneous multi-disk failure — use mixed firmware versions
Self-Assessment Discussion Prompts
- At what cluster utilization percentage would you block new VM provisioning to protect rebuild capacity?
- How does vSAN ESA change your disk failure response compared to OSA?
- What is the cost-benefit breakpoint for FTT=1 RAID-5 vs FTT=1 RAID-1 in a capacity-sensitive environment?
- How would you design disk failure resilience for a stretched vSAN cluster?
Extensions
vSAN RAID-5/6 Erasure Coding Recovery
Configure a storage policy with FTT=1 RAID-5 (erasure coding). Simulate disk failure and compare rebuild time and I/O impact with RAID-1 mirroring. Document the rebuild overhead difference: RAID-5 requires parity recalculation, making rebuild ~50% slower than RAID-1.
Multi-Disk Failure Scenario
Simulate simultaneous failure of 2 disks on different hosts in an FTT=2 cluster. Verify data remains accessible. Document the degraded capacity and rebuild timeline for dual-failure recovery.
Automated Disk Replacement with vSAN Proactive Operations
Configure Aria Operations to detect SMART warnings, automatically create a change ticket, and pre-stage the disk replacement workflow. Design the end-to-end automation from detection to replacement verification.
⚠ Known Pitfalls (from Community KB)
References
- vSAN 9.0 Administration Guide — Disk Management and Maintenance chapter
- vSAN 9.0 Troubleshooting Guide — Object Health and Repair
- VMware KB 2107713 — vSAN disk failure troubleshooting and replacement procedures
- vSAN ESA Architecture and Operations Guide
- VMware VCF 9.0 Storage Management — vSAN Lifecycle Operations