vSAN Failure & Recovery Scenario
Objectives
- Diagnose vSAN host failure impact on object availability and cluster capacity
- Interpret vSAN disk state transitions and rebuild timers in ESA architecture
- Execute vSAN rebuild rate tuning to balance recovery speed vs. VM performance
- Recover from vSAN network partition scenarios (split-brain detection and resolution)
- Design vSAN capacity alert thresholds to prevent data-at-risk emergencies
- Document vSAN failure recovery using RCAR methodology
Prerequisites
VCF lab with 4-host vSAN ESA cluster running production-like workloads
Prior labs: vcp-support-02
Required skills:
- esxcli vsan commands
- vSAN health check interpretation
- ESXi SSH access
Lab Environment
VCF workload domain with 4-host vSAN ESA cluster, FTT=1 RAID-1, mixed workloads
Tasks
Task 1 Host Failure Impact Analysis
Understand the cascading impact of a complete ESXi host failure on vSAN object availability, VM placement, and cluster capacity.
Document the pre-failure cluster state. For each host: hostname, number of VMs, vSAN disk count, vSAN capacity contribution, vSAN components hosted. Use: esxcli vsan debug disk list (or: esxcli vsan debug disk summary get) and vSphere Client Cluster Monitor vSAN Capacity view.
Simulate complete host failure. In Holodeck, power off one ESXi host VM abruptly. Observe: vSphere HA detects host failure (~30s), begins restarting VMs on surviving hosts, vSAN marks all components on the failed host as Absent, 60-minute repair timer starts.
Calculate capacity impact. With one host down: raw capacity drops 25%, remaining 3 hosts must accommodate all VMs. Verify HA admission control allows all VMs to restart. Document any VMs that cannot restart due to insufficient resources.
Monitor 60-minute absent timer and subsequent rebuild. After timer expires, vSAN rebuilds components from surviving mirrors. Track: esxcli vsan debug resync summary (rebuild progress), total data to rebuild, rebuild rate per host, estimated completion time.
Validation Gate
Check: Complete host failure simulation with impact analysis
Expected: Host failure simulated, HA restart verified, capacity impact calculated, rebuild timeline tracked.
Common Errors
Task 2 vSAN Network Partition & Split-Brain Resolution
Diagnose and recover from vSAN network partition scenarios where hosts lose connectivity but remain operational.
Understand vSAN network partition behavior. When vSAN network connectivity is lost between hosts: each host accesses local components, cannot access remote components. CMMDS detects partition, objects with majority components accessible remain available. Document the CMMDS quorum rules for FTT=1 RAID-1: 2 data components + 1 witness = 3 votes.
Simulate network partition in Holodeck. Isolate one ESXi host by disconnecting its vSAN VMkernel port. Observe: isolated host still runs VMs, objects with majority on remaining 3 hosts remain accessible, objects with both data and witness on isolated host become inaccessible on main cluster side.
Resolve the partition. Reconnect the vSAN VMkernel port. Monitor: CMMDS re-establishes cluster membership, vSAN merges partition with component resync, all objects return to healthy state. Track with esxcli vsan debug resync summary.
Design vSAN network resilience: redundant VMkernel ports on separate physical NICs, dedicated VLAN with QoS, network health monitoring with Aria Operations, alerts for latency >5ms and packet loss >0.01%.
Validation Gate
Check: Simulate and resolve vSAN network partition
Expected: Partition simulated, quorum behavior observed, partition resolved, network resilience documented.
Common Errors
Task 3 Capacity Emergency Procedures & Rebuild Rate Tuning
Handle vSAN capacity emergencies when failures push the cluster above safe thresholds.
Define vSAN capacity zones: Zone 1 (0-70%) normal; Zone 2 (70-80%) warning, plan expansion; Zone 3 (80-85%) critical, stop provisioning; Zone 4 (85-95%) emergency, reduce footprint; Zone 5 (>95%) crisis, immediate intervention.
Practice resync monitoring. View progress: esxcli vsan debug resync summary get. Note: 'esxcli vsan policy getdefault' shows default storage policy, not resync rates — OSA throttling is set in the vSphere Client (Resyncing Objects > Resync Throttling); ESA uses adaptive resync automatically. Monitor impact with esxtop disk view. Document tradeoff: 2x rebuild rate = ~50% faster rebuild but 15-25% VM latency increase.
Practice emergency capacity recovery options: (a) power off idle VMs (safe, immediate), (b) thin provision conversion (safe, slow), (c) FTT downgrade for non-critical VMs to FTT=0 (dangerous, immediate), (d) add storage/host (safe, requires procurement).
Calculate minimum cluster size: FTT=1 RAID-1 minimum 3 hosts (no rebuild capacity), operational minimum 4 hosts (rebuild after 1 failure). FTT=2 RAID-1 minimum 5 hosts. Formula: hosts_needed = FTT_minimum + rebuild_buffer + growth.
Validation Gate
Check: Complete capacity emergency procedures
Expected: Capacity zones defined, rebuild rate tuned, emergency procedures documented, sizing formula applied.
Common Errors
Task 4 vSAN Recovery Runbook & VCDX Defense
Create comprehensive recovery runbook and prepare VCDX defense for vSAN design decisions.
Build incident classification: Sev1 data loss/inaccessibility (immediate), Sev2 degraded/unprotected (1 hour), Sev3 single disk failure with FTT maintained (4 hours), Sev4 warning condition (24 hours). Document response times and escalation paths.
Create recovery flowchart with 4 entry points: disk failure, host failure, network partition, capacity emergency. Each path: detection, classification, impact assessment, mitigation, recovery, verification, post-incident review.
Calculate SLA impact: disk failure = 5h exposure at FTT=0, host failure = 7h exposure, network partition = variable. Map to SLA targets: 99.9% = 8.76h/year, 99.99% = 52.6min/year.
Prepare VCDX defense responses: (1) Why 4 hosts not 3 for FTT=1 — rebuild capacity after failure. (2) Monitoring strategy — five layers: SMART, object health, capacity, performance, network. (3) Stretched cluster failure handling — witness determines quorum. (4) Largest vSAN risk — silent data corruption from firmware bugs; mitigate with HCL compliance and integrity scrub.
Validation Gate
Check: Complete vSAN recovery runbook and VCDX defense
Expected: Incident classification, recovery flowchart, SLA analysis, and four defense responses completed.
Common Errors
Final Validation
Complete vSAN failure & recovery with advanced scenarios and VCDX defense
✓ Host failure analyzed → Capacity impact, HA interaction, rebuild timeline documented
✓ Network partition resolved → Split-brain understood, partition simulated and resolved
✓ Capacity emergency handled → Emergency procedures with rebuild rate tuning
✓ Recovery runbook created → Incident classification, flowchart, SLA analysis
✓ VCDX defense prepared → Four responses with quantitative backing
Cleanup / Restore
• Restore all hosts to powered-on state
• Verify vSAN cluster health green across all checks
• Confirm all objects compliant with storage policy
• Revert to snapshot if needed
Design Reflection (VCDX)
Advanced vSAN failure recovery demonstrates operational maturity. Design connects FTT policies to SLA requirements, implements multi-layer monitoring, and provides structured incident response.
Requirements
- R-001: vSAN cluster survives single host failure without data loss
- R-002: VM availability maintained during rebuild
- R-003: Recovery procedures executable by L2 support
Constraints
- Minimum 4 hosts for FTT=1 with rebuild capacity
- vSAN network requires redundant paths on dedicated VLAN
- Rebuild rate balances recovery speed with VM performance
Assumptions
- Aria Operations monitoring deployed
- Operations team trained on incident classification
- Spare hardware available within SLA window
Risks
- Second host failure during rebuild = data loss with FTT=1
- Network partition creates split-brain risk
- Capacity emergency during host failure requires immediate intervention
Self-Assessment Discussion Prompts
- At what cluster size does FTT=2 become cost-effective vs FTT=1 with monitoring?
- How would you design recovery differently for stretched cluster vs single-site?
- What is the operational impact of ESA vs OSA on failure recovery?
- How do you balance rebuild rate with production SLA during business hours?
Extensions
Stretched Cluster Failure Simulation
Deploy vSAN stretched cluster in Holodeck. Simulate site failure, verify automatic failover. Document witness behavior, quorum determination, and site restoration.
vSAN Data Integrity Verification
Enable data integrity scrub. Simulate bit-rot and verify vSAN detects and repairs corruption using mirror copies.
Automated vSAN Incident Response
Build automated workflow: vSAN alert -> classify severity -> execute diagnostics -> notify team -> open change ticket.
⚠ Known Pitfalls (from Community KB)
References
- vSAN 9.0 Troubleshooting Guide — Host Failure and Network Partition Recovery
- vSAN 9.0 Administration Guide — Rebuild Rate Tuning and Capacity Management
- VMware KB 2150769 — vSAN Object Health States and Recovery Procedures
- vSAN ESA Architecture Guide — Storage Pool Recovery
- VMware VCF 9.0 Operational Guide — vSAN Monitoring Best Practices