Lab: Monitor vSAN Resync After Host Failure
Objectives
- Simulate host failure, observe vSAN resync in action, monitor progress, verify data accessibility.
Prerequisites
VCF lab environment deployed and operational
Lab Environment
Standard VCF lab environment for vSphere Foundation 9.0 Support (Support Specialist)
Tasks
Task 1 Lab: Monitor vSAN Resync After Host Failure
Simulate host failure, observe vSAN resync in action, monitor progress, verify data accessibility.
vCenter → Cluster → Monitor → vSAN → Health (baseline green)
Identify one vSAN host to fail
Simulate failure: shutdown host or disconnect network
vCenter detects host unreachable within 30 seconds
vSAN health drops to yellow (component degraded)
Resync starts: Watch Resync Status % progress
Monitor network bandwidth consumed during resync (esxtop on remaining hosts)
Verify VMs continue to run (data accessible despite failure)
Host recovers: Power on, wait for rejoin
vSAN health returns to green
Validation Gate
Check: After simulating a host failure: verify vSAN resync starts automatically, monitor progress until completion, confirm all objects return to healthy state, and measure workload I/O impact during resync
Expected: Resync begins after CLOMD repair delay (default 60 min). Progress shows decreasing component count. After completion: all objects healthy with full FTT compliance. Workload I/O latency returned to pre-failure baseline.
Common Errors
Final Validation
Lab completed successfully
✓ All steps completed → No errors observed
Cleanup / Restore
• Revert to snapshot if needed
Design Reflection (VCDX)
Host failure and vSAN resync is the most common failure scenario VCDX panelists present. Know the timeline: failure detection (30s) → VM restart (2-5 min) → CLOMD delay (60 min) → resync (varies by data volume). Be able to walk through each phase.
Requirements
- Monitor vSAN resync after host failure
- Understand resync timing and impact on workloads
- Investigate host failure root cause before returning host to cluster
Constraints
- Do not initiate manual rebalance during active resync
- Resync consumes bandwidth — workload I/O may be impacted
- CLOMD repair delay is configurable but 60 min is default for good reason
Assumptions
- Cluster has N+1 capacity for host failure
- vSAN network bandwidth is sufficient for resync within acceptable timeframe
Risks
- Data loss from second failure during resync
- Persistent hardware issue causing repeated failures if host returned without investigation