Lab: vSAN Health Check & Resync Monitoring
Objectives
- Lab: vSAN Health Check & Resync Monitoring
Prerequisites
VCF lab environment deployed and operational
Lab Environment
Standard VCF lab environment for vSphere Foundation 9.0 Support (Support Specialist)
Tasks
Task 1 Lab: vSAN Health Check & Resync Monitoring
vCenter UI → Cluster → vSAN → Health Check → Run Now
Review results: Note any yellow/red alerts
Check network latency: Health Check → Network → Latency results (should be <2ms between hosts)
If MTU check fails: vCenter → Distributed Switch → vSAN port group → Edit → MTU field → Verify 9000
Simulate disk failure: Shutdown host → Remove one capacity disk from vSAN disk group (via web UI)
Monitor resync: vCenter → Cluster → Performance → Watch "Resync throughput" graph for >0 MB/s
After resync completes: Restore disk → vSAN auto-rebalances (watch throughput drop to 0)
Validation Gate
Check: Run vSAN health check, identify any warnings/errors, verify resync status (if any), and confirm all objects have full component count
Expected: vSAN health: all green (or yellow with known/documented exceptions). Resync: no active resyncs (or resync completing with expected ETA). Object health: all objects at full component count per storage policy.
Common Errors
Final Validation
Lab completed successfully
✓ All steps completed → No errors observed
Cleanup / Restore
• Revert to snapshot if needed
Design Reflection (VCDX)
vSAN operational health is a common VCDX defense topic. Panelists test whether you understand resync behavior, health check interpretation, and when manual intervention is appropriate vs when to let vSAN self-heal.
Requirements
- Monitor vSAN health checks continuously
- Understand resync behavior and timing
- Interpret object health status correctly
Constraints
- Do not intervene during active resync
- Resync consumes network bandwidth — may impact workload I/O
- Health checks must be reviewed proactively, not just during incidents
Assumptions
- vSAN health check service is running and alerts are forwarded
- Network bandwidth is sufficient for resync operations
Risks
- Data loss from second failure during reduced-availability state
- Resync impacting workload I/O during peak hours