Academy/vSphere Foundation 9.0 Support (2V0-18.25)/Lab: vSAN Health Check & Resync Monitoring
This lab targets VCF 9.0

Lab: vSAN Health Check & Resync Monitoring

VCF 9.0Intermediatevcp-foundation⏱ 105 min

Objectives

  • Lab: vSAN Health Check & Resync Monitoring

Prerequisites

VCF lab environment deployed and operational

Lab Environment

Standard VCF lab environment for vSphere Foundation 9.0 Support (Support Specialist)

Tasks

Task 1 Lab: vSAN Health Check & Resync Monitoring

vSAN health checks and resync monitoring are Day-2 operational essentials. Understanding what each health check tests, what resync means, and when to intervene vs when to let vSAN self-heal separates skilled operators from reactive firefighters.
Step 1
vCenter UI → Cluster → vSAN → Health Check → Run Now
Step 2

Review results: Note any yellow/red alerts

Step 3
Check network latency: Health Check → Network → Latency results (should be <2ms between hosts)
Step 4
If MTU check fails: vCenter → Distributed Switch → vSAN port group → Edit → MTU field → Verify 9000
Step 5
Simulate disk failure: Shutdown host → Remove one capacity disk from vSAN disk group (via web UI)
Step 6
Monitor resync: vCenter → Cluster → Performance → Watch "Resync throughput" graph for >0 MB/s
Step 7
After resync completes: Restore disk → vSAN auto-rebalances (watch throughput drop to 0)

Validation Gate

Check: Run vSAN health check, identify any warnings/errors, verify resync status (if any), and confirm all objects have full component count

Expected: vSAN health: all green (or yellow with known/documented exceptions). Resync: no active resyncs (or resync completing with expected ETA). Object health: all objects at full component count per storage policy.

Common Errors

Panicking during vSAN resync and attempting manual intervention
Fix: vSAN resync is automatic — after a host failure, disk failure, or policy change, vSAN rebuilds/moves components automatically. DO NOT add/remove hosts or change storage policies during resync. Monitor progress via: esxcli vsan debug resync summary. Typical resync: 10-60 minutes depending on data volume and network bandwidth.
Ignoring 'Objects with reduced availability' warning
Fix: This warning means some vSAN objects are running on fewer copies than the configured FTT. For RAID-1 FTT=1: objects should have 2 copies. If showing 1 copy, one mirror is missing — likely a host or disk failure is in progress. Check host status immediately. A second failure while in this state causes data loss.
Not understanding vSAN health check categories
Fix: vSAN health checks are grouped: Hardware compatibility (HCL check), Network (multicast, MTU), Data (object health, disk balance), Cluster (capacity, member health), Stretch cluster (witness, preferred domain). Focus on Data and Network categories — these indicate active issues. Hardware compatibility warnings can be addressed during the next maintenance window.
Running vSAN health check only during incidents instead of proactively
Fix: vSAN health check should run continuously (it does by default) with alerts forwarded to monitoring. Don't wait for an incident to check health. Schedule weekly health review: check disk balance, capacity utilization trending, and any persistent warnings. Address warnings before they escalate during a failure event.

Final Validation

Lab completed successfully

✓ All steps completed → No errors observed

Cleanup / Restore

• Revert to snapshot if needed

Design Reflection (VCDX)

vSAN operational health is a common VCDX defense topic. Panelists test whether you understand resync behavior, health check interpretation, and when manual intervention is appropriate vs when to let vSAN self-heal.

Requirements

  • Monitor vSAN health checks continuously
  • Understand resync behavior and timing
  • Interpret object health status correctly

Constraints

  • Do not intervene during active resync
  • Resync consumes network bandwidth — may impact workload I/O
  • Health checks must be reviewed proactively, not just during incidents

Assumptions

  • vSAN health check service is running and alerts are forwarded
  • Network bandwidth is sufficient for resync operations

Risks

  • Data loss from second failure during reduced-availability state
  • Resync impacting workload I/O during peak hours

⚠ Known Pitfalls (from Community KB)

Manually intervening during vSAN resync — adding/removing hosts or changing policies disrupts the automatic rebuild.
Ignoring 'reduced availability' warnings — this is the critical window where a second failure causes data loss.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.