Academy/VCF 9.0 Support (2V0-15.25)/vSAN Failure & Recovery Scenario
This lab targets VCF 9.0

vSAN Failure & Recovery Scenario

VCF 9.0Advancedvcp-foundationvcp-support⏱ 120 min

Advanced vSAN failure scenarios including host failure, network partition, and capacity emergency procedures.

Objectives

  • Diagnose vSAN host failure impact on object availability and cluster capacity
  • Interpret vSAN disk state transitions and rebuild timers in ESA architecture
  • Execute vSAN rebuild rate tuning to balance recovery speed vs. VM performance
  • Recover from vSAN network partition scenarios (split-brain detection and resolution)
  • Design vSAN capacity alert thresholds to prevent data-at-risk emergencies
  • Document vSAN failure recovery using RCAR methodology

Prerequisites

VCF lab with 4-host vSAN ESA cluster running production-like workloads

Prior labs: vcp-support-02

Required skills:

  • esxcli vsan commands
  • vSAN health check interpretation
  • ESXi SSH access

Lab Environment

VCF workload domain with 4-host vSAN ESA cluster, FTT=1 RAID-1, mixed workloads

Tasks

Task 1 Host Failure Impact Analysis

Builds from disk failure (Lab 02) to host failure. Panelists probe: What is the difference between host failure and disk failure for vSAN? How many hosts can you lose with FTT=1?

Understand the cascading impact of a complete ESXi host failure on vSAN object availability, VM placement, and cluster capacity.

Step 1

Document the pre-failure cluster state. For each host: hostname, number of VMs, vSAN disk count, vSAN capacity contribution, vSAN components hosted. Use: esxcli vsan debug disk list (or: esxcli vsan debug disk summary get) and vSphere Client Cluster Monitor vSAN Capacity view.

4-host cluster baseline: each host contributes ~25% of total capacity and hosts ~25% of total components. With FTT=1 RAID-1, each object has mirror copies on 2 different hosts plus a witness on a third.
Baseline documentation is critical for validating recovery completeness.
Step 2

Simulate complete host failure. In Holodeck, power off one ESXi host VM abruptly. Observe: vSphere HA detects host failure (~30s), begins restarting VMs on surviving hosts, vSAN marks all components on the failed host as Absent, 60-minute repair timer starts.

Two parallel recovery processes: vSphere HA restarts VMs (5-10 min) and vSAN starts 60-minute timer for component rebuild. VMs restart from mirror copies on surviving hosts. During 60-minute window, affected objects operate at FTT=0 (unprotected).
vSphere HA and vSAN recovery are independent. HA restarts VMs immediately. vSAN repair waits 60 minutes then rebuilds components.
Step 3

Calculate capacity impact. With one host down: raw capacity drops 25%, remaining 3 hosts must accommodate all VMs. Verify HA admission control allows all VMs to restart. Document any VMs that cannot restart due to insufficient resources.

3 remaining hosts must: run all VMs from the failed host (compute), have sufficient vSAN capacity for rebuild (storage), meet HA admission control requirements. If admission control tolerates 1 host failure, 33% of resources are reserved.
HA admission control must align with vSAN FTT policy. Both must be configured for the same failure tolerance level.
Step 4

Monitor 60-minute absent timer and subsequent rebuild. After timer expires, vSAN rebuilds components from surviving mirrors. Track: esxcli vsan debug resync summary (rebuild progress), total data to rebuild, rebuild rate per host, estimated completion time.

Rebuild begins after 60-minute timer. Data volume: all unique data from failed host (~25% of total). Rebuild distributes across remaining 3 hosts. Performance impact: 10-30% VM I/O latency increase during active rebuild.
Rebuild rate tuning: increase during off-peak for faster recovery, decrease during business hours for less VM impact.

Validation Gate

Check: Complete host failure simulation with impact analysis

Expected: Host failure simulated, HA restart verified, capacity impact calculated, rebuild timeline tracked.

Common Errors

Not accounting for HA admission control when calculating host failure tolerance
Confusing 60-minute absent timer with actual repair time (timer is just delay before repair starts)
Forgetting FTT=1 objects are unprotected during rebuild window
Not monitoring rebuild performance impact on running VMs

Task 2 vSAN Network Partition & Split-Brain Resolution

Advanced troubleshooting. Panelists probe: What is vSAN split-brain? How does vSAN handle network partitions? What role does the witness play?

Diagnose and recover from vSAN network partition scenarios where hosts lose connectivity but remain operational.

Step 1

Understand vSAN network partition behavior. When vSAN network connectivity is lost between hosts: each host accesses local components, cannot access remote components. CMMDS detects partition, objects with majority components accessible remain available. Document the CMMDS quorum rules for FTT=1 RAID-1: 2 data components + 1 witness = 3 votes.

Object remains accessible if partition with more than 50% of vote-carrying components can form quorum. Witness acts as tiebreaker. Partition with 2 of 3 votes wins quorum.
Witness placement is critical for partition resilience. Distribute witnesses across fault domains.
Step 2

Simulate network partition in Holodeck. Isolate one ESXi host by disconnecting its vSAN VMkernel port. Observe: isolated host still runs VMs, objects with majority on remaining 3 hosts remain accessible, objects with both data and witness on isolated host become inaccessible on main cluster side.

Network partition creates two sides: isolated (1 node) and remaining (3 nodes). Most objects have majority on 3-node side and remain accessible. vSAN prevents split-brain through quorum enforcement.
Network partition differs from host failure: isolated host is still running VMs, creating potential split-brain.
Step 3

Resolve the partition. Reconnect the vSAN VMkernel port. Monitor: CMMDS re-establishes cluster membership, vSAN merges partition with component resync, all objects return to healthy state. Track with esxcli vsan debug resync summary.

Partition resolution triggers resync of potentially diverged components. Short partitions (<5 min) resync in seconds. Long partitions with heavy I/O may take minutes to hours.
After resolving a partition, run full vSAN health check to verify no stale components remain.
Step 4

Design vSAN network resilience: redundant VMkernel ports on separate physical NICs, dedicated VLAN with QoS, network health monitoring with Aria Operations, alerts for latency >5ms and packet loss >0.01%.

Dual VMkernel ports on separate pNICs for NIC-level redundancy. Dedicated VLAN prevents broadcast storms. QoS prevents traffic starvation. Monitoring detects degradation before partition.
VCDX design: vSAN network design is as important as storage design. Network issues cause more outages than disk failures in production.

Validation Gate

Check: Simulate and resolve vSAN network partition

Expected: Partition simulated, quorum behavior observed, partition resolved, network resilience documented.

Common Errors

Confusing network partition with host failure
Not understanding CMMDS quorum rules
Reconnecting partitioned host without monitoring resync
Single vSAN VMkernel port with no redundancy

Task 3 Capacity Emergency Procedures & Rebuild Rate Tuning

Crisis management. Panelists ask: What do you do at 95% capacity after host failure? How do you prioritize VMs?

Handle vSAN capacity emergencies when failures push the cluster above safe thresholds.

Step 1

Define vSAN capacity zones: Zone 1 (0-70%) normal; Zone 2 (70-80%) warning, plan expansion; Zone 3 (80-85%) critical, stop provisioning; Zone 4 (85-95%) emergency, reduce footprint; Zone 5 (>95%) crisis, immediate intervention.

Five zones with escalating response actions mapped to vSAN health alerts.
Define zones BEFORE a crisis. When cluster hits 90%, follow pre-defined procedures.
Step 2

Practice resync monitoring. View progress: esxcli vsan debug resync summary get. Note: 'esxcli vsan policy getdefault' shows default storage policy, not resync rates — OSA throttling is set in the vSphere Client (Resyncing Objects > Resync Throttling); ESA uses adaptive resync automatically. Monitor impact with esxtop disk view. Document tradeoff: 2x rebuild rate = ~50% faster rebuild but 15-25% VM latency increase.

Rebuild rate controls I/O bandwidth for rebuilds. Higher rate = faster MTTR but more VM impact. Decision depends on time of day, failure severity, and SLA requirements.
Adaptive rate: increase during off-peak, decrease during business hours.
Step 3

Practice emergency capacity recovery options: (a) power off idle VMs (safe, immediate), (b) thin provision conversion (safe, slow), (c) FTT downgrade for non-critical VMs to FTT=0 (dangerous, immediate), (d) add storage/host (safe, requires procurement).

Options ranked by risk. FTT=0 downgrade is last resort; removes all data protection. Apply only to dev/test VMs with 48-hour deadline to restore.
FTT=0 must have documented risk acceptance and time limit.
Step 4

Calculate minimum cluster size: FTT=1 RAID-1 minimum 3 hosts (no rebuild capacity), operational minimum 4 hosts (rebuild after 1 failure). FTT=2 RAID-1 minimum 5 hosts. Formula: hosts_needed = FTT_minimum + rebuild_buffer + growth.

Cluster sizing with all variables. Key: minimum for FTT compliance differs from minimum for operational resilience.
VCDX: always design for N+1 operational capacity.

Validation Gate

Check: Complete capacity emergency procedures

Expected: Capacity zones defined, rebuild rate tuned, emergency procedures documented, sizing formula applied.

Common Errors

Running above 80% in steady state
Not tuning rebuild rate by time of day
FTT=0 for critical VMs
Sizing to exact FTT minimums without rebuild buffer

Task 4 vSAN Recovery Runbook & VCDX Defense

Synthesis task. Panelists expect structured runbook, SLA analysis, and design defense.

Create comprehensive recovery runbook and prepare VCDX defense for vSAN design decisions.

Step 1

Build incident classification: Sev1 data loss/inaccessibility (immediate), Sev2 degraded/unprotected (1 hour), Sev3 single disk failure with FTT maintained (4 hours), Sev4 warning condition (24 hours). Document response times and escalation paths.

4 severity levels mapped to response SLAs and escalation. Sev1 triggers emergency change management.
Align vSAN severities with organizational ITSM P1-P4 priorities.
Step 2

Create recovery flowchart with 4 entry points: disk failure, host failure, network partition, capacity emergency. Each path: detection, classification, impact assessment, mitigation, recovery, verification, post-incident review.

Flowchart with decision diamonds for escalation at each stage. Usable by L2 support without deep vSAN expertise.
Clear decision criteria and specific commands make this usable under pressure.
Step 3

Calculate SLA impact: disk failure = 5h exposure at FTT=0, host failure = 7h exposure, network partition = variable. Map to SLA targets: 99.9% = 8.76h/year, 99.99% = 52.6min/year.

SLA table showing single-failure tolerance within targets. Risk window = time between failure and rebuild completion at FTT=0.
Quantify the risk window for VCDX defense.
Step 4

Prepare VCDX defense responses: (1) Why 4 hosts not 3 for FTT=1 — rebuild capacity after failure. (2) Monitoring strategy — five layers: SMART, object health, capacity, performance, network. (3) Stretched cluster failure handling — witness determines quorum. (4) Largest vSAN risk — silent data corruption from firmware bugs; mitigate with HCL compliance and integrity scrub.

Four defense responses with quantitative justification.
Lead with business impact, then technical solution, then risk mitigation.

Validation Gate

Check: Complete vSAN recovery runbook and VCDX defense

Expected: Incident classification, recovery flowchart, SLA analysis, and four defense responses completed.

Common Errors

Recovery procedures requiring architect escalation for routine ops
Not quantifying FTT=0 exposure window
Ignoring network-related vSAN failures
No post-incident review process

Final Validation

Complete vSAN failure & recovery with advanced scenarios and VCDX defense

✓ Host failure analyzed → Capacity impact, HA interaction, rebuild timeline documented

✓ Network partition resolved → Split-brain understood, partition simulated and resolved

✓ Capacity emergency handled → Emergency procedures with rebuild rate tuning

✓ Recovery runbook created → Incident classification, flowchart, SLA analysis

✓ VCDX defense prepared → Four responses with quantitative backing

Cleanup / Restore

• Restore all hosts to powered-on state

• Verify vSAN cluster health green across all checks

• Confirm all objects compliant with storage policy

• Revert to snapshot if needed

Design Reflection (VCDX)

Advanced vSAN failure recovery demonstrates operational maturity. Design connects FTT policies to SLA requirements, implements multi-layer monitoring, and provides structured incident response.

Requirements

  • R-001: vSAN cluster survives single host failure without data loss
  • R-002: VM availability maintained during rebuild
  • R-003: Recovery procedures executable by L2 support

Constraints

  • Minimum 4 hosts for FTT=1 with rebuild capacity
  • vSAN network requires redundant paths on dedicated VLAN
  • Rebuild rate balances recovery speed with VM performance

Assumptions

  • Aria Operations monitoring deployed
  • Operations team trained on incident classification
  • Spare hardware available within SLA window

Risks

  • Second host failure during rebuild = data loss with FTT=1
  • Network partition creates split-brain risk
  • Capacity emergency during host failure requires immediate intervention

Self-Assessment Discussion Prompts

  1. At what cluster size does FTT=2 become cost-effective vs FTT=1 with monitoring?
  2. How would you design recovery differently for stretched cluster vs single-site?
  3. What is the operational impact of ESA vs OSA on failure recovery?
  4. How do you balance rebuild rate with production SLA during business hours?

Extensions

Stretched Cluster Failure Simulation

Deploy vSAN stretched cluster in Holodeck. Simulate site failure, verify automatic failover. Document witness behavior, quorum determination, and site restoration.

vSAN Data Integrity Verification

Enable data integrity scrub. Simulate bit-rot and verify vSAN detects and repairs corruption using mirror copies.

Automated vSAN Incident Response

Build automated workflow: vSAN alert -> classify severity -> execute diagnostics -> notify team -> open change ticket.

⚠ Known Pitfalls (from Community KB)

Not accounting for HA admission control interaction with vSAN host failure — VMs may not restart if compute resources not reserved
Ignoring network partition as vSAN failure scenario — network issues cause more production outages than disk failures
Running clusters above 80% capacity in steady state — no slack for rebuild after failure
Using FTT=0 downgrade as routine capacity tool — creates unacceptable data loss risk; only for documented emergencies

References

  • vSAN 9.0 Troubleshooting Guide — Host Failure and Network Partition Recovery
  • vSAN 9.0 Administration Guide — Rebuild Rate Tuning and Capacity Management
  • VMware KB 2150769 — vSAN Object Health States and Recovery Procedures
  • vSAN ESA Architecture Guide — Storage Pool Recovery
  • VMware VCF 9.0 Operational Guide — vSAN Monitoring Best Practices
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.