Multi-Site vSAN Stretched Cluster Design
Objectives
- Design a vSAN stretched cluster across two data center sites with witness placement
- Calculate network requirements for synchronous replication between sites
- Define site affinity policies and failure domain boundaries
- Analyze partition and failure scenarios with recovery procedures
- Size witness appliance using official Tiny/Medium/Large/XL tiers
Prerequisites
Access to VCF 9.0 multi-site architecture documentation
Prior labs: vcp-architect-02, vcp-architect-03
Required skills:
- vSAN storage architecture (ESA)
- Layer 2/Layer 3 networking between sites
- HA and DRS concepts
- Witness appliance fundamentals
Lab Environment
Design exercise — can be partially validated on Holodeck with nested vSAN stretched cluster (requires 3 distinct failure domains)
Tasks
Task 1 Stretched Cluster Network Requirements & Site Assessment
Validate that the inter-site network meets the strict requirements for vSAN stretched cluster before proceeding with design.
Document network requirements for vSAN stretched cluster (VCF 9.0):
- Latency requirements:
- vSAN data traffic: ≤5ms RTT between sites (hard requirement)
- Recommended: ≤1ms RTT for optimal performance
- Witness traffic: ≤200ms RTT to each site (witness is metadata only)
- vMotion: ≤150ms RTT (for cross-site vMotion during maintenance)
- Bandwidth requirements:
- vSAN replication: Calculate based on write rate
Formula: Write IOPS × average I/O size × FTT multiplier = bandwidth needed
Example: 10,000 write IOPS × 32KB × 2 (mirror) = 640 MB/s = ~5.1 Gbps
- Reserve: 2× calculated bandwidth for burst handling
- Minimum: 10 Gbps dedicated inter-site link (25 Gbps recommended)
- Network topology:
- vSAN VMkernel: Must be routable between sites (Layer 3 with or without L2 stretch)
- vMotion VMkernel: Must be routable between sites
- Management VMkernel: Routable (for vCenter, SDDC Manager communication)
- NSX Geneve overlay: ~54 bytes overhead; MTU minimum 1600, recommended 1700
Assess inter-site connectivity:
Site topology:
- Site A (Primary): Production data center
- Site B (Secondary): DR data center, 30km distance
- Site C (Witness): Corporate office or colocation (different fault domain)
Connectivity matrix:
| Path | Distance | Link Type | Bandwidth | Measured RTT | Requirement | Status |
|---|---|---|---|---|---|---|
| A↔B | 30km | Dark fiber | 100 Gbps | 0.8ms | ≤5ms | PASS |
| A↔C | 50km | MPLS | 1 Gbps | 2.5ms | ≤200ms | PASS |
| B↔C | 60km | MPLS | 1 Gbps | 3.0ms | ≤200ms | PASS |
Redundancy:
- A↔B: Dual path (primary dark fiber + backup MPLS)
- Failover: BGP-based, sub-second with BFD (default in NSX, 500ms×3 = 1.5s detection)
Document: If measured RTT exceeds 5ms, stretched cluster is NOT viable — use active-passive DR instead.
Design network segmentation for stretched cluster:
- VLAN design:
- vSAN traffic: Dedicated VLAN stretched across sites (or routed with static routes)
- vMotion traffic: Dedicated VLAN (can be L3 routed between sites)
- Management: Dedicated VLAN
- Workload overlay: NSX Geneve tunnels (no VLAN stretching needed)
- L2 stretching considerations:
- Option A: L2 stretch via VXLAN/OTV for vSAN VMkernel — simplifies IP management
Risk: STP issues, broadcast storms, split-brain
- Option B: L3 routing for vSAN VMkernel — each site has own subnet
Recommended: vSAN supports L3 routing between sites since vSAN 7.0u2
Configuration: Static routes or BGP for vSAN traffic
- Design decision: L3 routed vSAN (no L2 stretch)
Justification: Eliminates STP risk, supports future site additions, aligns with NSX overlay model
Document physical network design:
- ToR switches per site:
- 2× leaf switches per rack (redundancy)
- 25 GbE to each ESXi host (minimum 4 ports: 2 per leaf)
- Uplinks: 100 GbE to spine switches
- Inter-site links:
- Primary: 2× 100 GbE dark fiber (LAG/LACP)
- Secondary: 1× 10 GbE MPLS (failover only)
- Traffic engineering: vSAN traffic on primary; management/vMotion on secondary
- MTU configuration:
- End-to-end jumbo frames: MTU 9000 on physical switches
- vSAN VMkernel: MTU 9000
- NSX Geneve: MTU 1700 minimum (accounts for ~54 byte Geneve overhead on 1500-byte inner frame)
- Verify: No MTU mismatch at any hop (ping -s 8972 between vSAN VMkernel IPs)
Validation Gate
Check: Network assessment completed with latency measurements, bandwidth calculations, and L2/L3 routing decision documented
Expected: All inter-site paths verified against requirements with pass/fail status; bandwidth sized for vSAN write rate; L3 routing design justified
Common Errors
Task 2 Witness Appliance Sizing & Placement
Select the correct witness appliance size and place it in a failure domain independent of both data sites.
Document witness appliance sizing tiers (VCF 9.0 / vSAN 9.0):
The witness is deployed as an OVA appliance with predefined tiers:
| Tier | Max Components | vCPU | RAM | Storage | Use Case |
|---|---|---|---|---|---|
| Tiny | ≤750 | 2 | 8 GB | 15 GB | Small clusters, ≤10 hosts |
| Medium | ≤21,833 | 2 | 16 GB | 350 GB | Most production, ≤34 hosts |
| Large | ≤45,000 | 2 | 32 GB | 700 GB | Large clusters, ≤64 hosts |
| Extra Large | ≤64,000 | 4 | 64 GB | 3.3 TB | Maximum scale |
Component count estimation:
- Each VM object creates 2-4 components (depending on FTT)
- Formula: VMs × average_VMDKs × (FTT+1) × 2 (for witness metadata + data components)
- Example: 200 VMs × 2 VMDKs × 2 (FTT=1) × 2 = 1,600 components → Tiny sufficient - With growth: 200 × 1.2³ × 2 × 2 × 2 = 2,765 → Medium recommended
Design witness placement:
Placement requirements:
- Must be in a THIRD fault domain (not Site A or Site B)
- Purpose: In a site partition, witness casts the deciding vote for data availability
- If witness is co-located with Site A: Site A always wins the vote, Site B VMs become unavailable in any partition
- Placement options:
a) Corporate office (Site C) — Recommended if ≤200ms RTT to both data sites
b) Cloud (AWS/Azure/GCP) — Supported via Witness OVA on cloud VM; adds cloud cost
c) Colocation facility — Best for independence; higher cost
d) Third data center — Ideal but requires existing infrastructure
- Witness does NOT store VM data — only metadata (component placement, configuration)
- Bandwidth requirement: <5 Mbps typically (metadata only)
- Latency: ≤200ms RTT to each data site
- Witness HA:
- Deploy witness on infrastructure NOT managed by the stretched cluster it protects
- Run on separate vCenter or standalone ESXi at witness site
- If witness fails: Cluster continues operating normally until a site fails + witness is down simultaneously
Configure site affinity policies:
vSAN stretched cluster site affinity:
- Preferred site: Site A (primary)
- In a partition without witness, VMs on the preferred site continue running
- VMs on the non-preferred site are powered off to prevent split-brain
- VM-site affinity rules:
- Site A affinity group: VMs that should primarily run on Site A hosts
- Site B affinity group: VMs that should primarily run on Site B hosts
- No affinity: VMs that can run on either site (DRS decides)
- Data locality:
- vSAN places replica components on the same site as the VM (site-aware)
- Read-local policy: VMs read from local replica (no cross-site reads for normal I/O)
- Writes: Synchronous to both sites (this is the RTT-sensitive path)
- Design for 600-VM environment:
- Site A: 350 VMs (Tier-1 primary + Tier-2 subset)
- Site B: 250 VMs (Tier-2 remainder + Tier-3)
- Cross-site vMotion: Used only for planned maintenance, not routine DRS
Design stretched cluster capacity for single-site failure:
Capacity rule: Each site must have enough capacity to run ALL VMs when the other site fails.
Site A capacity:
- 8 hosts × 64 cores × 512 GB RAM = 512 cores, 4 TB RAM
- Under normal operation: 350 VMs (Tier-1 + Tier-2 subset)
- During Site B failure: Must accommodate 250 additional VMs from Site B
- Total: 600 VMs requiring ~640 vCPU, ~4.5 TB RAM (with Tier-2/3 overcommit)
- Fit check: 512 cores with applicable overcommit → sufficient
Site B capacity:
- 8 hosts × 64 cores × 512 GB RAM = 512 cores, 4 TB RAM
- Must accommodate 350 additional VMs from Site A during failure
Storage:
- FTT=1 site mirroring: Each site stores a full copy of all data
- Usable capacity per site must fit total dataset
- With failure: No protection degradation (data still mirrored, just missing one site copy)
- Rebuild: When failed site recovers, full resync occurs (plan for network impact)
Total host count: 16 data hosts + 1 witness = 17 hosts for stretched cluster
Validation Gate
Check: Witness sized using official OVA tiers, placed in independent fault domain, site affinity configured
Expected: Witness tier selected with component count calculation, placement in third site with <200ms RTT verified, site affinity rules defined, single-site failure capacity validated
Common Errors
Task 3 Partition & Failure Scenario Analysis
Analyze every meaningful failure scenario in a stretched cluster and document the expected behavior and recovery procedure.
Build a failure scenario matrix:
| # | Failure | Sites Available | Witness | Preferred Site | VM Behavior | Data Status |
|---|---|---|---|---|---|---|
| 1 | Site A fails | B + Witness | Up | A (down) | Site B VMs stay up; Site A VMs restart on Site B via HA | Data accessible (B has full copy) |
| 2 | Site B fails | A + Witness | Up | A (up) | Site A VMs stay up; Site B VMs restart on Site A via HA | Data accessible (A has full copy) |
| 3 | Witness fails | A + B | Down | N/A | All VMs continue normally on both sites | Data accessible (both copies intact); NO new protection until witness recovers |
| 4 | Site A + Witness fail | B only | Down | A (down) | Site B VMs stay up; Site A VMs CANNOT restart (no quorum for A's components) | Partial — only Site B local data accessible |
| 5 | Network partition A↔B (witness reachable from both) | A + Witness, B + Witness | Up | A | Site A VMs continue; Site B VMs continue; Writes quorum via witness | Data accessible on both sides via witness |
| 6 | Network partition A↔B (witness reachable from A only) | A + Witness | Up (A side) | A | Site A VMs continue; Site B VMs powered off (no quorum) | Site A data accessible; Site B isolated |
| 7 | All sites fail | None | Down | N/A | Full outage | Data intact on disk; recovery requires at least 1 site + witness |
Deep-dive Scenario 4 (Worst case: Site A + Witness fail simultaneously):
This is the hardest VCDX defense question. Walk through step by step:
- Site A fails: Hosts power off, local data copies unavailable
- Witness fails simultaneously: No quorum vote possible
- Site B hosts detect:
- Lost connectivity to Site A hosts (HA isolation detection)
- Lost connectivity to witness (cannot determine partition vs failure)
- vSAN cannot confirm whether Site A is truly down or just partitioned
- Result:
- VMs that were running on Site B hosts: Continue running (local compute still works)
- Data components: Only Site B copies available; vSAN marks Site A components as absent
- WITHOUT quorum: vSAN cannot promote Site B components to primary → data READ-ONLY or INACCESSIBLE for objects where Site B doesn't hold the primary copy 5. Recovery: - Option A: Restore witness first → quorum restored → vSAN promotes Site B components → full read/write
- Option B: If witness cannot be restored, use vSAN force recovery (CMMDS partition repair) — RISK: potential data inconsistency
Document: This scenario requires witness HA at the witness site (redundant power, network). It is the architectural justification for witness placement in a highly-available facility.
Design recovery procedures for each scenario:
Recovery from Scenario 1 (Site A failure):
- Immediate: HA restarts Site A VMs on Site B (automatic)
- Capacity check: Verify Site B can handle combined workload (may need to power off non-critical Tier-3 VMs)
- Site A recovery: When Site A comes back online, vSAN detects stale components
4. Resync: vSAN resyncs all stale components from Site B → Site A
- Duration: Depends on data volume and inter-site bandwidth
- Example: 50 TB of data, 80% changed, 10 Gbps link = ~9 hours
- Post-resync: DRS rebalances VMs back to original site affinity
- Verify: vSAN health shows all components in compliance
Recovery from Scenario 6 (Network partition, witness on A side):
- Site B VMs powered off by vSAN (no quorum)
- Do NOT manually restart Site B VMs — this risks split-brain with conflicting writes
- Fix network partition
- vSAN automatically reconciles and restarts Site B VMs
- Verify: No object version conflicts in vSAN health
Create a stretched cluster failure runbook:
- Detection:
- Monitor: vSAN Health → Stretched Cluster → Site connectivity - Alert: Any site connectivity loss → Critical alert → page on-call
- Dashboard: Aria Operations stretched cluster widget showing site status
- Decision tree:
- Is it a site failure or network partition?
→ Check witness connectivity to both sites
→ Check out-of-band connectivity (IPMI/iLO) to failed site hosts
- Is data accessible?
→ Check vSAN health → Object Health
→ Count accessible vs inaccessible objects
- Can we fail back safely?
→ Resync complete? (vSAN health → Resync dashboard)
→ All components healthy?- Escalation:
- Level 1: On-call verifies scenario, checks vSAN health
- Level 2: VMware Support engagement (GSS SR, Severity 1 for data loss risk)
- Level 3: vSAN force recovery (only with VMware Support guidance)
- Post-incident:
- Root cause analysis within 48 hours
- Update risk register if scenario was not previously anticipated
- Test witness failover to verify future resilience
Validation Gate
Check: All 7 failure scenarios documented with expected behavior and recovery procedures
Expected: Failure matrix covers all combinations of site/witness failures, deep-dive on worst-case scenario (Site A + Witness), recovery runbook with decision tree and escalation path
Common Errors
Task 4 Stretched Cluster Design Decision & VCDX Defense
Consolidate the stretched cluster design into defensible design decisions and prepare for VCDX panel challenges.
Create design decision D-008 — Multi-Site Architecture Pattern:
Decision: vSAN Stretched Cluster (synchronous active-active)
Alternatives considered:
A) Active-Passive with SRM:
- RPO: Minutes to hours (asynchronous replication)
- RTO: 30-60 minutes (SRM recovery plan execution)
- Pro: Simpler, no inter-site latency sensitivity
- Con: Data loss up to RPO; manual/semi-automatic failover
- Rejected: Customer requires RPO=0 for Tier-1 applications
B) Active-Active with HCX:
- HCX provides VM mobility but not storage replication
- Would need separate storage replication (e.g., array-based)
- Rejected: Adds complexity, violates single-vendor constraint
C) Pilot Light DR:
- Minimal infrastructure at DR site, powered up on demand
- RTO: 2-4 hours (power up infrastructure + restore from backup)
- Pro: Lowest cost
- Con: High RTO, manual process
- Rejected: Does not meet 4-hour RTO for 600 VMs
D) VMware Site Recovery (vSphere Replication):
- RPO: 5 minutes to 24 hours (configurable)
- RTO: 30-60 minutes
- Pro: Simpler than stretched cluster, no latency requirements
- Acceptable alternative if RPO=0 is relaxed to RPO=5min
Justification: Stretched cluster is the only VCF-native option providing RPO=0 with automatic failover.
Document total cost comparison:
| Item | Stretched Cluster | Active-Passive SRM |
|---|---|---|
| Data site hosts | 16 (8+8) | 12 (8 primary + 4 DR) |
| Witness | 1 appliance | N/A |
| Inter-site bandwidth | 10-25 Gbps (dark fiber) | 1-10 Gbps (asynchronous) |
| VCF licensing | 16 × 64 cores = 1,024 cores | 12 × 64 cores = 768 cores |
| RPO | 0 (synchronous) | 5-60 minutes |
| RTO | Seconds (automatic HA) | 30-60 minutes (SRM plan) |
| Operational complexity | High (witness management, partition handling) | Medium (replication monitoring, DR testing) |
Cost premium for RPO=0:
- Additional hosts: 4 × host cost
- Additional licensing: 256 × per-core cost
- Network: Dark fiber lease vs MPLS
- Total estimated premium: 30-40% over active-passive
Document: Is RPO=0 worth the cost premium? Map back to business requirement.
Prepare VCDX defense for stretched cluster challenges:
Challenge 1: 'Stretched clusters add complexity. Why not just use SRM?'
Response: Customer requirement R-001 specifies RPO=0 for Tier-1 applications. SRM minimum RPO is 5 minutes. The business case (financial trading / healthcare) cannot tolerate any data loss. However, for Tier-2/3, I would recommend SRM or vSphere Replication as a cost-effective alternative.
Challenge 2: 'What if the inter-site link latency degrades beyond 5ms?'
Response: vSAN will continue operating but with degraded write performance (every write waits for cross-site acknowledgment). If sustained >10ms, I would recommend converting to asynchronous replication (SRM) and accepting RPO>0. The design includes monitoring alerts at 3ms (warning) and 5ms (critical) thresholds.
Challenge 3: 'Your witness site is a corporate office — what if it loses power?'
Response: The witness only provides quorum votes. If the witness fails while both data sites are healthy, the cluster continues normally (Scenario 3 in the failure matrix). The risk window is witness failure + simultaneous site failure. Mitigation: UPS at witness site with 4-hour runtime, plus documented procedure for deploying replacement witness from OVA.
Challenge 4: 'How do you handle a planned site-wide maintenance (e.g., power shutdown at Site A)?'
Response: Planned failover procedure: (1) DRS evacuates all VMs to Site B, (2) Enter maintenance mode on all Site A hosts with 'Ensure data accessibility', (3) Power down Site A, (4) During maintenance: Site B runs at full capacity with witness providing quorum, (5) Reverse process to restore. Total planned failover time: ~2 hours for 350 VMs.
Create the final stretched cluster design summary:
- Architecture:
- 2 data sites (8+8 hosts) + 1 witness site (Medium OVA)
- vSAN ESA with FTT=1 site mirroring (data copy at each site)
- L3 routed inter-site connectivity (no L2 stretch)
- NSX overlay for workload networking
- Network:
- Inter-site: 2×100 GbE dark fiber (primary) + 1×10 GbE MPLS (backup)
- Measured RTT: 0.8ms (well within 5ms requirement)
- MTU: 9000 end-to-end, 1700 for Geneve overlay
- Witness:
- Tier: Medium (supports up to 21,833 components)
- Placement: Corporate office (Site C), 50km from both sites
- Connectivity: 1 Gbps MPLS, 2.5ms RTT to Site A, 3.0ms to Site B
- Failure handling:
- 7 scenarios documented with expected behavior
- Preferred site: Site A
- Recovery runbook with decision tree and escalation path
- Monitoring:
- Aria Operations dashboards for inter-site latency, resync progress
- Alerts: RTT >3ms warning, >5ms critical
- Weekly witness health verification
Validation Gate
Check: Stretched cluster design consolidated with cost comparison, alternative analysis, defense responses, and summary
Expected: Design decision entry with 4+ alternatives evaluated, cost comparison table, 4+ defense responses prepared, complete architecture summary
Common Errors
Final Validation
Complete multi-site vSAN stretched cluster design with network validation, witness sizing, failure analysis, and defensible design decisions
✓ Network requirements validated with measurements → RTT, bandwidth, and MTU verified for all inter-site paths
✓ Witness sized and placed correctly → Official OVA tier selected with component count math; third-site placement verified
✓ All failure scenarios documented → 7-scenario matrix with expected behavior, recovery procedures, and escalation path
✓ Design decision defensible → Alternatives evaluated, cost comparison documented, defense responses prepared
Cleanup / Restore
• Save all design documents and failure scenario matrix
• If using Holodeck: Remove stretched cluster configuration and revert to base snapshot
Design Reflection (VCDX)
Stretched clusters are the most challenging multi-site design pattern to defend in VCDX. Panelists test three things: (1) Do you understand the failure modes? The scenario matrix is your strongest tool — walk through each scenario with specific outcomes. (2) Can you justify the cost? Be ready with numbers showing the premium over SRM and why RPO=0 is worth it. (3) Do you know when NOT to use it? If the customer's RPO can relax to 5 minutes, SRM is simpler and cheaper. Architectural judgment means recommending the right pattern, not the most complex one.
Requirements
- RPO=0 for Tier-1 applications (no data loss on site failure)
- RTO in seconds (automatic failover, not manual SRM plan)
- 600 VMs across two sites with full failover capacity at each site
- Data sovereignty: Both copies must be within the same jurisdiction
Constraints
- Inter-site RTT must be ≤5ms for vSAN synchronous replication
- Witness must be in a third fault domain
- Each site must size for full workload (2× host count vs single-site)
- VCF 9.0 requires vSAN ESA for new deployments
Assumptions
- Dark fiber available between sites with <1ms RTT
- Corporate office suitable as witness site (<200ms RTT to both sites)
- Budget approved for 30-40% cost premium over active-passive DR
- Inter-site latency stable (no significant jitter or degradation during business hours)
Risks
- Inter-site link degradation causes vSAN write latency spike affecting all VMs
- Witness site failure coinciding with data site failure causes data unavailability
- Resync after site recovery saturates inter-site link for hours
- Stretched cluster complexity leads to operational errors during failure recovery
Self-Assessment Discussion Prompts
- At what inter-site distance does stretched cluster become impractical, and what alternative would you recommend?
- How would you design a hybrid approach using stretched cluster for Tier-1 and SRM for Tier-2/3?
- If the customer adds a third data site, can vSAN stretched cluster support 3 data sites? If not, what architecture would you use?
- What is the maximum number of VMs you would put in a single stretched cluster, and what drives that limit?
Extensions
Hybrid Multi-Site: Stretched Cluster + SRM
Design a hybrid approach where Tier-1 VMs use stretched cluster (RPO=0) and Tier-2/3 VMs use SRM with asynchronous replication (RPO=15min). Document the separate workload domains, independent vSAN datastores, and unified Aria Operations monitoring across both patterns.
Stretched Cluster with vSAN File Services
Add vSAN File Services (NFS/SMB shares) to the stretched cluster design. Evaluate how file service failover differs from VM failover, document the impact on witness component count, and design the client reconnection procedure.
Holodeck Stretched Cluster Simulation
Build a 6-node (3+3) stretched cluster on Holodeck with a witness VM. Simulate network partition by disabling vmnic between sites, observe vSAN behavior, and validate the failure scenario matrix from Task 3 against actual results.
⚠ Known Pitfalls (from Community KB)
References
- vSAN 9.0 Stretched Cluster Guide — Network Requirements and Witness Sizing
- VMware KB 2131662 — vSAN Stretched Cluster Best Practices
- vSAN Witness Appliance Deployment Guide — OVA Sizing Tiers
- VMware VCF 9.0 Multi-Site Design Considerations
- VMware KB 2108285 — Troubleshooting vSAN Network Performance