Academy/VCP-VCF 9.0 Architect (2V0-13.25)/Multi-Site vSAN Stretched Cluster Design
This lab targets VCF 9.0

Multi-Site vSAN Stretched Cluster Design

VCF 9.0Advancedvcp-foundation⏱ 90 min

VCF 9.0 stretched cluster — witness placement, site affinity, partition handling, and failover design

Objectives

  • Design a vSAN stretched cluster across two data center sites with witness placement
  • Calculate network requirements for synchronous replication between sites
  • Define site affinity policies and failure domain boundaries
  • Analyze partition and failure scenarios with recovery procedures
  • Size witness appliance using official Tiny/Medium/Large/XL tiers

Prerequisites

Access to VCF 9.0 multi-site architecture documentation

Prior labs: vcp-architect-02, vcp-architect-03

Required skills:

  • vSAN storage architecture (ESA)
  • Layer 2/Layer 3 networking between sites
  • HA and DRS concepts
  • Witness appliance fundamentals

Lab Environment

Design exercise — can be partially validated on Holodeck with nested vSAN stretched cluster (requires 3 distinct failure domains)

Tasks

Task 1 Stretched Cluster Network Requirements & Site Assessment

Stretched cluster designs fail most often due to network assumptions. VCDX panelists will challenge latency, bandwidth, and L2 stretching decisions — have the numbers ready.

Validate that the inter-site network meets the strict requirements for vSAN stretched cluster before proceeding with design.

Step 1

Document network requirements for vSAN stretched cluster (VCF 9.0):

  1. Latency requirements:
  • vSAN data traffic: ≤5ms RTT between sites (hard requirement)
  • Recommended: ≤1ms RTT for optimal performance
  • Witness traffic: ≤200ms RTT to each site (witness is metadata only)
  • vMotion: ≤150ms RTT (for cross-site vMotion during maintenance)
  1. Bandwidth requirements:
  • vSAN replication: Calculate based on write rate

Formula: Write IOPS × average I/O size × FTT multiplier = bandwidth needed
Example: 10,000 write IOPS × 32KB × 2 (mirror) = 640 MB/s = ~5.1 Gbps

  • Reserve: 2× calculated bandwidth for burst handling
  • Minimum: 10 Gbps dedicated inter-site link (25 Gbps recommended)
  1. Network topology:
  • vSAN VMkernel: Must be routable between sites (Layer 3 with or without L2 stretch)
  • vMotion VMkernel: Must be routable between sites
  • Management VMkernel: Routable (for vCenter, SDDC Manager communication)
  • NSX Geneve overlay: ~54 bytes overhead; MTU minimum 1600, recommended 1700
Step 2

Assess inter-site connectivity:

Site topology:

  • Site A (Primary): Production data center
  • Site B (Secondary): DR data center, 30km distance
  • Site C (Witness): Corporate office or colocation (different fault domain)

Connectivity matrix:

PathDistanceLink TypeBandwidthMeasured RTTRequirementStatus
A↔B30kmDark fiber100 Gbps0.8ms≤5msPASS
A↔C50kmMPLS1 Gbps2.5ms≤200msPASS
B↔C60kmMPLS1 Gbps3.0ms≤200msPASS

Redundancy:

- A↔B: Dual path (primary dark fiber + backup MPLS)
  • Failover: BGP-based, sub-second with BFD (default in NSX, 500ms×3 = 1.5s detection)

Document: If measured RTT exceeds 5ms, stretched cluster is NOT viable — use active-passive DR instead.

Step 3

Design network segmentation for stretched cluster:

  1. VLAN design:
  • vSAN traffic: Dedicated VLAN stretched across sites (or routed with static routes)
  • vMotion traffic: Dedicated VLAN (can be L3 routed between sites)
  • Management: Dedicated VLAN
  • Workload overlay: NSX Geneve tunnels (no VLAN stretching needed)
  1. L2 stretching considerations:
  • Option A: L2 stretch via VXLAN/OTV for vSAN VMkernel — simplifies IP management

Risk: STP issues, broadcast storms, split-brain

  • Option B: L3 routing for vSAN VMkernel — each site has own subnet

Recommended: vSAN supports L3 routing between sites since vSAN 7.0u2
Configuration: Static routes or BGP for vSAN traffic

  1. Design decision: L3 routed vSAN (no L2 stretch)

Justification: Eliminates STP risk, supports future site additions, aligns with NSX overlay model

Step 4

Document physical network design:

  1. ToR switches per site:
  • 2× leaf switches per rack (redundancy)
  • 25 GbE to each ESXi host (minimum 4 ports: 2 per leaf)
  • Uplinks: 100 GbE to spine switches
  1. Inter-site links:
  • Primary: 2× 100 GbE dark fiber (LAG/LACP)
  • Secondary: 1× 10 GbE MPLS (failover only)
  • Traffic engineering: vSAN traffic on primary; management/vMotion on secondary
  1. MTU configuration:
  • End-to-end jumbo frames: MTU 9000 on physical switches
  • vSAN VMkernel: MTU 9000
  • NSX Geneve: MTU 1700 minimum (accounts for ~54 byte Geneve overhead on 1500-byte inner frame)
  • Verify: No MTU mismatch at any hop (ping -s 8972 between vSAN VMkernel IPs)

Validation Gate

Check: Network assessment completed with latency measurements, bandwidth calculations, and L2/L3 routing decision documented

Expected: All inter-site paths verified against requirements with pass/fail status; bandwidth sized for vSAN write rate; L3 routing design justified

Common Errors

Assuming inter-site latency is always acceptable — even 30km fiber can have 2-5ms RTT depending on number of hops and equipment
Forgetting to size bandwidth for vSAN burst writes — average write rate ≠ peak; size for 2× average to handle burst
Stretching L2 for vSAN without evaluating STP risks — a single misconfigured switch can take down the stretched cluster
Not testing MTU end-to-end before deployment — a single hop with 1500 MTU will fragment vSAN traffic and cause performance degradation

Task 2 Witness Appliance Sizing & Placement

Witness sizing is a frequent VCDX challenge question. Candidates who cite generic specs instead of the official OVA tiers lose credibility.

Select the correct witness appliance size and place it in a failure domain independent of both data sites.

Step 1

Document witness appliance sizing tiers (VCF 9.0 / vSAN 9.0):

The witness is deployed as an OVA appliance with predefined tiers:

TierMax ComponentsvCPURAMStorageUse Case
Tiny≤75028 GB15 GBSmall clusters, ≤10 hosts
Medium≤21,833216 GB350 GBMost production, ≤34 hosts
Large≤45,000232 GB700 GBLarge clusters, ≤64 hosts
Extra Large≤64,000464 GB3.3 TBMaximum scale

Component count estimation:

  • Each VM object creates 2-4 components (depending on FTT)
  • Formula: VMs × average_VMDKs × (FTT+1) × 2 (for witness metadata + data components)
- Example: 200 VMs × 2 VMDKs × 2 (FTT=1) × 2 = 1,600 components → Tiny sufficient
- With growth: 200 × 1.2³ × 2 × 2 × 2 = 2,765 → Medium recommended
Step 2

Design witness placement:

Placement requirements:

  1. Must be in a THIRD fault domain (not Site A or Site B)
  • Purpose: In a site partition, witness casts the deciding vote for data availability
  • If witness is co-located with Site A: Site A always wins the vote, Site B VMs become unavailable in any partition
  1. Placement options:

a) Corporate office (Site C) — Recommended if ≤200ms RTT to both data sites
b) Cloud (AWS/Azure/GCP) — Supported via Witness OVA on cloud VM; adds cloud cost
c) Colocation facility — Best for independence; higher cost
d) Third data center — Ideal but requires existing infrastructure

  1. Witness does NOT store VM data — only metadata (component placement, configuration)
  • Bandwidth requirement: <5 Mbps typically (metadata only)
  • Latency: ≤200ms RTT to each data site
  1. Witness HA:
  • Deploy witness on infrastructure NOT managed by the stretched cluster it protects
  • Run on separate vCenter or standalone ESXi at witness site
  • If witness fails: Cluster continues operating normally until a site fails + witness is down simultaneously
Step 3

Configure site affinity policies:

vSAN stretched cluster site affinity:

  1. Preferred site: Site A (primary)
  • In a partition without witness, VMs on the preferred site continue running
  • VMs on the non-preferred site are powered off to prevent split-brain
  1. VM-site affinity rules:
  • Site A affinity group: VMs that should primarily run on Site A hosts
  • Site B affinity group: VMs that should primarily run on Site B hosts
  • No affinity: VMs that can run on either site (DRS decides)
  1. Data locality:
  • vSAN places replica components on the same site as the VM (site-aware)
  • Read-local policy: VMs read from local replica (no cross-site reads for normal I/O)
  • Writes: Synchronous to both sites (this is the RTT-sensitive path)
  1. Design for 600-VM environment:
  • Site A: 350 VMs (Tier-1 primary + Tier-2 subset)
  • Site B: 250 VMs (Tier-2 remainder + Tier-3)
  • Cross-site vMotion: Used only for planned maintenance, not routine DRS
Step 4

Design stretched cluster capacity for single-site failure:

Capacity rule: Each site must have enough capacity to run ALL VMs when the other site fails.

Site A capacity:

  • 8 hosts × 64 cores × 512 GB RAM = 512 cores, 4 TB RAM
  • Under normal operation: 350 VMs (Tier-1 + Tier-2 subset)
  • During Site B failure: Must accommodate 250 additional VMs from Site B
  • Total: 600 VMs requiring ~640 vCPU, ~4.5 TB RAM (with Tier-2/3 overcommit)
- Fit check: 512 cores with applicable overcommit → sufficient

Site B capacity:

  • 8 hosts × 64 cores × 512 GB RAM = 512 cores, 4 TB RAM
  • Must accommodate 350 additional VMs from Site A during failure

Storage:

  • FTT=1 site mirroring: Each site stores a full copy of all data
  • Usable capacity per site must fit total dataset
  • With failure: No protection degradation (data still mirrored, just missing one site copy)
  • Rebuild: When failed site recovers, full resync occurs (plan for network impact)

Total host count: 16 data hosts + 1 witness = 17 hosts for stretched cluster

Validation Gate

Check: Witness sized using official OVA tiers, placed in independent fault domain, site affinity configured

Expected: Witness tier selected with component count calculation, placement in third site with <200ms RTT verified, site affinity rules defined, single-site failure capacity validated

Common Errors

Using generic witness specs (2vCPU/4GB) instead of the official OVA tiers — panelists will check this
Placing witness at the same site as one of the data sites — defeats the purpose of the witness vote
Forgetting that stretched cluster effectively doubles host count — each site needs full failover capacity
Not configuring a preferred site — without it, a partition without witness causes unpredictable behavior

Task 3 Partition & Failure Scenario Analysis

VCDX panelists WILL ask 'what happens when...' questions about stretched clusters. Having a scenario matrix with specific outcomes is the strongest possible defense.

Analyze every meaningful failure scenario in a stretched cluster and document the expected behavior and recovery procedure.

Step 1

Build a failure scenario matrix:

#FailureSites AvailableWitnessPreferred SiteVM BehaviorData Status
1Site A failsB + WitnessUpA (down)Site B VMs stay up; Site A VMs restart on Site B via HAData accessible (B has full copy)
2Site B failsA + WitnessUpA (up)Site A VMs stay up; Site B VMs restart on Site A via HAData accessible (A has full copy)
3Witness failsA + BDownN/AAll VMs continue normally on both sitesData accessible (both copies intact); NO new protection until witness recovers
4Site A + Witness failB onlyDownA (down)Site B VMs stay up; Site A VMs CANNOT restart (no quorum for A's components)Partial — only Site B local data accessible
5Network partition A↔B (witness reachable from both)A + Witness, B + WitnessUpASite A VMs continue; Site B VMs continue; Writes quorum via witnessData accessible on both sides via witness
6Network partition A↔B (witness reachable from A only)A + WitnessUp (A side)ASite A VMs continue; Site B VMs powered off (no quorum)Site A data accessible; Site B isolated
7All sites failNoneDownN/AFull outageData intact on disk; recovery requires at least 1 site + witness
Step 2

Deep-dive Scenario 4 (Worst case: Site A + Witness fail simultaneously):

This is the hardest VCDX defense question. Walk through step by step:

  1. Site A fails: Hosts power off, local data copies unavailable
  2. Witness fails simultaneously: No quorum vote possible
  3. Site B hosts detect:
  • Lost connectivity to Site A hosts (HA isolation detection)
  • Lost connectivity to witness (cannot determine partition vs failure)
  • vSAN cannot confirm whether Site A is truly down or just partitioned
  1. Result:
  • VMs that were running on Site B hosts: Continue running (local compute still works)
  • Data components: Only Site B copies available; vSAN marks Site A components as absent
   - WITHOUT quorum: vSAN cannot promote Site B components to primary → data READ-ONLY or INACCESSIBLE for objects where Site B doesn't hold the primary copy
5. Recovery:
   - Option A: Restore witness first → quorum restored → vSAN promotes Site B components → full read/write
  • Option B: If witness cannot be restored, use vSAN force recovery (CMMDS partition repair) — RISK: potential data inconsistency

Document: This scenario requires witness HA at the witness site (redundant power, network). It is the architectural justification for witness placement in a highly-available facility.

Step 3

Design recovery procedures for each scenario:

Recovery from Scenario 1 (Site A failure):

  1. Immediate: HA restarts Site A VMs on Site B (automatic)
  2. Capacity check: Verify Site B can handle combined workload (may need to power off non-critical Tier-3 VMs)
  3. Site A recovery: When Site A comes back online, vSAN detects stale components
4. Resync: vSAN resyncs all stale components from Site B → Site A
  • Duration: Depends on data volume and inter-site bandwidth
  • Example: 50 TB of data, 80% changed, 10 Gbps link = ~9 hours
  1. Post-resync: DRS rebalances VMs back to original site affinity
  2. Verify: vSAN health shows all components in compliance

Recovery from Scenario 6 (Network partition, witness on A side):

  1. Site B VMs powered off by vSAN (no quorum)
  2. Do NOT manually restart Site B VMs — this risks split-brain with conflicting writes
  3. Fix network partition
  4. vSAN automatically reconciles and restarts Site B VMs
  5. Verify: No object version conflicts in vSAN health
Step 4

Create a stretched cluster failure runbook:

  1. Detection:
   - Monitor: vSAN Health → Stretched Cluster → Site connectivity
   - Alert: Any site connectivity loss → Critical alert → page on-call
  • Dashboard: Aria Operations stretched cluster widget showing site status
  1. Decision tree:
  • Is it a site failure or network partition?
     → Check witness connectivity to both sites
     → Check out-of-band connectivity (IPMI/iLO) to failed site hosts
   - Is data accessible?
     → Check vSAN health → Object Health
     → Count accessible vs inaccessible objects
   - Can we fail back safely?
     → Resync complete? (vSAN health → Resync dashboard)
     → All components healthy?
  1. Escalation:
  • Level 1: On-call verifies scenario, checks vSAN health
  • Level 2: VMware Support engagement (GSS SR, Severity 1 for data loss risk)
  • Level 3: vSAN force recovery (only with VMware Support guidance)
  1. Post-incident:
  • Root cause analysis within 48 hours
  • Update risk register if scenario was not previously anticipated
  • Test witness failover to verify future resilience

Validation Gate

Check: All 7 failure scenarios documented with expected behavior and recovery procedures

Expected: Failure matrix covers all combinations of site/witness failures, deep-dive on worst-case scenario (Site A + Witness), recovery runbook with decision tree and escalation path

Common Errors

Assuming HA automatically handles all stretched cluster failures — HA can only restart VMs if data is accessible (quorum)
Manually force-starting VMs on the isolated site during a partition — this causes split-brain and can lead to data corruption
Not planning for resync bandwidth — a full site recovery with 50+ TB of stale data can saturate the inter-site link for hours
Ignoring the witness-down-during-partition scenario — this is the only scenario where a stretched cluster can lose data availability

Task 4 Stretched Cluster Design Decision & VCDX Defense

Stretched cluster is a polarizing design choice. VCDX panelists may challenge whether it is the right pattern at all — be prepared to defend or propose alternatives.

Consolidate the stretched cluster design into defensible design decisions and prepare for VCDX panel challenges.

Step 1

Create design decision D-008 — Multi-Site Architecture Pattern:

Decision: vSAN Stretched Cluster (synchronous active-active)
Alternatives considered:
A) Active-Passive with SRM:

  • RPO: Minutes to hours (asynchronous replication)
  • RTO: 30-60 minutes (SRM recovery plan execution)
  • Pro: Simpler, no inter-site latency sensitivity
  • Con: Data loss up to RPO; manual/semi-automatic failover
  • Rejected: Customer requires RPO=0 for Tier-1 applications

B) Active-Active with HCX:

  • HCX provides VM mobility but not storage replication
  • Would need separate storage replication (e.g., array-based)
  • Rejected: Adds complexity, violates single-vendor constraint

C) Pilot Light DR:

  • Minimal infrastructure at DR site, powered up on demand
  • RTO: 2-4 hours (power up infrastructure + restore from backup)
  • Pro: Lowest cost
  • Con: High RTO, manual process
  • Rejected: Does not meet 4-hour RTO for 600 VMs

D) VMware Site Recovery (vSphere Replication):

  • RPO: 5 minutes to 24 hours (configurable)
  • RTO: 30-60 minutes
  • Pro: Simpler than stretched cluster, no latency requirements
  • Acceptable alternative if RPO=0 is relaxed to RPO=5min

Justification: Stretched cluster is the only VCF-native option providing RPO=0 with automatic failover.

Step 2

Document total cost comparison:

ItemStretched ClusterActive-Passive SRM
Data site hosts16 (8+8)12 (8 primary + 4 DR)
Witness1 applianceN/A
Inter-site bandwidth10-25 Gbps (dark fiber)1-10 Gbps (asynchronous)
VCF licensing16 × 64 cores = 1,024 cores12 × 64 cores = 768 cores
RPO0 (synchronous)5-60 minutes
RTOSeconds (automatic HA)30-60 minutes (SRM plan)
Operational complexityHigh (witness management, partition handling)Medium (replication monitoring, DR testing)

Cost premium for RPO=0:

  • Additional hosts: 4 × host cost
  • Additional licensing: 256 × per-core cost
  • Network: Dark fiber lease vs MPLS
  • Total estimated premium: 30-40% over active-passive

Document: Is RPO=0 worth the cost premium? Map back to business requirement.

Step 3

Prepare VCDX defense for stretched cluster challenges:

Challenge 1: 'Stretched clusters add complexity. Why not just use SRM?'
Response: Customer requirement R-001 specifies RPO=0 for Tier-1 applications. SRM minimum RPO is 5 minutes. The business case (financial trading / healthcare) cannot tolerate any data loss. However, for Tier-2/3, I would recommend SRM or vSphere Replication as a cost-effective alternative.

Challenge 2: 'What if the inter-site link latency degrades beyond 5ms?'
Response: vSAN will continue operating but with degraded write performance (every write waits for cross-site acknowledgment). If sustained >10ms, I would recommend converting to asynchronous replication (SRM) and accepting RPO>0. The design includes monitoring alerts at 3ms (warning) and 5ms (critical) thresholds.

Challenge 3: 'Your witness site is a corporate office — what if it loses power?'
Response: The witness only provides quorum votes. If the witness fails while both data sites are healthy, the cluster continues normally (Scenario 3 in the failure matrix). The risk window is witness failure + simultaneous site failure. Mitigation: UPS at witness site with 4-hour runtime, plus documented procedure for deploying replacement witness from OVA.

Challenge 4: 'How do you handle a planned site-wide maintenance (e.g., power shutdown at Site A)?'
Response: Planned failover procedure: (1) DRS evacuates all VMs to Site B, (2) Enter maintenance mode on all Site A hosts with 'Ensure data accessibility', (3) Power down Site A, (4) During maintenance: Site B runs at full capacity with witness providing quorum, (5) Reverse process to restore. Total planned failover time: ~2 hours for 350 VMs.

Step 4

Create the final stretched cluster design summary:

  1. Architecture:
  • 2 data sites (8+8 hosts) + 1 witness site (Medium OVA)
  • vSAN ESA with FTT=1 site mirroring (data copy at each site)
  • L3 routed inter-site connectivity (no L2 stretch)
  • NSX overlay for workload networking
  1. Network:
  • Inter-site: 2×100 GbE dark fiber (primary) + 1×10 GbE MPLS (backup)
  • Measured RTT: 0.8ms (well within 5ms requirement)
  • MTU: 9000 end-to-end, 1700 for Geneve overlay
  1. Witness:
  • Tier: Medium (supports up to 21,833 components)
  • Placement: Corporate office (Site C), 50km from both sites
  • Connectivity: 1 Gbps MPLS, 2.5ms RTT to Site A, 3.0ms to Site B
  1. Failure handling:
  • 7 scenarios documented with expected behavior
  • Preferred site: Site A
  • Recovery runbook with decision tree and escalation path
  1. Monitoring:
  • Aria Operations dashboards for inter-site latency, resync progress
  • Alerts: RTT >3ms warning, >5ms critical
  • Weekly witness health verification

Validation Gate

Check: Stretched cluster design consolidated with cost comparison, alternative analysis, defense responses, and summary

Expected: Design decision entry with 4+ alternatives evaluated, cost comparison table, 4+ defense responses prepared, complete architecture summary

Common Errors

Not having a cost comparison ready — panelists will ask whether stretched cluster is worth the premium over SRM
Unable to articulate when stretched cluster is NOT appropriate — shows lack of design judgment
Forgetting planned maintenance procedures — stretched cluster maintenance is more complex than single-site
Not documenting the threshold for converting stretched cluster to SRM — know when to recommend a simpler pattern

Final Validation

Complete multi-site vSAN stretched cluster design with network validation, witness sizing, failure analysis, and defensible design decisions

✓ Network requirements validated with measurements → RTT, bandwidth, and MTU verified for all inter-site paths

✓ Witness sized and placed correctly → Official OVA tier selected with component count math; third-site placement verified

✓ All failure scenarios documented → 7-scenario matrix with expected behavior, recovery procedures, and escalation path

✓ Design decision defensible → Alternatives evaluated, cost comparison documented, defense responses prepared

Cleanup / Restore

• Save all design documents and failure scenario matrix

• If using Holodeck: Remove stretched cluster configuration and revert to base snapshot

Design Reflection (VCDX)

Stretched clusters are the most challenging multi-site design pattern to defend in VCDX. Panelists test three things: (1) Do you understand the failure modes? The scenario matrix is your strongest tool — walk through each scenario with specific outcomes. (2) Can you justify the cost? Be ready with numbers showing the premium over SRM and why RPO=0 is worth it. (3) Do you know when NOT to use it? If the customer's RPO can relax to 5 minutes, SRM is simpler and cheaper. Architectural judgment means recommending the right pattern, not the most complex one.

Requirements

  • RPO=0 for Tier-1 applications (no data loss on site failure)
  • RTO in seconds (automatic failover, not manual SRM plan)
  • 600 VMs across two sites with full failover capacity at each site
  • Data sovereignty: Both copies must be within the same jurisdiction

Constraints

  • Inter-site RTT must be ≤5ms for vSAN synchronous replication
  • Witness must be in a third fault domain
  • Each site must size for full workload (2× host count vs single-site)
  • VCF 9.0 requires vSAN ESA for new deployments

Assumptions

  • Dark fiber available between sites with <1ms RTT
  • Corporate office suitable as witness site (<200ms RTT to both sites)
  • Budget approved for 30-40% cost premium over active-passive DR
  • Inter-site latency stable (no significant jitter or degradation during business hours)

Risks

  • Inter-site link degradation causes vSAN write latency spike affecting all VMs
  • Witness site failure coinciding with data site failure causes data unavailability
  • Resync after site recovery saturates inter-site link for hours
  • Stretched cluster complexity leads to operational errors during failure recovery

Self-Assessment Discussion Prompts

  1. At what inter-site distance does stretched cluster become impractical, and what alternative would you recommend?
  2. How would you design a hybrid approach using stretched cluster for Tier-1 and SRM for Tier-2/3?
  3. If the customer adds a third data site, can vSAN stretched cluster support 3 data sites? If not, what architecture would you use?
  4. What is the maximum number of VMs you would put in a single stretched cluster, and what drives that limit?

Extensions

Hybrid Multi-Site: Stretched Cluster + SRM

Design a hybrid approach where Tier-1 VMs use stretched cluster (RPO=0) and Tier-2/3 VMs use SRM with asynchronous replication (RPO=15min). Document the separate workload domains, independent vSAN datastores, and unified Aria Operations monitoring across both patterns.

Stretched Cluster with vSAN File Services

Add vSAN File Services (NFS/SMB shares) to the stretched cluster design. Evaluate how file service failover differs from VM failover, document the impact on witness component count, and design the client reconnection procedure.

Holodeck Stretched Cluster Simulation

Build a 6-node (3+3) stretched cluster on Holodeck with a witness VM. Simulate network partition by disabling vmnic between sites, observe vSAN behavior, and validate the failure scenario matrix from Task 3 against actual results.

⚠ Known Pitfalls (from Community KB)

Using generic witness specs instead of official OVA tiers — the witness appliance is an OVA with Tiny/Medium/Large/XL sizing tiers based on component count, not arbitrary CPU/RAM specs
Placing witness at the same physical location as one data site — this nullifies the quorum benefit; witness MUST be in an independent fault domain
Not sizing each site for full failover capacity — stretched cluster requires each site to run ALL VMs during single-site failure, effectively doubling the host count
Assuming stretched cluster works at any distance — the 5ms RTT requirement limits practical distance to ~100km with dark fiber; beyond that, use asynchronous replication (SRM)

References

  • vSAN 9.0 Stretched Cluster Guide — Network Requirements and Witness Sizing
  • VMware KB 2131662 — vSAN Stretched Cluster Best Practices
  • vSAN Witness Appliance Deployment Guide — OVA Sizing Tiers
  • VMware VCF 9.0 Multi-Site Design Considerations
  • VMware KB 2108285 — Troubleshooting vSAN Network Performance
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.