Dual-Site Holodeck Deployment
Objectives
- Design a non-overlapping IP addressing scheme for a dual-site Holodeck topology
- Configure the Holodeck Toolkit to deploy Site B management domain with cross-site routing
- Establish HoloRouter-A to HoloRouter-B peering using BGP or static routes
- Validate bidirectional management plane reachability and DNS cross-resolution
- Map a dual-site lab topology to production VCDX disaster recovery and stretched cluster architecture using RCAR framework
Prerequisites
Single Holodeck 9.0.2 management domain fully deployed and healthy (holodeck-02 completed). Physical host must have minimum 768 GB RAM, 64+ logical cores, 2+ TB SSD (due to doubled resource footprint for Site B). Site A instance must be in a stable state — snapshot taken. No existing dual-site configuration.
Prior labs: holodeck-02
Required skills:
- VCF management domain architecture (SDDC Manager, vCenter, NSX)
- IP address planning and CIDR notation
- BGP routing concepts (or comfort with static route configuration)
- vCenter Enhanced Linked Mode concepts
- NSX federation and multi-site networking
- Disaster recovery and RTO/RPO analysis
Lab Environment
Dual physical Holodeck instances on a single ESXi host OR on two separate physical hosts. Site A (production) uses 10.1.0.0/20 management CIDR with 4 ESXi hosts in 10.1.1.0/24. Site B (DR) uses 10.2.0.0/20 management CIDR with 4 ESXi hosts in 10.2.1.0/24. HoloRouter-A and HoloRouter-B are configured in dual-site mode with BGP peering (optional) or static route redistribution to achieve cross-site reachability. Both sites reach the shared HoloRouter's routing interfaces on 10.1.16.1 (Site A) and 10.2.16.1 (Site B).
graph TB
PHY[Single ESXi Host - 768 GB RAM] --> HR[HoloRouter Dual-Site 10.1.16.1 / 10.2.16.1]
HR -->|BGP Peering| BGP[BGP: Site A ASN 65001 | Site B ASN 65002]
subgraph SiteA["Site A - Production (10.1.0.0/20)"]
HR1[HoloRouter-A 10.1.16.1]
HR1 --> ESXiA1[ESXi-01 10.1.1.101]
HR1 --> ESXiA2[ESXi-02 10.1.1.102]
HR1 --> ESXiA3[ESXi-03 10.1.1.103]
HR1 --> ESXiA4[ESXi-04 10.1.1.104]
HR1 --> CBA[Cloud Builder 10.1.0.200]
HR1 --> SDDC_A[SDDC Manager 10.1.0.4]
HR1 --> VC_A[vCenter-A 10.1.0.6]
HR1 --> NSX_A[NSX Manager-A 10.1.0.10]
end
subgraph SiteB["Site B - DR (10.2.0.0/20)"]
HR2[HoloRouter-B 10.2.16.1]
HR2 --> ESXiB1[ESXi-01 10.2.1.101]
HR2 --> ESXiB2[ESXi-02 10.2.1.102]
HR2 --> ESXiB3[ESXi-03 10.2.1.103]
HR2 --> ESXiB4[ESXi-04 10.2.1.104]
HR2 --> CBB[Cloud Builder 10.2.0.200]
HR2 --> SDDC_B[SDDC Manager 10.2.0.4]
HR2 --> VC_B[vCenter-B 10.2.0.6]
HR2 --> NSX_B[NSX Manager-B 10.2.0.10]
end
BGP -->|Cross-site IP reachability| HR1
BGP -->|Cross-site IP reachability| HR2IP Addressing
| Network | Purpose | VLAN |
|---|---|---|
10.1.0.0/20 | Site A management network supernet | VLAN 1644 |
10.1.1.0/24 | Site A ESXi host management VMkernel | VLAN 1644 |
10.1.2.0/24 | Site A vMotion (if stretched cluster) | VLAN 1645 |
10.1.3.0/24 | Site A vSAN (if stretched vSAN) | VLAN 1646 |
10.1.4.0/24 | Site A NSX Host TEP | VLAN 1647 |
10.1.5.0/24 | Site A NSX Edge TEP | VLAN 1648 |
10.2.0.0/20 | Site B management network supernet | VLAN 1704 |
10.2.1.0/24 | Site B ESXi host management VMkernel | VLAN 1704 |
10.2.2.0/24 | Site B vMotion (if stretched cluster) | VLAN 1705 |
10.2.3.0/24 | Site B vSAN (if stretched vSAN) | VLAN 1706 |
10.2.4.0/24 | Site B NSX Host TEP | VLAN 1707 |
10.2.5.0/24 | Site B NSX Edge TEP | VLAN 1708 |
Credentials
| System | Username | Password |
|---|---|---|
| Physical ESXi Host | root | Set during ESXi installation |
| HoloConsole (Webtop) | holodeck | Default Holodeck password from Toolkit docs |
| Site A ESXi Hosts | root | Set by Holodeck config.json esxiPassword field |
| Site B ESXi Hosts | root | Set by Holodeck Site B config.json esxiPassword field |
| Cloud Builder A | admin@local | Set by Holodeck config.json cbPassword field |
| Cloud Builder B | admin@local | Set by Holodeck Site B config.json cbPassword field |
| SDDC Manager A | administrator@vsphere.local | Set in VCF bring-up spec JSON |
| SDDC Manager B | administrator@vsphere.local | Set in VCF bring-up spec JSON |
| vCenter-A | administrator@vsphere.local | Same as SDDC Manager A SSO password |
| vCenter-B | administrator@vsphere.local | Same as SDDC Manager B SSO password |
| NSX Manager-A | admin | Set in VCF bring-up spec JSON |
| NSX Manager-B | admin | Set in VCF bring-up spec JSON. API password is VMware123!VMware123! (doubled) |
Tasks
Task 1 Plan the dual-site topology and IP addressing scheme
availabilityIn VCDX design, topology and addressing are foundational. A panelist will probe your understanding of failure domain boundaries, network isolation, and resource allocation across sites. Dual-site planning surfaces the architectural trade-offs: single physical host = no site-level HA, but identical infrastructure = simpler operational model. This task mirrors the VCF Planning and Preparation Workbook exercise for multi-site deployments.
Document your dual-site topology in a design worksheet (text or spreadsheet). Start with Site A baseline (from holodeck-02): management CIDR 10.1.0.0/20, ESXi cluster 10.1.1.0/24, vMotion 10.1.2.0/24, vSAN 10.1.3.0/24, NSX Host TEP 10.1.4.0/24, NSX Edge TEP 10.1.5.0/24, VLAN range 1644-1648.
Design Site B network configuration using non-overlapping IP ranges. Recommended: Site B management CIDR 10.2.0.0/20 (no overlap with 10.1.0.0/20), ESXi cluster 10.2.1.0/24, vMotion 10.2.2.0/24, vSAN 10.2.3.0/24, NSX Host TEP 10.2.4.0/24, NSX Edge TEP 10.2.5.0/24, VLAN range 1704-1708 (no overlap with 1644-1648). This design allows symmetric topologies while avoiding overlap.
Document cross-site routing requirements. Both sites must reach each other's management plane (SDDC Manager, vCenter, NSX). Create a routing table for each direction: Site A to Site B routes (10.2.0.0/20 via Site B HoloRouter gateway) and Site B to Site A routes (10.1.0.0/20 via Site A HoloRouter gateway). Note the gateway IPs: HoloRouter-A will have interfaces in both 10.1.16.1/24 and an externally-accessible IP; HoloRouter-B will have interfaces in both 10.2.16.1/24 and an externally-accessible IP.
Analyze physical resource requirements. Site A already allocated ~16 vCPU + 384 GB RAM (4 ESXi hosts at 12vCPU/96GB each) from your physical host. Site B will need the same: ~16 vCPU + 384 GB RAM. Total dual-site footprint: ~32 vCPU + 768 GB RAM (plus 32 GB for management VMs). Verify your physical host has at least 768 GB RAM and 64 logical cores. If using a single physical host, document resource contention risks and mitigations.
Create a failure domain analysis. In this dual-site topology on a single physical host, what is the failure domain boundary? Answer: Both Site A and Site B share the same single point of failure (the physical host). In production, you would have separate physical hosts or data centers. Document how this affects your RTO/RPO strategy (covered in Task 5).
Validation Gate
Check: Review your design worksheet: (1) Site A and Site B CIDRs are non-overlapping, (2) VLAN ranges are non-overlapping, (3) cross-site routing table is complete, (4) resource allocation is verified against physical host capacity, (5) failure domain analysis is documented.
Expected: All five validation points pass. You are ready to implement the dual-site Holodeck configuration.
Common Errors
Task 2 Deploy Site B Holodeck instance with cross-site network configuration
availabilityExecuting the dual-site deployment is the operational foundation for VCDX multi-site architecture discussions. This task surfaces the orchestration complexity: two independent VCF bring-ups running on shared infrastructure, with potential resource contention and deployment ordering dependencies. Understanding what can fail independently vs what is tightly coupled is critical VCDX knowledge.
On the HoloConsole or management workstation, open PowerShell and navigate to the Holodeck runtime directory: cd C:\Holodeck\holodeck-runtime
Create a new Holodeck configuration for Site B. Use: New-HoloDeckConfig -Description 'holodeck-06-site-b' -TargetHost <your-physical-esxi-host> -Username root -Password <esxi-root-password>. Capture the generated ConfigID (e.g., 'a1b2').
Configure Site A network (if not already configured from holodeck-02). Run: New-HoloDeckNetworkConfig -Site a -MasterCIDR 10.1.0.0/20 -VLANRangeStart 10
Configure Site B network. Run: New-HoloDeckNetworkConfig -Site b -MasterCIDR 10.2.0.0/20 -VLANRangeStart 40. This generates a separate network config for Site B with CIDR 10.2.0.0/20 and VLANs starting at 1704 (40+1664).
Configure the HoloRouter for dual-site operation. Run: Set-HoloRouter -dualsite. This configures BGP peering between Site A and Site B router interfaces and sets up VLAN interfaces for both sites on the physical host.
Edit the Site B configuration file to customize hostname, IP assignments, and resource sizing. Open: notepad ./templates/config.json. Locate the 'site' or 'siteId' field and set it to 'b' or '2' (depending on Holodeck version). Verify managementCIDR is 10.2.0.0/20. Optionally adjust ESXi host resource sizing to avoid contention (e.g., reduce from 12vCPU/96GB to 10vCPU/80GB per host if physical host is constrained).
Prepare the Site B Holodeck instance. Run: New-HoloDeckInstance -ConfigFile ./templates/config.json -Verbose. This stages Site B artifacts, validates host compatibility, and configures HoloRouter-B interfaces. Monitor for errors.
Verify Site B HoloRouter is configured. From HoloConsole, SSH to the HoloRouter and run: vtysh -c 'show bgp summary'. You should see two neighbors (Site A and Site B) with established connections, or if using static routing, verify the routing table with 'show ip route'.
Start the Site B VCF deployment. Run: Start-HoloDeckInstance -InstanceID <site-b-config-id> -Verbose. This orchestrates nested ESXi provisioning, Cloud Builder deployment, and full VCF bring-up for Site B. Monitor the verbose output for progress.
Validation Gate
Check: Run: Get-HoloDeckInstance | Format-Table -Property InstanceId, State, VcfVersion, Site. Verify both Site A and Site B instances show 'Running' state.
Expected: Two Holodeck instances listed: Site A (or Instance 1) and Site B (or Instance 2), both in Running state, both VCF 9.0.2
Common Errors
Task 3 Configure cross-site networking and validate bidirectional reachability
manageabilityCross-site networking is where the architecture comes alive. In VCDX terms, this task probes your understanding of Layer 3 network design, failure domain isolation, and multi-site management traffic engineering. A panelist will ask: 'How do you ensure vCenter-A can manage ESXi-B?' or 'What happens if the cross-site link fails?' This task makes those questions concrete.
Verify HoloRouter BGP or static routes are configured for cross-site routing. From HoloConsole, SSH into the HoloRouter (ssh admin@10.1.1.1 or whichever is the main router IP) and run: vtysh -c 'show bgp ipv4 unicast' or if static routing: vtysh -c 'show ip route'. You should see routes for both 10.1.0.0/20 and 10.2.0.0/20 advertised or installed.
Test reachability from Site A to Site B management plane. From HoloConsole, ping Site B SDDC Manager: ping 10.2.0.4. Then ping Site B vCenter: ping 10.2.0.6. Both should reply with <1ms latency (same physical host).
Test reachability from Site B to Site A management plane. SSH to Site B HoloConsole or a Site B ESXi host (ping from 10.2.x.x range) and ping Site A SDDC Manager: ping 10.1.0.4. Then ping Site A vCenter: ping 10.1.0.6. Both should reply.
Configure DNS cross-resolution. From Site A HoloConsole, add Site B FQDNs to the HoloRouter DNS forwarder. SSH to HoloRouter and edit /etc/dnsmasq.conf to add: address=/sddc-manager-b.site-a.vcf.lab/10.2.0.4, address=/vcenter-b.site-a.vcf.lab/10.2.0.6, address=/nsx-manager-b.site-a.vcf.lab/10.2.0.10. Restart DNS: systemctl restart dnsmasq. Repeat for Site B (add Site A FQDNs).
Validate SDDC Manager awareness of both sites. From Site A SDDC Manager UI (https://10.1.0.4), navigate to Administration > Licensing or Inventory > Workload Domains. Verify Site A is listed. Then attempt to configure Site B workload domain addition (if your VCF licensing allows multi-domain). If multi-domain is not enabled, document the licensing constraint.
If deploying NSX federation (multi-site), configure NSX-A to peer with NSX-B. From NSX Manager-A UI (https://10.1.0.10), navigate to System > Federation. Click 'Join Federation' and enter NSX Manager-B IP (10.2.0.10) and credentials. Complete the federation handshake. Verify bidirectional federation status: both NSX managers should show 'Federation Status: Connected'.
Test management plane operations from both sites. From Site A vCenter, verify that you can see Site B ESXi hosts by navigating to Inventory > Hosts. If Enhanced Linked Mode is configured (vCenter-A linked to vCenter-B), you should see both sites' objects. Create a test VM on Site A cluster and verify it's managed by Site A vCenter.
Validation Gate
Check: Perform three checks: (1) ping 10.2.0.4 and 10.2.0.6 from Site A, (2) ping 10.1.0.4 and 10.1.0.6 from Site B, (3) nslookup vcenter-b.site-a.vcf.lab from Site A resolves to 10.2.0.6 and vice versa.
Expected: All three checks pass. Bidirectional IP and DNS reachability confirmed between sites.
Common Errors
Task 4 Validate dual-site management plane operations
manageabilityPost-deployment validation proves that the dual-site architecture is operational and ready for VCDX-level design discussions. This task operationalizes the abstract concepts: Can you manage both sites from a single pane of glass (SDDC Manager or vCenter Linked Mode)? What visibility does each site have into the other? How do you know if the cross-site link is healthy? These are the questions a VCDX panelist will ask during design defense.
From HoloConsole webtop, open Firefox and navigate to SDDC Manager-A: https://10.1.0.4. Log in with administrator@vsphere.local. Navigate to Inventory > Workload Domains. Verify the management domain shows Status: Active with 4 hosts, 1 cluster.
Open a new tab and navigate to SDDC Manager-B: https://10.2.0.4. Log in with administrator@vsphere.local (Site B credentials). Verify Site B management domain shows Status: Active with 4 hosts, 1 cluster.
Open a third tab and navigate to vCenter-A: https://10.1.0.6. Log in with administrator@vsphere.local. Navigate to Hosts and Clusters. Verify the management cluster shows 4 connected ESXi hosts, vSAN health green, DRS enabled, HA enabled.
Open a fourth tab and navigate to vCenter-B: https://10.2.0.6. Log in with administrator@vsphere.local (Site B credentials). Navigate to Hosts and Clusters. Verify the management cluster shows 4 connected ESXi hosts, vSAN health green, DRS enabled, HA enabled.
Configure vCenter Enhanced Linked Mode (optional but recommended for VCDX). In vCenter-A, navigate to Administration > System Configuration > Enhanced Linked Mode. Click 'Join'. Enter vCenter-B FQDN (vcenter-b.site-a.vcf.lab or IP 10.2.0.6) and Site B administrator@vsphere.local credentials. Complete the pairing. Repeat from vCenter-B to establish bidirectional link.
Navigate to NSX Manager-A: https://10.1.0.10. Log in with admin. Navigate to System > Fabric > Nodes > Host Transport Nodes. Verify all 4 Site A ESXi hosts show Configuration State: Success and Transport Node Status: Up. Repeat for NSX Manager-B to verify Site B hosts.
Test cross-site management from Site A vCenter. In vCenter-A Linked Mode view, verify you can see both Site A and Site B objects (datacenters, clusters, hosts). Create a test folder in Site A datacenter. From Site B vCenter, verify the folder is visible (if Linked Mode is working). This proves bidirectional management plane visibility.
Document the dual-site operational readiness. Create a status report showing: (a) Both SDDC Managers healthy and licensed, (b) Both vCenters healthy and Linked Mode connected, (c) Both NSX Managers healthy and federation status (if applicable), (d) Cross-site ping reachability confirmed, (e) Management traffic can traverse between sites without errors.
Validation Gate
Check: Verify: (1) SDDC Manager-A and SDDC Manager-B both show Active status, (2) vCenter-A and vCenter-B both show 4 healthy hosts with green vSAN/HA/DRS, (3) vCenter Linked Mode connected (if deployed), (4) NSX-A and NSX-B both show 4 host transport nodes up.
Expected: All management components dual-site healthy. Cross-site operations fully functional.
Common Errors
Task 5 Design exercise — Map dual-site Holodeck to production DR architecture using RCAR
recoverabilityThis is the VCDX defense moment. You have a working dual-site lab. Now articulate the architectural decisions, trade-offs, and risk mitigations using the RCAR framework. A VCDX panelist will probe: 'Why is this design suitable for DR?' 'What are the RPO and RTO assumptions?' 'How does your stretched cluster design differ from active-active?' 'What failover automation have you considered?' This task forces you to think like an architect, not just a lab technician.
Define the production DR scenario that your dual-site Holodeck represents. Example: 'Site A is the primary production datacenter (Houston). Site B is the DR datacenter (Dallas). Business applications run on Site A vSAN clusters. Site B is initially cold (powered off or minimal resources). RPO target: 1 hour (hourly snapshots + vSphere Replication). RTO target: 2 hours (manual failover via SDDC Manager + NSX policy failover).' Document your scenario in a design worksheet.
Create RCAR analysis for your design: Requirements, Constraints, Assumptions, Risks.
REQUIREMENTS: What must the design achieve? List at least 5 requirements. Example: (a) Ability to replicate application VMs from Site A to Site B, (b) Automatic network failover via NSX (traffic redirects to Site B IPs), (c) SDDC Manager can manage both sites, (d) RPO <= 1 hour, RTO <= 2 hours, (e) Cross-site management traffic must be encrypted and authenticated.
CONSTRAINTS: What limits your design? List at least 4. Example: (a) Single physical host in the lab (vs two separate data centers in production), (b) Holodeck 2.1.x supports dual-site but not NSX stretched security groups, (c) vSAN stretched cluster requires L2 adjacency (vMotion VLAN must span sites — not possible in this lab setup), (d) Licensing: dual-site failover requires additional per-host licenses, (e) Network latency: <5ms RTT required for stretched vSAN (lab has <2ms, but real WAN would not).
ASSUMPTIONS: What are you assuming will be true in production? List at least 4. Example: (a) Assume vSphere Replication licenses are available, (b) Assume cross-site network has <5ms latency and 1 Gbps+ bandwidth, (c) Assume vCenter Enhanced Linked Mode is deployed and functioning, (d) Assume Site B is pre-provisioned and ready for failover (not a cold DR site requiring hours to boot), (e) Assume SDDC Manager can manage both sites without license overhead.
RISKS: What can go wrong? List at least 6 with impact and mitigation. Example: (a) Risk: Cross-site network link fails. Impact: Site B becomes unreachable; applications on Site A continue running but replication stops. Mitigation: Monitor cross-site latency and packet loss continuously. If link fails, manual failover is required (RTO extends to 4+ hours). (b) Risk: vCenter Linked Mode breaks during failover (split-brain scenario). Impact: Operator confusion, potential dual VM registration. Mitigation: Document failover runbook explicitly stating which vCenter (A or B) is authoritative post-failover. (c) Risk: RPO not met due to replication lag. Impact: Data loss up to max lag period. Mitigation: Monitor replication lag per VM; alert if any VM exceeds RPO threshold. (d) Risk: Stretched vSAN cluster experiences latency-induced failures. Impact: vSAN rebuild operations slow down, increasing window of vulnerability. Mitigation: Use synchronous replication only for critical VMs; asynchronous for bulk workloads. (e) Risk: DNS failover not automated — Site B FQDNs not automatically updated. Impact: Applications connect to stale Site A IPs post-failover. Mitigation: Use dynamic DNS updates via DHCP or implement application-level connection pooling with fallback. (f) Risk: Disaster recovery drill reveals missing backup of SDDC Manager or vCenter configuration. Impact: Can't re-provision management domain on Site B if Site A fails completely. Mitigation: Weekly backup of SDDC Manager DB and vCenter; store offsite.
Map VCDX Design Quality dimensions to your dual-site architecture. For each of the 5 dimensions (Availability, Manageability, Performance, Recoverability, Security), describe how your design addresses it and where trade-offs were made. Example: Availability: Stretched vSAN with RAID-1 provides N+1 node failure tolerance. Trade-off: RAID-1 uses 50% of capacity. Performance: Asynchronous vSphere Replication avoids cross-site latency impact on Site A; RPO trade-off is 1 hour instead of near-zero. Manageability: Dual SDDC Managers (one per site) simplifies independent lifecycle operations but increases operational burden. Security: Cross-site replication traffic is encrypted; NSX Federation provides separate policy domains per site for workload isolation.
Create a failover decision tree. Starting from 'Primary site failure detected', walk through the decision logic: (a) Is it a network link failure only (Site A still operational but unreachable)? Decision: Do NOT failover; fix the link. (b) Is it a complete Site A power loss? Decision: Initiate failover (RTO 2 hours). (c) Is it a partial failure (e.g., vCenter-A down but ESXi-A hosts up)? Decision: SDDC Manager can manage from Site B; no failover needed. Document the tree as a flowchart or table.
Simulate a failover scenario. Assume Site A vCenter has become unavailable (you can disconnect its network or shut it down). Document the steps an operator would follow to: (a) Confirm Site A is truly down, (b) Update DNS to point to Site B vCenter (or confirm NSX failover has redirected traffic), (c) Open Site B vCenter and verify all replicated VMs are present and can be powered on, (d) Power on critical application VMs on Site B, (e) Verify applications are accessible on Site B IPs, (f) Document the actual RTO achieved and compare to target (2 hours). Do NOT actually perform the simulation in this lab — instead, write out the step-by-step runbook and estimate timings.
Validation Gate
Check: Verify all 9 steps completed: (1) Scenario documented, (2) RCAR worksheet created, (3)-(6) RCAR sections filled (Requirements, Constraints, Assumptions, Risks), (7) Design Quality matrix completed for all 5 dimensions, (8) Failover decision tree documented, (9) Failover runbook written with RTO estimate.
Expected: Complete design exercise documentation suitable for VCDX oral defense. A panelist would be able to ask follow-up questions on any section and receive detailed, architectural responses.
Common Errors
Final Validation
A complete dual-site Holodeck deployment (Site A and Site B) is fully operational with validated cross-site networking, management plane visibility, and a production-ready DR architecture design documented using RCAR framework. All 5 design qualities (Availability, Manageability, Performance, Recoverability, Security) have been analyzed with explicit trade-offs articulated. This lab is the gateway to VCDX multi-site architecture discussions.
✓ Two Holodeck instances running: Site A (10.1.x.x) and Site B (10.2.x.x), both state = Running → Get-HoloDeckInstance shows two instances with different IP ranges
✓ Cross-site IP reachability: ping 10.2.0.4 from Site A succeeds; ping 10.1.0.4 from Site B succeeds → Both pings reply with <2ms latency
✓ Cross-site DNS: nslookup vcenter-b.site-a.vcf.lab from Site A resolves to 10.2.0.6; vice versa → Both DNS lookups resolve correctly
✓ SDDC Manager-A and SDDC Manager-B both show management domain Active status → Inventory > Workload Domains shows Active with 4 hosts each
✓ vCenter-A and vCenter-B both show 4 healthy ESXi hosts, vSAN green, HA/DRS enabled → Hosts and Clusters view shows green health indicators
✓ NSX Manager-A and NSX Manager-B both show 4 host transport nodes in Success/Up state → System > Fabric > Nodes shows 4/4 transport nodes up per site
✓ vCenter Enhanced Linked Mode configured and connected (optional) → Administration > System Configuration shows Linked Mode Status: Connected (if deployed)
✓ RCAR design exercise completed: Requirements, Constraints, Assumptions, Risks, Quality Matrix, Failover Decision Tree, Failover Runbook → Design documentation comprehensive enough for VCDX oral defense
Cleanup / Restore
Snapshot: holodeck-06-complete
• Take snapshots of all Holodeck VMs from both Site A and Site B: Get-VM -Name 'Holo-*' | New-Snapshot -Name 'holodeck-06-complete' -Description 'Post dual-site deployment — both sites healthy, cross-site networking validated' -Memory:$false -Quiesce:$false
• Document dual-site topology: Create or update your lab notebook with Site A and Site B IP addressing, FQDN mappings, and cross-site routing configuration
• Archive RCAR design exercise: Save your design documentation (RCAR, Design Quality Matrix, Failover Runbook) to a central location for future reference and VCDX preparation
Design Reflection (VCDX)
Dual-site is the VCDX differentiator. A VCDX panelist examining a multi-site VCF design probes these dimensions: (1) Why dual-site vs single-site? (Answer: DR, business continuity, geographic diversity for compliance). (2) How do you decouple Site A from Site B operationally? (Answer: Separate SDDC Managers, separate vCenter clusters, ESXi hosts per-site, potential NSX federation for policy isolation). (3) What assumptions are baked into your cross-site latency budget? (Answer: <5ms RTT for stretched vSAN, <50ms RTT for asynchronous replication).
(4) How do you handle split-brain scenarios during failover? (Answer: DNS failover, explicit failover runbook, SDDC Manager authority per site). (5) What is your actual RTO and RPO, and how was it validated? (Answer: Lab simulation, failover runbook, documented assumptions). Be prepared to defend every architectural choice: 'Why stretched vSAN instead of separate clusters?' (Answer: Seamless VM migration for maintenance, unified capacity pool, but introduces L2 requirement and latency sensitivity).
'Why not active-active?' (Answer: Requires Byzantine consensus algorithms, complex DNS, split-brain risk; active-passive with fast failover is simpler and more operationally sound for this business profile).
Requirements
- Ability to replicate application VMs from Site A to Site B with RPO <= 1 hour
- Automatic or semi-automatic network failover via NSX (failover policies or manual SDDC Manager workload domain failover)
- SDDC Manager can manage workload domain resources at both sites for provisioning and lifecycle
- vCenter visibility across both sites (Enhanced Linked Mode or separate SDDC Manager UIs)
- Cross-site management traffic fully encrypted and authenticated
- Failover RTO <= 2 hours (includes detection, DNS update, VM power-on, application startup)
- Dual-site architecture must be operationally sustainable (runbooks, monitoring, alerting in place)
Constraints
- Single physical ESXi host in lab (vs separate data centers in production) — shared failure domain
- Holodeck 2.1.x dual-site support does not include stretched vSAN (requires separate vSAN clusters per site or async replication of vSAN objects)
- vMotion VLAN spanning both sites would require L2 adjacency (not available in pure BGP/Layer 3 design of this lab). Workaround: Live migration disabled; use vSphere Replication instead
- NSX multi-site federation available but Distributed Firewall policies are not automatically synced (must be manually deployed per site or via NSX policy templates)
- Licensing: vSphere Replication per-VM costs scale with replica count. NSX Federation requires multi-site add-on license
- Cross-site replication bandwidth: 1 Gbps network limit assumed (lab has virtual link; production WAN may be congested)
- SDDC Manager single-instance per site (no HA for SDDC Manager itself in this lab). Production would likely deploy redundant SDDC Managers per site
Assumptions
- Cross-site network latency <5ms RTT (met in lab; real WAN would be 50-150ms, requiring async replication only)
- Cross-site network availability >99.9% (SLA on WAN link or redundant paths)
- vSphere Replication licenses available for all application VMs
- Site B is pre-provisioned with compute capacity (4 ESXi hosts ready) — not a cold DR site requiring provisioning time
- NSX Federation licenses available if multi-site NSX is required
- DNS failover is manual (ops team updates DNS to point to Site B post-failover); no automated DNS update
- vCenter Enhanced Linked Mode is desired for operational convenience; not required for failover success
- Business tolerance for some manual failover steps (not fully automated) — acceptable due to infrequent failover events
Risks
- Cross-site network link failure or high latency (>5ms RTT). — Impact: vSphere Replication lag increases beyond RPO tolerance. If stretched vSAN were used, split-brain risk. Site A applications remain up; Site B replication stalls. — Mitigation: Continuous monitoring of cross-site ping latency and replication lag. Alert if latency >5ms or replication lag >15 minutes (exceeds RPO). Failover decision: if link down >30 min, manual failover to Site B. If latency increased but link operational, continue replication.
- vCenter Linked Mode breaks during failover (split-brain: both vCenters claim authority). — Impact: Operator confusion. Dual VM registration possible if apps powered on at both sites. Data corruption risk. — Mitigation: Failover runbook explicitly states: 'After failing over, Site B vCenter is sole authority. Site A vCenter must be isolated (network disconnect) until recovery.' Document this in vSphere alerting.
- RPO not met due to replication lag or network congestion. — Impact: Data loss up to max lag. For critical applications, unacceptable. — Mitigation: Implement per-VM replication lag alerting. Prioritize critical VMs for replication (sync replication). Bulk workloads use async. Weekly DR drill to measure actual RPO achieved.
- Site A complete power loss (not just network failure). — Impact: All Site A management domain components offline. SDDC Manager cannot manage. Failover to Site B mandatory. RTO extends if Site B vCenter takes 30+ minutes to discover all replicated VMs. — Mitigation: Power on Site B vCenter and SDDC Manager immediately post-failure. Run inventory scan to re-discover replicated VMs. Pre-create VM placement policies on Site B to speed power-on sequencing. Test in DR drill quarterly.
- Disaster recovery drill reveals missing backup of SDDC Manager configuration or vCenter settings. — Impact: Cannot re-provision management domain on Site B if Site A is unrecoverable. RTO becomes days, not hours. — Mitigation: Weekly backup of SDDC Manager PostgreSQL database. Weekly backup of vCenter config (using native backup tools or vSphere APIs). Store backups offsite (cloud storage or secondary site). Test restore quarterly.
- Stretched vSAN split-brain: network partition between Site A and Site B vSAN clusters. — Impact: Both partitions remain online, potentially writing conflicting data. Cluster corruption. — Mitigation: If using stretched vSAN (advanced extension), deploy VMware Site Recovery Manager (SRM) with vSAN synchronous replication and Byzantine agreement (requires witness VM). Preferred: DO NOT use stretched vSAN; use separate vSAN clusters per site with async replication instead.
- Operator error during failover: powers on VMs on both Site A and Site B simultaneously (not checking Site A status first). — Impact: Duplicate VMs, split-brain applications, data corruption. — Mitigation: Automate failover as much as possible: SDDC Manager API-driven workload domain failover, NSX policy automation for network redirection. Runbook includes verification step: 'Confirm all Site A VMs are powered off before powering on Site B replicas.' Implement VM registration locking in vCenter to prevent dual registration.
Self-Assessment Discussion Prompts
- Why is dual-site architecture critical for VCDX candidates to understand, and how does it differ from a single-site HA deployment?
- In your dual-site design, what is the failure domain boundary? What shared components exist between Site A and Site B that could cause correlated failures?
- How does your design handle network partitioning (split-brain) between Site A and Site B? What prevents two vCenters from independently managing the same VM?
- Compare two DR strategies: (a) stretched vSAN with synchronous replication, (b) separate vSAN clusters with async vSphere Replication. What are the RPO/RTO/cost trade-offs?
- If your cross-site network latency is 50ms (typical WAN) instead of <2ms (lab), how would your design change? What components could no longer be stretched?
- How would you automate failover detection and DNS failover in a production dual-site VCF deployment? What monitoring and alerting would you implement?
- In vCenter Enhanced Linked Mode across sites, a network partition occurs (Site A and Site B lose connectivity). Both vCenters are still running independently. What is the risk, and how do you recover?
- Your business requires RPO of 15 minutes and RTO of 1 hour. Is your dual-site Holodeck lab design capable of meeting these targets? What changes would be needed?
- NSX Federation provides separate policy domains per site. How does this affect network failover behavior? If Site A NSX is down, can Site B NSX automatically take over Site A workloads?
- Document the exact steps and timing for a failover scenario: Site A datacenter power loss at 2 PM. When are applications fully operational on Site B? What manual interventions are required?
Extensions
Deploy stretched vSAN cluster across Site A and Site B
Configure a vSAN cluster that spans both sites using synchronous replication and a witness VM. This advanced configuration is the foundation for stretched cluster designs in VCDX architecture. Document the latency sensitivity, failure domain implications, and how vSAN rebuild operations differ in stretched vs separate cluster designs.
much harderImplement NSX Federation and test multi-site security policies
Deploy NSX Federation between NSX Manager-A and NSX Manager-B. Create separate Tier-0 and Tier-1 gateways per site. Define security policies that differ between sites (e.g., Site A allows L7 inspection, Site B does not). Test cross-site VM communication with policies enforced. Document how NSX federation differs from stretched NSX (not supported in Holodeck).
much harderDeploy VMware Site Recovery Manager (SRM) for automated failover
If SRM licenses are available, deploy SRM on Site A and Site B. Configure protection groups for application VMs. Create recovery plans with VM boot sequencing and network remapping. Test failover automation: execute a full failover plan from Site A to Site B, measure actual RTO, document the orchestration workflow.
much harderSimulate a complete Site A failure and document recovery procedures
Disconnect or shut down all Site A VMs and management domain components. From Site B, execute a full failover: update DNS, power on application VMs, restore vCenter configuration from backup if needed. Measure actual RTO and RPO achieved. Document deviations from planned failover and lessons learned. This is the closest simulation to a real DR scenario.
harderDeploy vSphere Replication across sites with per-VM RPO targets
Configure vSphere Replication between Site A and Site B. Set different RPO targets per VM: critical apps = 15 min, standard = 1 hour, dev/test = 4 hours. Monitor replication lag and verify RPO compliance. Test recovery of individual VMs. Document how RPO variation affects cross-site network bandwidth utilization and cost.
harderDesign and implement multi-site workload domain in SDDC Manager
Using SDDC Manager, design a workload domain that spans Site A and Site B ESXi clusters. If license permits, add a workload domain to Site B and configure it for failover readiness. Document how SDDC Manager manages lifecycle operations (patching, upgrades) across sites.
harderConduct a full DR drill with metrics and sign-off
Schedule a formal DR drill: intentionally fail over from Site A to Site B. Measure and document: (1) Detection time, (2) Failover execution time, (3) Application startup time, (4) Application availability verification time, (5) Total RTO. Compare actual RTO to design target. Have stakeholders sign off on the drill results. This is the most realistic VCDX preparation exercise.
same⚠ Known Pitfalls (from Community KB)
References
- VMware Cloud Foundation 9.0 Multi-Site Deployment GuideTier 1 — Official
Official Broadcom documentation — covers multi-site topology, cross-site networking, SDDC Manager federation, and failover procedures. - Holodeck Toolkit Dual-Site Deployment DocumentationTier 1 — Official
Holodeck GitHub documentation — details the commands and workflow for deploying dual-site instances (New-HoloDeckNetworkConfig, Set-HoloRouter -dualsite). - VCF Community Forum — Holodeck Dual-Site DiscussionsTier 1 — Official
Broadcom community forum with 63 threads on Holodeck deployment, including dual-site IP config and vCenter issues. Threads 10, 19, and 30 are directly relevant to this lab. - vSphere Replication Administration GuideTier 1 — Official
Official guide for asynchronous VM replication across sites. RPO/RTO calculations, bandwidth requirements, and failover procedures. - NSX Federation Deployment GuideTier 1 — Official
NSX Federation architecture and deployment for multi-site network management and policy consistency. - VMware Site Recovery Manager (SRM) 8.x DocumentationTier 1 — Official
Automated failover orchestration, recovery plan execution, and RTO optimization techniques. - William Lam — VCF Multi-Site and DR Design PatternsTier 3 — Expert Blog
Deep dives on multi-site VCF, stretched cluster designs, and automation techniques. Authoritative community voice. - Cormac Hogan — vSAN Stretched Cluster ArchitectureTier 3 — Expert Blog
vSAN stretched cluster configurations, latency sensitivity, failure domain analysis, and recovery scenarios. - VCDX Design Principles — Design Quality Matrix FrameworkTier 2 — VMware Press
Reference for the 5 design quality dimensions (Availability, Manageability, Performance, Recoverability, Security) and how to articulate trade-offs. - VCF Planning and Preparation Workbook — Multi-Site Architecture SectionTier 1 — Official
Official planning document for multi-site deployments, including IP planning, failover strategy, and operational procedures.