Dual-Site Holodeck Deployment (Holodeck 9.1)
Objectives
- Design a non-overlapping IP addressing scheme for a dual-site Holodeck 9.1 topology, starting from the auto-generated Site A + Site B config pair
- Deploy the Site B management domain with cross-site routing on Holodeck 9.1
- Establish and validate cross-site routing between the Site A and Site B router functions (single-router vs VNA model: verify on live 9.1)
- Validate bidirectional management plane reachability and DNS cross-resolution using the 9.1 Technitium DNS service
- Map a dual-site lab topology to production VCDX disaster recovery and stretched cluster architecture using the RCAR framework
Prerequisites
Single Holodeck 9.1 management domain (VCF 9.1.0.0) fully deployed and healthy (holodeck91-02 completed). Physical host must have doubled resource headroom for Site B (9.0 baseline: minimum 768 GB RAM, 64+ logical cores, 2+ TB SSD — verify 9.1 sizing). Site A instance in a stable state — snapshot taken. No existing dual-site configuration active.
Prior labs: holodeck91-02
Required skills:
- VCF management domain architecture (SDDC Manager, vCenter, NSX)
- IP address planning and CIDR notation
- BGP routing concepts (or comfort with static route configuration)
- vCenter Enhanced Linked Mode concepts
- NSX federation and multi-site networking
- Disaster recovery and RTO/RPO analysis
Lab Environment
Dual Holodeck 9.1 instances on a single physical host. Site A (production) and Site B (DR) use the non-overlapping CIDR/VLAN allocations from the auto-generated 9.1 config pair (9.0 baseline pattern: Site A 10.1.0.0/20, Site B 10.2.0.0/20 — verify the 9.1 auto-generated defaults). Cross-site routing is provided by the HoloRouter dual-site function; on 9.1, evaluate whether VNA cluster networking changes the router topology.
graph TB
PHY[Single Physical ESXi Host] --> RTR[HoloRouter 9.1 dual-site or VNA cluster — verify]
subgraph SiteA[Site A - Production]
SDDC_A[SDDC Manager A]
VC_A[vCenter A]
NSX_A[NSX Manager A]
end
subgraph SiteB[Site B - DR]
SDDC_B[SDDC Manager B]
VC_B[vCenter B]
NSX_B[NSX Manager B]
end
RTR -->|Cross-site routing| SiteA
RTR -->|Cross-site routing| SiteBIP Addressing
| Network | Purpose | VLAN |
|---|---|---|
10.1.0.0/20 | Site A management supernet (9.0 baseline default — verify 9.1 auto-generated allocation) | 9.0 baseline VLAN 1644-1648 range — verify 9.1 defaults (VLAN range handling fixed in 9.1) |
10.2.0.0/20 | Site B management supernet (9.0 baseline default — verify 9.1 auto-generated allocation) | 9.0 baseline VLAN 1704-1708 range — verify 9.1 defaults (Site B port-group issues fixed in 9.1) |
Credentials
| System | Username | Password |
|---|---|---|
| Physical ESXi Host | root | Set during ESXi installation |
| HoloRouter 9.1 services (Webtop :30000 now authenticated; Authentik SSO :9443) | verify against live 9.1 toolkit | Webtop is behind authentication in 9.1; Authentik defaults are an open question — capture on first live deployment |
| SDDC Manager / vCenter (both sites) | administrator@vsphere.local | Set in the VCF bring-up spec per site |
| NSX Manager (both sites) | admin | Set in the VCF bring-up spec per site. 9.0-era doubled-password API quirk: verify whether it persists on 9.1 |
Tasks
Task 1 Plan the dual-site topology and IP addressing scheme
availabilityIn VCDX design, topology and addressing are foundational. 9.1 changes the starting point: the toolkit hands you an auto-generated Site A + Site B config pair, so planning becomes a review-and-align exercise rather than a from-scratch design — which makes collision analysis on shared hosts MORE important, not less. This task mirrors the VCF Planning and Preparation Workbook exercise for multi-site deployments.
Document your Site A baseline (from holodeck91-02) in a design worksheet: management CIDR, per-function subnets, and VLAN allocations as actually deployed.
Review the auto-generated Site B configuration produced by New-HoloDeckConfig (confirmed 9.1 behavior: both sites generated in one pass). Verify every Site B CIDR and VLAN is non-overlapping with Site A AND with anything else on the physical network.
Document cross-site routing requirements: both sites must reach each other's management plane (SDDC Manager, vCenter, NSX). Record the gateway/router interfaces for each direction. How the 9.1 router (single dual-site HoloRouter vs VNA cluster) exposes these interfaces: verify against live 9.1 toolkit.
Analyze physical resource requirements: Site B roughly doubles the footprint (docs-published 9.1 dual-site minimums: 96 logical cores, 1536 GB RAM, 5 TB NVMe total). Verify actual per-site consumption (plus GitOps/VNA overhead) and confirm the physical host has headroom; document contention risks if not.
Create a failure domain analysis: on a single physical host, both sites share one failure domain. Document how this differs from production dual-site and how it affects the RTO/RPO strategy developed in Task 5.
Validation Gate
Check: Review the worksheet: (1) Site A/B CIDRs non-overlapping, (2) VLAN ranges non-overlapping, (3) cross-site routing table complete, (4) resources verified against capacity, (5) failure domain analysis documented
Expected: All five points pass. Ready to deploy Site B.
Common Errors
Task 2 Deploy the Site B Holodeck instance with cross-site network configuration
availabilityExecuting the dual-site deployment is the operational foundation for VCDX multi-site architecture discussions: two independent VCF bring-ups on shared infrastructure, with resource contention and ordering dependencies. On 9.1 the exact cmdlet sequence is unverified — capture it live and note what the dual-site auto-generation removed from the 9.0-era manual workflow.
Open your management session (PowerShell workstation or the authenticated 9.1 Webtop) and navigate to the Holodeck runtime directory.
Locate the auto-generated Site B configuration from your New-HoloDeckConfig run (see holodeck91-01/91-02). Whether a separate per-site config-generation step still exists in 9.1: verify against live 9.1 toolkit.
Confirm the Site A network configuration in use matches what was deployed in holodeck91-02 (no drift).
Apply any Site B adjustments decided in Task 1 (CIDR/VLAN alignment, resource sizing). Note: Site B port-group issues were fixed in 9.1 — if you still hit port-group errors, capture them as new 9.1 findings rather than assuming 9.0 causes.
Enable the router's dual-site operation. The 9.0-era command was Set-HoloRouter -dualsite; whether 9.1 retains this cmdlet, changes its semantics under the new services stack, or replaces it with VNA cluster configuration: verify against live 9.1 toolkit.
Review the Site B configuration one final time: site identifier, management CIDR, resource sizing within the contention envelope planned in Task 1.
Run the prepare/staging phase for Site B (PowerShell path: New-HoloDeckInstance with the Site B config; GitOps path: trigger the Site B pipeline from the UI — exact 9.1 flow: verify).
Verify the router's cross-site routing state. 9.0-era method was vtysh BGP inspection on the HoloRouter; the equivalent on the 9.1 router (and whether FRR is still the routing daemon): verify against live 9.1 toolkit.
Start the Site B VCF 9.1.0.0 deployment and monitor to completion, watching physical host memory/CPU for contention with Site A throughout (Academy-measured 9.0 baseline: 90-150 minutes; official docs publish only 9.0.0.0 timings — ManagementOnly 4-5 h — so treat Academy numbers as environment-specific; verify 9.1 timing).
Validation Gate
Check: Query the toolkit for all instances: both Site A and Site B should report Running with version 9.1.0.0 (exact query cmdlet/fields: verify against live 9.1 toolkit)
Expected: Two Holodeck instances in Running state, both VCF 9.1.0.0
Common Errors
Task 3 Configure cross-site networking and validate bidirectional reachability
manageabilityCross-site networking is where the architecture comes alive: Layer 3 design, failure domain isolation, and management traffic engineering. On 9.1 the DNS layer changed (Technitium replaces the dnsmasq-era path) and VNA clusters may change where routing lives — so this task doubles as the re-baselining exercise for every 9.0-era networking assumption.
Verify cross-site routes on the 9.1 router function: both site supernets must be reachable in each direction. Inspection method (vtysh/FRR vs a new mechanism): verify against live 9.1 toolkit.
Test reachability from Site A to the Site B management plane: ping the Site B SDDC Manager and vCenter IPs.
Test reachability from Site B to the Site A management plane: ping the Site A SDDC Manager and vCenter IPs.
Configure DNS cross-resolution between sites. On 9.1 this is done in Technitium DNS (admin UI on :5380, behind the HTTPS reverse proxy) — the 9.0-era method of editing dnsmasq configuration does NOT carry forward. Exact zone/record workflow: verify against live 9.1 toolkit, then add each site's management FQDNs to the other site's resolution path.
Validate SDDC Manager awareness of both sites from each site's UI (Inventory > Workload Domains). Document any licensing constraints on multi-domain configuration.
Optional: configure NSX federation between the two sites' NSX Managers (System > Federation), if licensing allows. NSX version pairing on VCF 9.1.0.0: verify against live deployment.
Test management plane operations across sites (e.g. vCenter inventory visibility, test object creation) to prove bidirectional operational reachability.
Validation Gate
Check: Three checks: (1) Site A pings Site B management IPs, (2) Site B pings Site A management IPs, (3) cross-site FQDNs resolve in both directions via Technitium
Expected: All three pass. Bidirectional IP and DNS reachability confirmed.
Common Errors
Task 4 Validate dual-site management plane operations
manageabilityPost-deployment validation proves the dual-site architecture is operational and ready for VCDX-level defense: single-pane-of-glass management, per-site visibility, and cross-site link health. These UI-driven checks are largely version-agnostic; note any 9.1 UI differences as enrichment material.
From the (authenticated) Webtop or your workstation browser, open SDDC Manager A and verify the management domain shows Active with all hosts and clusters healthy.
Open SDDC Manager B and verify the Site B management domain shows Active.
Open vCenter A > Hosts and Clusters: verify all nested hosts connected, vSAN health green, HA/DRS enabled.
Open vCenter B > Hosts and Clusters: verify the same for Site B.
Optional: configure vCenter Enhanced Linked Mode between the two sites using FQDNs (requires the Task 3 DNS cross-resolution to be working).
Open each site's NSX Manager > System > Fabric > Nodes and verify all host transport nodes show Success/Up.
Test cross-site management visibility (Linked Mode inventory or per-site UIs) by creating a test object at Site A and confirming visibility per your chosen management model.
Document dual-site operational readiness: SDDC Managers healthy, vCenters healthy, NSX healthy, cross-site ping and DNS confirmed, management traffic flowing. Add a 9.1-specific appendix noting every place the live experience diverged from this stub.
Validation Gate
Check: Verify: (1) both SDDC Managers Active, (2) both vCenters healthy with green vSAN/HA/DRS, (3) Linked Mode connected if deployed, (4) all NSX transport nodes up at both sites
Expected: All management components dual-site healthy; cross-site operations functional
Common Errors
Task 5 Design exercise — map dual-site Holodeck to production DR architecture using RCAR
recoverabilityThis is the VCDX defense moment: articulate the architectural decisions, trade-offs, and risk mitigations of the working dual-site lab using the RCAR framework. The 9.1 twist worth defending: dual-site is now the toolkit default — discuss what 'DR rehearsal out of the box' changes about lab-to-production mapping, and where VNA distributed networking would alter the failure domain story.
Define the production DR scenario your dual-site lab represents: primary site, DR site, business objectives, RPO/RTO targets.
Create the RCAR analysis skeleton: Requirements, Constraints, Assumptions, Risks.
REQUIREMENTS: list at least 5 (replication capability, network failover, cross-site management, RPO/RTO, encrypted management traffic).
CONSTRAINTS: list at least 4 (single physical host, lab latency vs WAN reality, licensing, L2 adjacency limits; add 9.1-specific unknowns such as VNA sizing minimums).
ASSUMPTIONS: list at least 4 (replication licensing, cross-site latency/bandwidth, Site B readiness, manual vs automated DNS failover).
RISKS: list at least 6 with impact and mitigation (cross-site link failure, split-brain vCenter authority, RPO breach from replication lag, full site power loss, missing management-plane backups, operator error during failover).
Map the 5 VCDX design quality dimensions (Availability, Manageability, Performance, Recoverability, Security) to your dual-site design, including trade-offs. For Security, incorporate the 9.1 posture: HTTPS-everywhere reverse proxy, Vault-managed secrets, Authentik SSO — and contrast with the 9.0-era open-services model.
Create a failover decision tree covering at least: link-only failure (do not fail over), complete site power loss (fail over), and partial management-plane failure (manage from surviving components).
Write (do not execute) a failover runbook with step timings and estimated total RTO; compare against your target.
Validation Gate
Check: All 9 steps completed: scenario, RCAR worksheet, four RCAR sections, quality matrix, decision tree, runbook with RTO estimate
Expected: Design documentation comprehensive enough for VCDX oral defense with detailed follow-up answers available for any section
Common Errors
Final Validation
A complete dual-site Holodeck 9.1 deployment (Site A and Site B, both VCF 9.1.0.0) is operational with validated cross-site routing, Technitium-based DNS cross-resolution, management plane visibility, and a production-ready DR design documented via RCAR.
✓ Two Holodeck instances running (Site A and Site B) with non-overlapping ranges → Toolkit query shows two Running instances, both VCF 9.1.0.0
✓ Cross-site IP reachability in both directions → Management-plane pings succeed Site A <-> Site B
✓ Cross-site DNS via Technitium → FQDNs for each site resolve from the other site
✓ Both SDDC Managers show management domain Active → Inventory > Workload Domains healthy at both sites
✓ Both vCenters healthy; NSX transport nodes up at both sites → Green health indicators across both management planes
✓ RCAR design exercise completed → Requirements, Constraints, Assumptions, Risks, quality matrix, decision tree, and runbook documented
Cleanup / Restore
Snapshot: holodeck91-06-complete
• Take snapshots of all Holodeck VMs from both sites, named 'holodeck91-06-complete'
• Document the dual-site topology as actually deployed: IP addressing, FQDN mappings, cross-site routing, and every point where 9.1 diverged from this stub
• Archive the RCAR design exercise output for VCDX preparation
• If on a shared host: record the Site B allocations in the shared registry and communicate the router state changes to other instance owners
Design Reflection (VCDX)
Dual-site is the VCDX differentiator. A panelist probes: why dual-site vs single-site; how sites are decoupled operationally; latency-budget assumptions; split-brain handling; validated RTO/RPO. The 9.1-specific angles to defend: (1) dual-site auto-generation makes DR rehearsal a default rather than an advanced configuration — what does that change about your design process and your collision-avoidance controls? (2) VNA distributed networking offers an alternative to the single-router failure domain — when would you model with VNA clusters instead, and what does that buy in availability terms?
Requirements
- Replicate application VMs Site A to Site B within a defined RPO
- Cross-site management plane reachability (IP + DNS) in both directions
- Per-site SDDC Manager lifecycle independence
- Failover RTO within the defined target including detection, DNS update, and power-on sequencing
- Encrypted, authenticated cross-site management traffic (9.1 HTTPS-everywhere posture)
Constraints
- Single physical host = shared failure domain for both sites (lab-only constraint)
- 9.1 router/VNA dual-site mechanics unverified until live deployment
- Cross-site latency in the lab (<2ms) does not represent WAN reality (50-150ms)
- Licensing limits on NSX federation, replication, and multi-domain configuration
- Shared-host router state changes require coordination with other instance owners
Assumptions
- The auto-generated Site B config can be adjusted safely before deployment
- 9.0 baseline dual-site resource footprint (~768 GB total) approximates 9.1 until verified
- Technitium DNS supports the cross-site record configuration this design requires
- Manual DNS failover is acceptable to the business scenario chosen
Risks
- Cross-site link failure stalls replication — MITIGATION: latency/lag monitoring with explicit failover thresholds
- Split-brain vCenter authority post-failover — MITIGATION: runbook declares the sole authoritative site and isolation steps
- Auto-generated Site B collides with shared-host tenants — MITIGATION: pre-deploy allocation review and registry
- 9.0-era router workarounds applied blindly to the 9.1 stack corrupt router state — MITIGATION: treat all 9.0 fixes as unverified; re-diagnose from live logs
- Resource exhaustion during Site B bring-up destabilizes Site A — MITIGATION: continuous memory monitoring, conservative sizing, pause/resume plan
- Missing management-plane backups discovered during a drill — MITIGATION: scheduled SDDC Manager and vCenter backups stored off-instance
Self-Assessment Discussion Prompts
- 9.1 auto-generates the Site B configuration. What does 'DR by default' change about your design review process compared to 9.0's manual dual-site workflow?
- Where does the failure domain boundary sit in this lab vs production, and how would VNA distributed networking move it?
- How do you prevent two vCenters from independently managing the same VM during a network partition?
- Compare stretched vSAN synchronous replication vs separate clusters with async replication for this scenario's RPO/RTO/cost trade-offs
- If cross-site latency were 50ms instead of <2ms, which components could no longer be stretched and why?
- The 9.1 router stack is HTTPS-everywhere with authenticated services. How does that change the security section of your dual-site defense compared to the 9.0-era open-services router?
References
- HoloDeck 9.1 Release Announcement (captured verbatim)
Holodeck-9.1/01-release-announcement.mdlocal fileTier 1 — Official
Confirms dual-site auto-generation, Site B port-group fixes, VLAN range fixes, VNA cluster support, and the 9.1 services stack. - HoloDeck Toolkit 9.1 — Structured Reference & Changelog
Holodeck-9.1/02-structured-changelog.mdlocal fileTier 1 — Official
Structured 9.1 capture including the AMITH01 shared-host cautions for dual-site alignment. - Holodeck Toolkit GitHub RepositoryTier 1 — Official
Verify 9.1 dual-site documentation and cmdlet reference here once published. - VMware Cloud Foundation Multi-Site DocumentationTier 1 — Official
Locate the VCF 9.1 multi-site deployment guide for production-side design mapping. - William Lam — VCF Multi-Site and DR Design PatternsTier 3 — Expert Blog
Community deep dives on multi-site VCF and stretched cluster designs.