Academy/Holodeck Lab Setup & Operations/Dual-Site Holodeck Deployment (Holodeck 9.1)
This lab targets VCF 9.1.0.0

Dual-Site Holodeck Deployment (Holodeck 9.1)

VCF 9.1.0.0Advancedarchitectvcdx⏱ 180 min

Holodeck 9.1 deploys VCF 9.1.0.0, VVF 9.1.0.0 and VCF 5.2.4. Supported range: VCF 9.1.0.0 down to VCF 5.2 (9.1.0.0, 9.0.2.0, 9.0.1.0, 9.0.0.0, 5.2.4, 5.2.3, 5.2.2, 5.2.1, 5.2). Minimum nested ESX 8.0 U3. Docs dated May 2026; GA blog dated 2026-07-01 (conflict unresolved). No GitHub releases are published for vmware/Holodeck — vmware.github.io/Holodeck/9.1/ is the sole source of truth. Exact toolkit build number and PowerShell module version to be captured during the live W5 deployment.

Objectives

  • Design a non-overlapping IP addressing scheme for a dual-site Holodeck 9.1 topology, starting from the auto-generated Site A + Site B config pair
  • Deploy the Site B management domain with cross-site routing on Holodeck 9.1
  • Establish and validate cross-site routing between the Site A and Site B router functions (single-router vs VNA model: verify on live 9.1)
  • Validate bidirectional management plane reachability and DNS cross-resolution using the 9.1 Technitium DNS service
  • Map a dual-site lab topology to production VCDX disaster recovery and stretched cluster architecture using the RCAR framework

Prerequisites

Single Holodeck 9.1 management domain (VCF 9.1.0.0) fully deployed and healthy (holodeck91-02 completed). Physical host must have doubled resource headroom for Site B (9.0 baseline: minimum 768 GB RAM, 64+ logical cores, 2+ TB SSD — verify 9.1 sizing). Site A instance in a stable state — snapshot taken. No existing dual-site configuration active.

Prior labs: holodeck91-02

Required skills:

  • VCF management domain architecture (SDDC Manager, vCenter, NSX)
  • IP address planning and CIDR notation
  • BGP routing concepts (or comfort with static route configuration)
  • vCenter Enhanced Linked Mode concepts
  • NSX federation and multi-site networking
  • Disaster recovery and RTO/RPO analysis
📸 Starting State: S3-91 — Management Domain Deployed (VCF 9.1.0.0)

Lab Environment

Dual Holodeck 9.1 instances on a single physical host. Site A (production) and Site B (DR) use the non-overlapping CIDR/VLAN allocations from the auto-generated 9.1 config pair (9.0 baseline pattern: Site A 10.1.0.0/20, Site B 10.2.0.0/20 — verify the 9.1 auto-generated defaults). Cross-site routing is provided by the HoloRouter dual-site function; on 9.1, evaluate whether VNA cluster networking changes the router topology.

graph TB
  PHY[Single Physical ESXi Host] --> RTR[HoloRouter 9.1 dual-site or VNA cluster — verify]
  subgraph SiteA[Site A - Production]
    SDDC_A[SDDC Manager A]
    VC_A[vCenter A]
    NSX_A[NSX Manager A]
  end
  subgraph SiteB[Site B - DR]
    SDDC_B[SDDC Manager B]
    VC_B[vCenter B]
    NSX_B[NSX Manager B]
  end
  RTR -->|Cross-site routing| SiteA
  RTR -->|Cross-site routing| SiteB

IP Addressing

NetworkPurposeVLAN
10.1.0.0/20Site A management supernet (9.0 baseline default — verify 9.1 auto-generated allocation)9.0 baseline VLAN 1644-1648 range — verify 9.1 defaults (VLAN range handling fixed in 9.1)
10.2.0.0/20Site B management supernet (9.0 baseline default — verify 9.1 auto-generated allocation)9.0 baseline VLAN 1704-1708 range — verify 9.1 defaults (Site B port-group issues fixed in 9.1)

Credentials

SystemUsernamePassword
Physical ESXi HostrootSet during ESXi installation
HoloRouter 9.1 services (Webtop :30000 now authenticated; Authentik SSO :9443)verify against live 9.1 toolkitWebtop is behind authentication in 9.1; Authentik defaults are an open question — capture on first live deployment
SDDC Manager / vCenter (both sites)administrator@vsphere.localSet in the VCF bring-up spec per site
NSX Manager (both sites)adminSet in the VCF bring-up spec per site. 9.0-era doubled-password API quirk: verify whether it persists on 9.1

Tasks

Task 1 Plan the dual-site topology and IP addressing scheme

availability

In VCDX design, topology and addressing are foundational. 9.1 changes the starting point: the toolkit hands you an auto-generated Site A + Site B config pair, so planning becomes a review-and-align exercise rather than a from-scratch design — which makes collision analysis on shared hosts MORE important, not less. This task mirrors the VCF Planning and Preparation Workbook exercise for multi-site deployments.

Step 1

Document your Site A baseline (from holodeck91-02) in a design worksheet: management CIDR, per-function subnets, and VLAN allocations as actually deployed.

Worksheet showing the live Site A network baseline
Step 2

Review the auto-generated Site B configuration produced by New-HoloDeckConfig (confirmed 9.1 behavior: both sites generated in one pass). Verify every Site B CIDR and VLAN is non-overlapping with Site A AND with anything else on the physical network.

Site B allocation reviewed and confirmed non-overlapping
Overlapping CIDRs were a documented 9.0-era failure mode. Even though 9.1 auto-generates Site B, on shared hosts the generated defaults can still collide with OTHER instances — always verify.
Step 3

Document cross-site routing requirements: both sites must reach each other's management plane (SDDC Manager, vCenter, NSX). Record the gateway/router interfaces for each direction. How the 9.1 router (single dual-site HoloRouter vs VNA cluster) exposes these interfaces: verify against live 9.1 toolkit.

Routing table document showing bidirectional routes and gateway IPs
Step 4

Analyze physical resource requirements: Site B roughly doubles the footprint (docs-published 9.1 dual-site minimums: 96 logical cores, 1536 GB RAM, 5 TB NVMe total). Verify actual per-site consumption (plus GitOps/VNA overhead) and confirm the physical host has headroom; document contention risks if not.

Resource allocation table: Site A usage, Site B projection, total vs physical capacity
Step 5

Create a failure domain analysis: on a single physical host, both sites share one failure domain. Document how this differs from production dual-site and how it affects the RTO/RPO strategy developed in Task 5.

Failure domain matrix: failure mode, blast radius, detection, recovery, mitigation

Validation Gate

Check: Review the worksheet: (1) Site A/B CIDRs non-overlapping, (2) VLAN ranges non-overlapping, (3) cross-site routing table complete, (4) resources verified against capacity, (5) failure domain analysis documented

Expected: All five points pass. Ready to deploy Site B.

Common Errors

Auto-generated Site B ranges collide with another instance on a shared physical host
Cause: 9.1 auto-generation cannot know about other tenants' allocations
Fix: Adjust the Site B config before deploying; coordinate an allocation registry with other shared-host users.
Physical host memory exhaustion projected during planning
Cause: Insufficient buffer between Site A actuals and total host capacity
Fix: Reduce Site B nested host count or per-host sizing; or defer Site B to a second physical host.

Task 2 Deploy the Site B Holodeck instance with cross-site network configuration

availability

Executing the dual-site deployment is the operational foundation for VCDX multi-site architecture discussions: two independent VCF bring-ups on shared infrastructure, with resource contention and ordering dependencies. On 9.1 the exact cmdlet sequence is unverified — capture it live and note what the dual-site auto-generation removed from the 9.0-era manual workflow.

Step 1

Open your management session (PowerShell workstation or the authenticated 9.1 Webtop) and navigate to the Holodeck runtime directory.

Session ready at the Holodeck runtime location
Step 2

Locate the auto-generated Site B configuration from your New-HoloDeckConfig run (see holodeck91-01/91-02). Whether a separate per-site config-generation step still exists in 9.1: verify against live 9.1 toolkit.

Site B configuration identified; ConfigID/reference recorded
Step 3

Confirm the Site A network configuration in use matches what was deployed in holodeck91-02 (no drift).

Site A config confirmed against the live instance
Step 4

Apply any Site B adjustments decided in Task 1 (CIDR/VLAN alignment, resource sizing). Note: Site B port-group issues were fixed in 9.1 — if you still hit port-group errors, capture them as new 9.1 findings rather than assuming 9.0 causes.

Site B config updated with reviewed values
Step 5

Enable the router's dual-site operation. The 9.0-era command was Set-HoloRouter -dualsite; whether 9.1 retains this cmdlet, changes its semantics under the new services stack, or replaces it with VNA cluster configuration: verify against live 9.1 toolkit.

Router dual-site (or VNA) configuration completes successfully
On shared hosts the 9.0-era Set-HoloRouter was destructive to global router state. Re-validate this behavior on 9.1 BEFORE running it on any shared router — coordinate with other instance owners.
Step 6

Review the Site B configuration one final time: site identifier, management CIDR, resource sizing within the contention envelope planned in Task 1.

Site B config verified ready for deployment
Step 7

Run the prepare/staging phase for Site B (PowerShell path: New-HoloDeckInstance with the Site B config; GitOps path: trigger the Site B pipeline from the UI — exact 9.1 flow: verify).

Prepare completes; content validated; Site B router interfaces configured
Step 8

Verify the router's cross-site routing state. 9.0-era method was vtysh BGP inspection on the HoloRouter; the equivalent on the 9.1 router (and whether FRR is still the routing daemon): verify against live 9.1 toolkit.

Cross-site routes/peerings present for both site supernets
Step 9

Start the Site B VCF 9.1.0.0 deployment and monitor to completion, watching physical host memory/CPU for contention with Site A throughout (Academy-measured 9.0 baseline: 90-150 minutes; official docs publish only 9.0.0.0 timings — ManagementOnly 4-5 h — so treat Academy numbers as environment-specific; verify 9.1 timing).

Site B management domain deployment completes; instance reports Running

Validation Gate

Check: Query the toolkit for all instances: both Site A and Site B should report Running with version 9.1.0.0 (exact query cmdlet/fields: verify against live 9.1 toolkit)

Expected: Two Holodeck instances in Running state, both VCF 9.1.0.0

Common Errors

Router dual-site configuration hangs or fails
Cause: 9.0-era causes were FRR/iptables timing on the HoloRouter; 9.1 router internals differ (services stack behind reverse proxy) — causes must be re-baselined
Fix: Capture router logs on the live 9.1 appliance and diagnose fresh; do not assume 9.0 restart sequences apply. Verify against live 9.1 toolkit.
Site B deployment fails on storage allocation
Cause: Physical datastore exhausted after Site A allocation, or datastore name mismatch
Fix: Verify free space and datastore naming; clean up orphaned VMs from failed attempts before retrying.
Site B vCenter deployment fails mid-bring-up
Cause: 9.0-era pattern: memory oversubscription or Site B DNS not resolving; 9.1 DNS is served by Technitium — re-validate the DNS chain
Fix: Check physical host memory utilization; verify Site B DNS resolution against the 9.1 Technitium service (admin UI on :5380) rather than 9.0-era dnsmasq assumptions.

Task 3 Configure cross-site networking and validate bidirectional reachability

manageability

Cross-site networking is where the architecture comes alive: Layer 3 design, failure domain isolation, and management traffic engineering. On 9.1 the DNS layer changed (Technitium replaces the dnsmasq-era path) and VNA clusters may change where routing lives — so this task doubles as the re-baselining exercise for every 9.0-era networking assumption.

Step 1

Verify cross-site routes on the 9.1 router function: both site supernets must be reachable in each direction. Inspection method (vtysh/FRR vs a new mechanism): verify against live 9.1 toolkit.

Both site CIDRs present with valid next-hops in the routing state
Step 2

Test reachability from Site A to the Site B management plane: ping the Site B SDDC Manager and vCenter IPs.

Replies from both Site B management IPs with low latency (same physical host)
Step 3

Test reachability from Site B to the Site A management plane: ping the Site A SDDC Manager and vCenter IPs.

Replies from both Site A management IPs
Step 4

Configure DNS cross-resolution between sites. On 9.1 this is done in Technitium DNS (admin UI on :5380, behind the HTTPS reverse proxy) — the 9.0-era method of editing dnsmasq configuration does NOT carry forward. Exact zone/record workflow: verify against live 9.1 toolkit, then add each site's management FQDNs to the other site's resolution path.

Cross-site FQDNs resolve in both directions
Capture the Technitium workflow in detail during the live run — it replaces one of the most-referenced 9.0-era fixes (dnsmasq upstream forwarders) and the KB has zero 9.1 content on it yet.
Step 5

Validate SDDC Manager awareness of both sites from each site's UI (Inventory > Workload Domains). Document any licensing constraints on multi-domain configuration.

Each SDDC Manager shows its own management domain Active; constraints documented
Step 6

Optional: configure NSX federation between the two sites' NSX Managers (System > Federation), if licensing allows. NSX version pairing on VCF 9.1.0.0: verify against live deployment.

Federation connected in both directions, or the licensing constraint documented
Step 7

Test management plane operations across sites (e.g. vCenter inventory visibility, test object creation) to prove bidirectional operational reachability.

Management operations succeed against both sites from a single operator session

Validation Gate

Check: Three checks: (1) Site A pings Site B management IPs, (2) Site B pings Site A management IPs, (3) cross-site FQDNs resolve in both directions via Technitium

Expected: All three pass. Bidirectional IP and DNS reachability confirmed.

Common Errors

Cross-site ping fails with destination unreachable
Cause: Cross-site routing not established on the 9.1 router function
Fix: Re-inspect routing state on the live 9.1 router; capture the working diagnostic sequence for the KB (9.0-era vtysh commands: verify they still apply).
Cross-site DNS resolution fails (NXDOMAIN/timeout)
Cause: Cross-site records not configured in Technitium, or clients still pointing at a stale DNS IP
Fix: Verify records in the Technitium admin UI (:5380) and confirm the DNS server IPs handed to nested components; document the correct 9.1 DNS chain.

Task 4 Validate dual-site management plane operations

manageability

Post-deployment validation proves the dual-site architecture is operational and ready for VCDX-level defense: single-pane-of-glass management, per-site visibility, and cross-site link health. These UI-driven checks are largely version-agnostic; note any 9.1 UI differences as enrichment material.

Step 1

From the (authenticated) Webtop or your workstation browser, open SDDC Manager A and verify the management domain shows Active with all hosts and clusters healthy.

SDDC Manager A dashboard: management domain Active
Step 2

Open SDDC Manager B and verify the Site B management domain shows Active.

SDDC Manager B dashboard: management domain Active
Step 3

Open vCenter A > Hosts and Clusters: verify all nested hosts connected, vSAN health green, HA/DRS enabled.

vCenter A cluster healthy
Step 4

Open vCenter B > Hosts and Clusters: verify the same for Site B.

vCenter B cluster healthy
Step 5

Optional: configure vCenter Enhanced Linked Mode between the two sites using FQDNs (requires the Task 3 DNS cross-resolution to be working).

Linked Mode connected in both directions, or documented as skipped
Step 6

Open each site's NSX Manager > System > Fabric > Nodes and verify all host transport nodes show Success/Up.

All transport nodes up at both sites
Step 7

Test cross-site management visibility (Linked Mode inventory or per-site UIs) by creating a test object at Site A and confirming visibility per your chosen management model.

Bidirectional management plane visibility demonstrated
Step 8

Document dual-site operational readiness: SDDC Managers healthy, vCenters healthy, NSX healthy, cross-site ping and DNS confirmed, management traffic flowing. Add a 9.1-specific appendix noting every place the live experience diverged from this stub.

Status report with all health metrics passing plus 9.1 divergence notes

Validation Gate

Check: Verify: (1) both SDDC Managers Active, (2) both vCenters healthy with green vSAN/HA/DRS, (3) Linked Mode connected if deployed, (4) all NSX transport nodes up at both sites

Expected: All management components dual-site healthy; cross-site operations functional

Common Errors

Linked Mode pairing fails to establish
Cause: Cross-site DNS or certificate validation failing (FQDN mismatch)
Fix: Verify Technitium cross-site records resolve; use FQDNs (not IPs) for pairing; accept self-signed certificates when prompted.
Some Site B transport nodes show Install Failed
Cause: NSX VIB installation timeout under resource contention
Fix: Use the NSX Resolve action on the failed hosts; check Site B host memory and network connectivity if it recurs.

Task 5 Design exercise — map dual-site Holodeck to production DR architecture using RCAR

recoverability

This is the VCDX defense moment: articulate the architectural decisions, trade-offs, and risk mitigations of the working dual-site lab using the RCAR framework. The 9.1 twist worth defending: dual-site is now the toolkit default — discuss what 'DR rehearsal out of the box' changes about lab-to-production mapping, and where VNA distributed networking would alter the failure domain story.

Step 1

Define the production DR scenario your dual-site lab represents: primary site, DR site, business objectives, RPO/RTO targets.

Scenario document with quantified RPO/RTO targets
Step 2

Create the RCAR analysis skeleton: Requirements, Constraints, Assumptions, Risks.

RCAR worksheet created
Step 3

REQUIREMENTS: list at least 5 (replication capability, network failover, cross-site management, RPO/RTO, encrypted management traffic).

Requirements list with at least 5 quantified items
Step 4

CONSTRAINTS: list at least 4 (single physical host, lab latency vs WAN reality, licensing, L2 adjacency limits; add 9.1-specific unknowns such as VNA sizing minimums).

Constraints list with at least 4 items
Step 5

ASSUMPTIONS: list at least 4 (replication licensing, cross-site latency/bandwidth, Site B readiness, manual vs automated DNS failover).

Assumptions list with at least 4 items
Step 6

RISKS: list at least 6 with impact and mitigation (cross-site link failure, split-brain vCenter authority, RPO breach from replication lag, full site power loss, missing management-plane backups, operator error during failover).

Risks list with at least 6 items, each with impact and mitigation
Step 7

Map the 5 VCDX design quality dimensions (Availability, Manageability, Performance, Recoverability, Security) to your dual-site design, including trade-offs. For Security, incorporate the 9.1 posture: HTTPS-everywhere reverse proxy, Vault-managed secrets, Authentik SSO — and contrast with the 9.0-era open-services model.

Design quality matrix with design choice, trade-off, and rationale per dimension
Step 8

Create a failover decision tree covering at least: link-only failure (do not fail over), complete site power loss (fail over), and partial management-plane failure (manage from surviving components).

Decision tree with at least 3 failure scenarios and actions
Step 9

Write (do not execute) a failover runbook with step timings and estimated total RTO; compare against your target.

Failover runbook with estimated RTO vs target

Validation Gate

Check: All 9 steps completed: scenario, RCAR worksheet, four RCAR sections, quality matrix, decision tree, runbook with RTO estimate

Expected: Design documentation comprehensive enough for VCDX oral defense with detailed follow-up answers available for any section

Common Errors

RCAR sections are vague or generic (unquantified 'high availability' requirements)
Cause: Not connecting the lab topology to specific production constraints
Fix: Map each deployed component to a specific RCAR entry with concrete numbers (latency budgets, capacity, license counts).
Failover decision tree is not actionable
Cause: Missing partial-failure scenarios (e.g. vCenter down but hosts up)
Fix: Expand to at least 5 scenarios covering per-component and link-level failures, each with an explicit failover/no-failover decision.

Final Validation

A complete dual-site Holodeck 9.1 deployment (Site A and Site B, both VCF 9.1.0.0) is operational with validated cross-site routing, Technitium-based DNS cross-resolution, management plane visibility, and a production-ready DR design documented via RCAR.

✓ Two Holodeck instances running (Site A and Site B) with non-overlapping ranges → Toolkit query shows two Running instances, both VCF 9.1.0.0

✓ Cross-site IP reachability in both directions → Management-plane pings succeed Site A <-> Site B

✓ Cross-site DNS via Technitium → FQDNs for each site resolve from the other site

✓ Both SDDC Managers show management domain Active → Inventory > Workload Domains healthy at both sites

✓ Both vCenters healthy; NSX transport nodes up at both sites → Green health indicators across both management planes

✓ RCAR design exercise completed → Requirements, Constraints, Assumptions, Risks, quality matrix, decision tree, and runbook documented

Cleanup / Restore

Snapshot: holodeck91-06-complete

• Take snapshots of all Holodeck VMs from both sites, named 'holodeck91-06-complete'

• Document the dual-site topology as actually deployed: IP addressing, FQDN mappings, cross-site routing, and every point where 9.1 diverged from this stub

• Archive the RCAR design exercise output for VCDX preparation

• If on a shared host: record the Site B allocations in the shared registry and communicate the router state changes to other instance owners

Design Reflection (VCDX)

Dual-site is the VCDX differentiator. A panelist probes: why dual-site vs single-site; how sites are decoupled operationally; latency-budget assumptions; split-brain handling; validated RTO/RPO. The 9.1-specific angles to defend: (1) dual-site auto-generation makes DR rehearsal a default rather than an advanced configuration — what does that change about your design process and your collision-avoidance controls? (2) VNA distributed networking offers an alternative to the single-router failure domain — when would you model with VNA clusters instead, and what does that buy in availability terms?

Requirements

  • Replicate application VMs Site A to Site B within a defined RPO
  • Cross-site management plane reachability (IP + DNS) in both directions
  • Per-site SDDC Manager lifecycle independence
  • Failover RTO within the defined target including detection, DNS update, and power-on sequencing
  • Encrypted, authenticated cross-site management traffic (9.1 HTTPS-everywhere posture)

Constraints

  • Single physical host = shared failure domain for both sites (lab-only constraint)
  • 9.1 router/VNA dual-site mechanics unverified until live deployment
  • Cross-site latency in the lab (<2ms) does not represent WAN reality (50-150ms)
  • Licensing limits on NSX federation, replication, and multi-domain configuration
  • Shared-host router state changes require coordination with other instance owners

Assumptions

  • The auto-generated Site B config can be adjusted safely before deployment
  • 9.0 baseline dual-site resource footprint (~768 GB total) approximates 9.1 until verified
  • Technitium DNS supports the cross-site record configuration this design requires
  • Manual DNS failover is acceptable to the business scenario chosen

Risks

  • Cross-site link failure stalls replication — MITIGATION: latency/lag monitoring with explicit failover thresholds
  • Split-brain vCenter authority post-failover — MITIGATION: runbook declares the sole authoritative site and isolation steps
  • Auto-generated Site B collides with shared-host tenants — MITIGATION: pre-deploy allocation review and registry
  • 9.0-era router workarounds applied blindly to the 9.1 stack corrupt router state — MITIGATION: treat all 9.0 fixes as unverified; re-diagnose from live logs
  • Resource exhaustion during Site B bring-up destabilizes Site A — MITIGATION: continuous memory monitoring, conservative sizing, pause/resume plan
  • Missing management-plane backups discovered during a drill — MITIGATION: scheduled SDDC Manager and vCenter backups stored off-instance

Self-Assessment Discussion Prompts

  1. 9.1 auto-generates the Site B configuration. What does 'DR by default' change about your design review process compared to 9.0's manual dual-site workflow?
  2. Where does the failure domain boundary sit in this lab vs production, and how would VNA distributed networking move it?
  3. How do you prevent two vCenters from independently managing the same VM during a network partition?
  4. Compare stretched vSAN synchronous replication vs separate clusters with async replication for this scenario's RPO/RTO/cost trade-offs
  5. If cross-site latency were 50ms instead of <2ms, which components could no longer be stretched and why?
  6. The 9.1 router stack is HTTPS-everywhere with authenticated services. How does that change the security section of your dual-site defense compared to the 9.0-era open-services router?

References

  • HoloDeck 9.1 Release Announcement (captured verbatim) Holodeck-9.1/01-release-announcement.mdlocal fileTier 1 — Official
    Confirms dual-site auto-generation, Site B port-group fixes, VLAN range fixes, VNA cluster support, and the 9.1 services stack.
  • HoloDeck Toolkit 9.1 — Structured Reference & Changelog Holodeck-9.1/02-structured-changelog.mdlocal fileTier 1 — Official
    Structured 9.1 capture including the AMITH01 shared-host cautions for dual-site alignment.
  • Holodeck Toolkit GitHub RepositoryTier 1 — Official
    Verify 9.1 dual-site documentation and cmdlet reference here once published.
  • VMware Cloud Foundation Multi-Site DocumentationTier 1 — Official
    Locate the VCF 9.1 multi-site deployment guide for production-side design mapping.
  • William Lam — VCF Multi-Site and DR Design PatternsTier 3 — Expert Blog
    Community deep dives on multi-site VCF and stretched cluster designs.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.