Academy/Holodeck Lab Setup & Operations/Dual-Site Auto-Generation on Holodeck 9.1 — One-Pass Site A + Site B Configuration and Second-Site Deployment
This lab targets VCF 9.1.0.0

Dual-Site Auto-Generation on Holodeck 9.1 — One-Pass Site A + Site B Configuration and Second-Site Deployment

VCF 9.1.0.0Advancedarchitectvcdx⏱ 180 min

Holodeck 9.1 supports deployment of VCF 9.1.0.0, VVF 9.1.0.0, and the VCF 5.2.4 maintenance line (full range VCF 9.1.0.0 down to 5.2). This lab targets a VCF 9.1.0.0 dual-site deployment. 9.0.x twin: holodeck-06 (Dual-Site Holodeck Deployment) — the 9.0.x per-site manual workflow described there is superseded by 9.1 one-pass auto-generation but remains valid background reading. Do NOT mix 9.0.x and 9.1 toolkit modules in the same PowerShell session.

Objectives

  • Generate a complete dual-site (Site A + Site B) Holodeck configuration in a single New-HoloDeckConfig pass and explain how one-pass generation differs from the 9.0.x per-site workflow
  • Inspect and verify the auto-generated Site A and Site B configuration files for non-overlapping CIDR, VLAN, and port-group allocation
  • Switch the active PowerShell session context between site configurations using Import-HoloDeckConfig with the 9.1 -Site parameter
  • Run pre-deployment validation for both sites and deploy the Site B management domain alongside an existing Site A instance
  • Validate cross-site reachability and management-plane health across the dual-site topology
  • Articulate, in VCDX defense terms, why automated configuration generation reduces configuration drift, and defend the choice between dual-site DR and stretched-cluster topologies

Prerequisites

Holodeck 9.1 tooling installed and a Site A VCF 9.1.0.0 management domain deployed and healthy (holodeck91-02 completed), OR a clean physical host if you intend to deploy both sites in this lab. Physical host sizing for dual-site: minimum 768 GB RAM, 64+ logical cores, 2+ TB SSD (the Site B footprint doubles the single-site allocation). Nested ESX template must be 8.0 U3 or later — Holodeck 9.1 enforces this minimum. Snapshot of the Site A instance taken before starting.

Prior labs: holodeck91-02

Required skills:

  • Holodeck 9.1 lifecycle cmdlets (New-HoloDeckConfig, Import-HoloDeckConfig, New-HoloDeckInstance, Start-HoloDeckPrecheck) from holodeck91-01/02
  • IP address planning and CIDR notation
  • VCF management domain architecture (SDDC Manager, vCenter, NSX) — note the 9.1 unified management services runtime
  • BGP routing concepts or comfort with static route inspection (vtysh)
  • Disaster recovery concepts: RTO/RPO, failure domains, active-passive vs stretched topologies
📸 Starting State: S3-91 — Management Domain Deployed (VCF 9.1.0.0)

Lab Environment

Single physical ESXi host (or modest vSphere cluster) running two Holodeck 9.1 instances. Site A (production role) uses the 10.1.0.0/20 management supernet; Site B (DR role) uses 10.2.0.0/20. Unlike 9.0.x — where each site's network configuration was generated manually and independently (a documented source of overlap errors) — Holodeck 9.1 generates both site configurations from a single New-HoloDeckConfig invocation, guaranteeing deterministic, non-overlapping CIDR/VLAN allocation. The HoloRouter provides routing between sites plus the 9.1 built-in service stack: Technitium DNS (port 5380), HashiCorp Vault (8200), Authentik SSO (9443), and an authenticated Webtop (30000), all behind an HTTPS reverse proxy.

graph TB
  PHY[Physical ESXi Host - 768+ GB RAM] --> HR[HoloRouter 9.1<br/>HTTPS reverse proxy<br/>Technitium DNS :5380 / Vault :8200 / Authentik :9443]
  HR -->|One-pass config generation| CFG[New-HoloDeckConfig<br/>Site A + Site B configs generated together]
  subgraph SiteA["Site A - Production (10.1.0.0/20)"]
    ESXA[4x Nested ESX 8.0U3+ 10.1.1.10x]
    SDDCA[SDDC Manager-A 10.1.0.4]
    VCA[vCenter-A 10.1.0.6]
    NSXA[NSX Manager-A 10.1.0.10]
  end
  subgraph SiteB["Site B - DR (10.2.0.0/20)"]
    ESXB[4x Nested ESX 8.0U3+ 10.2.1.10x]
    SDDCB[SDDC Manager-B 10.2.0.4]
    VCB[vCenter-B 10.2.0.6]
    NSXB[NSX Manager-B 10.2.0.10]
  end
  HR --> SiteA
  HR --> SiteB
  SiteA <-->|Cross-site routing via HoloRouter| SiteB

IP Addressing

NetworkPurposeVLAN
10.1.0.0/20Site A management supernet (auto-generated by New-HoloDeckConfig)Site A VLAN block (auto-assigned — record actual IDs from generated config)
10.1.1.0/24Site A nested ESX host managementSite A management VLAN
10.2.0.0/20Site B management supernet (auto-generated in the same pass)Site B VLAN block (auto-assigned, guaranteed non-overlapping with Site A)
10.2.1.0/24Site B nested ESX host managementSite B management VLAN
172.16.0.0/12NSX overlay ranges (both sites, segmented per site by the generated configs)Overlay / Geneve

Credentials

SystemUsernamePassword
HoloRouter SSHrootSet at HoloRouter OVA deployment (9.1 default documented in toolkit docs)
SDDC Manager A / Badministrator@vsphere.localSet in the generated bring-up spec. API calls use the doubled master password convention (e.g. VMware123!VMware123!) — verify this convention still holds on your 9.1 build
vCenter A / Badministrator@vsphere.localSame SSO password as the site's SDDC Manager
NSX Manager A / BadminSet in generated bring-up spec; doubled password for API auth
HoloRouter web services (Vault/Authentik/Technitium/Webtop)admin (service-specific)9.1 puts all HoloRouter services behind Authentik SSO + HTTPS reverse proxy — retrieve initial credentials per the 9.1 toolkit docs; Authentik defaults were an open question at release, verify on your build

Tasks

Task 1 Generate the dual-site configuration in a single pass

manageability

This is the headline 9.1 workflow change (rn-9-1-018). In 9.0.x you ran New-HoloDeckNetworkConfig once per site and manually kept CIDRs and VLAN ranges from colliding — the community KB documents real failures from exactly that manual step. In 9.1, one New-HoloDeckConfig invocation emits both site configurations with deterministic, non-overlapping allocation. A VCDX panelist will recognize this as the classic 'automate the error-prone step' pattern: the tool now owns cross-site consistency instead of the operator.

Step 1

On the management workstation (or HoloConsole), open PowerShell 7.x and import the Holodeck 9.1 toolkit module. Confirm the module version reports a 9.1.x build: Get-Module HoloDeckToolkit | Select-Object Name, Version

Module loaded, version shows 9.1.x. If a 9.0.x module version loads, stop — remove the old module path before proceeding.
Do not run 9.0.x and 9.1 toolkit modules in the same session. Cmdlet names overlap but parameter sets differ (notably the new -Site behavior on Import-HoloDeckConfig).
Step 2

Generate the dual-site configuration. Run New-HoloDeckConfig with your target host parameters (TargetHost, Username, Password, datastore/network selections as in holodeck91-01). In 9.1 this single invocation generates TWO configuration files: one for Site A and one for Site B.

Success output referencing two generated configurations — a Site A config (loaded into the current session) and a Site B config (generated on disk, not loaded). Record both ConfigId values.
The 9.1 release notes state the dual-generation exists to eliminate race conditions when deploying dual-site in parallel: both configs are computed from the same allocation state at the same moment, so neither site can grab a range the other already claimed.
Version Notes: 9.0.x behavior for comparison: New-HoloDeckConfig generated one config; per-site network configs came from separate New-HoloDeckNetworkConfig -Site a / -Site b runs. That two-step flow is where CIDR overlap mistakes crept in (KB thread 10).
Step 3

Verify which site context is loaded. Run Get-HoloDeckConfig and confirm the session shows the Site A configuration as active. Note the ConfigId, InstanceID, and site indicator fields.

Get-HoloDeckConfig returns the Site A configuration metadata. Site B's config is visible on disk but is not the active session context.
Step 4

Inspect both generated configuration files on disk (location per your 9.1 install, typically under the holodeck runtime config directory — verify the exact path against your live 9.1 toolkit). Compare side by side: management CIDR (expect 10.1.0.0/20 vs 10.2.0.0/20 or your chosen supernets), VLAN ranges, port-group names, and reserved component IPs for SDDC Manager, vCenter, and NSX on each site.

Two config files with symmetric structure and strictly non-overlapping CIDRs, VLAN IDs, and port-group names. Component IP plans mirror each other across the 10.1.x / 10.2.x split.
Diff the two files programmatically (Compare-Object on parsed JSON) and save the diff as a design artifact — it is a compact, defensible record of exactly what differs between your two sites: addresses only, not architecture.
Step 5

Practice the site context switch. Load Site B: Import-HoloDeckConfig -ConfigId <site-b-config-id> -Site b. Run Get-HoloDeckConfig again and confirm the active context now reports Site B. Then switch back to Site A the same way.

Active configuration context flips to Site B (10.2.x addressing visible in queries), then back to Site A. No errors on either switch.
Every subsequent cmdlet in this lab is site-context-sensitive. Adopt the habit now: before ANY deployment or query action, run Get-HoloDeckConfig and confirm which site you are pointed at. Running a destructive action against the wrong site context is the 9.1 equivalent of the classic 'wrong terminal window' incident.

Validation Gate

Check: Two configuration files exist (Site A + Site B) from a single New-HoloDeckConfig pass. Diff confirms non-overlapping CIDRs/VLANs/port groups. Import-HoloDeckConfig -Site b successfully switches session context and back.

Expected: Dual-site configuration generated and verified; operator can deliberately control site context.

Common Errors

Only one configuration file is generated
Cause: You are running a 9.0.x toolkit module, or a 9.1 pre-GA build without the dual-generation behavior
Fix: Verify module version is 9.1 GA or later. Remove stale module paths ($env:PSModulePath) and re-import. Re-run New-HoloDeckConfig.
Get-HoloDeck* queries return Site A data even though you intended to work on Site B
Cause: Site B config was generated but never loaded — Site A is auto-loaded by default in 9.1
Fix: Run Import-HoloDeckConfig -ConfigId <site-b-id> -Site b, then re-run the query. Confirm with Get-HoloDeckConfig before proceeding.
New-HoloDeckConfig warns about template config integrity
Cause: The template config.json was hand-modified (behavior introduced in 9.0.1 template integrity validation, still present in 9.1)
Fix: Do not edit the template file. Restore the pristine template from the toolkit distribution and pass customizations via cmdlet parameters instead.
📋 KB: 9.0.1 GA announcement — template integrity validation

Task 2 Precheck both sites and verify the 9.1 dual-site fixes

availability

9.1 hardened exactly the failure points that broke 9.0.x dual-site deployments: VLAN range handling, Site B port groups, and VCF installer validation (rn-9-1-020). Prechecking both sites BEFORE committing hours of deployment time is the operational discipline a panel expects — 'how do you know the second site will deploy cleanly?' should have a better answer than 'we tried it'.

Step 1

With the Site A context loaded, run the pre-deployment validation: Start-HoloDeckPrecheck -Site a. Review every check: binary/content validation, target host reachability, networking (port groups, VLANs), storage headroom.

All Site A prechecks pass. If Site A is already deployed (from holodeck91-02), the precheck should reflect a healthy existing allocation.
Step 2

Switch context to Site B (Import-HoloDeckConfig -ConfigId <site-b-id> -Site b) and run Start-HoloDeckPrecheck -Site b.

All Site B prechecks pass, including the Site B port-group checks that were failure-prone on 9.0.x. Distinct port groups per site should be validated automatically (behavior introduced in 9.0.1, fixed and hardened in 9.1).
Step 3

Verify storage and memory headroom for the second site on the physical host BEFORE deploying. From the physical ESXi host client (or esxtop), record current consumed memory and datastore free space. Site B needs approximately the same footprint Site A consumed (~384 GB RAM and ~500+ GB disk for a 4-host management domain — verify against your actual Site A consumption).

Documented headroom calculation: physical capacity minus Site A actuals leaves enough for Site B plus a safety buffer (target: keep projected total below ~90% of physical RAM).
The dominant dual-site failure mode is not networking — it is memory exhaustion mid-deployment of Site B, which destabilizes the already-running Site A. KB thread 19 (Site B vCenter not deploying) is this exact failure. If headroom is marginal, reduce Site B nested host sizing before deploying, not after it fails.
Step 4

Verify nested ESX template version. Holodeck 9.1 enforces a minimum nested ESX of 8.0 U3. Confirm the template/ISO version the toolkit will use for Site B hosts.

Nested ESX version is 8.0 U3 or later. If older, update the template before deployment — the 9.1 installer validation will reject it (validation hardened in 9.1).

Validation Gate

Check: Start-HoloDeckPrecheck passes for both -Site a and -Site b. Headroom calculation documented. Nested ESX template confirmed at 8.0 U3+.

Expected: Both sites cleared for deployment with documented capacity evidence.

Common Errors

Site B precheck fails on port group validation
Cause: Manually created or leftover port groups from a previous 9.0.x dual-site attempt collide with the auto-generated Site B port-group plan
Fix: Remove stale Holodeck port groups from the physical host (verify nothing is attached first), then re-run the precheck. Do not hand-create Site B port groups on 9.1 — the toolkit owns them.
📋 KB: rn-9-1-020: Site B port group issues resolved in 9.1
Precheck passes but you doubt the VLAN allocation
Cause: Legacy habit from 9.0.x where VLAN range handling had defects
Fix: 9.1 fixed VLAN range handling (rn-9-1-020), but trust-and-verify: dump both site configs' VLAN assignments and confirm zero intersection. File a toolkit issue if you find overlap — that would be a regression.

Task 3 Deploy the Site B management domain alongside Site A

availability

Deploying a second full VCF instance next to a live one is a controlled-risk exercise in shared-infrastructure thinking: the two sites are logically independent (separate SDDC Manager, vCenter, NSX, vSAN) but physically entangled (same host, same HoloRouter). VCDX panels probe exactly this distinction — 'what is actually isolated here, and what only looks isolated?'

Step 1

Take a snapshot of all Site A VMs before starting (Get-VM pattern from earlier labs). Label it clearly (e.g. 'pre-siteb-deploy').

Snapshot exists for every Site A Holodeck VM. This is your rollback line if Site B deployment destabilizes the host.
Step 2

Confirm the Site B context is loaded (Get-HoloDeckConfig shows Site B), then start the deployment: New-HoloDeckInstance with the Site B configuration, followed by (or including, per your 9.1 workflow) Start-HoloDeckInstance for the Site B InstanceID. Use -Verbose for a progress stream. Alternative path: 9.1 also supports driving this deployment from the GitOps/UI pipeline (rn-9-1-016) — this lab uses the cmdlet path for repeatability; note the alternative in your runbook.

Site B deployment proceeds through nested ESX provisioning, installer validation (hardened in 9.1), management appliance deployment, and VCF bring-up. Expect 90-150 minutes; longer under resource contention with Site A.
Monitor physical host memory every 10-15 minutes during deployment (esxtop). If utilization exceeds ~95%, Site B components will start failing in confusing ways (vCenter services crashing, installer timeouts) and Site A becomes collateral damage.
Step 3

While deployment runs, verify Site A remains healthy: SDDC Manager-A UI reachable, vCenter-A shows all hosts connected, vSAN health green. Record any degradation with timestamps.

Site A stays fully operational throughout Site B deployment. Any Site A degradation correlates with physical resource pressure and should be documented as a shared-infrastructure risk observation for the design exercise.
Step 4

When deployment completes, query both instances. Switch contexts as needed and run Get-HoloDeckInstance for each site's InstanceID.

Two instances reported: Site A and Site B, both Running, both VCF 9.1.0.0.
Step 5

Log in to the Site B management plane: SDDC Manager-B (10.2.0.4 or the address from your generated config), vCenter-B, NSX Manager-B. Verify the management domain shows Active with all hosts commissioned. Note: 9.1's unified management services runtime (rn-9-1-001) means the SDDC Manager / VCF Operations footprint differs from what the 9.0 labs describe — record what you actually observe.

Site B management domain Active: hosts connected, vSAN green, NSX transport nodes up. Addresses match the auto-generated Site B config exactly.

Validation Gate

Check: Get-HoloDeckInstance shows Site A and Site B both Running on VCF 9.1.0.0. Site B SDDC Manager/vCenter/NSX all healthy. Site A suffered no persistent degradation.

Expected: Two independent VCF 9.1 management domains operational on the dual-site topology.

Common Errors

Site B deployment actions execute against Site A objects
Cause: Session context was still Site A when deployment cmdlets ran
Fix: Stop immediately. Verify damage scope on Site A (snapshot from step 1 is your recovery line). Re-import the Site B config with -Site b and re-run. Add a context-check step to your personal runbook.
Site B vCenter deployment fails; vmware-vpostgres or other services crash on start
Cause: Physical host memory oversubscription, or Site B DNS not resolving during vCenter firstboot
Fix: Check physical host memory (esxtop); reduce Site B nested host memory or count if >95%. Verify Site B DNS against the HoloRouter's Technitium DNS service — confirm Site B zones exist (Task 4).
📋 KB: Thread 19: Holodeck dual-site — Site B vCenter not deploying
Installer validation rejects the Site B environment
Cause: 9.1's hardened VCF installer validation catching a real problem earlier than 9.0.x would have (nested ESX below 8.0 U3, storage shortfall, network mismatch)
Fix: Read the validation failure precisely — 9.1 validation failing fast is a feature. Fix the flagged precondition rather than looking for a bypass.

Task 4 Validate the dual-site topology — cross-site routing, DNS, and management-plane independence

manageability

A dual-site lab is only useful for DR rehearsal if the sites can actually reach each other AND fail independently. This task validates both properties. The 9.1 twist: DNS is now served by Technitium on the HoloRouter behind the HTTPS reverse proxy (rn-9-1-017), which changes how you inspect and extend cross-site name resolution compared to the 9.0.x dnsmasq procedures.

Step 1

Verify cross-site IP reachability. From a Site A source (HoloConsole or a Site A component), ping Site B SDDC Manager and vCenter (10.2.0.4, 10.2.0.6 or your generated addresses). Repeat in the reverse direction from Site B.

Bidirectional replies, <2ms latency (same physical host). Both site supernets are routed via the HoloRouter.
Step 2

Inspect cross-site routing on the HoloRouter. SSH to the HoloRouter and review the routing state for both site supernets (vtysh 'show ip route' if FRR-based routing is in use on your build, or the routing method your 9.1 router uses — verify against live 9.1 toolkit, as the 9.1 router service stack changed substantially from 9.0.x).

Routes for both 10.1.0.0/20 and 10.2.0.0/20 present and pointing at the correct interfaces.
Step 3

Validate cross-site DNS with Technitium. Open the Technitium DNS console on the HoloRouter management IP, port 5380 (behind the HTTPS reverse proxy — authenticate as required on your build). Verify both Site A and Site B zones/records were provisioned by the dual-site generation. Test resolution both ways: from Site A, resolve a Site B FQDN; from Site B, resolve a Site A FQDN.

Technitium shows record sets for both sites. nslookup of cross-site FQDNs resolves correctly from both directions.
Version Notes: 9.0.x labs (holodeck-06 Task 3) had you hand-edit dnsmasq configs for cross-site resolution. On 9.1, verify whether Technitium fully replaces dnsmasq on your build (this was an open question at release) before applying any legacy DNS workaround — the old dnsmasq fix likely does not apply.
Step 4

Verify management-plane independence. Confirm that Site A and Site B each have their own SDDC Manager, vCenter, SSO domain, NSX management cluster, and vSAN datastore — i.e., no shared management state between sites other than the HoloRouter transit/DNS layer.

Component inventory table for both sites showing fully parallel, independent management stacks.
Step 5

Perform a controlled independence test: restart a non-critical Site B management service (or briefly disconnect a Site B nested host NIC), and confirm Site A registers zero impact. Restore immediately.

Site A health remains green throughout the Site B disturbance. Document this as evidence of logical failure-domain separation — while noting the shared physical host remains a common failure domain.

Validation Gate

Check: Bidirectional cross-site ping and DNS resolution pass. Both site supernets routed. Component inventory confirms parallel independent management stacks. Site B disturbance produced no Site A impact.

Expected: Dual-site topology validated: connected where intended, isolated where intended.

Common Errors

Cross-site DNS fails but cross-site ping works
Cause: Technitium zones for the second site missing or the client is pointed at a stale DNS address from a previous 9.0.x deployment
Fix: Check client resolv.conf/DNS settings point at the HoloRouter DNS IP. Inspect Technitium zone list on :5380. Re-sync the instance config if records are missing.
📋 KB: Thread 30: DNS resolution in Holodeck dual-site deployment (9.0.x-era; verify applicability on 9.1 Technitium stack)
Cannot reach Technitium/Vault/Authentik consoles on their documented ports
Cause: 9.1 places all HoloRouter services behind the HTTPS reverse proxy with authentication — direct-port habits from 9.0.x may not apply
Fix: Use https:// against the HoloRouter management IP with the documented port, accept the proxy certificate, authenticate via Authentik. Consult the 9.1 toolkit docs for the exact ingress paths on your build.

Task 5 Design exercise — auto-generation, drift, and the dual-site vs stretched decision

recoverability

This is the defense-facing payload of the lab. You now have first-hand evidence of what one-pass configuration generation buys you (deterministic non-overlap, no race conditions, symmetric sites) and what it does not (shared physical failure domain, no automatic DR orchestration). Convert that evidence into VCDX-grade positions using RCAR.

Step 1

Write a one-page position: 'Why automated dual-site configuration generation reduces configuration drift.' Ground it in what you observed: (a) both sites derive from one allocation computation — no human transcription between sites; (b) the config diff you produced in Task 1 shows addresses differ, architecture does not; (c) the 9.0.x failure modes this eliminates (KB thread 10 CIDR overlap, Site B port group collisions) were all human-transcription errors; (d) race-condition elimination when generating/deploying sites in parallel.

Position paper connecting the tooling change to the drift/error classes it eliminates, with the Task 1 diff attached as evidence
Step 2

Define the production scenario your lab models: Site A primary, Site B DR, active-passive, with target RPO/RPO. State explicitly which lab properties transfer to production (symmetric config generation, independent management stacks, cross-site DNS) and which do not (shared physical host, <2ms inter-site latency, single HoloRouter as shared transit).

Scenario statement with an honest lab-vs-production applicability table
Step 3

Build the dual-site vs stretched-cluster comparison. For each topology document: failure domain boundaries, latency requirements (stretched vSAN needs ~<5ms RTT and L2 adjacency; dual-site async replication tolerates WAN latency), RPO achievable (near-zero sync vs minutes-hours async), operational complexity (witness placement and split-brain handling vs failover runbooks), and licensing/cost. Conclude with which business profiles each fits.

Comparison matrix plus a defensible recommendation logic: 'choose stretched when..., choose dual-site DR when...'
Step 4

Complete an RCAR worksheet for the dual-site design (minimum: 5 requirements, 4 constraints, 4 assumptions, 6 risks with impact and mitigation). Include at least one risk unique to 9.1 tooling: e.g., wrong-site-context operator error enabled by fast context switching.

Completed RCAR worksheet
Step 5

Prepare panel-response drills: answer aloud, in under two minutes each — (a) 'Your two sites were generated from one template. What happens to your symmetry claim after two years of independent Day-2 change?' (b) 'Why not stretch the cluster instead?' (c) 'Where exactly does your failure-domain boundary sit in this lab, and where would it sit in production?'

Recorded or written answers demonstrating command of drift-over-time, topology selection, and failure-domain reasoning

Validation Gate

Check: All five artifacts exist: drift position paper with diff evidence, scenario + applicability table, dual-site vs stretched matrix, RCAR worksheet, panel drill answers.

Expected: Design documentation strong enough to survive follow-up probing in a VCDX defense.

Final Validation

A dual-site Holodeck 9.1 deployment is operational from a single-pass auto-generated configuration: Site A and Site B configs produced together with guaranteed non-overlap, Site B deployed alongside a live Site A, cross-site routing and Technitium DNS validated bidirectionally, and management-plane independence demonstrated. The design exercise converts the tooling change into defense material: automation of the cross-site consistency problem, and a reasoned dual-site vs stretched topology position.

✓ Task 1: Single New-HoloDeckConfig pass produced Site A + Site B configs; context switching via Import-HoloDeckConfig -Site b verified → Two config files, non-overlapping allocation confirmed by diff, session context controllable

✓ Task 2: Prechecks pass for both sites with documented capacity headroom → Start-HoloDeckPrecheck green for -Site a and -Site b; ESX 8.0 U3+ confirmed

✓ Task 3: Site B deployed alongside live Site A → Get-HoloDeckInstance shows both instances Running on VCF 9.1.0.0; Site A undisturbed

✓ Task 4: Cross-site reachability, DNS, and independence validated → Bidirectional ping + Technitium DNS resolution; independent management stacks; Site B disturbance invisible to Site A

✓ Task 5: Design exercise complete → Drift position, topology comparison, RCAR, and panel drills documented

Cleanup / Restore

Snapshot: holodeck91-05-complete

• Snapshot all VMs of BOTH sites with a clear label (e.g. 'holodeck91-05-complete') — subsequent Day-2 labs (holodeck91-08/09) can run against either site

• Archive the Task 1 config diff, RCAR worksheet, and topology comparison to your VCDX evidence folder

• If physical resources are constrained, Stop-HoloDeckInstance the Site B instance (do NOT remove it) to reclaim RAM for later labs — record which site is stopped in your lab notebook

• If you must fully reclaim Site B: use Remove-HoloDeckInstance against the SITE B context only, and verify the site context twice before executing. Do not use -ResetHoloRouter on a shared or dual-site router — it tears down global routing/DNS state for everything on the host

Design Reflection (VCDX)

Two defense threads converge in this lab. THREAD 1 — configuration drift as an availability risk: panels increasingly probe operational consistency, not just steady-state architecture. The 9.1 one-pass generation is a concrete instance of the strongest anti-drift pattern: derive all instances from one computation rather than validating N hand-built instances against each other. Be ready to generalize: the same argument justifies infrastructure-as-code for production VCF (and the toolkit's own GitOps path), golden-image host provisioning, and declarative NSX policy.

Also be ready for the counterpunch: generation guarantees symmetry at time zero only — without ongoing enforcement (drift detection, redeploy-from-source discipline), two years of independent Day-2 change erodes the symmetry claim. THREAD 2 — topology selection: dual-site active-passive DR vs stretched cluster is a classic panel fork. Dual-site (this lab): independent failure domains, tolerant of WAN latency, asynchronous replication RPO in minutes, explicit failover with runbooks — simpler to reason about, slower to recover.

Stretched: near-zero RPO and transparent HA restart across sites, but demands <5ms RTT, L2 adjacency or overlay stretch, witness placement, and rigorous split-brain design — a harder operational commitment. The VCDX move is to bind the choice to stated business requirements (RPO/RTO, distance between sites, network budget, operational maturity) rather than to technology preference.

Finally, name the lab's honesty caveat before the panel does: both 'sites' share one physical host and one HoloRouter, so the lab rehearses logical topology and operational workflow, not genuine site-level failure isolation.

Requirements

  • Two complete, independently manageable VCF 9.1 management domains (Site A production role, Site B DR role)
  • Deterministic, non-overlapping network allocation across sites, produced without manual cross-site transcription
  • Bidirectional cross-site management-plane reachability (IP and DNS) for replication and failover tooling
  • Site B deployable next to a live Site A without disrupting Site A operations
  • Operator ability to target either site unambiguously (session context control) for all lifecycle actions
  • Design documentation sufficient to defend the dual-site vs stretched decision against stated RPO/RTO targets

Constraints

  • Single physical host: both sites share one hardware failure domain and one HoloRouter transit/DNS layer — site-level disaster isolation cannot be genuinely rehearsed
  • Physical capacity: dual-site requires ~768 GB RAM / 64+ cores / 2+ TB SSD; Site B sizing may need reduction on smaller hosts
  • Holodeck 9.1 minimum nested ESX 8.0 U3 — older templates are rejected by the hardened installer validation
  • Lab latency (<2ms) is unrepresentative of production inter-site links; latency-sensitive conclusions (stretched vSAN feasibility) cannot be validated here
  • 9.1 HoloRouter service stack (Technitium/Vault/Authentik behind HTTPS proxy) invalidates 9.0.x DNS/service procedures; exact ingress paths and defaults must be verified per build

Assumptions

  • Site A instance (holodeck91-02) is healthy and snapshotted before Site B deployment begins
  • The auto-generated Site B allocation does not collide with any OTHER instance on a shared physical host (coordinate InstanceIDs and CIDRs with colleagues on shared hardware)
  • The doubled-password API convention from 9.0.x persists on 9.1 builds (verify on first API call)
  • Cmdlet-driven deployment path chosen for repeatability; GitOps/UI path (rn-9-1-016) assumed equivalent in outcome but not exercised here
  • Production mapping assumes real dual-site would add: independent physical infrastructure per site, WAN latency, replication tooling, and failover orchestration not present in the lab

Risks

  • Physical memory exhaustion during Site B deployment destabilizes live Site A — IMPACT: both sites degraded, possible total lab loss. MITIGATION: pre-deployment headroom calculation (Task 2), esxtop monitoring during deploy, Site A snapshots as rollback line, reduce Site B sizing when marginal.
  • Wrong-site-context operator error: destructive cmdlet executed against the unintended site — IMPACT: damage to the healthy site. MITIGATION: Get-HoloDeckConfig context check before every action; personal runbook rule; snapshots before any destructive step.
  • Stale 9.0.x artifacts (port groups, DNS records, module versions) contaminate the 9.1 dual-site deployment — IMPACT: precheck failures or subtle routing/DNS misbehavior. MITIGATION: clean the host of prior-version artifacts; never mix toolkit module versions in one session.
  • Symmetry erosion over time: Day-2 changes applied to one site only — IMPACT: DR site no longer representative; failover fails when needed. MITIGATION: change-management rule that Day-2 changes apply to both sites or are logged as accepted asymmetry; periodic config re-diff against the generated baseline.
  • Shared HoloRouter as hidden single point of failure for BOTH sites' transit and DNS — IMPACT: one router fault severs cross-site operations and name resolution everywhere. MITIGATION: acknowledge explicitly in the design; in production map this to redundant WAN/DNS infrastructure; see holodeck91-10 for the distributed VNA alternative to single-router networking.
  • Treating lab DR rehearsal as production DR validation — IMPACT: false confidence in RPO/RTO claims. MITIGATION: applicability table (Task 5 step 2) states exactly which conclusions transfer; production targets validated only by production-representative testing.

Self-Assessment Discussion Prompts

  1. Explain precisely how one-pass dual-site generation eliminates the CIDR-overlap failure class. What error classes does it NOT eliminate?
  2. Your two sites are symmetric today because one tool generated both. Describe a governance mechanism that keeps them defensibly symmetric after two years of Day-2 operations.
  3. The 9.1 release notes say dual generation 'eliminates race conditions when deploying dual-site in parallel.' What was the race, and what does its existence tell you about the 9.0.x allocation model?
  4. Defend choosing dual-site active-passive DR over a stretched cluster for a business with RPO 1 hour / RTO 4 hours and 40ms inter-site RTT. Now invert it: what requirement changes would flip your recommendation to stretched?
  5. In this lab, what is the true failure-domain boundary between Site A and Site B? List every shared component and classify each as acceptable-in-lab-only or acceptable-in-production.
  6. A panelist asks: 'Auto-generation is convenient, but you cannot explain an allocation you did not make. Walk me through why Site B got the VLANs it got.' How do you answer without hand-waving?
  7. How would you extend this dual-site foundation toward a rehearsable failover: what replication, DNS failover, and runbook elements are missing between this lab and a testable DR capability?

Extensions

Drive the Site B Deployment via the GitOps Pipeline

Repeat the Site B deployment using the 9.1 Click-to-Deploy GitOps path (HoloRouter OVA with GitOps enabled, GitLab-driven pipeline, rn-9-1-016) instead of cmdlets. Compare: auditability of the pipeline vs verbose cmdlet logs, rollback story, and how the GitOps path handles the dual-site configuration pair. Document which path you would standardize on for a team lab and why.

same

Dual-Site Workload Domain Rehearsal

With both management domains up, deploy a small VI workload domain on each site (see holodeck-05 methodology, adapted to 9.1) and document the full dual-site estate: per-site licensing, capacity, and the management-plane inventory a DR runbook would need. Note where 9.1's unified management services runtime changes the component inventory vs the 9.0 twin lab.

harder

Cross-Site Failover Runbook and Tabletop Exercise

Author a complete Site A to Site B failover runbook (detection, decision tree, DNS cutover via Technitium, service verification, RTO accounting) and run it as a tabletop exercise without executing destructive steps. Time each phase against a 2-hour RTO target and identify the longest-lead-time step.

harder

Drift Injection and Detection

Deliberately introduce three configuration drifts on Site B (e.g., an extra port group, a changed NTP source, a modified DNS record) and build a detection script that diffs live state against the generated configuration baseline. This turns the drift position paper from Task 5 into a working control.

same

⚠ Known Pitfalls (from Community KB)

Holodeck 9.0.2 dual-site IP config SUPERSEDED_IN_9_1
Problem: 9.0.x-era failure: Site A and Site B manually configured with overlapping CIDR ranges, causing routing conflicts after both sites deployed; HoloRouter reboot was part of the resolution.
Resolution: Root cause class (manual per-site config transcription) is eliminated by 9.1 one-pass dual generation. Remains relevant as the motivating failure story for Task 5's drift discussion, and the HoloRouter-state lesson still applies.
Holodeck 9.0.2 Dual Site B vCenter not deploying STILL_RELEVANT
Problem: Site B vCenter deployment failed due to insufficient remaining memory on the shared physical host after Site A allocation.
Resolution: Capacity, not tooling — fully applicable on 9.1. Mitigate with the Task 2 headroom calculation and live esxtop monitoring during Site B deployment; reduce Site B sizing when marginal.
Issue with DNS resolution in Holodeck dual-site deployment VERIFY_ON_9_1
Problem: Cross-site FQDN resolution failures in 9.0.x dual-site setups, resolved via dnsmasq forwarder configuration.
Resolution: 9.1 moves DNS to Technitium behind the HTTPS reverse proxy — the dnsmasq fix likely does not apply. Re-diagnose via the Technitium console (:5380) per Task 4; do not blindly apply the 9.0.x workaround.
Holodeck vcf 9 - Cannot select datastore STILL_RELEVANT
Problem: Deployment fails at storage selection due to full or misnamed datastore — commonly hit when deploying a second site into space already consumed by the first.
Resolution: Verify >500 GB free before Site B deployment and that the datastore name in the generated config matches physical inventory. Clean orphaned VMs/snapshots from failed prior attempts.

References

  • Announcing the General Availability of Holodeck 9.1 — VMware Cloud Foundation BlogTier 1 — Official
    Official GA announcement: VCF 9.1.0.0/VVF 9.1.0.0/5.2.4 support, dual-site auto-generation, VNA cluster support, Day-2 operations, HoloRouter service stack.
  • Holodeck 9.1 DocumentationTier 1 — Official
    Official 9.1 docs — authoritative for New-HoloDeckConfig dual-generation behavior, Import-HoloDeckConfig -Site parameter, and precheck workflow. Verify all cmdlet parameters here before deployment.
  • Simplify Multisite PoCs Using Holodeck — VMware Cloud Foundation BlogTier 1 — Official
    Multisite deployment patterns with Holodeck; written against 9.0.2 but the topology reasoning carries into 9.1.
  • Holodeck-9.1/01-release-announcement.md + 02-structured-changelog.md (repo-local) ../../Holodeck-9.1/02-structured-changelog.mdlocal fileTier 2 — VMware Press
    Captured release announcement and structured changelog used as the content baseline for this lab, including open questions to resolve against official release notes.
  • VCF Holodeck Toolkit — Broadcom Community ForumTier 1 — Official
    Threads 10 (dual-site IP config), 19 (Site B vCenter not deploying), 30 (dual-site DNS) document the 9.0.x failure modes this lab's 9.1 workflow addresses. Note corpus is 9.0.x-era — verify applicability on 9.1.
  • academy-v2/content/releases/vcf-9.1.json — rn-9-1-018, rn-9-1-020 ../releases/vcf-9.1.jsonlocal fileTier 2 — VMware Press
    Canonical changelog entries this lab implements: dual-site config auto-generation (rn-9-1-018) and the VLAN/Site-B/installer fixes (rn-9-1-020).
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.