Academy/Holodeck Lab Setup & Operations/Holodeck Instance Lifecycle Management (Holodeck 9.1)
This lab targets VCF 9.1.0.0

Holodeck Instance Lifecycle Management (Holodeck 9.1)

VCF 9.1.0.0Intermediateadminarchitectvcdx⏱ 90 min

Holodeck 9.1 deploys VCF 9.1.0.0, VVF 9.1.0.0 and VCF 5.2.4. Supported range: VCF 9.1.0.0 down to VCF 5.2 (9.1.0.0, 9.0.2.0, 9.0.1.0, 9.0.0.0, 5.2.4, 5.2.3, 5.2.2, 5.2.1, 5.2). Minimum nested ESX 8.0 U3. Docs dated May 2026; GA blog dated 2026-07-01 (conflict unresolved). No GitHub releases are published for vmware/Holodeck — vmware.github.io/Holodeck/9.1/ is the sole source of truth. Exact toolkit build number and PowerShell module version to be captured during the live W5 deployment.

Objectives

  • Design and execute a multi-checkpoint snapshot strategy for Holodeck 9.1 lab environments
  • Execute Remove-HoloDeckInstance on a 9.1 instance and verify complete cleanup of nested infrastructure (VMs, port groups, router state)
  • Redeploy from a fresh configuration pair (or GitOps pipeline re-run) and measure redeploy time vs snapshot revert time
  • Version-control the 9.1 Site A + Site B configuration pair and demonstrate config overlay patterns across the supported version range
  • Deploy and compare side-by-side VCF 5.2.4 and VCF 9.1.0.0 instances to articulate version-specific architectural differences
  • Automate Holodeck lifecycle operations via PowerShell and/or the GitOps pipeline, and document best practices for lab reset and scheduling

Prerequisites

Holodeck 9.1 pod with VCF 9.1.0.0 management domain deployed and healthy (holodeck91-02 completion snapshot available). Sufficient datastore space for 2+ snapshots per VM (9.0 baseline guidance: minimum 200 GB free — verify against the larger 9.1 footprint, 2.2 TB single-site baseline). PowerShell access to the Holodeck 9.1 runtime, and GitLab UI access if the GitOps driver was chosen in holodeck91-01.

Prior labs: holodeck91-02

Required skills:

  • Holodeck 9.1 instance deployment and basic PowerShell scripting
  • vSphere VM snapshot creation and revert operations
  • Understanding of VCF component versioning and binary staging structure
  • File system navigation and JSON editing
📸 Starting State: S3-91 — Management Domain Deployed (VCF 9.1.0.0)

Lab Environment

Primary Holodeck pod (VCF 9.1.0.0) from holodeck91-02. Secondary instance (VCF 5.2.4) deployed in Task 4 on the same physical host using a separate, non-overlapping configuration. Isolation via HoloRouter networking (9.1 defaults: Site A 10.1.0.0/20, Site B 10.2.0.0/20 — the auto-generated Site B config exists even for single-site deployments).

graph TB
  PHY[Physical ESXi Host — 9.1 envelope: 48 CPU / 650 GB / 2.2 TB single site] --> HR[HoloRouter 9.1 — services stack behind HTTPS proxy]
  HR --> MGMT-A[VCF 9.1.0.0 Instance]
  HR --> MGMT-B[VCF 5.2.4 Instance — Task 4, non-overlapping CIDR/VLAN]
  MGMT-A --> SNAP1[Snapshot: lcm-post-deploy-9.1.0.0]
  MGMT-A --> SNAP2[Snapshot: lcm-post-validation-9.1.0.0]
  HR -.->|GitOps path| GL[GitLab pipeline + runner]

IP Addressing

NetworkPurposeVLAN
10.1.0.0/20Site A primary instance (VCF 9.1.0.0) — 9.1 default MasterCIDR9.1 default Site A VLANs 0, 10-25 — verify actuals from the deployed instance
10.2.0.0/20Site B default supernet (auto-generated config pair). Task 4 secondary instance must use a non-overlapping custom /20 (e.g. 10.3.0.0/20) if Site B ranges are reserved9.1 default Site B VLANs 40-58 — verify; VLAN range handling was fixed in 9.1

Credentials

SystemUsernamePassword
Physical ESXi HostrootConfigured during ESXi installation
HoloRouter 9.1 (SSH) and services stack (Vault :8200, Authentik :9443, Technitium :5380, Webtop :30000, GitLab)root (SSH); per-service accounts per the 9.1 docs referenceRoot password set during OVA deployment. Service defaults documented in holodeck-official-docs-reference-9.1.md (master password convention) — confirm on the live deployment; Authentik admin username has a docs discrepancy (akadmin vs admin)
SDDC Manager / vCenteradministrator@vsphere.localMaster password from the VCF bring-up spec (9.1 docs publish a single master-password convention for all nested components)

Tasks

Task 1 Design and execute a multi-checkpoint snapshot strategy

recoverability

Snapshots are your lab reset button. In production they are not allowed as a protection mechanism, but in a lab they enable rapid rollback after experiments fail. VCDX panelists expect you to articulate when to snapshot (post-major-milestone), what to snapshot (which VMs), and the trade-offs (memory vs non-memory, disk overhead, consolidation timing). The mechanics are PowerCLI and largely version-agnostic; the 9.1 deltas are the larger VM inventory and disk footprint.

Step 1

Open your management session (PowerShell workstation or the authenticated 9.1 Webtop) and navigate to the Holodeck 9.1 runtime location (9.0 baseline: the holodeck-runtime directory; exact 9.1 layout on your chosen console: verify against live 9.1 toolkit).

Session ready at the Holodeck 9.1 runtime location
Step 2

Review the current instance state: Get-HoloDeckInstance (signature documented unchanged in 9.1). Record the instance ID and confirm the version reports 9.1.0.0 (exact output field names: verify against live 9.1 toolkit).

Instance in a running/healthy state with version 9.1.0.0
Step 3

List all Holodeck VMs currently deployed via PowerCLI (9.0 baseline naming: Get-VM -Name '<InstanceID>' or 'Holo-'). Record the actual 9.1 VM inventory — it may include additional appliances vs 9.0 (e.g. VNA cluster nodes if the distributed networking option was chosen).

Inventory captured: HoloRouter, nested ESX hosts, VCF Installer, SDDC Manager, vCenter, NSX Manager, plus any 9.1-specific appliances
Step 4

Check datastore free space to estimate snapshot overhead: Get-Datastore | Format-Table -Property Name, FreeSpaceGB, CapacityGB. The 9.1 footprint is larger than 9.0 (2.2 TB single-site baseline) — size snapshot headroom accordingly.

Free space recorded; sufficient headroom for a multi-snapshot strategy (9.0 baseline guidance: >200 GB — verify against 9.1 actuals)
Step 5

Create a non-memory snapshot at the post-deployment milestone: New-Snapshot across the instance VMs with name 'lcm-post-deploy-9.1.0.0', -Memory:$false -Quiesce:$false.

Snapshots created on all instance VMs
Non-memory snapshots revert in seconds vs minutes for memory-inclusive ones, at the cost of a guest reboot after revert. For lab checkpoints, non-memory is preferred — unchanged guidance from the 9.0 twin.
Step 6

Verify snapshot creation: Get-Snapshot across the instance VMs, sorted by name, showing Created and SizeGB.

All VMs show lcm-post-deploy-9.1.0.0 with a creation timestamp
Step 7

Perform a non-disruptive operational check (e.g. document current SDDC Manager state), then take a second checkpoint named 'lcm-post-validation-9.1.0.0'.

Second snapshot set created, representing a savepoint at a different phase
Step 8

Examine snapshot metadata and total growth: Get-Snapshot -Name 'lcm-*' | Measure-Object SizeGB -Sum. Keep total snapshot consumption below 50% of free datastore space.

Two snapshots per VM; total size within the safety envelope
Step 9

Document your naming convention in SNAPSHOT_NAMING.txt: '<component>-<phase>-<version>' (examples: lcm-post-deploy-9.1.0.0, lcm-post-wld-9.1.0.0). Version-scope every snapshot name so 9.0-era and 9.1 checkpoints can never be confused on a shared host.

Naming convention documented with 9.1-scoped examples

Validation Gate

Check: Verify: (1) at least 2 snapshots exist per VM, (2) names follow <component>-<phase>-<version> with the 9.1.0.0 version scope, (3) all snapshots are non-memory, (4) total snapshot size < 50% of free datastore space

Expected: Multi-checkpoint snapshot strategy in place with documented naming and growth tracking

Common Errors

New-Snapshot fails with 'Insufficient free space'
Cause: Datastore headroom consumed — more likely on 9.1 given the larger instance footprint
Fix: Check Get-Datastore free space; remove stale snapshots; if chronic, revisit the deployment sizing against the 9.1 envelope (soft validation permits under-provisioned hosts, which shows up here first).
Quiesced snapshots hang or time out
Cause: VMware Tools unresponsive in nested VMs — nested I/O latency makes quiescence unreliable
Fix: Use -Quiesce:$false in nested environments (unchanged guidance from the 9.0 twin).

Task 2 Destroy and redeploy an instance to verify cleanup and measure redeploy time

recoverability

The destroy-redeploy cycle is the ultimate validation that your instance is reproducible and the toolkit cleans up after itself. VCDX panelists probe: how do you ensure repeatability, detect orphaned resources, and automate this? On 9.1 there are two redeploy drivers to time (PowerShell vs GitOps pipeline re-run), and the one-config-one-deployment rule means a redeploy starts from a FRESH config pair.

Step 1

Before destroying, back up instance metadata: Get-HoloDeckInstance | ConvertTo-Json | Out-File instance-backup-before-destroy.json. Also record the deployment parameters actually used (Version, InstanceID, CIDR, VLANRangeStart, networking switches) — you need them for an identical redeploy.

Instance metadata and deployment parameters captured
Step 2

Stop the instance gracefully: Stop-HoloDeckInstance -InstanceID <id> (documented unchanged from 9.0.2 — graceful ordered shutdown). Prefer this over blunt per-VM power-off.

Ordered shutdown completes; all instance VMs powered off
Step 3

Record pre-destruction datastore state: capture FreeSpaceGB for the target datastore(s) for later comparison.

Datastore metrics recorded
Step 4

Execute instance removal: Remove-HoloDeckInstance (documented unchanged in 9.1; optional -ResetHoloRouter clears router networking with an automatic reboot). Decide DELIBERATELY whether to include -ResetHoloRouter: on a shared host it resets global router state for every tenant, and its behavior under the 9.1 services stack (Technitium, reverse proxy, GitLab runner) must be re-validated before use.

Removal completes: VM deletion, port group cleanup, state cleanup (exact 9.1 progress output: verify against live 9.1 toolkit)
Do not interrupt removal mid-run. On shared hosts, coordinate with other instance owners BEFORE any -ResetHoloRouter — the 9.0-era destructiveness caution carries forward until re-baselined on 9.1.
Step 5

Verify cleanup: (a) no instance VMs remain (Get-VM count 0 for the instance prefix), (b) instance-specific port groups removed, (c) datastore space reclaimed vs the step-3 baseline (expect a larger reclaim than 9.0 given the 9.1 footprint).

All three cleanup checks pass; reclaimed GB recorded
Step 6

Generate a FRESH configuration for the redeploy: New-HoloDeckConfig -TargetHost <host> -UserName <u> -Password <p>. Confirmed 9.1 behavior: this creates TWO config files (Site A + Site B) and auto-loads Site A into the session. Do NOT reuse the consumed config from the destroyed deployment (one-config-one-deployment rule).

New Site A + Site B config pair created; Site A loaded; ConfigID recorded
Step 7

Start the timed redeploy. PowerShell path: New-HoloDeckInstance -Version 9.1.0.0 -InstanceID <id> with the recorded parameters (start a stopwatch first). GitOps path: trigger the deployment pipeline from the GitLab UI and use the pipeline duration as the measurement.

Deployment progresses to completion; elapsed time recorded
Step 8

Compare redeploy time vs snapshot revert time. Reference points: snapshot revert ~5-10 minutes; 9.0-era Academy baseline for a mgmt-only redeploy was 90-150 minutes, while the official docs still publish only 9.0.0.0 timings (ManagementOnly 4-5 h). No 9.1-specific timings are published — your measurement IS the enrichment data for this stub.

Time comparison documented with actual 9.1 numbers

Validation Gate

Check: After destroy: (1) no instance VMs remain, (2) port groups cleaned, (3) datastore space recovered. After redeploy: (1) instance healthy at 9.1.0.0, (2) redeploy time measured and logged, (3) fresh config pair used (not the consumed one)

Expected: Complete lifecycle cycle demonstrated: destroy validates cleanup, redeploy validates reproducibility, timing captured for the 9.1 KB

Common Errors

Remove-HoloDeckInstance hangs during VM removal
Cause: VMs locked by snapshots or in-flight operations (9.0-era pattern; 9.1 behavior to re-baseline)
Fix: Remove stale snapshots, force power-off the stuck VM, retry. Capture the exact 9.1 failure signature for the KB — the corpus has zero 9.1 content.
Redeploy fails immediately at config load
Cause: Attempted reuse of the consumed config, or Import-HoloDeckConfig called without the now-mandatory -Site parameter
Fix: Generate a fresh pair with New-HoloDeckConfig; when re-importing in a new session always use Import-HoloDeckConfig -ConfigID <id> -Site a.
GitOps pipeline unreachable or fails to trigger
Cause: GitLab not accessible — known first-boot issue GH#139 (iptables 80/443 rules below the DROP rule) or runner not registered
Fix: Apply the GH#139 workaround (move the port 80/443 ACCEPT rules above the DROP rule), verify the runner registration, retry from the GitLab UI.

Task 3 Version-control the 9.1 configuration pair and demonstrate overlay patterns

manageability

Configuration management at scale: template reuse, inheritance, and override patterns instead of copy-paste. The 9.1 twist: every New-HoloDeckConfig run now yields a Site A + Site B pair, so your version-control scheme must track pairs, and version-scoped naming matters more because one toolkit drives targets from 5.2 to 9.1.0.0.

Step 1

Locate the 9.1 configuration artifacts on the HoloRouter: the template at /holodeck-runtime/templates/config.json (confirmed in the 9.1 docs) and per-deployment configs under /holodeck-runtime/config/. Confirmed: 9.1 creates TWO files per config (Site A + Site B); the naming scheme of the pair: verify against live 9.1 toolkit.

Template and per-deployment config pair located; pair naming documented
Step 2

Create a version-control directory structure for configs: configs/{base,overlays,deployed,archive} in your working area.

Directory structure created
Step 3

Copy the 9.1 template as your base: config-base-9.1.0.0.json. Capture the live template's default nested-host sizing while you are here — 9.1 defaults are unpublished (9.0 baseline: 4 hosts x 12 vCPU / 96 GB) and this is a W5 capture item.

Base config file created; 9.1 default sizing recorded for the lab notebook
Step 4

Record the valid 9.1 -Version strings in your convention doc: '9.1.0.0', '9.0.2.0', '9.0.1.0', '9.0.0.0', '5.2.4', '5.2.3', '5.2.2', '5.2.1', '5.2'. Version-scope every config filename (e.g. config-vcf9100-mgmt-only.json, config-vcf524-mgmt-only.json) so cross-version drift is impossible to miss.

Version matrix documented; filenames version-scoped
Step 5

Build an overlay example: load the base config into a PowerShell object, apply a scenario delta (e.g. reduced nested host count or a mgmt-only profile), and save as a deployed config. Use object manipulation (ConvertFrom-Json / ConvertTo-Json -Depth 20), never hand-edited JSON.

Overlay-derived config saved and parseable; modified values verified
Step 6

Document the naming convention in CONFIG_NAMING.md, including the new pair convention: every generated config is a Site A + Site B pair; archive them together; note which member was consumed by a deployment (one config = one deployment).

Naming convention documented including pair handling
Step 7

Create a VERSION_HISTORY.txt log recording each config's creation date, target version, deployment outcome, and measured deploy time.

Version history log created and populated with at least the holodeck91-02 and Task 2 entries
Step 8

Write a pre-deployment validation function: parse the config, assert the version string is in the valid 9.1 set, assert the binaries folder /holodeck-runtime/bin/<version>/ is staged, and assert the CIDR does not collide with any recorded allocation on the shared host. Note for session recovery: Import-HoloDeckConfig -ConfigID <id> -Site a (the -Site parameter is mandatory in 9.1).

Validation function runs clean against your deployed config

Validation Gate

Check: Verify: (1) base config exists at configs/base/, (2) at least 2 version-scoped deployed configs, (3) naming convention documents the Site A/B pair handling, (4) version history log populated, (5) validation function passes

Expected: Config management framework in place, adapted to the 9.1 pair-generation and multi-version reality

Common Errors

Config edits break JSON parsing
Cause: Manual editing introduced syntax errors
Fix: Manipulate configs as PowerShell objects and re-serialize with ConvertTo-Json -Depth 20 (unchanged guidance from the 9.0 twin).
Deployment picks up stale values despite an edited config
Cause: 9.0-era behavior cached host specs after first prepare; whether 9.1 retains this: verify against live 9.1 toolkit
Fix: Assume the safe pattern: destroy, generate a fresh config pair, redeploy.

Task 4 Deploy side-by-side VCF 5.2.4 and 9.1.0.0 instances and document version differences

manageability

This is VCDX gold. The 9.1 toolkit's supported range (VCF 9.1.0.0 down to 5.2) turns upgrade-path rehearsal into a lab exercise: run the 5.2 maintenance line beside 9.1 and articulate what changed (bring-up orchestrator, component versions, UI, licensing, services). Panelists probe exactly these version-boundary design decisions when upgrades are on the table.

Step 1

Confirm the version plan: the 9.1 toolkit deploys VCF 9.1.0.0 / VVF 9.1.0.0 / VCF 5.2.4 (official docs also list 5.2.3). Choose 5.2.4 as the comparison target and confirm your entitlement covers the 5.2.4 content set.

Comparison plan documented: existing 9.1.0.0 instance + new 5.2.4 instance
Step 2

Stage the VCF 5.2.4 binaries to /holodeck-runtime/bin/5.2.4/ (5.2.x deployments need the ESX ISO + Cloud Builder OVA; 9.x uses the VCF Installer OVA instead). Verify checksums — do not rename files to force a match.

5.2.4 binaries staged with matching checksums
GH#151 documents 9.1-era depot file-set growth and checksum failures from renamed binaries. Stage exactly what the manifest expects.
Step 3

Plan non-overlapping networking for the second instance: a custom /20 CIDR (e.g. 10.3.0.0/20) and a -VLANRangeStart clear of the primary instance (Site A needs >= 16 consecutive VLAN IDs). Check existing allocations first (Get-HoloDeckSubnet -Site a and Get-HolodeckServiceIPPools -Site a; signatures documented unchanged in 9.1).

Non-overlapping CIDR/VLAN plan recorded, validated against live allocations
Step 4

Resource-check before committing: the VCF 9.1 single-site envelope is already 48 CPU / 650 GB / 2.2 TB. Concurrent 5.2.4 + 9.1.0.0 instances will likely exceed a shared host — the recommended pattern is sequential: snapshot the 9.1 instance (Task 1 checkpoints), deploy 5.2.4, compare, then remove it. 9.1 soft validation will warn rather than block if you push past recommendations; treat warnings as design data, not noise.

Resource decision documented: concurrent vs sequential, with the numbers that justified it
Step 5

Deploy the 5.2.4 instance: New-HoloDeckInstance -Version 5.2.4 -InstanceID <unique-id> with the planned CIDR/VLAN parameters (5.2.x-specific switches and behavior on the 9.1 toolkit: verify against live 9.1 toolkit).

5.2.4 bring-up progresses; note where the Cloud Builder-driven sequence diverges from the 9.1 VCF Installer sequence
Step 6

Once deployed, list both instances via Get-HoloDeckInstance and confirm two healthy instances at different versions.

Two instances reported: one 9.1.0.0, one 5.2.4
Step 7

Access both management planes and document differences: (a) bring-up orchestrator (Cloud Builder for 5.2.x vs VCF Installer for 9.x, including the Installer compatibility matrix), (b) SDDC Manager UI and Day-2 menus, (c) component versions (vCenter, NSX 4.x vs 9.x-aligned), (d) evaluation licensing (5.2.x: 60-day trial; 9.x: 90-day trial, License Later mode), (e) the surrounding services experience (9.1 HoloRouter services stack vs the 5.2-era access model).

Notes or screenshots covering all five comparison axes
Step 8

Write the comparison document with at least 5 concrete architectural or operational differences, each tied to a design implication (upgrade windows, capacity, automation surface, security posture).

COMPARISON_5.2.4-vs-9.1.0.0.md created with 5+ differences and design implications
Step 9

Close with a VCDX design reflection: what the 5.2-to-9.x boundary means for a customer's upgrade design — orchestrator replacement, component re-versioning, licensing change, and operational-model shift — and which of those you can now demonstrate live rather than assert.

Design reflection written

Validation Gate

Check: Verify: (1) both instances deployed and healthy (sequentially or concurrently per the documented resource decision), (2) both management planes accessed, (3) comparison document with 5+ differences and design implications, (4) VCDX reflection written

Expected: Side-by-side comparison complete with architectural insights ready for panelist discussion

Common Errors

5.2.4 deployment fails at content validation
Cause: Binary set incomplete or checksums mismatched (renamed files), or manifest not aligned to 5.2.4
Fix: Re-stage the exact 5.2.4 file set per the manifest; never rename binaries (GH#151 pattern). Exact 9.1-toolkit error strings for 5.2.x targets: verify against live 9.1 toolkit.
Both instances compete for resources; one stalls
Cause: Physical host capacity exceeded — the 9.1 envelope alone is much larger than 9.0-era sizing
Fix: Fall back to the sequential pattern: snapshot the 9.1 instance, remove or stop it, run the 5.2.4 comparison, restore afterwards.

Task 5 Automate Holodeck operations and build a lab reset capability

manageability

Operational maturity means repeatability and automation. On 9.1 you have two automation substrates: your own PowerShell module (imperative, flexible) and the built-in GitOps pipeline (declarative runs with a durable audit trail). VCDX-level thinking builds the PowerShell framework AND evaluates when the pipeline is the better control plane.

Step 1

Create a PowerShell module (HoloDeckOps.psm1) with a deployment wrapper that logs every run: timestamped log file, the New-HoloDeckConfig call (capturing BOTH generated config IDs), the New-HoloDeckInstance call with -Version 9.1.0.0, and error capture. Exact 9.1 cmdlet output to parse: verify against live 9.1 toolkit.

Module created with a logging deployment function
Step 2

Add checkpoint functions: New-HoloDeckCheckpoint / Revert-HoloDeckCheckpoint wrapping PowerCLI snapshot create/revert for the instance VM set (version-agnostic mechanics, carried from the 9.0 twin).

Checkpoint functions added and exported
Step 3

Create Reset-Lab.ps1 that reverts all instance VMs to a named snapshot (default 'holodeck91-02-complete'), with confirmation prompt and per-VM missing-snapshot handling, then powers the VMs back on.

Lab reset script created and readable
Step 4

Create ScheduledSnapshot.ps1: weekly non-memory checkpoint named auto-weekly-<date> plus cleanup of snapshots older than 28 days. Schedule it via your workstation's task scheduler.

Scheduled snapshot script created; retention policy documented
Step 5

Add the Task 3 config validation function to the module (version string in the 9.1 valid set, binaries staged, CIDR collision check) so every automated deploy is pre-validated.

Config validation wired into the deployment wrapper
Step 6

Create Show-LabStatus.ps1: instance inventory (Get-HoloDeckInstance), datastore usage, and snapshot counts per VM in one dashboard output.

Status dashboard script created and producing output
Step 7

Evaluate the GitOps pipeline as the audit-grade driver: open the pre-loaded Holodeck repository in GitLab, read the pipeline definition, and review the pipeline run history from your deployments. Compare against your module on repeatability (identical runs), auditability (who/what/when per run), and control (intervention mid-run). Record which driver your lab standardizes on and why.

Driver comparison documented with a standardization decision
Step 8

Document the automation framework in AUTOMATION_README.md: function inventory, reset/schedule usage, idempotency notes (state on HoloRouter — re-run New-HoloDeckInstance to resume from failure, carried into 9.1), and the GitOps-vs-PowerShell decision.

Automation framework documented

Validation Gate

Check: Verify: (1) module with at least 3 functions, (2) Reset-Lab.ps1 and ScheduledSnapshot.ps1 created, (3) config validation wired in, (4) status dashboard works, (5) GitOps pipeline reviewed with a documented driver decision, (6) README complete

Expected: Repeatable, idempotent lab operations framework in place, with a deliberate choice of automation driver

Common Errors

Module import fails or functions not found
Cause: Module path/manifest issue
Fix: Dot-source the script as a fallback; keep the .psm1 in a folder matching the module name (unchanged guidance from the 9.0 twin).
Reset script cannot find the target snapshot on some VMs
Cause: Snapshot missing on VMs added after the checkpoint (e.g. Day-2 added hosts)
Fix: Handle missing snapshots per-VM with -ErrorAction SilentlyContinue and report the delta; re-baseline the checkpoint after any Day-2 expansion.

Final Validation

The Holodeck 9.1 instance lifecycle is fully managed: snapshots checkpoint progress, a destroy-redeploy cycle validated cleanup and reproducibility with measured 9.1 timings, the Site A + Site B config pair is version-controlled with overlay patterns, a 5.2.4 instance was compared side-by-side against 9.1.0.0 across the toolkit's full support range, and an automation framework exists with a deliberate PowerShell-vs-GitOps driver decision. 1 deployment (W5).

✓ Snapshots: at least 2 named checkpoints per VM → Names follow <component>-<phase>-<version> with 9.1.0.0 scoping; all non-memory

✓ Destroy-redeploy cycle completed with timings → Cleanup verified (VMs, port groups, datastore); redeploy from a FRESH config pair; actual 9.1 timings logged (no official 9.1 timing table exists)

✓ Config management framework → Version-scoped base/deployed/overlay configs; Site A/B pair convention documented; validation function passes

✓ Version comparison: 5.2.4 vs 9.1.0.0 → 5+ documented differences with design implications (orchestrator, components, licensing, UI, services stack)

✓ Automation: module + reset + schedule + dashboard + driver decision → All scripts runnable; GitOps pipeline reviewed; standardization decision documented

Cleanup / Restore

Snapshot: holodeck91-07-complete

• Take a final snapshot 'holodeck91-07-complete' across all instance VMs after all tasks complete

• Consolidate intermediate lcm-* snapshots (keep only holodeck91-02-complete and holodeck91-07-complete)

• Remove the 5.2.4 comparison instance if still deployed, and verify its cleanup (Task 2 checklist)

• Archive config files (including both members of each Site A/B pair) and automation scripts to your reference location

• If on a shared host: record allocation changes in the shared registry and note every place live 9.1 behavior diverged from this stub

Design Reflection (VCDX)

A VCDX panelist examining operational maturity would probe: how do you manage multiple lab scenarios simultaneously; can you roll back a failed change; how would you standardize deployments across a team; can you articulate snapshot-vs-redeploy trade-offs with measured numbers? The 9.1-specific angles: (1) the toolkit now hands you an audit-grade deployment driver (GitOps pipeline history) — defend when pipeline-driven lifecycle beats imperative cmdlets and vice versa; (2) the one-config-one-deployment rule plus auto-generated Site A/B pairs changes your configuration-management design — pairs must be tracked and consumed configs never reused; (3) the support range down to VCF 5.2 makes upgrade rehearsal demonstrable — use the side-by-side evidence, not doctrine.

Requirements

  • Support rapid environment reset for iterative design testing (snapshot revert in minutes)
  • Enable version-comparison testing across the toolkit's supported range (VCF 9.1.0.0 down to 5.2) on shared infrastructure
  • Provide reproducible, version-controlled configurations including the 9.1 Site A/B pair convention
  • Detect and clean up orphaned resources after instance removal
  • Automate routine operations with an auditable execution record (GitOps pipeline history or logged PowerShell runs)

Constraints

  • The 9.1 resource envelope (48 CPU / 650 GB / 2.2 TB single site) constrains concurrent multi-version instances far more than 9.0-era sizing did
  • One config = one deployment; consumed configs cannot drive redeploys
  • Snapshots are lab-only mechanisms — storage-intensive and not a production protection pattern
  • Nested performance does not represent production; timing data is directional only
  • 9.0-era KB and workarounds are unverified against the 9.1 services stack

Assumptions

  • Stop/Start/Remove-HoloDeckInstance behave per their unchanged documented signatures on 9.1
  • The GitOps pipeline can serve as a lifecycle driver equivalent to the cmdlet path for full deployments
  • 9.0 baseline snapshot guidance (non-memory, <50% free space) carries forward until 9.1 evidence says otherwise
  • Broadcom entitlement covers both 9.1.0.0 and 5.2.4 content sets

Risks

  • Snapshot chain corruption causes restore failures — MITIGATION: periodic consolidation and test restores
  • -ResetHoloRouter on a shared host destroys other tenants' networking — MITIGATION: coordinate teardown windows; re-validate its 9.1 behavior before first use
  • Resource exhaustion during side-by-side testing destabilizes the primary instance — MITIGATION: sequential pattern with checkpoints; treat soft-validation warnings as stop signals
  • Config pair mismanagement (reusing a consumed config, losing the Site B twin) breaks reproducibility — MITIGATION: pair-aware archiving and the pre-deploy validation function
  • Applying 9.0-era fixes to the 9.1 stack corrupts router state — MITIGATION: treat all 9.0 workarounds as unverified; re-diagnose from live logs

Self-Assessment Discussion Prompts

  1. When would you snapshot vs redeploy on 9.1, and how do your MEASURED timings change the answer vs the 9.0-era baseline?
  2. The GitOps pipeline gives you deployment audit history for free. What does that change about your automation design compared to hand-rolled PowerShell logging?
  3. One config = one deployment, and every config generation now yields a Site A + Site B pair. Design a configuration-management scheme that makes misuse structurally impossible.
  4. You have one physical host at the 9.1 resource envelope and need to compare 5.2.4, 9.1.0.0, and a future release. Sequence the work and justify the checkpoint strategy.
  5. Which parts of your 9.0-era reset/automation scripts survive contact with the 9.1 services stack unchanged, and how would you prove it?
  6. If you provisioned Holodeck 9.1 labs for 10 VCDX candidates, would you standardize on the GitOps driver or the cmdlet path? Defend both answers.

Extensions

Holodeck-as-Code on the Built-in Pipeline

Extend the pre-loaded GitLab pipeline: fork the deployment repo, parameterize it for your config overlay scheme, and add a validation stage that runs your Task 3 checks before any deploy. 9.1 makes this native — the extension is making it YOURS.

harder

Automated Health Checks and Alerting

Build a monitoring loop that checks instance health (SDDC Manager availability, NSX fabric status, datastore space, HoloRouter services stack endpoints) and triggers remediation (snapshot revert, VM restart) on anomalies.

harder

Multi-Tenant Lab Hosting Service

Design a system where multiple users provision independent Holodeck 9.1 instances on shared infrastructure: quota management, allocation registry (CIDR/VLAN), Authentik-backed RBAC, and automated cleanup of abandoned instances.

much harder

Full CI/CD Lab Testing Workflow

Wire config commits to automatic pipeline-driven deployment, post-deploy validation tests, and result reporting — turning the destroy-redeploy cycle of Task 2 into a scheduled regression test for the lab platform itself.

much harder

Multi-Host Holodeck Estate

Span the lifecycle framework across multiple physical hosts (or a DRS cluster target): coordinated snapshots, cross-host allocation registry, and estate-wide status dashboards.

much harder

⚠ Known Pitfalls (from Community KB)

Encounter Issues when Modified default ESXi Host CPU & Memory RESOLVED
Problem: 9.0-era record: config modifications did not apply consistently to all hosts after a partial deployment
Resolution: 9.0-era fix: customize config before first prepare; destroy and re-prepare otherwise. Whether 9.1 retains the caching behavior: verify against live 9.1 toolkit.
Start-HoloDeckInstance -InstanceID xxxxxx command failed. Cannot find instance OPEN
Problem: 9.0-era record: Start failed after successful Prepare due to missing Appliance property array in instance JSON
Resolution: 9.0-era guidance: verify Get-HoloDeckInstance lists the instance with populated appliances. Re-validate the failure signature on 9.1 before reuse.
Holodeck vcf 9 - Cannot select datastore RESOLVED
Problem: 9.0-era record: datastore selection failed despite adequate free space
Resolution: 9.0-era fix: verify datastore naming, clean orphaned VMs, rescan. Likely still relevant on 9.1 given the larger footprint — confirm on live deployment.
Waiting for FRR service LIKELY_RESOLVED
Problem: 9.0-era record: prepare hung waiting for the BGP daemon on the HoloRouter
Resolution: 9.0-era fix restarted FRR. The 9.1 router runs a different services stack behind a reverse proxy — do NOT assume the restart sequence applies; re-diagnose from live 9.1 logs.
VCF Installer is not ready yet RESOLVED
Problem: 9.0-era record: bring-up hung waiting for installer initialization, usually DNS-rooted
Resolution: 9.0-era fix checked DNS/NTP on the installer. On 9.1, DNS is served by Technitium (:5380) rather than dnsmasq — re-baseline the DNS diagnostic chain before applying any old fix.

References

  • HoloDeck 9.1 Release Announcement (captured verbatim) Holodeck-9.1/01-release-announcement.mdlocal fileTier 1 — Official
    Primary source for confirmed 9.1 facts: targets, support range, GitOps, services stack, dual-site auto-generation.
  • HoloDeck Toolkit 9.1 — Structured Reference & Changelog Holodeck-9.1/02-structured-changelog.mdlocal fileTier 1 — Official
    Structured capture with open questions (GitLab overhead, VNA sizing, Technitium vs dnsmasq, upgrade path).
  • Holodeck 9.1 Official Docs (versioned site)Tier 1 — Official
    Use the /9.1/ URLs only — un-versioned pages serve stale legacy content. Cmd reference confirms unchanged Stop/Start/Remove signatures and the Import-HoloDeckConfig -Site breaking change.
  • Holodeck Toolkit GitHub RepositoryTier 1 — Official
    Issue tracker for 9.1-era issues: GH#139 (GitLab first boot), GH#143 (Day-2 All-Apps Org), GH#151 (9.1 depot file set).
  • Broadcom Support Portal — VCF 9.1 / 5.2.4 DownloadsTier 1 — Official
    Official content source for both comparison targets. Requires entitlement; use the latest manifest for 9.1 depots.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.