Day-2 Operations on Holodeck 9.1 — Part 1: Adding Nested ESX Hosts and Deploying VCF Automation In-Toolkit
Objectives
- Baseline a deployed Holodeck 9.1 environment (inventory, capacity, health) before making any Day-2 change
- Provision additional nested ESX hosts into a running Holodeck 9.1 instance using New-HoloDeckESXiNodes
- Commission the new hosts into SDDC Manager and expand a cluster, validating vSAN and NSX integration end to end
- Deploy VCF Automation as a Day-2 operation via Update-HoloDeckInstance and validate the Automation consumption surface
- Apply capacity-first Day-2 discipline: measure before and after every change, and document the deltas
- Articulate the VCDX case for staged deployment (core stack first, consumption layer later) versus monolithic Day-1 deployment
Prerequisites
Healthy Holodeck 9.1 VCF 9.1.0.0 management domain deployed (holodeck91-02 completed), single site is sufficient. SDDC Manager, vCenter, and NSX Manager all reachable and green. Physical host must have spare capacity beyond current consumption: budget roughly one additional nested host's footprint (typically 8-12 vCPU / 64-96 GB RAM / 200+ GB disk per host at lab sizing — verify against your instance's host sizing) plus the VCF Automation appliance footprint. Snapshot 'holodeck91-02-complete' (or later) exists.
Prior labs: holodeck91-02
Required skills:
- Holodeck 9.1 lifecycle and config cmdlets (holodeck91-01/02)
- Environment discovery methodology (holodeck-08 — version-agnostic 4-phase discovery playbook)
- SDDC Manager host commissioning and cluster expansion concepts (holodeck-05 covers the 9.0 mechanics; UI details differ on 9.1)
- vSAN cluster expansion and rebalancing behavior
- REST API authentication against SDDC Manager (bearer token flow)
Lab Environment
Single-site Holodeck 9.1 instance: management domain with 4 nested ESX hosts, HoloRouter 9.1 (Technitium DNS :5380, Vault :8200, Authentik :9443, authenticated Webtop :30000, all behind HTTPS reverse proxy). This lab adds 1-2 nested ESX hosts Day-2 and then deploys VCF Automation onto the running stack. Component addressing follows the 9.1 generated configuration — the values below reflect the common Holodeck scheme; always read actuals from your generated config (Get-HoloDeckAppNetwork after Import-HoloDeckConfig).
Credentials
| System | Username | Password |
|---|---|---|
| SDDC Manager UI/API | administrator@vsphere.local | Master password; API auth uses the doubled-password convention from 9.0.x (verify on your 9.1 build) |
| vCenter | administrator@vsphere.local | Same SSO as SDDC Manager |
| NSX Manager | admin | Set in bring-up spec; doubled password for API |
| New nested ESX hosts | root | Set by the Holodeck config (esxiPassword field) — same as existing hosts |
| VCF Automation (post-deployment) | verify default admin account against live 9.1 toolkit | Assigned during Day-2 Automation deployment — record it in your lab notebook immediately |
Tasks
Task 1 Baseline the environment before any Day-2 change
manageabilityDay-2 discipline starts with knowing exactly what you have. Every capacity, inventory, and health metric you record now becomes the 'before' column in your change evidence. This mirrors production change management: a VCDX panelist asking 'how did you know the expansion succeeded?' expects a before/after delta, not a green dashboard screenshot.
Load the instance context: Import-HoloDeckConfig -ConfigId <id> (add -Site if dual-site). Run Get-HoloDeckInstance -InstanceID <id> and record: deployment version, host count, component health, state.
Authenticate to the SDDC Manager API (POST /v1/tokens with administrator@vsphere.local and the doubled password, exactly as drilled in holodeck-08 Phase 3 — the flow is unchanged in concept on 9.1). Then capture the inventory baseline: GET /v1/hosts, GET /v1/domains, GET /v1/clusters per domain. Save the JSON responses to files.
Capture the capacity baseline at three layers: (a) physical host — esxtop/host client: CPU, memory, datastore free; (b) nested cluster — vCenter cluster Summary: CPU/memory/vSAN capacity and current utilization; (c) vSAN — Monitor > vSAN > Capacity: usable capacity, consumed, slack.
Verify health gates before proceeding: vSAN health green (no resync in progress), no failed tasks in SDDC Manager task queue, NSX transport nodes all Success/Up, no expired credentials (GET /v1/credentials — check lastRotated/expiry).
Take a snapshot of all instance VMs labeled 'pre-day2-expansion'. Record snapshot names in your lab notebook.
Validation Gate
Check: Baseline artifacts exist: instance query output, API inventory JSON files, three-layer capacity table, health-gate checklist, pre-change snapshot.
Expected: Documented, healthy starting state with a rollback line.
Common Errors
Task 2 Provision additional nested ESX hosts with New-HoloDeckESXiNodes
manageabilityThis is the first half of rn-9-1-019: in-toolkit host addition. Before 9.1's Day-2 catalog matured, growing a nested lab often meant hand-building ESX VMs or redeploying the instance. New-HoloDeckESXiNodes automates nested host provisioning to match the instance's existing conventions (naming, networking, DNS records, disk layout for the instance's vSAN mode). The design lesson: expansion should be a product of the same automation that built the environment — hand-built additions are where drift and snowflakes enter.
Review the New-HoloDeckESXiNodes parameters in the 9.1 toolkit help (Get-Help New-HoloDeckESXiNodes -Full). Expect parameters covering CPU, memory, site, and vSAN mode alignment; confirm exact parameter names, count semantics, and defaults against the live 9.1 toolkit before running — do not rely on 9.0.2-era notes.
Re-check physical headroom against your Task 1 baseline: projected consumption = current + (new host footprint x host count). Keep projected physical RAM below ~90%.
Run New-HoloDeckESXiNodes for the target site/instance with your chosen sizing (CPU, memory, vSAN mode consistent with the instance — ESA or OSA as deployed). Use -Verbose and capture the transcript.
Verify the new host(s) at the infrastructure layer before touching SDDC Manager: (a) VM(s) running on the physical host; (b) ESX responds on its management IP (ping, then browse to the host client); (c) DNS forward and reverse resolution for the new host FQDN(s) via the HoloRouter's Technitium DNS; (d) NTP synced on the new host(s).
Confirm the new host(s) meet VCF commissioning prerequisites: root credentials as configured, no vSAN partitions already claimed, correct VLAN reachability from SDDC Manager. Run a manual connectivity test from the SDDC Manager appliance (SSH, then curl/ping toward the new host).
Validation Gate
Check: New nested ESX host(s) provisioned by the toolkit, running 8.0 U3+, reachable, DNS-registered (forward + reverse), NTP-synced, commissioning prerequisites verified.
Expected: Host(s) ready for SDDC Manager commissioning with zero hand-patched configuration.
Common Errors
Task 3 Commission the new hosts and expand the cluster
availabilityProvisioned is not the same as productive. Commissioning into SDDC Manager and expanding the cluster exercises the full VCF integration chain: host validation, cluster membership, vSAN disk claiming and rebalancing, NSX transport node preparation. This is the same Day-2 motion as production hardware expansion — the panel question it rehearses is 'walk me through adding capacity to a live domain without a maintenance window'.
In SDDC Manager, commission the new host(s): Inventory > Hosts > add/commission workflow, entering FQDN, credentials, and network pool association. Note: 9.1's unified management services runtime (rn-9-1-001) reorganized parts of the management UI relative to 9.0 — record the actual navigation path you use, and treat 9.0 lab screenshots as approximate.
Expand the management cluster (or a workload cluster if you deployed one) with the new host(s): select the domain > cluster > add host workflow. Choose the commissioned host(s) and confirm.
Monitor vSAN expansion: vCenter > cluster > Monitor > vSAN > Capacity and Resyncing Objects. Watch the new capacity appear and rebalancing traffic distribute objects onto the new host.
Verify NSX integration: NSX Manager > System > Fabric > Nodes > Host Transport Nodes. The new host(s) must show Configuration State: Success and Node Status: Up.
Re-run the Task 1 API inventory queries (GET /v1/hosts, /v1/clusters) and produce the before/after delta: host count, cluster membership, vSAN capacity, transport node count.
Functional proof: vMotion an existing VM onto a new host, and confirm DRS considers the new host in its placement (cluster > Monitor > DRS).
Validation Gate
Check: New host(s) Active and assigned in SDDC Manager, in the vCenter cluster with vSAN contributing and healthy, NSX transport nodes Success/Up, before/after delta documented, vMotion onto new host proven.
Expected: Cluster expanded end to end with quantified capacity uplift and zero downtime to running workloads.
Common Errors
Task 4 Deploy VCF Automation as a Day-2 operation
manageabilitySecond half of rn-9-1-019: the consumption layer added after the fact. The 9.1 pattern — deploy the core stack, verify resource consumption and stability, THEN add VCF Automation — is exactly the staged-deployment discipline architects argue for in production: sequence deployments so each layer's resource claim and health is proven before the next layer lands on top of it.
Check capacity headroom for the Automation appliance against your running baseline (Task 3 delta). Verify against the live 9.1 toolkit docs what footprint the VCF Automation deployment claims in a lab-sized instance — do not assume production sizing figures.
Review the Day-2 Automation deployment operation: Get-Help Update-HoloDeckInstance -Full. Identify the parameter set for VCF Automation deployment (in 9.0.2 this family included switches such as -AddVcfAutomationAllAppsOrg; confirm the exact 9.1 parameter set names against the live toolkit). Remember: one parameter set per invocation — Automation deployment cannot be combined with other Day-2 operations in a single call.
Run the Automation deployment via Update-HoloDeckInstance with the Automation parameter set for your site/instance, with -Verbose, capturing the transcript.
Validate the Automation deployment: (a) appliance VM(s) running and healthy in vCenter; (b) Automation UI reachable at its FQDN (record the URL and admin account from the deployment output); (c) Automation registered against the VCF instance (visible from the management console); (d) DNS records for Automation components present in Technitium.
Exercise the consumption surface minimally: log in to Automation, locate the default organization/project structure (if your build deployed an All Apps Org, note it), and confirm the environment's inventory/resources are visible to Automation.
Complete the capacity story: re-measure the three-layer capacity table and produce the final before/after ledger for the whole lab — baseline, post-host-expansion, post-Automation.
Validation Gate
Check: VCF Automation deployed Day-2 via Update-HoloDeckInstance, UI reachable and registered, DNS records present, and the three-stage capacity ledger completed.
Expected: Consumption layer operational on top of a proven core stack, with the full resource cost of the day's changes quantified.
Common Errors
Final Validation
A running Holodeck 9.1 environment was expanded in place — nested ESX hosts added with New-HoloDeckESXiNodes, commissioned and integrated through SDDC Manager into vSAN and NSX, then VCF Automation deployed as a Day-2 operation via Update-HoloDeckInstance — all without redeploying the instance and with a quantified before/after capacity ledger at every stage. This demonstrates the 9.1 shift from bring-up tool to Day-2 lab platform (rn-9-1-019) and rehearses the production discipline of staged, measured, reversible change.
✓ Task 1: Baseline captured (inventory, three-layer capacity, health gates, snapshot) → Documented starting state with rollback line
✓ Task 2: Nested host(s) provisioned via New-HoloDeckESXiNodes → Host(s) booted on 8.0 U3+, DNS forward+reverse, NTP synced, commissioning-ready
✓ Task 3: Cluster expanded end to end → Hosts Active in SDDC Manager, vSAN contributing and green, NSX transport nodes Up, vMotion proven, delta documented
✓ Task 4: VCF Automation deployed Day-2 and validated → Automation UI reachable and registered; three-stage capacity ledger complete
Cleanup / Restore
Snapshot: holodeck91-08-complete
• Snapshot all instance VMs including the new hosts and Automation appliance as 'holodeck91-08-complete' — holodeck91-09 builds directly on this state
• Archive the capacity ledger, API baseline/delta JSON, and deployment transcripts to your VCDX evidence folder
• Do NOT decommission the added hosts or remove Automation — they are prerequisites for holodeck91-09 (Supervisor deployment needs both the capacity and, for some paths, the Automation layer)
• If physical capacity is now marginal, note it prominently in the lab notebook: holodeck91-09 adds Supervisor footprint on top of this state
Design Reflection (VCDX)
Three defensible positions come out of this lab. (1) EXPAND-IN-PLACE vs REDEPLOY: 9.1's Day-2 catalog makes in-place expansion cheap in the lab, but the panel-grade reasoning is economic and risk-based: redeploy gives you a provably clean, generation-consistent state at the cost of downtime and lost in-flight work; expansion preserves running state at the cost of accumulating history.
The mature answer is conditional — expand when the change is additive and automatable by the same tooling that built the environment (as here); redeploy when drift or partial failures have made state untrustworthy. (2) STAGED DEPLOYMENT: deploying the core stack first, measuring, then adding VCF Automation is a sequencing argument: each layer's resource claim and failure modes are isolated in time, so attribution is trivial when something breaks. Monolithic Day-1 deployment of everything is faster on the calendar but couples all failure domains into a single debugging session.
The 9.1 release notes explicitly endorse the staged pattern ('deploy the rest of the stack and check resource consumption and availability before deploying VCF Automation'). (3) MEASUREMENT AS CHANGE EVIDENCE: the three-stage capacity ledger is the lab-scale version of production change validation — before/after deltas at every layer, tied to specific changes. A panelist probing 'how do you know?' should always land on recorded deltas, not dashboards.
Be ready for the counter-question on the toolkit dependency: 'your expansion is only as consistent as the tool — what is your process when the toolkit cannot express the change you need?' (Answer: that is precisely when change control gets heavier, because you are now creating state the automation does not know about — document it as deliberate, tracked asymmetry or extend the automation.)
Requirements
- Add compute capacity to a running VCF 9.1 lab instance without downtime to existing workloads
- New hosts must be provisioned by the toolkit to instance conventions (naming, DNS, networking, vSAN mode) — no hand-built hosts
- Full integration chain validated: commissioning, cluster membership, vSAN contribution, NSX transport node preparation
- VCF Automation deployed as a discrete, post-core Day-2 operation with its own validation gate
- Every change bracketed by measured before/after capacity and inventory deltas
- All changes reversible: snapshot line before the sequence, documented cleanup path for partial failures
Constraints
- Physical host capacity bounds the expansion: each nested host and the Automation appliance claim real RAM/CPU/disk from a finite pool
- Update-HoloDeckInstance parameter sets are mutually exclusive — Day-2 operations must be sequenced, not batched
- Nested ESX minimum 8.0 U3 applies to added hosts exactly as to Day-1 hosts
- 9.1 UI navigation differs from 9.0 labs due to the unified management services runtime — procedures reference capabilities, not pixel paths
- Exact cmdlet parameter names for New-HoloDeckESXiNodes and the Automation parameter set must be verified against the live 9.1 toolkit (this lab flags them rather than fabricating specifics)
Assumptions
- The instance was deployed by Holodeck 9.1 (not upgraded from 9.0.x mid-flight), so Day-2 cmdlets operate on state they generated
- vSAN mode (ESA/OSA) of the instance is known and matched by the new hosts' provisioning parameters
- Doubled-password API convention persists on the 9.1 build (verified at first token call)
- Lab-sized Automation footprint fits remaining capacity after host expansion — verified, not assumed, in Task 4 step 1
- Single operator; on shared physical hosts, capacity coordination with colleagues happens before provisioning
Risks
- Capacity exhaustion cascade: expansion + Automation overshoots physical RAM and destabilizes the whole instance — IMPACT: everything degrades at once. MITIGATION: headroom gates before each operation (90% ceiling), staged sequencing so each claim is measured before the next.
- Snowflake host: manual fixes applied to a toolkit-provisioned host — IMPACT: drift, commissioning failures, untrustworthy automation state. MITIGATION: reprovision instead of patch; treat hand-edits as forbidden on toolkit-owned objects.
- Partial Automation deployment leaves orphaned artifacts — IMPACT: retries fail, cleanup burden grows. MITIGATION: never interrupt; on failure capture logs, clean completely, then retry; rehearse the cleanup path (feeds holodeck91-09 teardown discipline).
- DNS gaps (especially PTR) silently break commissioning or Automation reachability — IMPACT: confusing mid-workflow failures. MITIGATION: explicit forward+reverse verification in Technitium before commissioning and after Automation deploy.
- Toolkit/parameter mismatch from stale 9.0.2-era notes — IMPACT: wrong invocations, wasted cycles. MITIGATION: Get-Help against the live module is the source of truth; version-scoped docs kept separate per the repo's coexistence policy.
- Expansion validated only by dashboards, not deltas — IMPACT: unnoticed partial integration (e.g., host in cluster but vSAN not contributing). MITIGATION: before/after API inventory diffs and the three-layer capacity ledger as mandatory artifacts.
Self-Assessment Discussion Prompts
- Defend expand-in-place over redeploy for this lab's host addition. Now give the scenario in which you would refuse to expand and insist on redeploy.
- Why does it matter that the new hosts were provisioned by New-HoloDeckESXiNodes rather than hand-built to the same spec? What exactly does 'same tooling' buy you that 'same spec' does not?
- The 9.1 Day-2 contract is one operation per Update-HoloDeckInstance invocation. As a design principle, argue for and against batched multi-operation change APIs.
- Your capacity ledger shows three deltas. A panelist asks: 'Which single number in that ledger would make you halt the next planned change, and why that threshold?'
- Staged deployment (core first, Automation later) lengthens the calendar. Justify the added elapsed time to a project sponsor in risk terms.
- In production VCF, what is the equivalent of this lab's 'pre-day2-expansion' snapshot line, given you cannot snapshot a physical estate? What actually constitutes your rollback capability there?
- How would you detect, six months from now, that someone hand-patched one of these nested hosts outside the toolkit? What control makes that detectable by design?
Extensions
Scale-In: Remove a Day-2 Host Cleanly
Reverse the expansion: evacuate, decommission, and remove one of the added hosts through SDDC Manager, verify vSAN health through the contraction (FTT implications at the new host count), and then remove the nested VM via the toolkit. Compare the effort and risk profile of scale-in vs scale-out and document why contraction is operationally harder.
sameDay-2 Expansion via the API Only
Repeat the commissioning and cluster-expansion flow driving SDDC Manager purely through its REST API (token, host commissioning payloads, cluster expansion, task polling) with no UI. Produces a reusable expansion script and exercises the 9.1 API-first story (rn-9-1-002).
harderAutomation Consumption Pilot
Build one minimal catalog item in the Day-2-deployed VCF Automation (e.g., a single-VM blueprint onto the expanded cluster), request it as a consumer, and trace the full request-to-provision path. Establishes the consumption baseline that the Supervisor scenarios in holodeck91-09 extend.
sameFailure Injection: Break Commissioning on Purpose
On a spare provisioned host, deliberately induce each classic commissioning failure (remove the PTR record, skew NTP, wrong root password) one at a time, run commissioning validation, and catalog the exact error signature each produces. Turns this lab's warnings into a personal troubleshooting index.
same⚠ Known Pitfalls (from Community KB)
References
- Announcing the General Availability of Holodeck 9.1 — VMware Cloud Foundation BlogTier 1 — Official
Official GA announcement: Day-2 operations catalog (add nested ESX hosts, deploy VCF Automation, Supervisors in VNA/Edge modes). - Holodeck 9.1 DocumentationTier 1 — Official
Authoritative for New-HoloDeckESXiNodes parameters and Update-HoloDeckInstance Day-2 parameter sets — verify all specifics here before running. - Add Clusters and Hosts in VCF 9.0 using Holodeck — VMware Cloud Foundation BlogTier 1 — Official
The 9.0-era walk-through of host/cluster addition with Holodeck — the direct ancestor of this lab's workflow; read for the SDDC Manager-side mechanics. - academy-v2/content/releases/vcf-9.1.json — rn-9-1-019, rn-9-1-001, rn-9-1-002
../releases/vcf-9.1.jsonlocal fileTier 2 — VMware Press
Canonical changelog entries: Day-2 operations (rn-9-1-019), unified management runtime affecting UI navigation (rn-9-1-001), API-first (rn-9-1-002) for the API extension. - holodeck-08 — Holodeck Environment Discovery Playbook (9.0 twin)
holodeck-08.jsonlocal fileTier 2 — VMware Press
The 4-phase discovery methodology used for Task 1's baseline; version-agnostic in approach, 9.0.x in specifics. - William Lam — VCF 9.1 Automated Nested Lab DeploymentTier 3 — Expert Blog
Alternative nested-lab automation for VCF 9.1 — useful comparative perspective on what 'toolkit-owned provisioning' looks like outside Holodeck.