Academy/Holodeck Lab Setup & Operations/Day-2 Operations on Holodeck 9.1 — Part 2: Deploying vSphere Supervisors in Edge (Centralized) and VNA (Distributed) Modes, with Teardown Discipline
This lab targets VCF 9.1.0.0

Day-2 Operations on Holodeck 9.1 — Part 2: Deploying vSphere Supervisors in Edge (Centralized) and VNA (Distributed) Modes, with Teardown Discipline

VCF 9.1.0.0Advancedadminarchitectvcdx⏱ 240 min

9.0.x twin: holodeck-09 (Command Reference — its Supervisor coverage is limited to the 9.0.2-era -DeploySupervisor* switches and Day-0 deployment only; Day-2 Supervisor deployment and the Centralized/Distributed mode choice did not exist). SupervisorDeploymentMode 'Distributed' requires VCF 9.1.0.0 or later per the official 9.1 docs. The Distributed path has been community-verified at Day-0 (GitHub #143 deployed -DeploySupervisorMgmtDomain Distributed alongside -VnaClusterMgmtDomain successfully); the Day-2 Update-HoloDeckInstance -DeploySupervisor set is documented but not yet community-verified — every step against it flags verify_live accordingly. Read holodeck91-08 first: it covers Day-2 part 1 (host addition + VCF Automation) and its capacity ledger and cleanup lessons are prerequisites here.

Objectives

  • Gate a Supervisor deployment behind explicit capacity and health checks, extending the holodeck91-08 three-layer capacity ledger
  • Deploy a vSphere Supervisor as a Day-2 operation in Centralized (NSX Edge-backed) mode via Update-HoloDeckInstance and validate its control plane and networking path end to end
  • Tear the Centralized Supervisor down cleanly, inventorying and clearing every residual artifact (control plane VMs, NSX objects, content library state, DNS records) before the next deployment
  • Deploy a Supervisor in Distributed (VNA-backed) mode on a VNA-enabled domain and contrast its networking path against the Centralized deployment
  • Produce a mode-comparison artifact: resource footprint, networking path, dependency chain, and failure domain for Edge vs VNA Supervisor placement
  • Codify teardown discipline as a reusable runbook: pre-change snapshot line, one-operation-per-invocation sequencing, post-teardown residue audit, and evidence capture

Prerequisites

Healthy Holodeck 9.1 VCF 9.1.0.0 environment as left by holodeck91-08: management domain expanded to 5-6 nested hosts, VCF Automation deployed and registered, snapshot 'holodeck91-08-complete' taken, capacity ledger current. The management domain must have an NSX Edge cluster for the Centralized deployment (present if the instance was deployed with -NsxEdgeClusterMgmtDomain; verify before starting). The Distributed deployment (Task 4) additionally requires a VNA cluster in the target domain — a Day-0 decision: if your instance was not deployed with -VnaClusterMgmtDomain, Task 4 runs against the VNA-enabled instance built in holodeck91-10, or is executed as a documented-failure experiment (both paths are legitimate and spelled out in the task). Physical headroom: budget for Supervisor control plane VMs (typically three) plus per-mode networking components on top of the post-91-08 consumption — measured in Task 1, not assumed.

Prior labs: holodeck91-02, holodeck91-08

Required skills:

  • Day-2 discipline from holodeck91-08: baselining, capacity ledgers, mutually exclusive Update-HoloDeckInstance parameter sets, partial-failure cleanup
  • vSphere Supervisor / Workload Management concepts: control plane VMs, namespaces, content libraries, workload networks
  • NSX constructs consumed by Supervisors: Tier-1 gateways, load balancers, segments, ingress/egress CIDRs
  • SDDC Manager and vCenter UI navigation on VCF 9.1 (expect deltas from 9.0 screenshots per rn-9-1-001)
  • kubectl basics for post-deployment Supervisor validation
📸 Starting State: S3-91 — Management Domain Deployed (VCF 9.1.0.0)

Lab Environment

Single-site Holodeck 9.1 instance carried forward from holodeck91-08: expanded management domain (5-6 nested ESX hosts), VCF Automation deployed, HoloRouter 9.1 service stack (Technitium DNS :5380, Vault :8200, Authentik :9443, Webtop :30000 behind the HTTPS reverse proxy). This lab adds, validates, and removes Supervisor deployments in two modes. Known 9.1 naming from community-verified deployments: management-domain Supervisor 'supervisor-mgmt', zone 'zone-mgmt-a', cluster 'cluster-mgmt-01a' — verify against your generated configuration. Centralized mode consumes the domain's NSX Edge cluster for north-south and load balancing; Distributed mode consumes the domain's VNA cluster (distributed networking, no Tier-0 dependency — the community-observed deployment log explicitly reports 'Skipping T0 retrieval as distributed mode selected').

Credentials

SystemUsernamePassword
SDDC Manager UI/APIadministrator@vsphere.localMaster password; API auth uses the doubled-password convention (verify on your 9.1 build)
vCenteradministrator@vsphere.localSame SSO as SDDC Manager
NSX ManageradminSet in bring-up spec; doubled password for API
VCF Automationadmin (org: system)As recorded during the holodeck91-08 Day-2 deployment
Supervisor kubectl accessadministrator@vsphere.local (via kubectl vsphere login)SSO password; the Supervisor API endpoint IP is assigned at deployment — record it from the Workload Management UI

Tasks

Task 1 Gate the change: capacity, health, and the pre-Supervisor snapshot line

manageability

A Supervisor is the heaviest single Day-2 addition in the catalog: three control plane VMs plus per-mode networking components land on a nested estate that holodeck91-08 already grew twice. The gating discipline drilled there applies with higher stakes — and this lab adds a twist: because you will deploy, tear down, and deploy again, your baseline must be good enough to prove LATER that teardown actually returned the environment to it. The baseline is both the 'before' column and the teardown acceptance criterion.

Step 1

Load context and confirm starting state: Import-HoloDeckConfig -ConfigId <id> -Site a (the -Site parameter is mandatory in 9.1), then Get-HoloDeckInstance -InstanceID <id>. Confirm the post-91-08 state: expanded host count, VCF Automation present, all components healthy.

Instance reports VCF 9.1.0.0, Running, expanded host count, healthy components — matching the holodeck91-08 final ledger.
Step 2

Verify the mode prerequisites in NSX and the toolkit config: (a) confirm the management domain has an NSX Edge cluster (NSX Manager > System > Fabric > Nodes > Edge Transport Nodes, and the edge cluster object) — required for Centralized mode; (b) determine whether a VNA cluster exists in the domain (check your Day-0 New-HoloDeckInstance parameters for -VnaClusterMgmtDomain, and inventory vCenter for VNA-related appliances/components; their concrete form is characterized properly in holodeck91-10).

Edge cluster confirmed present. VNA presence determined and recorded — this decides whether Task 4 runs in-place, against the holodeck91-10 instance, or as a documented-failure experiment.
VNA cluster deployment is a Day-0 switch (-VnaClusterMgmtDomain / -VnaClusterWkldDomain on New-HoloDeckInstance). No documented Day-2 operation adds a VNA cluster to an existing domain — if you lack one, do not improvise; follow the Task 4 branching.
Step 3

Extend the three-layer capacity ledger from holodeck91-08 with a fresh 'pre-supervisor' column: physical host CPU/memory/datastore, cluster CPU/memory/vSAN utilization, and vSAN capacity/slack. Then estimate the Supervisor claim: three control plane VMs plus networking components. Exact lab-sized control plane VM sizing is not published for Holodeck 9.1 — verify against the live toolkit during deployment and record actuals.

Ledger extended with timestamped pre-supervisor column and a documented headroom decision (proceed / reduce scope). Keep projected physical RAM below the ~90% ceiling used throughout the Day-2 labs.
holodeck91-08's closing note warned that capacity might be marginal after expansion + Automation. If your ledger says marginal, the correct move is deliberate scope reduction (e.g. plan only one mode in-place and run the other against holodeck91-10's instance) — record the decision, do not push through the ceiling.
Step 4

Run the health gates: vSAN health green with no resync, no failed SDDC Manager tasks, NSX transport nodes (host AND edge) Success/Up, DNS forward+reverse spot checks against Technitium (:5380) for the components involved.

All gates green. Remediate anything amber BEFORE deploying — a Supervisor on a degraded base multiplies the cleanup burden this lab exists to teach.
Step 5

Establish the snapshot line: snapshot all instance VMs as 'pre-supervisor-91-09'. This is the rollback line for the entire two-deployment sequence and the reference state for teardown verification.

Snapshot line recorded in the lab notebook with VM list and timestamps.
Snapshot BEFORE the first deployment, not between the two. If Task 3's teardown fails to fully clean, you must be able to fall back to a provably pre-Supervisor state rather than stacking a second deployment on residue.
Step 6

Review the Day-2 Supervisor operation before running anything: Get-Help Update-HoloDeckInstance -Full. Locate the DeploySupervisor parameter set: -Site <a|b> -DeploySupervisor -VIDomain <Management|Workload> [-SupervisorDeploymentMode <Centralized|Distributed>]. Confirm on your build that the mode parameter defaults to Centralized and that Distributed is gated to VCF 9.1.0.0+.

Parameter set semantics recorded from the live module (the authoritative source), including any deviation from the documented syntax.
This is the same one-operation-per-invocation contract as holodeck91-08: the Supervisor set cannot be batched with anything else.

Validation Gate

Check: Instance healthy and context loaded; Edge cluster confirmed and VNA presence determined; capacity ledger extended with a headroom decision; health gates green; 'pre-supervisor-91-09' snapshot line established; live parameter set reviewed.

Expected: Deployment gated, reversible, and evidence-ready — with the Task 4 branching decision already made.

Common Errors

Import-HoloDeckConfig fails or Get-HoloDeck* queries return nothing
Cause: -Site omitted (mandatory in 9.1) or wrong ConfigId
Fix: Get-HoloDeckConfig to list configs, then Import-HoloDeckConfig -ConfigId <id> -Site a. Update any saved 9.0-era session scripts — the missing -Site is the classic 9.1 migration papercut.
No NSX Edge cluster in the management domain
Cause: Instance deployed without -NsxEdgeClusterMgmtDomain
Fix: Centralized mode cannot proceed without it. Options: redeploy the instance with the edge switch (weigh against holodeck91-08's expand-vs-redeploy reasoning), or run this lab's Centralized tasks against a domain that has one. Do not attempt to hand-build an edge cluster outside the toolkit for this purpose.
📋 KB: Carried 9.0-era edge deployment issues exist (GitHub #94) — nested edge clusters are resource-sensitive

Task 2 Deploy the Supervisor in Centralized (Edge) mode and validate end to end

availability

Centralized is the default and the architecturally familiar mode: Supervisor networking rides the domain's NSX Edge cluster — Tier-1 gateways, edge-hosted load balancers, ingress/egress through the edge. Deploying it first gives you the reference implementation against which Distributed mode is contrasted, and its validation chain (control plane up, API reachable, networking path traced) is the template you will reuse in Task 4.

Step 1

Run the deployment: Update-HoloDeckInstance -Site a -DeploySupervisor -VIDomain Management -SupervisorDeploymentMode Centralized (mode stated explicitly even though it is the default — commands in evidence logs should not rely on defaults). Capture the transcript with -Verbose if supported on your build.

Toolkit orchestrates Supervisor enablement against the management domain: expect content library setup, control plane VM deployment, and NSX object creation. Extended runtime on nested hardware — monitor rather than assume. Exact stage list is verify_live; record the stages your transcript actually shows.
Do not interrupt mid-deployment. holodeck91-08's lesson applies doubled: a partially enabled Supervisor spans vCenter AND NSX state and is the messiest partial failure in the Day-2 catalog. If it fails, capture logs first, then proceed to the Task 3 teardown methodology rather than retrying onto residue.
Step 2

Monitor from the vSphere side while the toolkit runs: vCenter > Workload Management shows the enablement progress; watch the three SupervisorControlPlaneVM machines appear and the cluster config state advance. Record timing per phase for the capacity/mode comparison.

Three control plane VMs deployed and running; Workload Management reaches a Ready/Running config status for the Supervisor (expect the name supervisor-mgmt and zone zone-mgmt-a per observed 9.1 naming — verify against your build).
Step 3

Validate the control plane: from the Workload Management UI record the control plane API endpoint IP, then log in: kubectl vsphere login --server <endpoint> -u administrator@vsphere.local. Run kubectl get nodes and kubectl get pods -A.

Login succeeds; control plane nodes Ready; system pods running. The Supervisor is alive as a Kubernetes API surface, not just as green icons.
Step 4

Trace the Centralized networking path in NSX Manager: identify the Supervisor-created objects — Tier-1 gateway(s), load balancer (hosting the control plane API VIP), segments for the workload networks, and the ingress/egress CIDR allocations. Diagram the path: workload -> segment -> Tier-1 -> Edge cluster (Tier-0) -> physical (HoloRouter).

NSX object inventory attributed to the Supervisor, with a drawn path showing the edge cluster as the north-south and load-balancing choke point. This diagram is half of the Task 5 comparison artifact.
Note WHERE the control plane VIP actually lives (edge-hosted LB). In Task 4 the equivalent question — where does the VIP live in Distributed mode? — is the sharpest single contrast between the modes.
Step 5

Exercise the consumption surface minimally: create one namespace via Workload Management, assign storage policy and permissions, and confirm it appears via kubectl get ns. If VCF Automation integration is part of your study track, note that the Supervisor is now visible to Automation (but defer All Apps Org work — see the known-issue warning).

Namespace created and visible from both UI and kubectl — proof of a working end-to-end consumption path.
If you later add a VCF Automation All Apps Org on top of a DISTRIBUTED-mode Supervisor, know the open 9.1 issue: the Day-2 org creation can fail with an IP Space / Gateway CIDR mismatch on the Distributed VLAN Connection (documented workaround: align the IP Space range with the gateway CIDR, clean the partial Region, re-run). Not applicable to this Centralized deployment, but it is the adjacent minefield.
Step 6

Close the deployment with measurements: re-run the capacity ledger (post-centralized column) and record the Supervisor's actual claim — control plane VM sizes, edge LB footprint, and the deployment wall-clock time from step 1.

Quantified cost of the Centralized Supervisor at all three ledger layers, plus phase timings. This is the left-hand side of the mode comparison.

Validation Gate

Check: Supervisor deployed Day-2 via the DeploySupervisor parameter set in Centralized mode; three control plane VMs running; kubectl login and node/pod checks pass; NSX object inventory and networking path diagrammed; test namespace round-trip done; post-centralized capacity column recorded.

Expected: Reference-implementation Supervisor validated end to end with its full resource and networking cost documented.

Common Errors

Update-HoloDeckInstance rejects the invocation with a parameter set error
Cause: Supervisor parameters mixed with another Day-2 operation, or -SupervisorDeploymentMode used with a value/casing the build rejects
Fix: One operation per invocation; use exactly Centralized or Distributed as string values. Verify spelling against Get-Help on the live module rather than notes.
Enablement stalls with control plane VMs deployed but config status stuck
Cause: The classic nested-lab trio: resource starvation, DNS gaps, or NSX edge health — control plane components wait on networking that is not coming up
Fix: Check edge transport node health first (a degraded edge blocks the LB/VIP path), then Technitium for missing records, then physical CPU/memory saturation. Give nested timings generous multiples of production expectations before declaring failure; capture the Workload Management event log either way.
📋 KB: Nested-platform failure class carried from the 9.0 era; edge sensitivity documented in GitHub #94
kubectl vsphere login fails against a Ready-looking Supervisor
Cause: Control plane VIP unreachable from your workstation (routing/DNS to the ingress range), not a Supervisor fault
Fix: Verify the endpoint IP is in a range your workstation routes to via the HoloRouter (holodeck-04/91-04 routing patterns); test from Webtop (in-pod bastion) to isolate workstation routing from Supervisor health.
TKC/workload cluster creation fails later on a pre-installed Supervisor
Cause: Known open issue class from the 9.0.2 era (GitHub #117) — status on 9.1 unverified
Fix: This lab stops at namespace level deliberately. If you extend into TKC creation and hit failures, check the issue's current state and capture your 9.1 reproduction — you would be producing the first 9.1 data point.
📋 KB: GitHub #117 (open at authoring time, 9.0.2-era)

Task 3 Tear down the Centralized Supervisor — the residue audit

recoverability

This is the heart of the lab's second theme. Supervisor enablement writes state into vCenter, NSX, storage, and DNS — and there is no documented Holodeck cmdlet that removes it: the toolkit's Day-2 catalog is additive. Teardown therefore happens on the vSphere side (Workload Management deactivation) followed by a manual residue audit against the Task 1 baseline. The production parallel is exact: platform teams that can deploy but cannot cleanly decommission accumulate cost and risk silently. An environment is only as manageable as its worst teardown.

Step 1

Before touching anything, inventory what you are about to remove: from Task 2's evidence, list every artifact the Supervisor created — control plane VMs, NSX objects (Tier-1, LB, segments, IP allocations), content library, namespaces, DNS records, storage policy usage. This list is your teardown checklist; removal without an inventory is deletion, not teardown.

Teardown checklist derived from observed deployment artifacts, each with its owning layer (vCenter / NSX / storage / DNS).
Step 2

Remove consumption first: delete the test namespace from Workload Management and confirm its NSX segment and any allocated VIPs are released. Consumption objects must go before the platform beneath them.

Namespace gone from UI and kubectl; associated NSX objects released.
Ordering is the discipline: consumption -> platform -> infrastructure. The same ordering argument appears in production decommissioning runbooks and in your defense answers about safe change reversal.
Step 3

Deactivate the Supervisor: vCenter > Workload Management > select the Supervisor > Disable/Remove (the exact control naming on VCF 9.1 may differ from 9.0 screenshots — record the actual path). Monitor the disablement tasks to completion.

Deactivation tasks complete; the three control plane VMs are deleted by the workflow.
There is no documented Update-HoloDeckInstance reverse operation for -DeploySupervisor, and whether the toolkit's internal state records the Supervisor as present after a vSphere-side removal is verify_live (checked in step 6). You are deliberately stepping outside the toolkit here — which is exactly why the audit trail matters.
Step 4

Run the residue audit against your step 1 checklist: (a) vCenter — control plane VMs gone, no orphaned VMs/folders/resource pool remnants; (b) NSX — Supervisor-created Tier-1, LB, segments, and IP allocations removed (search by the naming pattern captured in Task 2 step 4); (c) content library — decide deliberately whether to keep it (it is re-used by the next deployment and holding it saves re-download time; keeping it is a documented decision, not an omission); (d) Technitium — any Supervisor-related DNS records cleared.

Audit table: every checklist row marked REMOVED, KEPT-DELIBERATELY, or ORPHANED. Orphans get cleaned manually with the method recorded.
NSX residue is the classic Supervisor teardown failure: orphaned segments, stale LB VIPs, or IP-pool allocations that block the NEXT enablement with address-conflict errors. Audit NSX even if the vCenter side looks perfectly clean.
Step 5

Verify return-to-baseline: re-run the capacity ledger (post-teardown column) and diff against the Task 1 pre-supervisor column. CPU/memory should return to baseline within noise; storage may legitimately differ by the kept content library.

Post-teardown ledger column matching pre-supervisor within documented deltas — the quantitative proof that teardown succeeded. Any unexplained delta is an orphan you have not found yet.
Step 6

Check toolkit state coherence: run Get-HoloDeckInstance and any Supervisor-related state the toolkit exposes. Record whether the toolkit still believes a Supervisor exists after the vSphere-side removal, and whether a subsequent -DeploySupervisor invocation would be accepted.

Toolkit-state observation recorded (verify_live — undocumented territory and one of this lab's most valuable original findings). If state is stale, note it prominently: it determines whether Task 4 can run in-place or needs the snapshot line.
This is the general automation lesson: acting outside the tool that owns the state creates a divergence you must detect and reconcile deliberately. holodeck91-08 called hand-patched hosts snowflakes; a toolkit that believes in a Supervisor that no longer exists is the same disease in the opposite direction.

Validation Gate

Check: Teardown checklist built before removal; consumption-first ordering followed; Supervisor deactivated with control plane VMs removed; four-layer residue audit complete with every row dispositioned; ledger returned to baseline within documented deltas; toolkit state coherence recorded.

Expected: Provably clean environment ready for the second deployment — and a reusable teardown runbook with evidence.

Common Errors

Deactivation completes but NSX still shows Supervisor segments, LB, or IP allocations
Cause: Incomplete NSX cleanup by the disablement workflow — the most common Supervisor teardown residue
Fix: Remove the orphans manually in NSX Manager, working leaf-to-root (LB virtual servers/pools, then Tier-1 attachments, then segments, then IP allocations), recording each. If the next enablement later fails with address conflicts, this is where you missed one.
📋 KB: Well-known Supervisor disablement residue class; no 9.1-specific KB yet — your audit is the seed
Disable/Remove option missing or greyed out in Workload Management
Cause: Namespaces or workloads still present, or an in-flight task blocking state transition
Fix: Complete step 2 fully (all namespaces gone), wait for running tasks, retry. If genuinely wedged, capture the state and use the 'pre-supervisor-91-09' snapshot line — a wedged teardown is precisely what the snapshot exists for.
Capacity ledger does not return to baseline after apparent full teardown
Cause: Hidden residue: orphaned VMDKs, content library growth, or vSAN objects pending cleanup
Fix: Give vSAN time to complete deletions, then hunt the storage delta specifically (datastore file browser, content library size). Unexplained deltas are findings, not rounding errors.

Task 4 Deploy the Supervisor in Distributed (VNA) mode and contrast the networking path

availability

Distributed mode is the 9.1-new alternative: Supervisor networking rides the domain's VNA cluster instead of the NSX Edge cluster, with the community-observed deployment explicitly skipping Tier-0 retrieval in this mode. Running it immediately after a clean Centralized teardown — same domain, same baseline, same validation chain — turns a feature checkbox into a controlled A/B experiment. The branching honesty: this task requires a VNA cluster, which is a Day-0 property. All three legitimate paths are specified; pick the one your Task 1 finding dictates.

Step 1

Execute your Task 1 branching decision. PATH A (domain has a VNA cluster): proceed in-place. PATH B (no VNA here, holodeck91-10's VNA-enabled instance available): run this task there, after loading that instance's config context and confirming its own capacity gates. PATH C (no VNA anywhere yet): run the deployment command anyway as a documented-failure experiment — capture the exact rejection or failure the toolkit produces when Distributed mode lacks a VNA cluster, which is itself unpublished information; then return to this task after building holodeck91-10's instance.

Branch chosen and recorded with rationale. Path C produces a documented error signature instead of a Supervisor — a legitimate lab outcome that feeds the findings log.
HONESTY GATE: the dependency of Distributed mode on a VNA cluster is inferred from the community-verified deployment (which carried both -VnaClusterMgmtDomain and Distributed Supervisor together) and from the mode's definition — the official docs do not state the failure behavior when the VNA is absent. Whatever Path C shows you is primary evidence; record it as such.
Step 2

Deploy: Update-HoloDeckInstance -Site a -DeploySupervisor -VIDomain Management -SupervisorDeploymentMode Distributed. Capture the transcript and phase timings exactly as in Task 2 step 1.

Supervisor enablement proceeds against the VNA-backed networking path. The Day-2 Distributed path is documented but not yet community-verified — your transcript is potentially the first recorded run; note every deviation from the Centralized transcript.
Same non-interruption rule as Task 2. You now also have a proven teardown runbook (Task 3) as your failure recovery — which changes the risk calculus of the whole operation. Note that shift for the design reflection: rehearsed reversibility is what makes bolder changes acceptable.
Step 3

Validate the control plane identically to Task 2: control plane VMs running, Workload Management Ready, kubectl vsphere login, get nodes / get pods -A. Using the identical validation chain is what makes the comparison valid.

Same green results as Centralized — establishing that the modes are functionally equivalent at the API surface before you examine how differently they get there.
Step 4

Trace the Distributed networking path and diff it against the Task 2 diagram: in NSX (and the VNA components' own surface as discovered in holodeck91-10), determine where the control plane VIP lives, what replaced the edge-hosted LB, whether a Tier-1/Tier-0 chain exists at all (the observed deployment log said 'Skipping T0 retrieval as distributed mode selected'), and how ingress/egress reach the HoloRouter. Draw the second diagram.

Distributed-mode path diagram with the explicit deltas marked: what disappeared (edge dependency, T0 hop), what replaced it (VNA-hosted functions), and what stayed identical (segments, API surface).
Where the VIP lives is the sharpest question. In Centralized mode the answer was 'on the edge cluster'. If in Distributed mode the answer is 'distributed across the VNA cluster', you are looking at the SPOF-versus-scale-out argument of holodeck91-10 materialized in one address.
Step 5

Repeat the minimal consumption proof (one namespace, kubectl visibility) and the capacity measurement (post-distributed ledger column with phase timings).

Working consumption path and the right-hand side of the mode comparison: Distributed footprint and timings alongside Centralized.
If you proceed to a VCF Automation All Apps Org on THIS deployment, the known 9.1 issue applies directly: Day-2 org creation against a distributed-mode domain can fail with the IP Space vs Gateway CIDR mismatch (BAD_REQUEST on the Distributed VLAN Connection). The documented workaround is to align the IP Space range with the gateway CIDR, clean up the partial Region, and re-run. Attempt only with time budgeted.

Validation Gate

Check: Branch decision executed and recorded; Distributed deployment run (or failure experiment documented); identical validation chain passed; Distributed networking path diagrammed with explicit deltas from Centralized; consumption proof and post-distributed ledger column captured.

Expected: Both Supervisor modes deployed and validated with methodologically comparable evidence — or a documented, evidence-backed account of the VNA dependency.

Common Errors

Deployment rejected or fails early with Distributed mode selected
Cause: No VNA cluster in the target domain (Day-0 property absent), or VCF version below 9.1.0.0
Fix: This IS the Path C result — capture the exact error text and stage. Verify VCF version; then either move to a VNA-enabled instance (holodeck91-10) or accept the documented finding.
📋 KB: Distributed mode gated to VCF 9.1.0.0+ per official docs; VNA-absence behavior undocumented
Distributed Supervisor healthy, but a later VCF Automation All Apps Org creation fails with BAD_REQUEST on the Distributed VLAN Connection
Cause: Known open 9.1 issue: IP Space associated with the Distributed VLAN Connection must match the Gateway CIDR, and the generated values mismatch even at default CIDR
Fix: Apply the community workaround: adjust the management IP Space range to match the gateway CIDR, delete the partially created Region, re-run the org creation. Record whether your build reproduces it.
📋 KB: GitHub #143 (open, workaround provided)
Distributed deployment noticeably slower or resource-tighter than Centralized was
Cause: The VNA cluster's own footprint plus Supervisor components landing on an estate that has absorbed every prior Day-2 addition
Fix: Consult the ledger — if physical RAM is above the ceiling, this is the ledger telling you the host is done growing. Reduce scope rather than wait out thrashing; the comparison data is still valid with the caveat recorded.

Task 5 Final teardown, mode-comparison artifact, and the teardown runbook

recoverability

The lab closes by practicing what it preached: the Distributed Supervisor comes down through the same audited teardown, the two deployments become one comparison artifact, and the teardown method itself is promoted from 'steps I did' to a written runbook with acceptance criteria. That artifact set — comparison table, two path diagrams, residue-audit method — is direct defense material for both the Supervisor placement question and the broader 'how do you decommission safely?' question.

Step 1

Tear down the Distributed Supervisor using the Task 3 runbook verbatim: inventory, consumption-first removal, deactivation, four-layer residue audit, ledger return-to-baseline, toolkit-state check. Note every place the Distributed teardown differs from the Centralized one (different NSX/VNA residue is expected and is comparison data).

Second clean teardown with its own audit table; deltas between the two teardowns recorded — particularly which layer held residue in each mode.
Running the same runbook twice against different implementations is how you find out whether you wrote a runbook or a diary. Every step that needed improvisation the second time gets rewritten now.
Step 2

Build the mode-comparison artifact, one row per dimension, columns [Centralized (Edge), Distributed (VNA), Evidence]: resource footprint (ledger columns), deployment wall-clock, networking path (the two diagrams), VIP/LB placement, Tier-0 dependency (present vs skipped), failure domain (edge cluster vs VNA cluster), teardown residue profile, known issues touching each mode.

Complete comparison table, every cell citing captured evidence, ready to defend the placement recommendation it implies.
Step 3

Write the placement recommendation as you would in a design document: for a stated workload profile (choose one — e.g. lab/training platform vs production-adjacent staging), recommend a mode with justification from your table, and state the conditions under which you would choose the other.

One-paragraph, conditions-based recommendation. 'It depends' with explicit conditions is the architect answer; an unconditional preference is the one a panel dismantles.
Cross-reference holodeck91-10's design reflection: the Supervisor mode choice is the consumption-layer echo of the same single-router-versus-distributed argument at the infrastructure layer.
Step 4

Finalize the teardown runbook as a standalone artifact: preconditions (snapshot line, inventory), ordered steps, per-layer audit checklists, acceptance criteria (ledger return-to-baseline, zero undispositioned orphans, toolkit-state coherence), and the escalation path (snapshot revert) with its criteria.

A runbook a colleague could execute without you — the test of whether the discipline transferred out of your head.
Step 5

Close the lab state: verify the environment matches the pre-supervisor baseline (final ledger diff), take the 'holodeck91-09-complete' snapshot across instance VMs, and archive the artifact set (comparison table, diagrams, both audit tables, runbook, transcripts) to the VCDX evidence folder.

Environment at documented baseline; snapshot taken; evidence archived. The estate is clean for whichever lab follows — which is the whole point.

Validation Gate

Check: Distributed teardown completed via the runbook with its audit table; mode-comparison artifact complete with evidence citations; conditions-based placement recommendation written; standalone teardown runbook finalized; environment at baseline with 'holodeck91-09-complete' snapshot and archived evidence.

Expected: Two deployments, two clean teardowns, one defensible comparison — and a transferable teardown discipline.

Common Errors

Second teardown leaves residue the first did not
Cause: Mode-specific artifacts (VNA-side networking state) not covered by the Centralized-derived checklist
Fix: Extend the runbook's inventory step to include the VNA component surface discovered in holodeck91-10; disposition the new orphans and record them — this is exactly the delta the comparison artifact wants.
📋 KB: No 9.1 KB exists for Distributed-mode teardown; your audit is the first record
Final ledger will not reconcile to baseline after both cycles
Cause: Cumulative small residue: content library growth, log/support bundles, vSAN cleanup lag
Fix: Disposition each delta explicitly (KEPT-DELIBERATELY entries are fine; UNKNOWN entries are not). If genuinely irreconcilable, the snapshot line is the honest reset — record that the double-cycle exceeded clean-teardown capability and why.

Final Validation

A vSphere Supervisor was deployed as a Day-2 operation twice on the same Holodeck 9.1 domain — first in Centralized (NSX Edge-backed) mode, then, after a fully audited teardown, in Distributed (VNA-backed) mode — using the Update-HoloDeckInstance -DeploySupervisor parameter set with -SupervisorDeploymentMode. Both deployments passed an identical validation chain (control plane, kubectl, networking-path trace, namespace consumption), both came down through a residue-audited teardown with ledger-verified return to baseline, and the outputs were consolidated into a mode-comparison artifact, a conditions-based placement recommendation, and a standalone teardown runbook. This completes the rn-9-1-019 Day-2 catalog begun in holodeck91-08 and elevates teardown discipline to a first-class, evidenced skill.

✓ Task 1: Change gated — capacity column, health gates, VNA-presence branching decision, 'pre-supervisor-91-09' snapshot line → Reversible, evidence-ready starting state

✓ Task 2: Centralized Supervisor deployed and validated (control plane, kubectl, NSX path diagram, namespace, ledger column) → Reference implementation with quantified cost

✓ Task 3: Centralized teardown with four-layer residue audit and ledger return-to-baseline → Provably clean environment; toolkit-state coherence recorded

✓ Task 4: Distributed deployment per branching path (in-place / holodeck91-10 instance / documented-failure experiment) with identical validation and path diagram → Comparable evidence for the VNA-backed mode or primary evidence of the VNA dependency

✓ Task 5: Final teardown, mode-comparison artifact, placement recommendation, standalone runbook, 'holodeck91-09-complete' snapshot → Clean estate and complete defense artifact set

Cleanup / Restore

Snapshot: holodeck91-09-complete

• Snapshot all instance VMs as 'holodeck91-09-complete' AFTER the final teardown — this snapshot deliberately captures the clean, Supervisor-free baseline, not a deployed Supervisor

• Archive the artifact set to the VCDX evidence folder: mode-comparison table, both networking-path diagrams, both residue-audit tables, the teardown runbook, deployment transcripts, and the full capacity ledger

• If Task 4 ran against the holodeck91-10 instance, repeat the teardown and audit THERE too — do not leave a Supervisor running on an instance another lab owns

• Retain the content library only as a documented KEPT-DELIBERATELY item (it accelerates future Supervisor labs); delete it if storage headroom is marginal

• Update the release-intake notes: mark which rn-9-1-019 Supervisor claims are now field-verified on your build and which remain open (Day-2 Distributed path verification status, VNA-absence failure behavior)

• Delete the 'pre-supervisor-91-09' snapshot line only after 'holodeck91-09-complete' is boot-verified

Design Reflection (VCDX)

Two defense narratives come out of this lab, and the panel will test both past their comfortable versions. (1) SUPERVISOR PLACEMENT — EDGE vs VNA: Centralized mode concentrates Supervisor north-south and load balancing on the NSX Edge cluster: a mature, well-understood path with strong operational tooling, at the cost of making the edge cluster a shared choke point and failure domain for every Supervisor consumer — and of coupling Kubernetes API availability to edge health.

Distributed (VNA) mode spreads those functions across the VNA cluster: the choke point dissolves and the observed deployment even drops the Tier-0 dependency, but you buy a newer, less-proven data plane with a thinner operational corpus (its first public bug — the IP Space/gateway CIDR mismatch — appeared within weeks of GA) and a Day-0 prerequisite (the VNA cluster) that cannot be retrofitted by any documented Day-2 operation.

The panel-grade answer is conditions-based: Centralized where operational maturity, tooling familiarity, and edge capacity already exist; Distributed where edge concentration is the binding constraint — scale-out of consumers, blast-radius reduction, or edge-cluster contention. Then expect the follow-up: 'your Distributed mode depends on a Day-0 decision — what is your migration story for a Centralized estate that later needs Distributed?' The honest answer today is redeploy-and-migrate, and saying so cleanly scores better than inventing an in-place path.

(2) TEARDOWN AS A DESIGNED CAPABILITY: the toolkit's Day-2 catalog deploys Supervisors but does not remove them — an asymmetry that mirrors most real platforms, where build automation outruns decommission automation. This lab's position: reversibility is a designed property with acceptance criteria (residue audit, ledger return-to-baseline, state coherence), not a hope. Note the second-order effect you experienced directly — after Task 3 proved the teardown, Task 4's deployment carried less risk, because rehearsed reversibility changes the risk calculus of change itself.

That is the deep answer to 'how do you de-risk changes in production?': not fewer changes, but proven reversal paths. The counter-question to prepare for: 'your teardown went outside the toolkit — reconcile that with your own snowflake argument from the expansion lab.' Answer: the asymmetry is the toolkit's gap, not the operator's license; the mitigation is the audited runbook plus explicit toolkit-state verification, and the strategic fix is extending the automation, which you name as such.

Requirements

  • Supervisor deployable to a chosen domain as a discrete Day-2 operation in either Centralized or Distributed mode, without redeploying the instance
  • Both modes validated through an identical chain (control plane health, kubectl access, networking-path trace, namespace consumption) so the comparison is methodologically sound
  • Every deployment bracketed by capacity-ledger columns; every teardown proven by residue audit plus ledger return-to-baseline
  • The full sequence reversible at any point via the pre-supervisor snapshot line
  • Outputs consolidated into reusable artifacts: mode-comparison table, placement recommendation, standalone teardown runbook

Constraints

  • Update-HoloDeckInstance parameter sets are mutually exclusive — Supervisor deployment cannot be batched with any other Day-2 operation
  • SupervisorDeploymentMode 'Distributed' requires VCF 9.1.0.0+ and (by all available evidence) a VNA cluster in the target domain — a Day-0 property with no documented Day-2 retrofit
  • No documented toolkit operation removes a Supervisor: teardown is vSphere-side, creating a toolkit-state divergence that must be checked explicitly
  • Nested physical capacity is finite and pre-loaded by holodeck91-08's additions: control plane VMs plus per-mode networking must fit under the ~90% ceiling
  • The Day-2 Distributed deployment path is documented but not community-verified at authoring time; the adjacent VCF Automation org flow has a known open defect in distributed mode
  • Exact UI paths and stage lists may deviate from 9.0-era materials (unified management runtime); procedures reference capabilities, not pixel paths

Assumptions

  • The instance was deployed by Holodeck 9.1 with an NSX Edge cluster in the management domain (verified in Task 1, not assumed silently)
  • Observed 9.1 naming (supervisor-mgmt, zone-mgmt-a, cluster-mgmt-01a) applies to this build — confirmed against the generated config at deployment
  • vSphere-side Supervisor deactivation is the legitimate teardown path in the absence of a toolkit operation, provided the residue audit and state check follow
  • Content library retention between deployments is safe and accelerates the second enablement
  • Single operator; on shared hosts, both deployments and teardowns are announced to co-tenants since edge/VNA and capacity effects are shared

Risks

  • Partial Supervisor enablement spanning vCenter and NSX state — IMPACT: the messiest partial failure in the Day-2 catalog; retries onto residue compound it. MITIGATION: never interrupt; on failure go straight to the audited teardown, then the snapshot line if the audit cannot reconcile.
  • Teardown residue in NSX (segments, LB VIPs, IP allocations) — IMPACT: the NEXT enablement fails with address conflicts, misattributed as a deployment bug. MITIGATION: four-layer residue audit with per-row disposition; NSX audited even when vCenter looks clean.
  • Toolkit-state divergence after vSphere-side teardown — IMPACT: subsequent toolkit operations act on false state (refuse valid operations or duplicate objects). MITIGATION: explicit state-coherence check after every teardown; findings logged as primary evidence since behavior is undocumented.
  • Capacity exhaustion on the third Day-2 wave (hosts + Automation + Supervisor + mode networking) — IMPACT: estate-wide degradation misread as a Supervisor defect. MITIGATION: ledger gates before each deployment; deliberate scope reduction (single mode in-place) when the ceiling approaches.
  • Distributed-mode adjacency defects (IP Space vs gateway CIDR mismatch in the VCF Automation org flow) — IMPACT: Day-2 workflows fail after a healthy Supervisor deploy, burning time on misdiagnosis. MITIGATION: known-issue awareness with the published workaround; org-creation work explicitly time-budgeted or deferred.
  • Unverified Day-2 Distributed path — IMPACT: this lab may hit behavior no one has recorded. MITIGATION: transcript-everything discipline; deviations from the Centralized run treated as findings, not noise; Path C failure experiment defined in advance so even the negative result is evidence.

Self-Assessment Discussion Prompts

  1. Defend Centralized (Edge) mode for a production-adjacent staging environment, then defend Distributed (VNA) mode for the same environment. Which defense needed more conditions attached, and what does that tell you about the maturity gap?
  2. In Centralized mode the Supervisor API VIP lives on the edge cluster; you traced where it lives in Distributed mode. Walk a panel through what fails, and in what order, when each hosting layer degrades — and which failure is easier to detect from the consumer side.
  3. The VNA cluster is a Day-0 decision and Distributed mode depends on it. A customer on a Centralized estate now needs Distributed. Lay out the migration options, their downtime profiles, and the one you would recommend — including what you tell them about why no in-place path exists.
  4. Your teardown went outside the toolkit because the Day-2 catalog is additive. Reconcile that with the snowflake argument from holodeck91-08: when is out-of-tool action legitimate, and what obligations does it create?
  5. Task 3's proven teardown made Task 4's deployment less risky. Generalize: how does rehearsed reversibility change what changes you are willing to approve in production, and how would you evidence 'rehearsed' to a change advisory board?
  6. Your residue audit has three verdicts: REMOVED, KEPT-DELIBERATELY, ORPHANED. A panelist asks why KEPT-DELIBERATELY is not just residue with paperwork. Answer with the content library example and the general principle.
  7. The ledger returned to baseline after teardown. Which single unexplained delta would concern you most — CPU, memory, storage, or an NSX object count — and why?
  8. You may have produced the first recorded run of the Day-2 Distributed path. What obligations come with operating ahead of the community corpus, and how does your evidence discipline discharge them?

Extensions

Workload-Domain Supervisor A/B

Repeat both modes against a workload domain (-VIDomain Workload) on an instance that has one, comparing against the management-domain results: does domain placement change footprint, networking path, or teardown residue? Completes the four-quadrant picture (two modes x two domains) the placement recommendation ideally rests on.

harder

TKC on Top — the #117 Probe

On one deployed Supervisor, attempt a Tanzu Kubernetes cluster creation and validate node provisioning end to end. The 9.0.2-era failure (TKC creation from a pre-installed Supervisor) is still open at authoring time with no 9.1 data point — your run either verifies the fix or produces the first 9.1 reproduction.

harder

Supervisor Failure Injection

With the snapshot line proven, degrade each mode's hosting layer deliberately (power off an edge node under Centralized; a VNA node under Distributed) and record the Supervisor-side symptoms, detection latency, and recovery behavior. Converts the failure-domain rows of your comparison table from reasoning into measurement.

much harder

Automation-Consumed Supervisor

Wire the deployed Supervisor into VCF Automation consumption (namespace or All Apps Org path), deliberately walking into the known distributed-mode IP Space defect with the workaround pre-staged. Produces a documented end-to-end consumption chain plus a first-hand known-issue reproduction — both defense-grade artifacts.

harder

Teardown Runbook Peer Test

Hand your Task 5 runbook to a colleague (or execute it yourself after a two-week gap) against a fresh Supervisor deployment, recording every point of confusion or improvisation. Runbooks are only real after they survive a second operator; fold the findings back in.

same

⚠ Known Pitfalls (from Community KB)

[9.1 known issue] VCFA All Apps Org creation fails against a Distributed-mode domain OPEN (workaround provided)
Problem: Day-2 org creation fails with BAD_REQUEST: the IP Space associated with the Distributed VLAN Connection must match the Gateway CIDR — reproduced even with the default MasterCIDR when distributed mode is selected.
Resolution: Align the management IP Space range with the gateway CIDR, delete the partially created Region, re-run. (GitHub #143)
[Carried 9.0.2-era issue] TKC creation from a pre-installed Supervisor fails OPEN, 9.1 status unverified
Problem: Workload cluster creation on a toolkit-deployed Supervisor failed in the 9.0.2 era (GitHub #117); no 9.1 data point exists.
Resolution: This lab stops at namespace scope by design; the TKC extension doubles as the 9.1 verification probe.
[Carried 9.0-era issue] Nested NSX Edge cluster deployment and health sensitivity OPEN, 9.1 status unverified
Problem: Edge cluster problems (GitHub #94 era) directly gate Centralized-mode Supervisors: a degraded edge blocks the LB/VIP path and presents as a stuck enablement.
Resolution: Gate on edge transport node health before deploying (Task 1 step 4); check edge health FIRST when Centralized enablement stalls.
[Structural] Supervisor teardown has no toolkit operation BY DESIGN in 9.1 (Day-2 catalog is additive)
Problem: Deployment is in-toolkit; removal is vSphere-side, risking NSX residue and toolkit-state divergence.
Resolution: Use this lab's audited teardown runbook: consumption-first ordering, four-layer residue audit, ledger return-to-baseline, explicit toolkit-state coherence check.
[Capacity] Third Day-2 wave on a finite host STRUCTURAL
Problem: Supervisor components land on an estate already carrying the 91-08 host expansion and VCF Automation; overshoot degrades everything at once and masquerades as component bugs.
Resolution: Ledger gates before every deployment; deliberate scope reduction (one mode in-place, the other via holodeck91-10's instance) when headroom is marginal.

References

  • Holodeck 9.1 Documentation — Update-HoloDeckInstance Day-2 parameter setsTier 1 — Official
    Authoritative for the DeploySupervisor set: -Site, -DeploySupervisor, -VIDomain, -SupervisorDeploymentMode (Centralized = NSX Edge, Distributed = VNA; VCF 9.1.0.0+ only; default Centralized), and the one-operation-per-invocation contract. Use the versioned /9.1/ URLs.
  • Announcing the General Availability of Holodeck 9.1 — VMware Cloud Foundation BlogTier 1 — Official
    GA announcement of the Day-2 catalog including Supervisors in VNA/Edge modes.
  • VCF 9.1 Release Intake — rn-9-1-019 (Day-2 operations), rn-9-1-015 (VNA), rn-9-1-009 (VKS 3.6) academy-v2/content/releases/vcf-9.1.jsonlocal fileTier 1 — Official
    Canonical changelog entries: Day-2 Supervisor deployment (rn-9-1-019), the VNA networking model Distributed mode rides on (rn-9-1-015), and the VKS 3.6 scale context (rn-9-1-009).
  • Holodeck GitHub Issue #143 — Day-2 VCFA All Apps Org fails on Distributed-mode deploymentTier 1 — Official
    The community-verified 9.1 deployment combining -VnaClusterMgmtDomain with -DeploySupervisorMgmtDomain Distributed; source of the observed naming (supervisor-mgmt, zone-mgmt-a, cluster-mgmt-01a), the 'Skipping T0 retrieval' log line, and the IP Space / Gateway CIDR workaround.
  • Holodeck GitHub Issue #117 — TKC creation from pre-installed Supervisor fails (9.0.2-era, open)Tier 1 — Official
    Open issue relevant to any extension of this lab into TKC creation; no 9.1 status at authoring time.
  • holodeck91-08 — Day-2 Operations Part 1 (hosts + VCF Automation) holodeck91-08.jsonlocal fileTier 2 — VMware Press
    Direct predecessor: supplies the environment state, the capacity-ledger method, the one-operation contract, and the partial-failure cleanup lessons this lab builds into a full teardown discipline.
  • holodeck91-10 — VNA Cluster Networking holodeck91-10.jsonlocal fileTier 2 — VMware Press
    Companion lab characterizing the VNA cluster itself; provides the VNA-enabled instance for Task 4 Path B and the infrastructure-layer half of the Edge-vs-VNA argument.
  • vSphere Supervisor / Workload Management Documentation (Broadcom Techdocs)Tier 2 — VMware Press
    Reference for Supervisor enablement/disablement mechanics, control plane architecture, and namespace management on the vSphere side.
  • VCF Holodeck Toolkit — Broadcom Community ForumTier 1 — Official
    Community forum. The captured KB corpus remains 9.0-era; no thread on Day-2 Supervisor modes existed at authoring time — this lab's transcripts are candidate seed content.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.