Holodeck 9.1 Click-to-Deploy GitOps — Part 2: Full Site A VCF 9.1.0.0 Deployment from the UI
Objectives
- Read and interpret the auto-provisioned deployment pipeline definition before trusting it with a multi-hour build
- Trigger a full Site A VCF 9.1.0.0 deployment from the Click-to-Deploy UI without executing any manual PowerShell
- Monitor a running deployment across all three observation planes: GitLab pipeline jobs, vSphere inventory on the physical host, and the VCF installer UI
- Validate the deployed management domain end to end (SDDC Manager, vCenter, NSX) to the same health bar as the 9.0 track
- Treat the Git history and pipeline record as the deployment's evidence chain, and assess the GitOps loop for drift and repeatability
- Extend the verify-live findings log with pipeline-stage mapping and deployed component versions
Prerequisites
Holodeck 9.1 GitOps control plane fully provisioned per holodeck91-02: HoloRouter 9.1 running with GitLab, Vault, Authentik, Technitium, and Webtop live behind the HTTPS reverse proxy; Site A VCF 9.1.0.0 configuration committed to Git and diff-reviewed; snapshot 'holodeck91-02-complete' boot-verified. Physical host has full headroom for the nested estate (no other heavyweight VMs contending). Allow an uninterrupted window — the pipeline run is expected in the 2-3 hour range like the 9.0 bring-up, but the 9.1 duration is unmeasured until you measure it.
Prior labs: holodeck91-01, holodeck91-02
Required skills:
- Reading GitLab pipeline/job UIs and CI logs
- VCF management domain architecture and bring-up sequencing (VCF installer, SDDC Manager, vCenter, NSX) — as exercised on the 9.0 track
- vSphere Client inventory monitoring on the physical host
- Health validation of SDDC Manager, vCenter (vSAN/HA/DRS), and NSX transport nodes
- Disciplined evidence capture (screenshots, commit hashes, job IDs)
Lab Environment
Single physical ESXi host carrying the HoloRouter 9.1 (GitOps control plane) from holodeck91-02. This lab's pipeline run provisions the full nested Site A estate: nested ESX 8.0 U3+ hosts for the management cluster, the VCF installer appliance, and through it SDDC Manager, vCenter, and NSX for VCF 9.1.0.0. All nested components sit behind the HoloRouter on the management supernet committed in the Site A configuration (default 10.1.1.0/20 mirrored from the 9.0 track unless you changed it). Technitium DNS on the router provides name resolution for the estate.
graph TB WS[Operator Workstation] -->|HTTPS| GL[GitLab on HoloRouter 9.1<br/>Click-to-Deploy UI] GL -->|pipeline run| PIPE[Deployment Pipeline] PIPE --> ESX1[Nested ESX-01 10.1.1.101] PIPE --> ESX2[Nested ESX-02 10.1.1.102] PIPE --> ESX3[Nested ESX-03 10.1.1.103] PIPE --> ESX4[Nested ESX-04 10.1.1.104] PIPE --> INST[VCF Installer 10.1.1.200 - verify live] INST --> SDDC[SDDC Manager 10.1.1.4] INST --> VC[vCenter 10.1.1.6] INST --> NSX[NSX Manager 10.1.1.10] HR[HoloRouter 9.1<br/>Technitium DNS / gateway] --- ESX1 HR --- SDDC
IP Addressing
| Network | Purpose | VLAN |
|---|---|---|
10.1.1.0/20 | Management supernet (SDDC Manager, vCenter, NSX, VCF installer) — as committed in the Site A config from holodeck91-02 | Per committed config (9.0-era default VLAN 1644; 9.1 VLAN-range handling was fixed — trust the committed values) |
10.1.2.0/24 | vMotion | Per committed config |
10.1.3.0/24 | vSAN | Per committed config |
10.1.4.0/24 | NSX Host TEP | Per committed config |
10.1.5.0/24 | NSX Edge TEP | Per committed config |
Credentials
| System | Username | Password |
|---|---|---|
| GitLab / Click-to-Deploy UI | root (or bootstrap admin per holodeck91-02 findings log) | Mechanism recorded in your holodeck91-02 findings log |
| Nested ESX Hosts | root | Set in the Site A configuration committed in holodeck91-02 Task 5 |
| VCF Installer appliance | admin@local (verify_live for 9.1) | Set in the Site A configuration; record the 9.1 field name in your findings log |
| SDDC Manager | administrator@vsphere.local | SSO password from the Site A configuration |
| vCenter Server | administrator@vsphere.local | Same SSO password |
| NSX Manager | admin | Set in the Site A configuration |
Tasks
Task 1 Read the pipeline before you run it
manageabilityClicking deploy on a pipeline you have not read is operator faith, not engineering. This task maps the 9.1 pipeline stages onto the 9.0 mental model you already own (Prepare -> nested ESX -> installer -> bring-up -> validation) so that when a stage fails at hour two, you know exactly which 9.0-era failure class to reach for.
In GitLab, open the pipeline definition associated with the deployment repository (e.g. .gitlab-ci.yml or the equivalent the appliance generated — per your holodeck91-02 Task 4 findings). Read it top to bottom.
Build the mapping table in your findings log: [9.1 pipeline stage] -> [9.0 cmdlet-flow equivalent] -> [expected duration] -> [most likely 9.0-era failure class]. Use holodeck-02's milestones as the right-hand column: content validation, router config, nested ESX provisioning, installer deployment, bring-up, host commissioning, final validation.
Identify where the pipeline consumes the Site A configuration commit: which job reads it, and whether the pipeline pins a specific commit SHA or floats on the branch HEAD.
Identify the pipeline's validation stages. The 9.1 changelog states VCF installer validation was hardened; locate where in the pipeline validation happens relative to resource-consuming stages.
Confirm the run will target Site A only, per the parked-Site-B state you verified in holodeck91-02 Task 5. Locate whatever selector (variable, job parameter, config flag) scopes the run to one site.
Validation Gate
Check: Findings log contains: enumerated pipeline stages, the 9.1-to-9.0 mapping table, the exact commit the run will consume, the location of validation stages, and the confirmed Site A scoping.
Expected: You can narrate the entire upcoming run before it starts — which is exactly what a panelist means by 'walk me through your deployment'.
Common Errors
Task 2 Trigger the Site A deployment from the UI
manageabilityThis is the moment the 9.1 release headline describes: a full Holodeck deployment 'directly from the UI — no manual PowerShell required'. The discipline is in what you record at trigger time, because the trigger record IS the change record in a GitOps model.
Pre-trigger snapshot check: confirm 'holodeck91-02-complete' exists on the HoloRouter VM (your pipeline-never-ran restore point).
Open the Click-to-Deploy entry point you identified in holodeck91-02 Task 4 and initiate the Site A VCF 9.1.0.0 deployment.
Record the trigger evidence immediately: pipeline/run ID, the commit SHA it consumed (from Task 1 step 3), trigger timestamp, and the identity the UI attributed the trigger to.
Watch the first validation stage complete before walking away. The hardened installer validation should pass cleanly given your holodeck91-02 Task 5 review.
Validation Gate
Check: Pipeline run is in progress past its validation stage, and the findings log holds the full trigger record (run ID, commit SHA, timestamp, attributed identity).
Expected: Deployment running, evidence chain started, zero manual PowerShell executed.
Common Errors
Task 3 Monitor the run across three observation planes
availabilityA GitOps deployment gives you a new primary console (the pipeline) but the old planes still tell truths the pipeline cannot: vSphere shows what actually got provisioned, and the installer UI shows bring-up internals. Operators who watch only the pipeline discover failures late; operators who can triangulate across all three planes diagnose in minutes. This is the 9.1 version of holodeck-02's 'monitor via Verbose stream plus installer UI' guidance.
Plane 1 — GitLab: keep the pipeline graph open and follow the live job log of the active stage. Note each stage's start/end time against your Task 1 duration priors.
Plane 2 — vSphere Client on the physical host: watch the inventory as nested VMs appear. Expect nested ESX hosts first, then the VCF installer appliance, mirroring the 9.0 ordering.
Plane 3 — VCF installer UI: once the installer appliance is up (per your committed addressing, e.g. https://10.1.1.200 — verify the actual address from the pipeline log or committed config), open it from your workstation or the authenticated Webtop and follow the bring-up tracker.
During the long bring-up middle (SDDC Manager/vCenter/NSX deployment), sample the HoloRouter's resource utilization in vSphere and confirm Technitium DNS is resolving for the estate: from your workstation, resolve an estate FQDN against the router (e.g. nslookup sddc-manager.<your-site-a-domain> <holorouter-mgmt-ip>).
Record stage completion times as they land, building the measured 9.1 duration profile next to your 9.0 priors in the mapping table.
Wait for the pipeline to reach completion. Do not intervene on transient warnings; intervene only on hard stage failure (see common errors).
Validation Gate
Check: Pipeline green end to end; nested estate visible in vSphere (nested ESX hosts, installer, SDDC Manager, vCenter, NSX all powered on); installer UI reports bring-up complete; measured duration profile recorded; mid-run Technitium resolution check logged.
Expected: Full Site A VCF 9.1.0.0 deployment completed from a single UI trigger with zero manual PowerShell.
Common Errors
Task 4 Validate the deployed management domain end to end
availabilityA green pipeline is a claim, not proof. This task applies the same health bar the 9.0 track set in holodeck-02 Tasks 4-5 — SDDC Manager, vCenter cluster health, NSX fabric — so the two tracks are comparable artifact-for-artifact in your VCDX evidence pack.
From your workstation (or the authenticated Webtop), open SDDC Manager at its committed address (e.g. https://10.1.1.4 or its FQDN via Technitium) and log in with administrator@vsphere.local.
Verify Inventory > Workload Domains shows the management domain Active, and Inventory > Hosts shows all nested ESX hosts Active with no errors.
Open vCenter at its committed address and validate the management cluster: all hosts connected, vSAN health green, HA and DRS enabled.
Open NSX Manager at its committed address and validate: System > Fabric > Nodes shows all host transport nodes with Configuration State Success and Status Up.
Record the deployed component versions (SDDC Manager, vCenter, NSX, ESX build on nested hosts) from their respective About/summary pages into your findings log.
Confirm end-to-end name resolution and access: resolve and browse SDDC Manager, vCenter, and NSX by FQDN from your workstation using Technitium as resolver.
Validation Gate
Check: SDDC Manager: management domain Active with all hosts. vCenter: cluster healthy, vSAN green, HA/DRS on. NSX: all transport nodes Success/Up. Component version list recorded. FQDN access via Technitium confirmed for all three UIs.
Expected: Management domain healthy to the identical bar as the 9.0 track, with versions and access evidence logged.
Common Errors
Task 5 Audit the evidence chain and probe the GitOps loop for drift
manageabilityThis is where the GitOps promise is tested, not assumed. Repeatability, auditability, and drift-resistance are the three claims that justified moving off PowerShell — each gets a concrete probe here. The outputs are direct exhibits for a VCDX defense: 'show me your change record' has a two-click answer if this task is done honestly.
Assemble the deployment evidence chain in your findings log as a linked sequence: config commit SHA -> pipeline run ID -> stage log excerpts -> health screenshots -> component version list. Verify every link actually cross-references (the run really points at that SHA, etc.).
Drift probe: make a small, deliberate out-of-band change to the running estate that contradicts the committed config (e.g. rename something cosmetic or alter a non-critical setting directly in vCenter — choose something reversible and harmless).
Observe whether any part of the 9.1 toolkit detects or reports the divergence, and what a pipeline re-run would do about it (read the pipeline definition rather than re-running the full multi-hour build; re-run only if the pipeline supports a cheap validation-only mode).
Revert your out-of-band change in the estate, and log the drift-probe result: detected/undetected, reconciled/ignored.
Repeatability assessment on paper: using your evidence chain, write the exact minimal sequence a colleague would follow to reproduce this estate from a bare 8.0 U3 host (holodeck91-01 staging -> holodeck91-02 OVA + commit -> single UI trigger). Note every point where tribal knowledge (your findings log) is still required.
Take the completion snapshots: snapshot the HoloRouter and the nested estate VMs as 'holodeck91-03-complete'. PowerCLI on the physical host: Get-VM -Name '<your-91-instance-prefix>*' | New-Snapshot -Name 'holodeck91-03-complete' -Description 'Site A VCF 9.1.0.0 deployed via GitOps pipeline, validated healthy' -Memory:$false
Validation Gate
Check: Evidence chain assembled and cross-verified; drift probe executed, observed, reverted, and its conclusion logged; reproduction recipe written with tribal-knowledge gaps listed; 'holodeck91-03-complete' snapshots in place across router and estate.
Expected: GitOps claims tested against reality, evidence pack defense-ready, environment snapshotted for the labs that follow.
Common Errors
Final Validation
A complete Site A VCF 9.1.0.0 management domain is running, deployed entirely from the Click-to-Deploy UI by a GitOps pipeline with zero manual PowerShell. Health matches the 9.0-track bar (SDDC Manager Active, vCenter cluster green, NSX fabric up). The deployment's evidence chain — commit SHA to run ID to health screenshots — is assembled, the drift behavior of the toolkit has been probed and recorded, and the estate is snapshotted as 'holodeck91-03-complete'.
✓ Pipeline run: green end to end from a single UI trigger → No manual PowerShell anywhere in the run; interventions (if any) logged as findings
✓ SDDC Manager: Site A management domain Active with all nested hosts → Active domain, all hosts healthy
✓ vCenter: cluster with all hosts connected, vSAN green, HA/DRS enabled → All green, no alarms
✓ NSX Manager: all host transport nodes Success/Up → Fabric fully operational
✓ Evidence chain: commit SHA -> run ID -> logs -> health artifacts, cross-verified → No broken links; comparison against 9.0 evidence-availability written
✓ Drift probe executed and conclusion recorded (reconciling vs fire-and-forget) → Grounded, evidenced answer in the findings log
✓ Snapshots 'holodeck91-03-complete' across router and estate → Restore point for subsequent 9.1 labs
Cleanup / Restore
Snapshot: holodeck91-03-complete
• Snapshot router and estate as 'holodeck91-03-complete' (done in Task 5)
• Export or bookmark the pipeline run page and archive key job logs — GitLab log retention on the appliance is unverified (verify_live), so do not assume the run record persists indefinitely
• Update the findings log's measured-duration column and mark the deployment-path open question (GitOps vs cmdlets) as resolved-by-experience for your environment
• Close all bootstrap-admin sessions; note in the log that identity hardening (Authentik) is pending holodeck91-04
Design Reflection (VCDX)
The panel question this lab arms you for is: 'Your design says deployments are automated and auditable — prove it, and tell me what breaks.' You can now answer with a real evidence chain (commit -> run -> healthy estate) and a tested drift conclusion, instead of vendor language. Expect follow-ups on both flanks. Operational flank: what happens when the pipeline fails at stage four of seven — is partial state cleaned up, resumed, or orphaned, and how do you know? Governance flank: the UI trigger was attributed to a bootstrap admin — who SHOULD be able to click deploy, who approves the config commit, and where is separation of duties? A candidate who presents GitOps as 'better because newer' fails this exchange; a candidate who presents measured durations, a probed reconciliation model, and an honest tribal-knowledge gap list turns it into the strongest section of their defense.
Requirements
- Full Site A VCF 9.1.0.0 management domain deployable from a single UI action against a reviewed, committed configuration
- Deployment must produce a verifiable evidence chain linking configuration version to running state
- Post-deployment health must meet the same bar as the 9.0 cmdlet track (SDDC Manager Active, vCenter cluster green, NSX fabric up) for cross-track comparability
- Deployment process must be reproducible by another operator from documented steps alone
Constraints
- Pipeline internals, stage semantics, and log retention are undocumented — operational knowledge must be built empirically and recorded
- The GitOps control plane and the deployment target share one physical host; pipeline resource demand competes with the estate it is building
- Multi-hour unattended run requires a change freeze on the config repo (commit-consumption semantics verified in Task 1, not assumed)
- No 9.1 community KB exists; all troubleshooting maps 9.0-era failure classes onto new mechanisms
- Nested-lab performance figures (durations, resource profiles) do not transfer to production sizing — they inform lab operations only
Assumptions
- The committed Site A configuration from holodeck91-02 is complete and internally consistent (enforced partly by 9.1's hardened validation)
- Technitium DNS registers and serves estate records during bring-up without the manual forwarder intervention the 9.0 track required — verified live in Task 3
- Pipeline failure leaves the control plane (GitLab/Git history) intact, so diagnosis and re-run are always possible from the router
- The physical host remains dedicated to this instance for the duration of the run
Risks
- Mid-pipeline failure with partial estate state — IMPACT: orphaned nested VMs consuming resources and blocking clean re-runs; MITIGATION: 'holodeck91-02-complete' restore point, documented cleanup-before-rerun check, idempotency observations from Task 3 common errors
- Green-but-wrong: pipeline stage success criteria weaker than actual provisioning success — IMPACT: false confidence, late discovery; MITIGATION: three-plane monitoring with vSphere as ground truth, end-to-end health validation in Task 4
- Audit trail rot: GitLab run logs on the appliance may not persist — IMPACT: evidence chain breaks retroactively; MITIGATION: export/archive run records at completion (cleanup action)
- Fire-and-forget loop: if the toolkit does not reconcile drift, the Git repo silently stops describing reality — IMPACT: the audit trail becomes actively misleading over time; MITIGATION: drift probe conclusion drives either periodic re-validation runs or explicit 'point-in-time only' labeling of the config repo
- Bootstrap-identity triggering: all actions attributed to a shared root account — IMPACT: attribution without accountability; MITIGATION: Authentik SSO integration (holodeck91-04) before treating the trail as audit-grade
Self-Assessment Discussion Prompts
- Compare your measured 9.1 pipeline deployment against the holodeck-02 cmdlet deployment across: operator time invested, wall-clock time, evidence produced, and skill floor required. Where does each path win?
- Repeatability: the cmdlet path re-run depends on the operator re-executing steps identically; the pipeline re-run depends on the same commit producing the same result. Which failure modes does each dependence create, and which did you actually observe?
- Auditability: your evidence chain links commit to running estate. A panelist asks 'could an operator have made the estate diverge from that commit without leaving a trace?' Answer honestly from your drift-probe result, and design the control that closes whatever gap you found.
- Drift: is the 9.1 toolkit a reconciling GitOps system or a versioned trigger system, per your Task 5 observation? What operational practice must you add in the second case to keep Git truthful?
- The pipeline failed at the nested-ESX stage for a colleague (datastore misnamed in the committed config). Walk through the recovery: what gets fixed where, what gets cleaned up, and what does the Git history look like afterwards if done properly?
- Your reproduction recipe still contains tribal-knowledge items from the findings log. For each, decide: automate it, document it into the repo, or accept it — and justify the split as you would in a design defense.
- In production VCF, would you accept a deployment system whose control plane (GitLab) lives inside the failure domain it deploys (on the router of the estate)? Map this to production analogs and state your placement decision.
Extensions
Teardown-and-Redeploy Repeatability Proof
Destroy the nested estate (retaining the router and Git history), re-trigger the identical commit, and diff the resulting estate against your Task 4 validation artifacts. Measure whether the second run is faster, identical, and intervention-free. This converts the repeatability claim from argument to experiment.
samePipeline Failure Injection
Deliberately commit a plausible-but-wrong configuration (bad datastore name, undersized host) and observe where the hardened validation catches it versus where it slips through to a late stage failure. Document the failure surface map and the cleanup burden of each late failure. Mirrors the 9.0 track's failure-injection extension under the new driver.
harderSite B Dual-Site Activation
Activate the parked Site B configuration and drive a second-site deployment from the same control plane, managing IP/VLAN separation from Site A. Exercises 9.1's dual-site auto-generation end to end and sets up DR/stretched design discussion. On a shared host, coordinate first — this doubles the footprint.
much harderDay-2 Operations from the Toolkit
Use the 9.1 Day-2 capabilities announced for the toolkit — add a nested ESX host to the estate, then deploy VCF Automation or spin up a Supervisor in VNA/Edge mode — and observe whether Day-2 actions flow through the same Git/pipeline loop or bypass it. The answer determines whether the audit chain survives past day one.
harder⚠ Known Pitfalls (from Community KB)
References
- HoloDeck 9.1 Release Announcement (verbatim capture)
Holodeck-9.1/01-release-announcement.mdlocal fileTier 1 — Official
Primary source: 'trigger full Holodeck deployments directly from the UI — no manual PowerShell required'; VCF 9.1.0.0 support; hardened VCF installer validation. - HoloDeck Toolkit 9.1 — Structured Reference & Changelog
Holodeck-9.1/02-structured-changelog.mdlocal fileTier 1 — Official
9.0-to-9.1 orientation table (deployment driver: PowerShell -> GitOps/UI) and open-questions list tracked by this lab's findings log. - VCF 9.1 Release Intake — rn-9-1-016 (Click-to-Deploy GitOps)
academy-v2/content/releases/vcf-9.1.jsonlocal fileTier 1 — Official
Confirms both paths must be presented in lab guidance: GitOps/UI (this lab) and cmdlets (holodeck-02 remains the 9.0.x record). - VMware Cloud Foundation Documentation (Broadcom)Tier 1 — Official
Authoritative VCF documentation portal — consult the 9.1 deployment and bring-up guides for component-level detail as they publish. - VCF Holodeck Toolkit — Broadcom Community ForumTier 1 — Official
All captured KB threads are 9.0-era. Deployment-phase failure classes (datastore, legacy CPU, installer readiness) are referenced here as carried-over hypotheses, not confirmed 9.1 behavior. - GitLab CI/CD Pipelines DocumentationTier 2 — VMware Press
For reading pipeline graphs, job logs, and run metadata used throughout Tasks 1-3 and the evidence chain in Task 5. - OpenGitOps PrinciplesTier 2 — VMware Press
Framework for the Task 5 drift probe: distinguishing a reconciling GitOps system from a versioned trigger system.