Academy/Holodeck Lab Setup & Operations/Holodeck 9.1 Click-to-Deploy GitOps — Part 2: Full Site A VCF 9.1.0.0 Deployment from the UI
This lab targets VCF 9.1.0.0

Holodeck 9.1 Click-to-Deploy GitOps — Part 2: Full Site A VCF 9.1.0.0 Deployment from the UI

VCF 9.1.0.0Intermediateadminarchitectvcdx⏱ 300 min

Written for Holodeck Toolkit 9.1 GA deploying VCF 9.1.0.0 Site A. The same UI/pipeline flow applies to VVF 9.1.0.0 and down-rev targets to VCF 5.2 with different content selections. NSX and component build numbers ship inside the VCF 9.1.0.0 content and are not independently selected in this flow — record the deployed component versions during validation (verify_live). This is the 9.1 parallel twin of holodeck-02's deployment tasks; the 9.0.x cmdlet-driven deployment remains documented in holodeck-02 and is not superseded.

Objectives

  • Read and interpret the auto-provisioned deployment pipeline definition before trusting it with a multi-hour build
  • Trigger a full Site A VCF 9.1.0.0 deployment from the Click-to-Deploy UI without executing any manual PowerShell
  • Monitor a running deployment across all three observation planes: GitLab pipeline jobs, vSphere inventory on the physical host, and the VCF installer UI
  • Validate the deployed management domain end to end (SDDC Manager, vCenter, NSX) to the same health bar as the 9.0 track
  • Treat the Git history and pipeline record as the deployment's evidence chain, and assess the GitOps loop for drift and repeatability
  • Extend the verify-live findings log with pipeline-stage mapping and deployed component versions

Prerequisites

Holodeck 9.1 GitOps control plane fully provisioned per holodeck91-02: HoloRouter 9.1 running with GitLab, Vault, Authentik, Technitium, and Webtop live behind the HTTPS reverse proxy; Site A VCF 9.1.0.0 configuration committed to Git and diff-reviewed; snapshot 'holodeck91-02-complete' boot-verified. Physical host has full headroom for the nested estate (no other heavyweight VMs contending). Allow an uninterrupted window — the pipeline run is expected in the 2-3 hour range like the 9.0 bring-up, but the 9.1 duration is unmeasured until you measure it.

Prior labs: holodeck91-01, holodeck91-02

Required skills:

  • Reading GitLab pipeline/job UIs and CI logs
  • VCF management domain architecture and bring-up sequencing (VCF installer, SDDC Manager, vCenter, NSX) — as exercised on the 9.0 track
  • vSphere Client inventory monitoring on the physical host
  • Health validation of SDDC Manager, vCenter (vSAN/HA/DRS), and NSX transport nodes
  • Disciplined evidence capture (screenshots, commit hashes, job IDs)
📸 Starting State: S3-91 — Management Domain Deployed (VCF 9.1.0.0)

Lab Environment

Single physical ESXi host carrying the HoloRouter 9.1 (GitOps control plane) from holodeck91-02. This lab's pipeline run provisions the full nested Site A estate: nested ESX 8.0 U3+ hosts for the management cluster, the VCF installer appliance, and through it SDDC Manager, vCenter, and NSX for VCF 9.1.0.0. All nested components sit behind the HoloRouter on the management supernet committed in the Site A configuration (default 10.1.1.0/20 mirrored from the 9.0 track unless you changed it). Technitium DNS on the router provides name resolution for the estate.

graph TB
  WS[Operator Workstation] -->|HTTPS| GL[GitLab on HoloRouter 9.1<br/>Click-to-Deploy UI]
  GL -->|pipeline run| PIPE[Deployment Pipeline]
  PIPE --> ESX1[Nested ESX-01 10.1.1.101]
  PIPE --> ESX2[Nested ESX-02 10.1.1.102]
  PIPE --> ESX3[Nested ESX-03 10.1.1.103]
  PIPE --> ESX4[Nested ESX-04 10.1.1.104]
  PIPE --> INST[VCF Installer 10.1.1.200 - verify live]
  INST --> SDDC[SDDC Manager 10.1.1.4]
  INST --> VC[vCenter 10.1.1.6]
  INST --> NSX[NSX Manager 10.1.1.10]
  HR[HoloRouter 9.1<br/>Technitium DNS / gateway] --- ESX1
  HR --- SDDC

IP Addressing

NetworkPurposeVLAN
10.1.1.0/20Management supernet (SDDC Manager, vCenter, NSX, VCF installer) — as committed in the Site A config from holodeck91-02Per committed config (9.0-era default VLAN 1644; 9.1 VLAN-range handling was fixed — trust the committed values)
10.1.2.0/24vMotionPer committed config
10.1.3.0/24vSANPer committed config
10.1.4.0/24NSX Host TEPPer committed config
10.1.5.0/24NSX Edge TEPPer committed config

Credentials

SystemUsernamePassword
GitLab / Click-to-Deploy UIroot (or bootstrap admin per holodeck91-02 findings log)Mechanism recorded in your holodeck91-02 findings log
Nested ESX HostsrootSet in the Site A configuration committed in holodeck91-02 Task 5
VCF Installer applianceadmin@local (verify_live for 9.1)Set in the Site A configuration; record the 9.1 field name in your findings log
SDDC Manageradministrator@vsphere.localSSO password from the Site A configuration
vCenter Serveradministrator@vsphere.localSame SSO password
NSX ManageradminSet in the Site A configuration

Tasks

Task 1 Read the pipeline before you run it

manageability

Clicking deploy on a pipeline you have not read is operator faith, not engineering. This task maps the 9.1 pipeline stages onto the 9.0 mental model you already own (Prepare -> nested ESX -> installer -> bring-up -> validation) so that when a stage fails at hour two, you know exactly which 9.0-era failure class to reach for.

Step 1

In GitLab, open the pipeline definition associated with the deployment repository (e.g. .gitlab-ci.yml or the equivalent the appliance generated — per your holodeck91-02 Task 4 findings). Read it top to bottom.

A stage/job graph you can enumerate: stage names, ordering, and any manual gates.
The real stage names are undocumented in the release announcement — record what the live toolkit actually defines (verify_live) and resist renaming them to 9.0 terms in your notes; keep a mapping table instead.
Step 2

Build the mapping table in your findings log: [9.1 pipeline stage] -> [9.0 cmdlet-flow equivalent] -> [expected duration] -> [most likely 9.0-era failure class]. Use holodeck-02's milestones as the right-hand column: content validation, router config, nested ESX provisioning, installer deployment, bring-up, host commissioning, final validation.

A one-page mapping that turns the unfamiliar pipeline into familiar territory.
Expected durations for the 9.0 flow: Prepare 15-30 min, full bring-up 90-150 min. Use these as priors, then correct them with your measured 9.1 numbers after the run.
Step 3

Identify where the pipeline consumes the Site A configuration commit: which job reads it, and whether the pipeline pins a specific commit SHA or floats on the branch HEAD.

You can state precisely which commit the run will deploy.
If the pipeline floats on HEAD, any commit pushed mid-run could matter (or not, depending on when the config is read). Note the behavior — it defines your change-freeze discipline during deployments.
Step 4

Identify the pipeline's validation stages. The 9.1 changelog states VCF installer validation was hardened; locate where in the pipeline validation happens relative to resource-consuming stages.

Validation stage(s) identified, ideally early (fail-fast) rather than post-provisioning.
Early validation is the concrete payoff of the 9.1 hardening: 9.0-era deployments burned hours before failing on content mismatches that a pre-check could have caught.
Step 5

Confirm the run will target Site A only, per the parked-Site-B state you verified in holodeck91-02 Task 5. Locate whatever selector (variable, job parameter, config flag) scopes the run to one site.

Site scoping mechanism identified and confirmed to be Site A.
Dual-site auto-generation is a 9.1 feature; an accidental dual-site run doubles resource demand and, on a shared host, risks IP/VLAN collisions with any existing instance. Verify the scope BEFORE triggering.

Validation Gate

Check: Findings log contains: enumerated pipeline stages, the 9.1-to-9.0 mapping table, the exact commit the run will consume, the location of validation stages, and the confirmed Site A scoping.

Expected: You can narrate the entire upcoming run before it starts — which is exactly what a panelist means by 'walk me through your deployment'.

Common Errors

Pipeline definition is opaque wrapper scripts with no readable stages
Cause: The appliance may encapsulate logic in container images or bundled scripts rather than inline YAML
Fix: Read what is readable (stage names, job names, artifacts), note the opaque boundaries explicitly in your mapping table, and treat opaque stages as single failure units for monitoring purposes.
📋 KB: verify_live — pipeline internals undocumented; no 9.1 KB exists

Task 2 Trigger the Site A deployment from the UI

manageability

This is the moment the 9.1 release headline describes: a full Holodeck deployment 'directly from the UI — no manual PowerShell required'. The discipline is in what you record at trigger time, because the trigger record IS the change record in a GitOps model.

Step 1

Pre-trigger snapshot check: confirm 'holodeck91-02-complete' exists on the HoloRouter VM (your pipeline-never-ran restore point).

Snapshot present. If missing, take it now before triggering anything.
Step 2

Open the Click-to-Deploy entry point you identified in holodeck91-02 Task 4 and initiate the Site A VCF 9.1.0.0 deployment.

A pipeline run starts. The UI reflects a running deployment.
Do not open a PowerShell session 'just to help it along'. The entire point of this lab is to prove the UI-only path; any manual intervention invalidates the repeatability claim you will make in the design reflection. If the run genuinely requires manual intervention to succeed, that is a FINDING — record it prominently.
Step 3

Record the trigger evidence immediately: pipeline/run ID, the commit SHA it consumed (from Task 1 step 3), trigger timestamp, and the identity the UI attributed the trigger to.

Complete trigger record in your findings log.
Note WHO the system says triggered it. If everything is attributed to a bootstrap root account, your audit story has a hole that Authentik SSO (holodeck91-04) exists to close.
Step 4

Watch the first validation stage complete before walking away. The hardened installer validation should pass cleanly given your holodeck91-02 Task 5 review.

Validation stage green within the early minutes of the run.
If validation fails here, STOP and fix via a new commit through the UI — do not hand-patch anything on the router or bypass the gate. Failed-validation-plus-clean-fix-commit is a better audit trail than a mysteriously green second run.

Validation Gate

Check: Pipeline run is in progress past its validation stage, and the findings log holds the full trigger record (run ID, commit SHA, timestamp, attributed identity).

Expected: Deployment running, evidence chain started, zero manual PowerShell executed.

Common Errors

Validation stage fails on content or version checks
Cause: Staged content does not match the committed VCF 9.1.0.0 selection — the 9.1 hardened validation failing fast (working as designed)
Fix: Fix the content staging or the committed configuration via the UI (new commit), then re-trigger. Log both the failure and fix commits — this pair is a textbook auditability artifact.
📋 KB: 9.0-era manifest-mismatch failure class (e.g. 'fail to deploy vcf 9.0.2.0'), now caught early by 9.1 hardening
Trigger button/flow not found where expected
Cause: Entry-point assumption from holodeck91-02 was wrong or a UI update changed it
Fix: Re-locate via GitLab's pipeline pages (Pipelines > Run pipeline) as the fallback path, and correct your findings log entry.
📋 KB: verify_live

Task 3 Monitor the run across three observation planes

availability

A GitOps deployment gives you a new primary console (the pipeline) but the old planes still tell truths the pipeline cannot: vSphere shows what actually got provisioned, and the installer UI shows bring-up internals. Operators who watch only the pipeline discover failures late; operators who can triangulate across all three planes diagnose in minutes. This is the 9.1 version of holodeck-02's 'monitor via Verbose stream plus installer UI' guidance.

Step 1

Plane 1 — GitLab: keep the pipeline graph open and follow the live job log of the active stage. Note each stage's start/end time against your Task 1 duration priors.

Stages progressing in order; job logs streaming provisioning detail.
Pin the job-log view of the current stage rather than the graph — the logs surface warnings that never bubble up to stage status.
Step 2

Plane 2 — vSphere Client on the physical host: watch the inventory as nested VMs appear. Expect nested ESX hosts first, then the VCF installer appliance, mirroring the 9.0 ordering.

Nested ESX VMs power on and the installer appliance appears, consistent with pipeline progress.
If the pipeline claims progress but no VMs appear (or vice versa), you have a control-plane/data-plane disagreement — treat the vSphere view as ground truth and start diagnosis from the pipeline job that should have created the missing objects.
Step 3

Plane 3 — VCF installer UI: once the installer appliance is up (per your committed addressing, e.g. https://10.1.1.200 — verify the actual address from the pipeline log or committed config), open it from your workstation or the authenticated Webtop and follow the bring-up tracker.

Installer UI shows the bring-up task sequence for SDDC Manager, vCenter, and NSX with per-task status.
The installer's granular tracker remains the best diagnostic view during bring-up, exactly as on the 9.0 track. The Webtop now requires login (9.1 change) — credentials per your holodeck91-02 findings log.
Step 4

During the long bring-up middle (SDDC Manager/vCenter/NSX deployment), sample the HoloRouter's resource utilization in vSphere and confirm Technitium DNS is resolving for the estate: from your workstation, resolve an estate FQDN against the router (e.g. nslookup sddc-manager.<your-site-a-domain> <holorouter-mgmt-ip>).

Router utilization within reason; estate FQDNs resolving via Technitium as records are registered.
On the 9.0 track, DNS gaps (missing upstream forwarders in dnsmasq) caused mid-bring-up hangs. Technitium supersedes that path in 9.1 — but 'supersedes' is a hypothesis until this resolution check passes during a live bring-up. Log the result either way; holodeck91-04 goes deep on it.
Step 5

Record stage completion times as they land, building the measured 9.1 duration profile next to your 9.0 priors in the mapping table.

Mapping table gains a 'measured' column; total wall-clock tracked toward the final number.
This measured profile is original data. It also tells future-you when a 9.1 stage is genuinely stuck versus merely slow — the exact judgment the 9.0 KB threads spent pages arguing about.
Step 6

Wait for the pipeline to reach completion. Do not intervene on transient warnings; intervene only on hard stage failure (see common errors).

Pipeline completes green; installer UI shows bring-up complete; vSphere shows the full nested estate powered on.
Expected total is in the multi-hour range. Walking away is fine; intervening out of impatience is how clean audit trails die.

Validation Gate

Check: Pipeline green end to end; nested estate visible in vSphere (nested ESX hosts, installer, SDDC Manager, vCenter, NSX all powered on); installer UI reports bring-up complete; measured duration profile recorded; mid-run Technitium resolution check logged.

Expected: Full Site A VCF 9.1.0.0 deployment completed from a single UI trigger with zero manual PowerShell.

Common Errors

Pipeline stage fails during nested ESX provisioning
Cause: Carried-over 9.0 failure classes: datastore capacity/name issues ('Cannot select datastore', the most-discussed 9.0 KB thread), or CPU compatibility on older hardware (allowLegacyCPU boot.cfg class)
Fix: Diagnose exactly as on 9.0 (datastore free space and naming, physical CPU generation), but REMEDIATE through the GitOps loop: fix the committed configuration or host, then re-run the pipeline. Log whether the re-run is idempotent or requires cleanup of partial state.
📋 KB: 9.0-era Threads 51 (datastore) and 8 (legacy CPU) — mechanisms carry over, driver changed
Bring-up stalls with the installer waiting or unready
Cause: 9.0-era 'VCF Installer is not ready yet' class: installer appliance up but initialization blocked, historically DNS/NTP
Fix: Check DNS resolution against Technitium and time sync on the installer appliance. The DNS root cause SHOULD be structurally improved by Technitium in 9.1 — if you still hit it, that is a headline finding for your uplift notes.
📋 KB: 9.0-era Thread 5 ('VCF Installer is not ready yet') — re-validate under Technitium
Pipeline shows green for a stage whose artifacts are visibly absent in vSphere
Cause: Stage success criteria weaker than actual provisioning success (control-plane/data-plane disagreement)
Fix: Trust vSphere. Capture the job log of the lying stage, file it in your findings log as a 9.1 toolkit defect candidate, and do not proceed to dependent stages without manual verification of the missing artifacts.
📋 KB: No 9.1 KB — document rigorously

Task 4 Validate the deployed management domain end to end

availability

A green pipeline is a claim, not proof. This task applies the same health bar the 9.0 track set in holodeck-02 Tasks 4-5 — SDDC Manager, vCenter cluster health, NSX fabric — so the two tracks are comparable artifact-for-artifact in your VCDX evidence pack.

Step 1

From your workstation (or the authenticated Webtop), open SDDC Manager at its committed address (e.g. https://10.1.1.4 or its FQDN via Technitium) and log in with administrator@vsphere.local.

SDDC Manager dashboard loads showing the Site A management domain.
Step 2

Verify Inventory > Workload Domains shows the management domain Active, and Inventory > Hosts shows all nested ESX hosts Active with no errors.

Management domain Active; all hosts (4 by default sizing) Active, no red/yellow.
Screenshot this view with the timestamp visible — it pairs with the pipeline run ID as your deployment-complete evidence.
Step 3

Open vCenter at its committed address and validate the management cluster: all hosts connected, vSAN health green, HA and DRS enabled.

Cluster healthy: hosts connected, vSAN green, HA/DRS on — the same bar as holodeck-02 Task 4.
Step 4

Open NSX Manager at its committed address and validate: System > Fabric > Nodes shows all host transport nodes with Configuration State Success and Status Up.

All transport nodes Success/Up. NSX cluster stable.
NSX remains the most resource-sensitive component in nested labs; if transport nodes are degraded, check nested host memory pressure before suspecting the 9.1 toolkit.
Step 5

Record the deployed component versions (SDDC Manager, vCenter, NSX, ESX build on nested hosts) from their respective About/summary pages into your findings log.

Concrete component version list for VCF 9.1.0.0 as actually deployed.
The announcement names only the VCF train (9.1.0.0), not component builds. Your recorded list closes that gap and becomes the version table for the 9.1 study track.
Step 6

Confirm end-to-end name resolution and access: resolve and browse SDDC Manager, vCenter, and NSX by FQDN from your workstation using Technitium as resolver.

All three management UIs reachable by FQDN. Certificate warnings acceptable at this stage (trust workflow is a holodeck91-04 topic).
On the 9.0 track this required manual routes plus dnsmasq forwarder fixes (holodeck-04). Note in your log which parts 9.1 made unnecessary — that delta is design-reflection fuel.

Validation Gate

Check: SDDC Manager: management domain Active with all hosts. vCenter: cluster healthy, vSAN green, HA/DRS on. NSX: all transport nodes Success/Up. Component version list recorded. FQDN access via Technitium confirmed for all three UIs.

Expected: Management domain healthy to the identical bar as the 9.0 track, with versions and access evidence logged.

Common Errors

SDDC Manager UI unreachable shortly after pipeline completion
Cause: Services still initializing post-bring-up (carried over from 9.0 behavior), or DNS record not yet present in Technitium
Fix: Wait 10-15 minutes, verify the FQDN resolves against Technitium, then check service status on the appliance. Only then suspect deployment failure.
📋 KB: 9.0-era pattern (holodeck-02 Task 4 common errors)
vSAN health warnings in the nested cluster
Cause: Nested vSAN is timing-sensitive during initial resync; also expected-noise items differ per release
Fix: Re-test health after resync completes. Distinguish structural failures (missing disks, partition groups) from nested-lab noise before acting.
📋 KB: 9.0-era nested vSAN guidance carries over

Task 5 Audit the evidence chain and probe the GitOps loop for drift

manageability

This is where the GitOps promise is tested, not assumed. Repeatability, auditability, and drift-resistance are the three claims that justified moving off PowerShell — each gets a concrete probe here. The outputs are direct exhibits for a VCDX defense: 'show me your change record' has a two-click answer if this task is done honestly.

Step 1

Assemble the deployment evidence chain in your findings log as a linked sequence: config commit SHA -> pipeline run ID -> stage log excerpts -> health screenshots -> component version list. Verify every link actually cross-references (the run really points at that SHA, etc.).

A verifiable chain from 'what we said we wanted' to 'what is running', with no gaps.
Now perform the 9.0 thought experiment: which links would exist for a holodeck-02 cmdlet deployment? (Answer: roughly none — a transcript if you saved it, and an unversioned config.json.) Write that comparison down while it is vivid.
Step 2

Drift probe: make a small, deliberate out-of-band change to the running estate that contradicts the committed config (e.g. rename something cosmetic or alter a non-critical setting directly in vCenter — choose something reversible and harmless).

Estate now diverges from Git in one known, controlled way.
Keep the change trivially reversible and note it. You are testing whether the toolkit notices drift, not sabotaging your lab.
Step 3

Observe whether any part of the 9.1 toolkit detects or reports the divergence, and what a pipeline re-run would do about it (read the pipeline definition rather than re-running the full multi-hour build; re-run only if the pipeline supports a cheap validation-only mode).

A grounded answer to 'is this GitOps loop reconciling, or fire-and-forget?' — recorded with evidence.
OpenGitOps principle four is continuous reconciliation. Many 'GitOps' systems are actually versioned-trigger systems with no reconcile loop. Which the 9.1 toolkit is determines your drift-management design — do not guess, observe (verify_live).
Step 4

Revert your out-of-band change in the estate, and log the drift-probe result: detected/undetected, reconciled/ignored.

Estate back in conformance; drift-probe conclusion recorded.
Step 5

Repeatability assessment on paper: using your evidence chain, write the exact minimal sequence a colleague would follow to reproduce this estate from a bare 8.0 U3 host (holodeck91-01 staging -> holodeck91-02 OVA + commit -> single UI trigger). Note every point where tribal knowledge (your findings log) is still required.

A reproduction recipe with an honest list of remaining tribal-knowledge dependencies.
Each tribal-knowledge dependency is a gap between 'automated' and 'repeatable'. The shorter that list, the stronger your repeatability claim in front of a panel.
Step 6

Take the completion snapshots: snapshot the HoloRouter and the nested estate VMs as 'holodeck91-03-complete'. PowerCLI on the physical host: Get-VM -Name '<your-91-instance-prefix>*' | New-Snapshot -Name 'holodeck91-03-complete' -Description 'Site A VCF 9.1.0.0 deployed via GitOps pipeline, validated healthy' -Memory:$false

Snapshots present across the estate; skip memory snapshots to save space, as on the 9.0 track.
Adjust the VM name filter to your actual 9.1 instance naming (verify from vSphere inventory — the 9.1 naming convention may differ from the 9.0 'Holo-*' prefix).

Validation Gate

Check: Evidence chain assembled and cross-verified; drift probe executed, observed, reverted, and its conclusion logged; reproduction recipe written with tribal-knowledge gaps listed; 'holodeck91-03-complete' snapshots in place across router and estate.

Expected: GitOps claims tested against reality, evidence pack defense-ready, environment snapshotted for the labs that follow.

Common Errors

Evidence chain has a gap (e.g. cannot prove which commit a run consumed)
Cause: Trigger record from Task 2 incomplete, or the pipeline does not surface its consumed SHA
Fix: Recover what is recoverable from GitLab's run metadata now, and add the missing capture step to your reproduction recipe so the next run's chain is complete.
📋 KB: Process discipline, not a toolkit defect
Snapshot creation is slow or storage-heavy across the full estate
Cause: Full nested estate is many VMs with large disks
Fix: Snapshot without memory (as instructed), stagger the operations, and verify datastore headroom first. Consider snapshotting only router + management appliances if space-constrained, and record the reduced scope.
📋 KB: Carried-over 9.0 practice (holodeck-02 Task 6)

Final Validation

A complete Site A VCF 9.1.0.0 management domain is running, deployed entirely from the Click-to-Deploy UI by a GitOps pipeline with zero manual PowerShell. Health matches the 9.0-track bar (SDDC Manager Active, vCenter cluster green, NSX fabric up). The deployment's evidence chain — commit SHA to run ID to health screenshots — is assembled, the drift behavior of the toolkit has been probed and recorded, and the estate is snapshotted as 'holodeck91-03-complete'.

✓ Pipeline run: green end to end from a single UI trigger → No manual PowerShell anywhere in the run; interventions (if any) logged as findings

✓ SDDC Manager: Site A management domain Active with all nested hosts → Active domain, all hosts healthy

✓ vCenter: cluster with all hosts connected, vSAN green, HA/DRS enabled → All green, no alarms

✓ NSX Manager: all host transport nodes Success/Up → Fabric fully operational

✓ Evidence chain: commit SHA -> run ID -> logs -> health artifacts, cross-verified → No broken links; comparison against 9.0 evidence-availability written

✓ Drift probe executed and conclusion recorded (reconciling vs fire-and-forget) → Grounded, evidenced answer in the findings log

✓ Snapshots 'holodeck91-03-complete' across router and estate → Restore point for subsequent 9.1 labs

Cleanup / Restore

Snapshot: holodeck91-03-complete

• Snapshot router and estate as 'holodeck91-03-complete' (done in Task 5)

• Export or bookmark the pipeline run page and archive key job logs — GitLab log retention on the appliance is unverified (verify_live), so do not assume the run record persists indefinitely

• Update the findings log's measured-duration column and mark the deployment-path open question (GitOps vs cmdlets) as resolved-by-experience for your environment

• Close all bootstrap-admin sessions; note in the log that identity hardening (Authentik) is pending holodeck91-04

Design Reflection (VCDX)

The panel question this lab arms you for is: 'Your design says deployments are automated and auditable — prove it, and tell me what breaks.' You can now answer with a real evidence chain (commit -> run -> healthy estate) and a tested drift conclusion, instead of vendor language. Expect follow-ups on both flanks. Operational flank: what happens when the pipeline fails at stage four of seven — is partial state cleaned up, resumed, or orphaned, and how do you know? Governance flank: the UI trigger was attributed to a bootstrap admin — who SHOULD be able to click deploy, who approves the config commit, and where is separation of duties? A candidate who presents GitOps as 'better because newer' fails this exchange; a candidate who presents measured durations, a probed reconciliation model, and an honest tribal-knowledge gap list turns it into the strongest section of their defense.

Requirements

  • Full Site A VCF 9.1.0.0 management domain deployable from a single UI action against a reviewed, committed configuration
  • Deployment must produce a verifiable evidence chain linking configuration version to running state
  • Post-deployment health must meet the same bar as the 9.0 cmdlet track (SDDC Manager Active, vCenter cluster green, NSX fabric up) for cross-track comparability
  • Deployment process must be reproducible by another operator from documented steps alone

Constraints

  • Pipeline internals, stage semantics, and log retention are undocumented — operational knowledge must be built empirically and recorded
  • The GitOps control plane and the deployment target share one physical host; pipeline resource demand competes with the estate it is building
  • Multi-hour unattended run requires a change freeze on the config repo (commit-consumption semantics verified in Task 1, not assumed)
  • No 9.1 community KB exists; all troubleshooting maps 9.0-era failure classes onto new mechanisms
  • Nested-lab performance figures (durations, resource profiles) do not transfer to production sizing — they inform lab operations only

Assumptions

  • The committed Site A configuration from holodeck91-02 is complete and internally consistent (enforced partly by 9.1's hardened validation)
  • Technitium DNS registers and serves estate records during bring-up without the manual forwarder intervention the 9.0 track required — verified live in Task 3
  • Pipeline failure leaves the control plane (GitLab/Git history) intact, so diagnosis and re-run are always possible from the router
  • The physical host remains dedicated to this instance for the duration of the run

Risks

  • Mid-pipeline failure with partial estate state — IMPACT: orphaned nested VMs consuming resources and blocking clean re-runs; MITIGATION: 'holodeck91-02-complete' restore point, documented cleanup-before-rerun check, idempotency observations from Task 3 common errors
  • Green-but-wrong: pipeline stage success criteria weaker than actual provisioning success — IMPACT: false confidence, late discovery; MITIGATION: three-plane monitoring with vSphere as ground truth, end-to-end health validation in Task 4
  • Audit trail rot: GitLab run logs on the appliance may not persist — IMPACT: evidence chain breaks retroactively; MITIGATION: export/archive run records at completion (cleanup action)
  • Fire-and-forget loop: if the toolkit does not reconcile drift, the Git repo silently stops describing reality — IMPACT: the audit trail becomes actively misleading over time; MITIGATION: drift probe conclusion drives either periodic re-validation runs or explicit 'point-in-time only' labeling of the config repo
  • Bootstrap-identity triggering: all actions attributed to a shared root account — IMPACT: attribution without accountability; MITIGATION: Authentik SSO integration (holodeck91-04) before treating the trail as audit-grade

Self-Assessment Discussion Prompts

  1. Compare your measured 9.1 pipeline deployment against the holodeck-02 cmdlet deployment across: operator time invested, wall-clock time, evidence produced, and skill floor required. Where does each path win?
  2. Repeatability: the cmdlet path re-run depends on the operator re-executing steps identically; the pipeline re-run depends on the same commit producing the same result. Which failure modes does each dependence create, and which did you actually observe?
  3. Auditability: your evidence chain links commit to running estate. A panelist asks 'could an operator have made the estate diverge from that commit without leaving a trace?' Answer honestly from your drift-probe result, and design the control that closes whatever gap you found.
  4. Drift: is the 9.1 toolkit a reconciling GitOps system or a versioned trigger system, per your Task 5 observation? What operational practice must you add in the second case to keep Git truthful?
  5. The pipeline failed at the nested-ESX stage for a colleague (datastore misnamed in the committed config). Walk through the recovery: what gets fixed where, what gets cleaned up, and what does the Git history look like afterwards if done properly?
  6. Your reproduction recipe still contains tribal-knowledge items from the findings log. For each, decide: automate it, document it into the repo, or accept it — and justify the split as you would in a design defense.
  7. In production VCF, would you accept a deployment system whose control plane (GitLab) lives inside the failure domain it deploys (on the router of the estate)? Map this to production analogs and state your placement decision.

Extensions

Teardown-and-Redeploy Repeatability Proof

Destroy the nested estate (retaining the router and Git history), re-trigger the identical commit, and diff the resulting estate against your Task 4 validation artifacts. Measure whether the second run is faster, identical, and intervention-free. This converts the repeatability claim from argument to experiment.

same

Pipeline Failure Injection

Deliberately commit a plausible-but-wrong configuration (bad datastore name, undersized host) and observe where the hardened validation catches it versus where it slips through to a late stage failure. Document the failure surface map and the cleanup burden of each late failure. Mirrors the 9.0 track's failure-injection extension under the new driver.

harder

Site B Dual-Site Activation

Activate the parked Site B configuration and drive a second-site deployment from the same control plane, managing IP/VLAN separation from Site A. Exercises 9.1's dual-site auto-generation end to end and sets up DR/stretched design discussion. On a shared host, coordinate first — this doubles the footprint.

much harder

Day-2 Operations from the Toolkit

Use the 9.1 Day-2 capabilities announced for the toolkit — add a nested ESX host to the estate, then deploy VCF Automation or spin up a Supervisor in VNA/Edge mode — and observe whether Day-2 actions flow through the same Git/pipeline loop or bypass it. The answer determines whether the audit chain survives past day one.

harder

⚠ Known Pitfalls (from Community KB)

[9.0-era, re-validate on 9.1] Holodeck vcf 9 - Cannot select datastore RESOLVED (9.0)
Problem: Nested ESX provisioning failed on datastore capacity/naming — the most-discussed 9.0 KB thread. The failure class survives into the 9.1 pipeline's provisioning stage; only the driver changed.
Resolution: Under 9.1: verify datastore name and >1.5 TB free before triggering; remediate via a config commit and pipeline re-run, cleaning up partial VMs first.
[9.0-era, re-validate on 9.1] VCF Installer is not ready yet LIKELY_RESOLVED (9.0)
Problem: Bring-up hung on installer initialization, historically DNS/NTP-rooted. Technitium DNS in 9.1 should structurally improve the DNS root cause — unverified until observed.
Resolution: Check estate FQDN resolution against Technitium and installer time sync. Log the outcome as a 9.1 data point either way.
[9.0-era, re-validate on 9.1] Issue with legacy CPU support (boot.cfg) RESOLVED (9.0)
Problem: Older physical CPUs required allowLegacyCPU=TRUE in nested ESX boot.cfg. Whether and where the 9.1 pipeline exposes this knob is unknown.
Resolution: If provisioning fails on CPU compatibility, search the committed config schema and pipeline variables for a legacy-CPU option before hand-patching anything (verify_live).
[9.1 open question] Pipeline idempotency and partial-state cleanup OPEN
Problem: Behavior of a pipeline re-run over a partially built estate (orphan cleanup, resume, or collision) is undocumented.
Resolution: Prefer reverting to 'holodeck91-02-complete' and re-running clean until re-run semantics are observed and recorded.
[9.1 open question] Pipeline/run log retention on the appliance OPEN
Problem: Evidence chain depends on GitLab run records whose retention configuration on the embedded appliance is unverified.
Resolution: Archive run logs and metadata at completion (cleanup action) rather than trusting retention defaults.

References

  • HoloDeck 9.1 Release Announcement (verbatim capture) Holodeck-9.1/01-release-announcement.mdlocal fileTier 1 — Official
    Primary source: 'trigger full Holodeck deployments directly from the UI — no manual PowerShell required'; VCF 9.1.0.0 support; hardened VCF installer validation.
  • HoloDeck Toolkit 9.1 — Structured Reference & Changelog Holodeck-9.1/02-structured-changelog.mdlocal fileTier 1 — Official
    9.0-to-9.1 orientation table (deployment driver: PowerShell -> GitOps/UI) and open-questions list tracked by this lab's findings log.
  • VCF 9.1 Release Intake — rn-9-1-016 (Click-to-Deploy GitOps) academy-v2/content/releases/vcf-9.1.jsonlocal fileTier 1 — Official
    Confirms both paths must be presented in lab guidance: GitOps/UI (this lab) and cmdlets (holodeck-02 remains the 9.0.x record).
  • VMware Cloud Foundation Documentation (Broadcom)Tier 1 — Official
    Authoritative VCF documentation portal — consult the 9.1 deployment and bring-up guides for component-level detail as they publish.
  • VCF Holodeck Toolkit — Broadcom Community ForumTier 1 — Official
    All captured KB threads are 9.0-era. Deployment-phase failure classes (datastore, legacy CPU, installer readiness) are referenced here as carried-over hypotheses, not confirmed 9.1 behavior.
  • GitLab CI/CD Pipelines DocumentationTier 2 — VMware Press
    For reading pipeline graphs, job logs, and run metadata used throughout Tasks 1-3 and the evidence chain in Task 5.
  • OpenGitOps PrinciplesTier 2 — VMware Press
    Framework for the Task 5 drift probe: distinguishing a reconciling GitOps system from a versioned trigger system.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.