vLCM Cluster Image Build & Staged Upgrade (VCF 9.0)
Objectives
- Design and build a vLCM cluster image incorporating ESXi, firmware, and vendor-specific VIBs within VCF 9.0's Operations Depot paradigm
- Execute a staged, rolling host remediation workflow that respects DRS constraints and business availability windows
- Detect and remediate infrastructure drift (non-compliant hosts) using vLCM compliance baselines
- Compare VCF 9.0's Operations Depot-driven upgrade strategy against VCF 5.2.x's SDDC Manager image lifecycle
- Articulate the architectural implications of lifecycle management as a distributed platform service in a multi-domain VCF stack
Prerequisites
VCF 9.0 management and workload domains fully deployed and healthy. One workload domain cluster (4-6 ESXi hosts running 8.0 Update 3) accessible via SDDC Manager and vCenter. Hosts are compliant with baseline image as of deployment. Internet connectivity or offline depot configured. At least 100 GB free space on datastores for image staging.
Prior labs: holodeck-02
Required skills:
- vCenter and SDDC Manager UI navigation
- ESXi host lifecycle operations (power, remediation, snapshot)
- Understanding of vSphere DRS, vMotion, and resource reservations
- Basic API navigation (vLCM Cluster Image Spec)
- Knowledge of vendor-specific VIBs (NIC drivers, HBA firmware, GPU firmware)
Lab Environment
VCF 9.0 management domain (1 SDDC Manager, 1 vCenter) + 1 workload domain with 4-6 ESXi hosts (8.0 U3) in a single cluster. All hosts are identical hardware (same NICs, HBAs, firmware revisions). Workload cluster includes: DRS enabled (balances guest workloads), HA enabled (admission control configured), vSAN disabled (to simplify drift scenarios). At least 3-4 test VMs per host to simulate production workload distribution.
graph TB VC[vCenter 10.0.0.6] --> EsxiCluster[Workload Cluster<br/>4-6 ESXi 8.0 U3 hosts] SDDC[SDDC Manager 10.0.0.4] --> OpDepot[Operations Depot] OpDepot --> vLCM[vLCM Cluster Image] vLCM --> EsxiCluster VC --> VM1[VM workload-01] VC --> VM2[VM workload-02] VC --> VM3[VM workload-03] EsxiCluster --> Host1[ESXi-1] EsxiCluster --> Host2[ESXi-2] EsxiCluster --> Host3[ESXi-3] EsxiCluster --> Host4[ESXi-4]
Credentials
| System | Username | Password |
|---|---|---|
| SDDC Manager | administrator@vsphere.local | From VCF deployment |
| vCenter Server | administrator@vsphere.local | Same as SDDC Manager |
| ESXi Hosts | root | Set during host commissioning |
Tasks
Task 1 Assess baseline cluster image compliance and gather upgrade requirements
ManageabilityIn production, you never design an upgrade in isolation. You must inventory the current state, validate hardware heterogeneity, identify vendor-specific addons, and establish a compatibility matrix. This mirrors a VCDX design consultation where you'd spend 30-40% of your time on discovery before recommending an upgrade path.
Navigate to SDDC Manager (https://sddc-manager.vcf.sddc.lab). Go to Infrastructure > Workload Domains. Select your workload domain. Click 'Lifecycle' tab.
In the Lifecycle view, note the Current Cluster Image ID (format: example-image-v1). Click on it to view the image composition: ESXi version, firmware versions (BIOS, BMC, NIC, HBA), and VIBs included.
For each ESXi host in the cluster, verify hardware uniformity. In vCenter, go to Hosts and Clusters > [Cluster] > Hosts. For each host, click its name and navigate to Hardware. Verify CPU model, RAM quantity, NIC types (name), and HBA types match across all hosts.
SSH to one ESXi host and gather detailed VIB inventory. Run: esxcli software vib list | grep -E 'vib|vendorvendor|version' | head -20
Review the upgrade requirements document. In SDDC Manager, go to Cluster Image > Available Images. Identify the target ESXi version (8.0 U4) and its firmware/VIB composition. Compare to current state: list 3 major changes (e.g., NIC driver version bump, BIOS firmware update, new GPU VIB).
Document cluster upgrade constraints. In vCenter, check: (a) DRS enabled? (b) HA enabled + admission control setting (conservative, moderate, disabled)? (c) Any powered-on VMs with memory reservations > 10 GB? (d) vMotion network latency (ping between hosts < 5 ms expected).
Validation Gate
Check: Document captured: (1) baseline cluster image ID + composition, (2) hardware inventory for all hosts, (3) 3+ specific VIB/firmware changes in target image, (4) cluster constraints (DRS/HA settings, VM reservations, network RTT)
Expected: Pre-upgrade assessment complete. You have a clear change control document ready for a production change advisory board.
Common Errors
Task 2 Build and stage the target vLCM cluster image
AvailabilityvLCM cluster image creation is the core operational task that requires deep understanding of the image specification: not just ESXi version, but the dependency graph of VIBs, firmware, and OEM addons. In VCDX, panelists will ask: 'What happens if you include two conflicting VIBs?' and 'How would you validate an image before rolling it to 4,000 hosts?' This task builds that expertise.
In SDDC Manager, go to Infrastructure > Cluster Images. Click 'Create Cluster Image'.
Select ESXi 8.0 Update 4 as the base. SDDC Manager will show 'Recommended Components' for that ESXi version. Review the list: ESXi 8.0 U4 + default Broadcom NIC drivers + default Broadcom HBA drivers. Do NOT accept defaults yet.
In the wizard, click 'Customize Components'. For the NIC drivers, select the version corresponding to your current hardware. Example: if baseline had bnx2x 8.45.13, check if 8.46.21 is available. If a major version bump (8.45 -> 8.46), research release notes for breaking changes (e.g., https://compatibilityguide.broadcom.com/). Document any breaking changes.
Click 'Add VIB' to include any vendor-specific addons. Example: if you're using NVIDIA GPUs, add the nvidia-vgpu VIB. If using Broadcom BladeLogic or Dell iDRAC, add those management VIBs. Select the latest stable version available in the Operations Depot.
Review the final image composition. The image should show: (a) ESXi 8.0 U4 base, (b) minimum 2-3 vendor VIBs (NIC, HBA, optional GPU/management), (c) firmware updates (BIOS, BMC). Click 'Save Image as Draft'.
Now stage the image: with the draft image selected, click 'Stage Image to Cluster'. Select your workload domain cluster. This operation copies the image payload to the cluster's datastores and validates it against cluster compatibility. Monitor the staging progress in the Activity pane.
Validation Gate
Check: In SDDC Manager Cluster Images, the target image is listed with status 'Staged' and a green checkmark. Image composition includes ESXi 8.0 U4 + 3+ VIBs.
Expected: Target image built and staged. Cluster is ready for remediation.
Common Errors
Task 3 Execute a rolling host remediation with DRS-aware sequencing
AvailabilityThis is where theory meets practice. You will orchestrate the remediation of 4-6 hosts one at a time, respecting DRS constraints (vMotion workloads before entering remediation mode). In VCDX, you'll be asked: 'If you have a 6-host cluster with 10 VMs each and a 2-hour upgrade window, how do you sequence this to minimize guest downtime?' This task answers that question empirically.
In vCenter, navigate to your workload cluster. Right-click the cluster and select 'Expand' or click each host individually. For each host, note the current VM count. Create a remediation order starting with the host that has the fewest powered-on VMs (or manually vMotioned VMs to other hosts first). Target order: Host 1 (least VMs) -> Host 2 -> Host 3 -> Host 4.
Begin remediation of Host 1. In SDDC Manager, go to Infrastructure > Cluster Images. Select your staged image. Click 'Remediate Cluster' (or 'Remediate Host' if granular control is available — check SDDC Manager UI options).
In the remediation wizard, SDDC Manager will show the remediation workflow: (a) set host to maintenance mode, (b) vMotion all VMs off the host (DRS handles this), (c) upgrade ESXi via vLCM, (d) reboot host, (e) exit maintenance mode. Confirm the workflow sequence and start remediation.
While Host 1 is remediating, monitor in parallel: (a) In vCenter, watch the vMotion network traffic in the cluster. (b) Check DRS recommendations (home icon in cluster summary) — are migrations being scheduled? (c) Monitor SDDC Manager activity for remediation sub-task progress (firmware updates, VIB installations, etc.).
Host 1 remediation should complete in 15-25 minutes (ESXi reboot + VIB application time). Once Host 1 shows 'Compliant' status and is back online in vCenter, begin remediation of Host 2 using the same sequence.
After all hosts are remediated, verify cluster-wide compliance. In SDDC Manager Cluster Images, the cluster should show status 'All Hosts Compliant'. In vCenter, verify all hosts are online, all VMs are running, and DRS load balancing has normalized VM distribution across the cluster.
Validation Gate
Check: (1) SDDC Manager shows cluster image compliance = 100% (all hosts compliant), (2) vCenter cluster shows all hosts connected with no alarms, (3) all powered-on VMs are in Running state, (4) documented remediation duration per host (target: 15-25 min/host)
Expected: Cluster successfully upgraded from ESXi 8.0 U3 to 8.0 U4 with zero guest downtime (due to vMotion during maintenance mode).
Common Errors
Task 4 Detect and remediate infrastructure drift with vLCM compliance baselines
SecurityPost-upgrade, your cluster is compliant. But in production, hosts drift: a system administrator manually installs a NIC driver update, firmware is patched out-of-band, or a third-party tool drops a management VIB. vLCM's drift detection is your enforcement mechanism. This task simulates drift and shows you how to find it, report it, and fix it — critical for a VCDX architect designing governance.
Simulate drift on Host 2: SSH to that host. Check the current NIC driver version: esxcli software vib list | grep bnx2x (or equivalent for your NIC). Note the version, e.g., 8.46.21.
Manually 'downgrade' the driver to simulate an out-of-band patch. Run: esxcli software vib remove -n bnx2x (this removes the VIB). Then install the older version from a staging location (e.g., a USB stick or pre-cached bundle). In a real lab, you'd have older VIB bundles staged. For this simulation, manually edit the VIB version metadata (in /opt/vmware/ or similar) to show an older version, or just remove the VIB to trigger a 'missing component' drift.
In SDDC Manager, navigate to Infrastructure > Cluster Images. Select your staged cluster image. Click 'Compliance Check' or 'Drift Detection'. Select Host 2 (the host you just drifted).
Review the drift report. SDDC Manager should show: (a) the baseline cluster image specification, (b) the host's actual state, (c) differences highlighted. Document 2-3 specific mismatches.
Remediate the drift: in SDDC Manager, select Host 2 and click 'Remediate Host' or 'Force Compliance'. This will re-apply the correct VIB version. Monitor the remediation in vCenter (host enters maintenance mode, VMs migrate).
After remediation, verify Host 2 is back to compliant state. SSH to the host and re-run: esxcli software vib list | grep bnx2x. Confirm the version is now back to 8.46.21.
Validation Gate
Check: (1) Drift detected on Host 2 via SDDC Manager compliance check, (2) drift report correctly identified missing/wrong VIB versions, (3) Host 2 remediated and returned to compliant state, (4) SSH verification confirms correct version post-remediation
Expected: Drift detection and auto-remediation workflow demonstrated. This is a production-ready governance pattern.
Common Errors
Task 5 Compare VCF 9.0 vLCM workflow with VCF 5.2.x and document operational complexity reduction
ManageabilityThis is the VCDX perspective: you're not just executing steps, you're comparing two design paradigms. VCF 5.2.x had SDDC Manager managing lifecycle but with manual vLCM tasks. VCF 9.0 unified lifecycle under Operations Depot + vLCM Cluster Image APIs. Understanding this evolution—what was manual, what's now automated, where complexity shifted—is the difference between a VCP and a VCDX.
Document the VCF 9.0 workflow you just executed: (1) Task 1: cluster assessment (manual, ~30 min), (2) Task 2: image build + stage (SDDC Manager API, ~20 min), (3) Task 3: rolling remediation (SDDC Manager orchestrated, ~90 min for 4 hosts), (4) Task 4: drift detection (automated, ~10 min). Total operator time: ~2 hours. Total wall-clock time: ~2.5 hours (overlap during staging/remediation).
Now review the VCF 5.2.x approach (consult the VCF 5.2 Planning & Preparation guide or community forum threads). In 5.2, the workflow is: (a) SDDC Manager > Lifecycle > Cluster Images (same API, similar UX), (b) Remediation is SDDC Manager-driven but requires operator to pre-stage vMotion/maintenance mode logic manually (not automated), (c) Drift detection is via vLCM baseline comparison, but remediation is not as tightly integrated into SDDC Manager. Document 3 manual steps in 5.2 that are now automated in 9.0.
Quantify operator effort: VCF 9.0 required ~2 hours of hands-on time (largely monitoring). VCF 5.2 would require ~3-4 hours (more manual vMotion orchestration, separate remediation validation, manual drift remediation sequencing). Document the hours saved per cluster upgrade (Delta: ~1-2 hours) and extrapolate to a multi-cluster environment (e.g., 10 workload clusters = 10-20 hours saved per upgrade cycle).
Identify where complexity shifted. In 9.0, complexity shifted LEFT (earlier): image design and composition (vib selection, vendor compatibility validation) happens up-front. In 5.2, you'd discover issues during remediation. Document this: 'VCF 9.0 moves image validation left, reducing remediation failures by ~40%, but requires deeper pre-flight analysis.'
For a VCDX panelist, articulate: 'What are the architectural trade-offs?' Prompt: If you had 50 workload domains with mixed hardware, would you prefer VCF 5.2's per-cluster flexibility (custom images) or VCF 9.0's centralized Operations Depot (consistency, but less flexibility)?' Document your answer and reasoning (2-3 paragraphs).
Finally, measure upgrade time empirically. Document: (1) Time from 'Start Remediation Host 1' to 'Host 1 Back Online', (2) Average time per host, (3) Total cluster remediation time. In a real production run, this data informs your change advisory board's risk assessment ('If we have a 4-hour maintenance window and 6 hosts, can we complete the upgrade? What's our rollback decision point?').
Validation Gate
Check: (1) VCF 9.0 workflow timeline documented (all 4 tasks + wall-clock time), (2) VCF 5.2 workflow comparison with 3+ specific differences, (3) Operator time delta calculated for 1 cluster + extrapolated to 10 clusters, (4) Complexity shift analysis (left-shifting image validation), (5) VCDX-level trade-off analysis (consistency vs flexibility), (6) Empirically measured host remediation times captured
Expected: You can now articulate to a VCDX panelist: 'VCF 9.0's Operations Depot + vLCM integration reduced lifecycle operator overhead by ~30-40% vs 5.2, but requires stronger upfront image design governance. Here are the measured timelines and the architectural trade-offs.'
Common Errors
Final Validation
You have successfully designed, built, staged, and executed a rolling ESXi cluster upgrade using VCF 9.0's vLCM cluster image paradigm. You've detected and remediated infrastructure drift, and you've articulated the architectural implications vs VCF 5.2. All hosts are compliant, the cluster is healthy, and you have measured data to inform production change control processes.
✓ Baseline cluster image assessment: documented current state, target state, and change list → 3+ specific VIB/firmware changes identified
✓ Target cluster image built and staged: ESXi 8.0 U4 + 3+ vendor VIBs, staged to cluster datastore → Image status = Staged, ready for remediation
✓ All 4-6 hosts remediated successfully: zero guest downtime (vMotion during maintenance mode) → All hosts online, compliant, running their original VM workload (or redistributed by DRS)
✓ Drift detection: manually introduced drift on 1 host, detected via vLCM compliance scan, remediated back to compliant → Host showed non-compliant, then compliant post-remediation
✓ Workflow comparison: VCF 9.0 vs 5.2 analyzed; operator time savings quantified; architectural trade-offs articulated → Documented timeline comparison + 2-3 paragraph VCDX-level analysis
Cleanup / Restore
Snapshot: workload-domain-post-vlcm-upgrade
• Take a snapshot of the workload domain cluster post-upgrade (all hosts compliant, all VMs running): snapshot all cluster VMs as 'workload-domain-post-vlcm-upgrade'
• Document the final cluster image ID and composition for reference in subsequent labs
• If you need to revert for future labs, restore to the 'workload-domain-baseline' snapshot (pre-upgrade state)
Design Reflection (VCDX)
A VCDX panelist examining a large-scale VCF lifecycle management design would probe: How do you handle image composition governance across 50+ workload domains with mixed hardware? What's your rollback strategy if drift remediation fails on 20% of hosts? How does vLCM integrate into your change management process, and who approves image changes? How would you design automated compliance checks that alert on drift within 24 hours? The vLCM cluster image is not just a technical feature—it's a policy enforcement layer. Be ready to argue whether centralized (one image per VCF instance) or distributed (one image per workload domain) image governance scales better, and defend the trade-offs with operational data from this lab.
Requirements
- Cluster image must incorporate ESXi base OS, vendor-specific drivers (NIC, HBA), firmware (BIOS, BMC), and optional VIBs (GPU, management agents)
- Remediation must be zero-downtime (vMotion-based) and respect DRS constraints; no guest VM downtime allowed
- Drift detection must identify non-compliant hosts within 1 hour of deviation and alert operators
- Lifecycle operations must be auditable: log every image change, remediation, and drift detection event to a centralized audit trail
Constraints
- vLCM images are cluster-scoped: a single image cannot span multiple hardware revisions or ESXi versions. Heterogeneous clusters require multiple images or host retirement.
- Remediation bandwidth is limited by vMotion network capacity and storage I/O (snapshot creation during maintenance mode). Typical throughput: 1 host per 15-30 minutes.
- Vendor VIB availability depends on Operations Depot sync status. Missing VIBs delay image staging by hours or days if depot is offline.
- DRS constraints (VM memory reservations, resource affinity rules) can block vMotion, causing remediation stalls. Operator must pre-validate cluster has enough free capacity.
Assumptions
- All hosts in a remediation cluster have identical hardware (CPU, NIC, HBA, firmware UEFI level)
- vMotion network is available and has > 1 Gbps capacity for concurrent VM migrations
- Datastore has sufficient free space: image size + (host count * 50 GB snapshot overhead) + 100 GB buffer
- No VMs have 'must stay on this host' affinity rules that would block remediation
- Operations Depot is reachable and syncs daily with upstream Broadcom depot
Risks
- Image composition error (conflicting VIB versions): if two VIBs depend on incompatible kernel ABIs, remediation fails mid-flight and host rolls back, causing remediation stall. MITIGATION: test image on 1-2 hosts before full-scale rollout; maintain rollback image.
- vMotion stall during maintenance mode: if remaining cluster capacity cannot absorb all VMs from remediating host, vMotion never completes and remediation times out. IMPACT: host stuck in maintenance mode, unavailable to guest workload. MITIGATION: pre-validate DRS has enough free capacity; migrate non-critical VMs off before starting.
- Datastore fills during multi-host remediation: concurrent snapshots for maintenance mode can exhaust datastore capacity, causing failures and data loss. IMPACT: cluster partially degraded, recovery requires manual snapshot deletion + rollback. MITIGATION: monitor datastore free space hourly during remediation cycle; pre-allocate headroom.
- Drift remediation cascades: if drift remediation fails on host A (say, VIB install error), the host stays non-compliant. If unchecked, subsequent hosts may inherit the bad image during re-staging. IMPACT: cluster-wide compliance failure. MITIGATION: image staging validates each host's compatibility independently; failed remediation on one host does not affect others, but requires manual investigation.
- Operator error: manually removing a critical VIB (as in Task 4 drift simulation) is easy and undetectable until drift scan runs hours later. In production, this could cause silent compliance gaps on 10% of hosts. MITIGATION: enforce vib removal/install via vLCM only, disable direct esxcli vib commands on production hosts.
Self-Assessment Discussion Prompts
- If you have 50 workload domains, each with slightly different hardware (vendor A's NICs in 20 domains, vendor B's in 30), how would you design your vLCM image strategy? One image per vendor? One per domain? What are the governance implications?
- Your change control process requires 2-week approval before any ESXi upgrade. But a critical NIC driver CVE is patched, and you need to update the cluster image within 48 hours. How do you fast-track the image change while maintaining governance?
- A host fails remediation mid-upgrade (VIB installation error, host panics during reboot). What's your rollback decision logic? Do you revert the host from snapshot, or do you re-remediate with a different image? What data do you need to decide?
- VCF Operations Depot is your single source of truth for VIB versions. If the Depot becomes unavailable (network outage, Broadcom service disruption), can you still remediate hosts? What's your contingency?
- After upgrading 4 hosts successfully, you discover a kernel bug in ESXi 8.0 U4 that affects your specific NIC (bnx2x 8.46.21). Hosts are crashing under load. What's your rollback strategy? Can you roll back 4 hosts without affecting the 2 remaining hosts on 8.0 U3?
- You have 1000 hosts across 20 clusters. You want to upgrade to ESXi 8.0 U4. Each cluster takes 8 hours to remediate. You have a 4-day maintenance window. How many clusters can you upgrade in parallel? What are the bottlenecks (vMotion bandwidth, datastore I/O, SDDC Manager API rate limits)?
Extensions
Build and test a heterogeneous cluster image strategy for mixed hardware
Repeat Tasks 2-3, but this time create TWO cluster images: one optimized for hosts with Intel NICs (i40en), one for Broadcom (bnx2x). Deploy the lab cluster with mixed hardware (some hosts have Intel, some Broadcom). Demonstrate that vLCM can selectively apply different images to different hosts based on hardware detection. Document the operational complexity of managing multiple images in production (versioning, testing, approval workflows).
harderSimulate a vLCM remediation failure and execute a rollback recovery
Deliberately corrupt the staged image on one datastore (e.g., truncate the ESXi ISO) to cause a remediation failure on Host 3. Document the failure mode, the error logs, and the recovery procedure. How long does rollback take? Is the host recoverable, or must it be re-imaged? What does this teach you about backup image strategies and rollback decision thresholds?
harderDesign a multi-cluster upgrade orchestration workflow
You have 5 workload clusters (each with 6 hosts). Design a vLCM upgrade strategy that maximizes parallelization while respecting business constraints: (a) max 1 cluster can be in remediation at a time (to avoid vMotion contention), (b) each cluster must complete within 12 hours, (c) you want minimum guest downtime. Build a timeline and sequencing plan. Validate it against vMotion bandwidth, datastore I/O, and SDDC Manager API concurrency limits. This is a production-scale orchestration problem.
harderCompare Operations Depot image distribution with offline vLCM bundle deployment
In Task 2, you used SDDC Manager's Operations Depot (online, synced from Broadcom). Simulate an air-gapped environment where you download vLCM bundles offline and stage them manually. Create a cluster image using the offline bundle. Compare: (a) time to build image (online vs offline), (b) validation strictness (does offline mode have fewer checks?), (c) compliance gap (if an offline bundle becomes outdated, how do you detect it?). This is relevant for financial/government VCF deployments where Broadcom connectivity is restricted.
sameReferences
- VMware Cloud Foundation 9.0 Lifecycle Management with vLCM Cluster ImagesTier 1 — Official
Official Broadcom documentation — authoritative reference for cluster image composition, staging, and remediation workflows - VCF 9.0 Operations Depot and vLCM IntegrationTier 1 — Official
Explains how Operations Depot syncs with vLCM and how vendor VIBs are versioned and made available - ESXi Lifecycle Manager API Reference (vLCM Cluster Image Spec)Tier 1 — Official
API documentation for programmatically creating, staging, and managing cluster images - William Lam — vLCM Cluster Image Deep Dive and Drift DetectionTier 3 — Expert Blog
Expert blog covering vLCM best practices, performance tuning, and troubleshooting. Includes hands-on drift remediation scenarios. - Cormac Hogan — VCF Lifecycle and Operational GovernanceTier 3 — Expert Blog
Architectural perspective on integrating vLCM into change control and governance processes at scale - Broadcom Community Forum — vLCM and Cluster ImagesTier 1 — Official
Broadcom-hosted forum with 40+ threads on vLCM image failures, drift detection, and multi-cluster orchestration