Academy/VCP-VCF 9.0 Administrator (2V0-17.25)/vLCM Cluster Image Build & Staged Upgrade (VCF 9.0)
This lab targets VCF 9.0

vLCM Cluster Image Build & Staged Upgrade (VCF 9.0)

VCF 9.0Advancedadminarchitectvcdx⏱ 180 min

Core vLCM workflow applies to VCF 9.0.x. VCF 5.2.x uses older SDDC Manager-driven lifecycle with manual vLCM — see Extension 1 for 5.2 comparison.

Objectives

  • Design and build a vLCM cluster image incorporating ESXi, firmware, and vendor-specific VIBs within VCF 9.0's Operations Depot paradigm
  • Execute a staged, rolling host remediation workflow that respects DRS constraints and business availability windows
  • Detect and remediate infrastructure drift (non-compliant hosts) using vLCM compliance baselines
  • Compare VCF 9.0's Operations Depot-driven upgrade strategy against VCF 5.2.x's SDDC Manager image lifecycle
  • Articulate the architectural implications of lifecycle management as a distributed platform service in a multi-domain VCF stack

Prerequisites

VCF 9.0 management and workload domains fully deployed and healthy. One workload domain cluster (4-6 ESXi hosts running 8.0 Update 3) accessible via SDDC Manager and vCenter. Hosts are compliant with baseline image as of deployment. Internet connectivity or offline depot configured. At least 100 GB free space on datastores for image staging.

Prior labs: holodeck-02

Required skills:

  • vCenter and SDDC Manager UI navigation
  • ESXi host lifecycle operations (power, remediation, snapshot)
  • Understanding of vSphere DRS, vMotion, and resource reservations
  • Basic API navigation (vLCM Cluster Image Spec)
  • Knowledge of vendor-specific VIBs (NIC drivers, HBA firmware, GPU firmware)

Lab Environment

VCF 9.0 management domain (1 SDDC Manager, 1 vCenter) + 1 workload domain with 4-6 ESXi hosts (8.0 U3) in a single cluster. All hosts are identical hardware (same NICs, HBAs, firmware revisions). Workload cluster includes: DRS enabled (balances guest workloads), HA enabled (admission control configured), vSAN disabled (to simplify drift scenarios). At least 3-4 test VMs per host to simulate production workload distribution.

graph TB
  VC[vCenter 10.0.0.6] --> EsxiCluster[Workload Cluster<br/>4-6 ESXi 8.0 U3 hosts]
  SDDC[SDDC Manager 10.0.0.4] --> OpDepot[Operations Depot]
  OpDepot --> vLCM[vLCM Cluster Image]
  vLCM --> EsxiCluster
  VC --> VM1[VM workload-01]
  VC --> VM2[VM workload-02]
  VC --> VM3[VM workload-03]
  EsxiCluster --> Host1[ESXi-1]
  EsxiCluster --> Host2[ESXi-2]
  EsxiCluster --> Host3[ESXi-3]
  EsxiCluster --> Host4[ESXi-4]

Credentials

SystemUsernamePassword
SDDC Manageradministrator@vsphere.localFrom VCF deployment
vCenter Serveradministrator@vsphere.localSame as SDDC Manager
ESXi HostsrootSet during host commissioning

Tasks

Task 1 Assess baseline cluster image compliance and gather upgrade requirements

Manageability

In production, you never design an upgrade in isolation. You must inventory the current state, validate hardware heterogeneity, identify vendor-specific addons, and establish a compatibility matrix. This mirrors a VCDX design consultation where you'd spend 30-40% of your time on discovery before recommending an upgrade path.

Step 1

Navigate to SDDC Manager (https://sddc-manager.vcf.sddc.lab). Go to Infrastructure > Workload Domains. Select your workload domain. Click 'Lifecycle' tab.

Lifecycle dashboard showing current ESXi version (8.0 Update 3) and cluster image compliance status
Step 2

In the Lifecycle view, note the Current Cluster Image ID (format: example-image-v1). Click on it to view the image composition: ESXi version, firmware versions (BIOS, BMC, NIC, HBA), and VIBs included.

Image details page showing ESXi 8.0 U3 + vendor-specific firmware/VIBs (note their versions for later comparison)
Screenshot this page — it's your baseline. In production, you'd attach this to your change management record.
Step 3

For each ESXi host in the cluster, verify hardware uniformity. In vCenter, go to Hosts and Clusters > [Cluster] > Hosts. For each host, click its name and navigate to Hardware. Verify CPU model, RAM quantity, NIC types (name), and HBA types match across all hosts.

All hosts show identical hardware specifications (e.g., 2x Intel Xeon, 384 GB RAM, 2x 10 Gbe NICs, 1x FC HBA)
Step 4

SSH to one ESXi host and gather detailed VIB inventory. Run: esxcli software vib list | grep -E 'vib|vendorvendor|version' | head -20

List of critical VIBs including NIC drivers (e.g., bnx2x, i40en), HBA drivers (e.g., lpfc), and Broadcom vendors
Step 5

Review the upgrade requirements document. In SDDC Manager, go to Cluster Image > Available Images. Identify the target ESXi version (8.0 U4) and its firmware/VIB composition. Compare to current state: list 3 major changes (e.g., NIC driver version bump, BIOS firmware update, new GPU VIB).

Documented list of 3+ changes between current and target image (example: bnx2x driver 8.45.13 -> 8.46.21, BIOS 12.5 -> 12.7, new nvidia-vgpu 16.1)
Step 6

Document cluster upgrade constraints. In vCenter, check: (a) DRS enabled? (b) HA enabled + admission control setting (conservative, moderate, disabled)? (c) Any powered-on VMs with memory reservations > 10 GB? (d) vMotion network latency (ping between hosts < 5 ms expected).

Constraint document listing DRS state, HA admission control, VM reservation totals, and network readiness

Validation Gate

Check: Document captured: (1) baseline cluster image ID + composition, (2) hardware inventory for all hosts, (3) 3+ specific VIB/firmware changes in target image, (4) cluster constraints (DRS/HA settings, VM reservations, network RTT)

Expected: Pre-upgrade assessment complete. You have a clear change control document ready for a production change advisory board.

Common Errors

Cluster Image details page shows 'Image not found' or empty composition
Cause: Operations Depot not fully synced post-deployment, or Cluster Image metadata was corrupted
Fix: In SDDC Manager, go to Infrastructure > Operations Depot. Click 'Resync'. Wait 5-10 minutes for depot metadata refresh. Then retry accessing Cluster Image details.
📋 KB: Community KB: 'VCF 9.0 Operations Depot Image Sync'
Hardware heterogeneity detected: two hosts have different NIC models (bnx2x vs ixgbe)
Cause: Hosts were commissioned from different hardware generations — common in phased expansions
Fix: You cannot use a single vLCM image for heterogeneous hosts. Create separate images per hardware class, or plan to replace the outlier host(s) with matching models before upgrade.
📋 KB: Community KB: 'Handling Heterogeneous Hardware in vLCM Upgrades'
HA admission control is disabled; cluster has memory reservations totaling 500 GB on 4 hosts (512 GB per host)
Cause: Over-provisioning + no HA safety margin
Fix: Before upgrade, migrate 2-3 VMs to reduce per-host reservation. Then enable HA admission control (conservative). Test remediation order — you can only remediate 1 host at a time in this scenario.

Task 2 Build and stage the target vLCM cluster image

Availability

vLCM cluster image creation is the core operational task that requires deep understanding of the image specification: not just ESXi version, but the dependency graph of VIBs, firmware, and OEM addons. In VCDX, panelists will ask: 'What happens if you include two conflicting VIBs?' and 'How would you validate an image before rolling it to 4,000 hosts?' This task builds that expertise.

Step 1

In SDDC Manager, go to Infrastructure > Cluster Images. Click 'Create Cluster Image'.

Cluster Image creation wizard opens with step 1: 'Select ESXi Release'
Step 2

Select ESXi 8.0 Update 4 as the base. SDDC Manager will show 'Recommended Components' for that ESXi version. Review the list: ESXi 8.0 U4 + default Broadcom NIC drivers + default Broadcom HBA drivers. Do NOT accept defaults yet.

Recommended components list shown
Step 3

In the wizard, click 'Customize Components'. For the NIC drivers, select the version corresponding to your current hardware. Example: if baseline had bnx2x 8.45.13, check if 8.46.21 is available. If a major version bump (8.45 -> 8.46), research release notes for breaking changes (e.g., https://compatibilityguide.broadcom.com/). Document any breaking changes.

VIB selection interface with version comparisons
In production, you'd run this new image on 2-3 test hosts in an isolated lab first. For this lab, we'll trust the image, but the concept is critical.
Step 4

Click 'Add VIB' to include any vendor-specific addons. Example: if you're using NVIDIA GPUs, add the nvidia-vgpu VIB. If using Broadcom BladeLogic or Dell iDRAC, add those management VIBs. Select the latest stable version available in the Operations Depot.

VIB added to image composition. Image now shows ESXi + NIC driver + HBA driver + custom VIBs
Step 5

Review the final image composition. The image should show: (a) ESXi 8.0 U4 base, (b) minimum 2-3 vendor VIBs (NIC, HBA, optional GPU/management), (c) firmware updates (BIOS, BMC). Click 'Save Image as Draft'.

Draft image created with name auto-generated (example: esxi-8.0u4-prod-workload-001) or custom name if provided
Step 6

Now stage the image: with the draft image selected, click 'Stage Image to Cluster'. Select your workload domain cluster. This operation copies the image payload to the cluster's datastores and validates it against cluster compatibility. Monitor the staging progress in the Activity pane.

Staging progress shown. Completion message: 'Image staged successfully. Ready for remediation.'
Staging takes 10-20 minutes depending on image size (typically 2-4 GB) and network bandwidth. Do NOT interrupt this process.

Validation Gate

Check: In SDDC Manager Cluster Images, the target image is listed with status 'Staged' and a green checkmark. Image composition includes ESXi 8.0 U4 + 3+ VIBs.

Expected: Target image built and staged. Cluster is ready for remediation.

Common Errors

VIB selection shows no 'recommended' VIBs for a particular vendor (e.g., NIC driver missing from dropdown)
Cause: Operations Depot hasn't synced that vendor's VIB bundle yet, or the specific version is not available for ESXi 8.0 U4 compatibility
Fix: In SDDC Manager, go to Infrastructure > Operations Depot > Depots. Manually import the missing VIB bundle (e.g., Broadcom bnx2x 8.46.21 driver). Refresh the Cluster Image wizard. If still missing, file a support request — VIB availability is a known friction point in nested lab environments.
Image staging fails with 'Not enough free space on datastore' after 15 minutes of progress
Cause: Datastore capacity was over-estimated; image copy is failing mid-flight due to insufficient free space
Fix: Check datastore capacity: in vCenter, right-click the datastore and select 'Properties'. If free space < image size + 20%, delete or vMotion old snapshots/images. Re-stage the image.
Staged image shows compatibility warning: 'One or more VIBs are not signed by trusted publishers'
Cause: Custom VIBs (e.g., in-house management tools) don't have valid Broadcom signatures — common in on-premises datacenters
Fix: This is informational in lab; in production, you'd either get the vendor to sign the VIB or accept the security risk at the CAB level. For now, proceed — the warning is just about trust, not functionality.

Task 3 Execute a rolling host remediation with DRS-aware sequencing

Availability

This is where theory meets practice. You will orchestrate the remediation of 4-6 hosts one at a time, respecting DRS constraints (vMotion workloads before entering remediation mode). In VCDX, you'll be asked: 'If you have a 6-host cluster with 10 VMs each and a 2-hour upgrade window, how do you sequence this to minimize guest downtime?' This task answers that question empirically.

Step 1

In vCenter, navigate to your workload cluster. Right-click the cluster and select 'Expand' or click each host individually. For each host, note the current VM count. Create a remediation order starting with the host that has the fewest powered-on VMs (or manually vMotioned VMs to other hosts first). Target order: Host 1 (least VMs) -> Host 2 -> Host 3 -> Host 4.

Remediation sequencing plan documented (e.g., 'Host-1: 2 VMs, Host-2: 3 VMs, Host-3: 4 VMs, Host-4: 3 VMs')
Step 2

Begin remediation of Host 1. In SDDC Manager, go to Infrastructure > Cluster Images. Select your staged image. Click 'Remediate Cluster' (or 'Remediate Host' if granular control is available — check SDDC Manager UI options).

Remediation wizard opens, asking to select hosts or to proceed with all hosts. Select Host 1 only.
Step 3

In the remediation wizard, SDDC Manager will show the remediation workflow: (a) set host to maintenance mode, (b) vMotion all VMs off the host (DRS handles this), (c) upgrade ESXi via vLCM, (d) reboot host, (e) exit maintenance mode. Confirm the workflow sequence and start remediation.

Remediation progress shown in SDDC Manager Activity pane. Host 1 transitions to Maintenance Mode.
Step 4

While Host 1 is remediating, monitor in parallel: (a) In vCenter, watch the vMotion network traffic in the cluster. (b) Check DRS recommendations (home icon in cluster summary) — are migrations being scheduled? (c) Monitor SDDC Manager activity for remediation sub-task progress (firmware updates, VIB installations, etc.).

vMotion activity observed (VMs migrating off Host 1), DRS making recommendations, SDDC Manager showing per-step progress
Step 5

Host 1 remediation should complete in 15-25 minutes (ESXi reboot + VIB application time). Once Host 1 shows 'Compliant' status and is back online in vCenter, begin remediation of Host 2 using the same sequence.

Host 1 shows green status (compliant), back in the cluster, all its VMs have migrated back or are running
Do NOT remediate multiple hosts in parallel initially — you need to demonstrate understanding of sequential, safety-conscious remediation. After the first 2 hosts, you can remediate Host 3 and Host 4 in parallel if time permits.
Step 6

After all hosts are remediated, verify cluster-wide compliance. In SDDC Manager Cluster Images, the cluster should show status 'All Hosts Compliant'. In vCenter, verify all hosts are online, all VMs are running, and DRS load balancing has normalized VM distribution across the cluster.

All hosts compliant, all VMs running, cluster back to normal state

Validation Gate

Check: (1) SDDC Manager shows cluster image compliance = 100% (all hosts compliant), (2) vCenter cluster shows all hosts connected with no alarms, (3) all powered-on VMs are in Running state, (4) documented remediation duration per host (target: 15-25 min/host)

Expected: Cluster successfully upgraded from ESXi 8.0 U3 to 8.0 U4 with zero guest downtime (due to vMotion during maintenance mode).

Common Errors

Host enters maintenance mode but vMotion stalls (VMs do not migrate off). Remediation times out after 30 minutes.
Cause: VM has a memory reservation >= available capacity on remaining hosts, or vMotion network is congested. DRS cannot successfully evacuate the host.
Fix: Cancel remediation. In vCenter, manually reduce VM memory reservations or migrate the blocking VM manually using cold migrate (powered off). Then retry remediation of that host.
📋 KB: Community KB: 'vMotion Stalls During vLCM Remediation'
After remediation completes, host shows status 'Compliant' in SDDC Manager but vCenter shows ESXi version still 8.0 U3
Cause: vCenter cache hasn't updated post-reboot. The host actually IS running 8.0 U4, but vCenter is showing stale data.
Fix: In vCenter, right-click the host > Refresh. If still shows old version, restart the vCenter services: systemctl restart vmware-vpxd on the vCenter VM.
Second or third host remediation fails with 'Image not accessible on datastore'
Cause: Datastore transitioned to read-only (snapshot space exhausted or unplanned rebalancing)
Fix: Check datastore free space. If < 50 GB, delete old snapshots (from previous labs). Re-stage the cluster image. Retry remediation.

Task 4 Detect and remediate infrastructure drift with vLCM compliance baselines

Security

Post-upgrade, your cluster is compliant. But in production, hosts drift: a system administrator manually installs a NIC driver update, firmware is patched out-of-band, or a third-party tool drops a management VIB. vLCM's drift detection is your enforcement mechanism. This task simulates drift and shows you how to find it, report it, and fix it — critical for a VCDX architect designing governance.

Step 1

Simulate drift on Host 2: SSH to that host. Check the current NIC driver version: esxcli software vib list | grep bnx2x (or equivalent for your NIC). Note the version, e.g., 8.46.21.

Current NIC driver version shown (baseline after remediation)
Step 2

Manually 'downgrade' the driver to simulate an out-of-band patch. Run: esxcli software vib remove -n bnx2x (this removes the VIB). Then install the older version from a staging location (e.g., a USB stick or pre-cached bundle). In a real lab, you'd have older VIB bundles staged. For this simulation, manually edit the VIB version metadata (in /opt/vmware/ or similar) to show an older version, or just remove the VIB to trigger a 'missing component' drift.

NIC driver version is now different from cluster image baseline (or missing)
Step 3

In SDDC Manager, navigate to Infrastructure > Cluster Images. Select your staged cluster image. Click 'Compliance Check' or 'Drift Detection'. Select Host 2 (the host you just drifted).

Drift detection scan runs. Result: Host 2 shows 'Non-Compliant' with a list of mismatched VIBs (e.g., 'bnx2x: expected 8.46.21, found [missing or older version]')
Step 4

Review the drift report. SDDC Manager should show: (a) the baseline cluster image specification, (b) the host's actual state, (c) differences highlighted. Document 2-3 specific mismatches.

Detailed drift report showing version deltas for at least 2 VIBs
Step 5

Remediate the drift: in SDDC Manager, select Host 2 and click 'Remediate Host' or 'Force Compliance'. This will re-apply the correct VIB version. Monitor the remediation in vCenter (host enters maintenance mode, VMs migrate).

Host 2 enters maintenance mode, remediates (5-10 minutes), then exits maintenance mode and shows 'Compliant'
Step 6

After remediation, verify Host 2 is back to compliant state. SSH to the host and re-run: esxcli software vib list | grep bnx2x. Confirm the version is now back to 8.46.21.

NIC driver version is back to baseline (8.46.21 or expected version)

Validation Gate

Check: (1) Drift detected on Host 2 via SDDC Manager compliance check, (2) drift report correctly identified missing/wrong VIB versions, (3) Host 2 remediated and returned to compliant state, (4) SSH verification confirms correct version post-remediation

Expected: Drift detection and auto-remediation workflow demonstrated. This is a production-ready governance pattern.

Common Errors

Drift detection runs but shows all hosts compliant even though you manually removed a VIB from Host 2
Cause: Drift detection is not enabled by default, or the removed VIB is optional (not part of the baseline image)
Fix: Ensure you removed a REQUIRED VIB (NIC driver, HBA driver) not an optional one. In SDDC Manager, enable 'Compliance Monitoring' in cluster image settings. Refresh and re-scan.
Manual VIB removal succeeds, but 'esxcli software vib list' still shows the old version
Cause: ESXi caches VIB metadata; you may need to reboot the host for the removal to be visible
Fix: Reboot Host 2 (or place in maintenance mode and trigger a reload). Then check vib list again.

Task 5 Compare VCF 9.0 vLCM workflow with VCF 5.2.x and document operational complexity reduction

Manageability

This is the VCDX perspective: you're not just executing steps, you're comparing two design paradigms. VCF 5.2.x had SDDC Manager managing lifecycle but with manual vLCM tasks. VCF 9.0 unified lifecycle under Operations Depot + vLCM Cluster Image APIs. Understanding this evolution—what was manual, what's now automated, where complexity shifted—is the difference between a VCP and a VCDX.

Step 1

Document the VCF 9.0 workflow you just executed: (1) Task 1: cluster assessment (manual, ~30 min), (2) Task 2: image build + stage (SDDC Manager API, ~20 min), (3) Task 3: rolling remediation (SDDC Manager orchestrated, ~90 min for 4 hosts), (4) Task 4: drift detection (automated, ~10 min). Total operator time: ~2 hours. Total wall-clock time: ~2.5 hours (overlap during staging/remediation).

Timeline document for VCF 9.0 workflow
Step 2

Now review the VCF 5.2.x approach (consult the VCF 5.2 Planning & Preparation guide or community forum threads). In 5.2, the workflow is: (a) SDDC Manager > Lifecycle > Cluster Images (same API, similar UX), (b) Remediation is SDDC Manager-driven but requires operator to pre-stage vMotion/maintenance mode logic manually (not automated), (c) Drift detection is via vLCM baseline comparison, but remediation is not as tightly integrated into SDDC Manager. Document 3 manual steps in 5.2 that are now automated in 9.0.

Comparison document: 3-5 VCF 5.2 vs 9.0 operational differences
Step 3

Quantify operator effort: VCF 9.0 required ~2 hours of hands-on time (largely monitoring). VCF 5.2 would require ~3-4 hours (more manual vMotion orchestration, separate remediation validation, manual drift remediation sequencing). Document the hours saved per cluster upgrade (Delta: ~1-2 hours) and extrapolate to a multi-cluster environment (e.g., 10 workload clusters = 10-20 hours saved per upgrade cycle).

Operator time savings calculation
Step 4

Identify where complexity shifted. In 9.0, complexity shifted LEFT (earlier): image design and composition (vib selection, vendor compatibility validation) happens up-front. In 5.2, you'd discover issues during remediation. Document this: 'VCF 9.0 moves image validation left, reducing remediation failures by ~40%, but requires deeper pre-flight analysis.'

Complexity shift analysis
Step 5

For a VCDX panelist, articulate: 'What are the architectural trade-offs?' Prompt: If you had 50 workload domains with mixed hardware, would you prefer VCF 5.2's per-cluster flexibility (custom images) or VCF 9.0's centralized Operations Depot (consistency, but less flexibility)?' Document your answer and reasoning (2-3 paragraphs).

Architectural trade-off analysis (consistency vs flexibility in multi-domain lifecycle management)
Step 6

Finally, measure upgrade time empirically. Document: (1) Time from 'Start Remediation Host 1' to 'Host 1 Back Online', (2) Average time per host, (3) Total cluster remediation time. In a real production run, this data informs your change advisory board's risk assessment ('If we have a 4-hour maintenance window and 6 hosts, can we complete the upgrade? What's our rollback decision point?').

Measured upgrade timeline: Host remediation duration, total cluster time, rollback threshold (e.g., 'If not 3 of 4 hosts done by 2-hour mark, we abort and revert').

Validation Gate

Check: (1) VCF 9.0 workflow timeline documented (all 4 tasks + wall-clock time), (2) VCF 5.2 workflow comparison with 3+ specific differences, (3) Operator time delta calculated for 1 cluster + extrapolated to 10 clusters, (4) Complexity shift analysis (left-shifting image validation), (5) VCDX-level trade-off analysis (consistency vs flexibility), (6) Empirically measured host remediation times captured

Expected: You can now articulate to a VCDX panelist: 'VCF 9.0's Operations Depot + vLCM integration reduced lifecycle operator overhead by ~30-40% vs 5.2, but requires stronger upfront image design governance. Here are the measured timelines and the architectural trade-offs.'

Common Errors

VCF 5.2 documentation is hard to find or outdated
Cause: VCF 5.2 reached end-of-support in late 2024 (now April 2026). Documentation moved to archives.
Fix: Consult the Broadcom Knowledge Base or the community forum threads referenced in the Extensions section. Look for blog posts from 2023-2024 era comparing 5.2 to 9.0.

Final Validation

You have successfully designed, built, staged, and executed a rolling ESXi cluster upgrade using VCF 9.0's vLCM cluster image paradigm. You've detected and remediated infrastructure drift, and you've articulated the architectural implications vs VCF 5.2. All hosts are compliant, the cluster is healthy, and you have measured data to inform production change control processes.

✓ Baseline cluster image assessment: documented current state, target state, and change list → 3+ specific VIB/firmware changes identified

✓ Target cluster image built and staged: ESXi 8.0 U4 + 3+ vendor VIBs, staged to cluster datastore → Image status = Staged, ready for remediation

✓ All 4-6 hosts remediated successfully: zero guest downtime (vMotion during maintenance mode) → All hosts online, compliant, running their original VM workload (or redistributed by DRS)

✓ Drift detection: manually introduced drift on 1 host, detected via vLCM compliance scan, remediated back to compliant → Host showed non-compliant, then compliant post-remediation

✓ Workflow comparison: VCF 9.0 vs 5.2 analyzed; operator time savings quantified; architectural trade-offs articulated → Documented timeline comparison + 2-3 paragraph VCDX-level analysis

Cleanup / Restore

Snapshot: workload-domain-post-vlcm-upgrade

• Take a snapshot of the workload domain cluster post-upgrade (all hosts compliant, all VMs running): snapshot all cluster VMs as 'workload-domain-post-vlcm-upgrade'

• Document the final cluster image ID and composition for reference in subsequent labs

• If you need to revert for future labs, restore to the 'workload-domain-baseline' snapshot (pre-upgrade state)

Design Reflection (VCDX)

A VCDX panelist examining a large-scale VCF lifecycle management design would probe: How do you handle image composition governance across 50+ workload domains with mixed hardware? What's your rollback strategy if drift remediation fails on 20% of hosts? How does vLCM integrate into your change management process, and who approves image changes? How would you design automated compliance checks that alert on drift within 24 hours? The vLCM cluster image is not just a technical feature—it's a policy enforcement layer. Be ready to argue whether centralized (one image per VCF instance) or distributed (one image per workload domain) image governance scales better, and defend the trade-offs with operational data from this lab.

Requirements

  • Cluster image must incorporate ESXi base OS, vendor-specific drivers (NIC, HBA), firmware (BIOS, BMC), and optional VIBs (GPU, management agents)
  • Remediation must be zero-downtime (vMotion-based) and respect DRS constraints; no guest VM downtime allowed
  • Drift detection must identify non-compliant hosts within 1 hour of deviation and alert operators
  • Lifecycle operations must be auditable: log every image change, remediation, and drift detection event to a centralized audit trail

Constraints

  • vLCM images are cluster-scoped: a single image cannot span multiple hardware revisions or ESXi versions. Heterogeneous clusters require multiple images or host retirement.
  • Remediation bandwidth is limited by vMotion network capacity and storage I/O (snapshot creation during maintenance mode). Typical throughput: 1 host per 15-30 minutes.
  • Vendor VIB availability depends on Operations Depot sync status. Missing VIBs delay image staging by hours or days if depot is offline.
  • DRS constraints (VM memory reservations, resource affinity rules) can block vMotion, causing remediation stalls. Operator must pre-validate cluster has enough free capacity.

Assumptions

  • All hosts in a remediation cluster have identical hardware (CPU, NIC, HBA, firmware UEFI level)
  • vMotion network is available and has > 1 Gbps capacity for concurrent VM migrations
  • Datastore has sufficient free space: image size + (host count * 50 GB snapshot overhead) + 100 GB buffer
  • No VMs have 'must stay on this host' affinity rules that would block remediation
  • Operations Depot is reachable and syncs daily with upstream Broadcom depot

Risks

  • Image composition error (conflicting VIB versions): if two VIBs depend on incompatible kernel ABIs, remediation fails mid-flight and host rolls back, causing remediation stall. MITIGATION: test image on 1-2 hosts before full-scale rollout; maintain rollback image.
  • vMotion stall during maintenance mode: if remaining cluster capacity cannot absorb all VMs from remediating host, vMotion never completes and remediation times out. IMPACT: host stuck in maintenance mode, unavailable to guest workload. MITIGATION: pre-validate DRS has enough free capacity; migrate non-critical VMs off before starting.
  • Datastore fills during multi-host remediation: concurrent snapshots for maintenance mode can exhaust datastore capacity, causing failures and data loss. IMPACT: cluster partially degraded, recovery requires manual snapshot deletion + rollback. MITIGATION: monitor datastore free space hourly during remediation cycle; pre-allocate headroom.
  • Drift remediation cascades: if drift remediation fails on host A (say, VIB install error), the host stays non-compliant. If unchecked, subsequent hosts may inherit the bad image during re-staging. IMPACT: cluster-wide compliance failure. MITIGATION: image staging validates each host's compatibility independently; failed remediation on one host does not affect others, but requires manual investigation.
  • Operator error: manually removing a critical VIB (as in Task 4 drift simulation) is easy and undetectable until drift scan runs hours later. In production, this could cause silent compliance gaps on 10% of hosts. MITIGATION: enforce vib removal/install via vLCM only, disable direct esxcli vib commands on production hosts.

Self-Assessment Discussion Prompts

  1. If you have 50 workload domains, each with slightly different hardware (vendor A's NICs in 20 domains, vendor B's in 30), how would you design your vLCM image strategy? One image per vendor? One per domain? What are the governance implications?
  2. Your change control process requires 2-week approval before any ESXi upgrade. But a critical NIC driver CVE is patched, and you need to update the cluster image within 48 hours. How do you fast-track the image change while maintaining governance?
  3. A host fails remediation mid-upgrade (VIB installation error, host panics during reboot). What's your rollback decision logic? Do you revert the host from snapshot, or do you re-remediate with a different image? What data do you need to decide?
  4. VCF Operations Depot is your single source of truth for VIB versions. If the Depot becomes unavailable (network outage, Broadcom service disruption), can you still remediate hosts? What's your contingency?
  5. After upgrading 4 hosts successfully, you discover a kernel bug in ESXi 8.0 U4 that affects your specific NIC (bnx2x 8.46.21). Hosts are crashing under load. What's your rollback strategy? Can you roll back 4 hosts without affecting the 2 remaining hosts on 8.0 U3?
  6. You have 1000 hosts across 20 clusters. You want to upgrade to ESXi 8.0 U4. Each cluster takes 8 hours to remediate. You have a 4-day maintenance window. How many clusters can you upgrade in parallel? What are the bottlenecks (vMotion bandwidth, datastore I/O, SDDC Manager API rate limits)?

Extensions

Build and test a heterogeneous cluster image strategy for mixed hardware

Repeat Tasks 2-3, but this time create TWO cluster images: one optimized for hosts with Intel NICs (i40en), one for Broadcom (bnx2x). Deploy the lab cluster with mixed hardware (some hosts have Intel, some Broadcom). Demonstrate that vLCM can selectively apply different images to different hosts based on hardware detection. Document the operational complexity of managing multiple images in production (versioning, testing, approval workflows).

harder

Simulate a vLCM remediation failure and execute a rollback recovery

Deliberately corrupt the staged image on one datastore (e.g., truncate the ESXi ISO) to cause a remediation failure on Host 3. Document the failure mode, the error logs, and the recovery procedure. How long does rollback take? Is the host recoverable, or must it be re-imaged? What does this teach you about backup image strategies and rollback decision thresholds?

harder

Design a multi-cluster upgrade orchestration workflow

You have 5 workload clusters (each with 6 hosts). Design a vLCM upgrade strategy that maximizes parallelization while respecting business constraints: (a) max 1 cluster can be in remediation at a time (to avoid vMotion contention), (b) each cluster must complete within 12 hours, (c) you want minimum guest downtime. Build a timeline and sequencing plan. Validate it against vMotion bandwidth, datastore I/O, and SDDC Manager API concurrency limits. This is a production-scale orchestration problem.

harder

Compare Operations Depot image distribution with offline vLCM bundle deployment

In Task 2, you used SDDC Manager's Operations Depot (online, synced from Broadcom). Simulate an air-gapped environment where you download vLCM bundles offline and stage them manually. Create a cluster image using the offline bundle. Compare: (a) time to build image (online vs offline), (b) validation strictness (does offline mode have fewer checks?), (c) compliance gap (if an offline bundle becomes outdated, how do you detect it?). This is relevant for financial/government VCF deployments where Broadcom connectivity is restricted.

same

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.