Academy/VCP-VCF 9.0 Administrator (2V0-17.25)/vLCM Cluster Image Management: Design, Build, and Staged Remediation
This lab targets VCF 9.0

vLCM Cluster Image Management: Design, Build, and Staged Remediation

VCF 9.0Intermediateadminarchitectvcdx⏱ 120 min

vLCM (vSphere Lifecycle Manager) is the unified cluster image management system in VCF 9.0.x, replacing the vSAN 8.x image builder. Applies to ESXi 8.0 Update 2 and later. For VCF 5.2.x with ESXi 7.0.x, see Extension 1.

Objectives

  • Understand vLCM architecture and the cluster image concept in VCF 9.0
  • Design and build a cluster image incorporating base ESXi, drivers, firmware, and vSAN-ready components
  • Assign a cluster image to a management domain cluster and interpret precheck results
  • Execute a staged, rolling remediation across 4 ESXi hosts with zero workload downtime
  • Troubleshoot common vLCM failures (VIB conflicts, firmware incompatibility, staging errors)
  • Validate compliance of remediated hosts against design specifications (driver versions, firmware levels)

Prerequisites

VCF 9.0 lab environment with fully deployed management domain (SDDC Manager, vCenter, NSX Manager, 4 ESXi hosts in a single cluster running ESXi 8.0.0 or 8.0.1). At least 100 GB free space in /boot on each ESXi host for image staging. vSAN is configured and healthy. Internet connectivity to Broadcom Support Portal (or offline ESXi ISO pre-staged in /tmp/esxi on one ESXi host). HA and DRS enabled on the cluster.

Prior labs: holodeck-02 or equivalent VCF deployment, vcp-admin-01 (recommended for understanding VOM and monitoring)

Required skills:

  • ESXi host management and VMkernel networking
  • Understanding of VCF cluster architecture and lifecycle operations
  • Comfort with SDDC Manager console and vCenter Server
  • Basic understanding of VIBs (vSphere Installation Bundles), firmware updates, and driver compatibility
  • SSH access to ESXi hosts and ability to read system logs (/var/log/vLCM.log)

Lab Environment

VCF Management Domain: 1 management cluster with 4 ESXi hosts (ESXi-01 through ESXi-04) running 8.0.0 or 8.0.1. All hosts connected to shared vSAN datastore (vsan-md-cluster). HA enabled with admission control = 1 host failure tolerated. DRS enabled with default aggressiveness (3). NSX is deployed on all hosts. Single vCenter instance managing the cluster. SDDC Manager accessible at https://sddc-manager.vcf.local:443.

IP Addressing

NetworkPurposeVLAN
10.0.0.0/20Management services (SDDC Manager, vCenter, NSX, vLCM)1644
10.1.1.0/24ESXi host management VMkernel1644
10.1.2.0/24vMotion1645
10.1.3.0/24vSAN1646

Credentials

SystemUsernamePassword
SDDC Manageradministrator@vsphere.localVCF admin password from deployment spec
vCenter Serveradministrator@vsphere.localSame as SDDC Manager
ESXi HostsrootESXi root password from Holodeck config (default: generated during deployment)

Tasks

Task 1 Understand vLCM Architecture and Cluster Image Design

Manageability

vLCM is a paradigm shift from traditional 'patch then reboot' to 'image-based cluster remediation'. Instead of applying patches individually, you define a Desired Image (base ESXi + validated components), stage it to hosts, then remediate in a coordinated rollout. This is the VCDX-critical architectural pattern: declaring desired state vs. imperative patching. A panelist will probe: 'Why image-based lifecycle management? What are the advantages vs. traditional patching?'

Step 1

Read the vLCM architecture overview: Navigate to SDDC Manager > Lifecycle > Images > Help (or online docs: https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-0.html). Document 5 key concepts: (a) Cluster Image definition, (b) Image Staging vs. Remediation, (c) Precheck validation, (d) HA-aware remediation (rolling reboot), (e) Compliance checking post-remediation.

You can articulate the vLCM workflow: Design Image → Create Image → Assign to Cluster → Precheck → Stage → Remediate → Validate
Step 2

Access SDDC Manager and navigate to Lifecycle > Images > Cluster Images. Review any pre-existing images. Note the columns: Image Name, Base ESXi Version, Components (VIBs, drivers, firmware), Created Date, Compatible Clusters.

You see a list of available images (may be empty if this is the first lab)
Step 3

Review the current state of your management cluster: SDDC Manager > Inventory > Clusters > management-cluster. Click 'Image Compliance'. Note the current Desired Image (likely 'None' or a baseline from deployment). Take a screenshot showing the current state.

Screenshot shows the current image state of the management cluster
Image Compliance is the 'design state vs. actual state' view. If Desired Image is set but actual hosts are not running that image, the cluster is out of compliance.
Step 4

Design your target cluster image by creating a design spec document (or add to lab notes):

'Cluster Image Design Specification: vcp-admin-02-management-cluster

Target Image: ESXi-8.0-U3-vSAN-Ready-20240501
Base ESXi: 8.0 Update 3 (ESXi 8.0.3.0.xxx)
VIB Components:

  1. vSAN Driver VIB (VIBs): vsan-vsanmgmt-8.0.3.x, vsan-esx-8.0.3.x (auto-selected by vLCM)
  2. Network Drivers: Intel NIC driver (ixgbe or i40e, version X.X.x)
  3. Storage Drivers: RAID controller driver (LSI MegaRAID, version Y.Y.y)

Firmware:

  1. NIC Firmware: Intel X550 (version Z.Z.z)
  2. RAID Controller Firmware: LSI SAS3x08 (version A.A.a)

Design Intent:

  • All 4 management domain hosts must run the same image for consistency
  • No community/third-party VIBs; all components are Broadcom-certified
  • Image will be staged to all hosts, then remediated in rolling fashion (1 host at a time, HA-aware)
  • Estimated remediation time: 4 hosts × 15 min (boot + health check) = 60 minutes
  • Zero workload downtime expected (HA handles VM migration during host reboot)

Compliance Checkpoints:

  • Post-remediation: All 4 hosts running image version X
  • vSAN rebalance completion (data rebuild after each host reboot)
  • NSX transport node health confirmed
  • SDDC Manager health check passes'
Design spec document is created and includes image components, remediation timeline, and compliance checkpoints
This design spec is the 'architecture' part of the task. You're not just updating hosts; you're declaring the desired state and building a remediation plan.
Step 5

Verify prerequisites for image building:

(a) ESXi ISO availability: Check if ESXi 8.0.3 ISO is available in SDDC Manager's content library. Navigate to SDDC Manager > Lifecycle > Images > Available Components > Base ESXi. If not listed, you'll need to upload it.

(b) Driver/firmware sources: Document where you'll source NIC drivers and RAID firmware:

  • Broadcom Support Portal (https://support.broadcom.com/)
  • Vendor websites (Intel for NIC, LSI/Broadcom for RAID)
  • Pre-staged in your lab environment (/tmp/drivers on the management appliance)

(c) Staging space: SSH into one ESXi host and verify /boot has >100 GB free: ssh root@esxi-01.vcf.local, then 'df -h /boot'

You've confirmed ESXi ISO is available, identified driver sources, and verified staging space
Step 6

Document your vLCM remediation plan (add to design spec or lab notes):

'Remediation Plan:
Phase 1 (Pre-remediation): Verify HA & DRS ready, take snapshot of all hosts, confirm vSAN health
Phase 2 (Staged rollout):

  - Host 1 (ESXi-01): Enter maintenance mode (HA evacuates VMs) → Stage image → Remediate (reboot) → Validate ESXi version, vSAN status, NSX TN health
  • Host 2 (ESXi-02): Repeat
  • Host 3 (ESXi-03): Repeat
  • Host 4 (ESXi-04): Repeat

Phase 3 (Post-remediation): Verify all hosts at target version, re-enable HA admission control, document time/resources used

Risk Mitigation:

  • If remediation fails on a host, revert by exiting maintenance mode and running a different image (rollback image)
  • If vSAN gets unhealthy during host reboots, pause the rollout and allow vSAN to rebuild
  • Estimated total time: 4 hours (4 hosts × 60 min per host including vSAN rebuild time)'
Remediation plan is documented with phases, risks, and timeline

Validation Gate

Check: You can navigate SDDC Manager > Lifecycle > Images and articulate: (1) what a cluster image is, (2) the 5-step vLCM workflow, (3) your target image components, (4) your remediation plan with HA considerations, (5) compliance checkpoints

Expected: Design spec and remediation plan documents exist and are detailed enough for a peer to follow and execute

Common Errors

ESXi 8.0.3 ISO is not available in SDDC Manager's content library
Cause: The ISO was not staged during VCF deployment, or it's in a different SDDC Manager instance
Fix: Download ESXi 8.0.3 ISO from Broadcom Support Portal. Use SDDC Manager > Lifecycle > Images > Upload Component to add the ISO. Wait 5-10 minutes for SDDC Manager to process and index it.
You're confused about the difference between 'Desired Image' and 'Current Image'
Cause: vLCM uses declarative language: Desired = goal state, Current = actual state. This is different from traditional imperative patching.
Fix: Think of it like this: in vSphere HA, you declare 'I want 4 hosts online' (Desired), and HA ensures you have 4 hosts (Current). Same with vLCM: you declare Desired Image, and vLCM ensures all hosts remediate to that image.
You create an image but don't understand what VIBs are included
Cause: vLCM auto-selects VIBs based on the base ESXi version and the cluster's hardware profile. You may not see all VIBs until the image is created.
Fix: After creating the image (Task 2), open it and review 'Image Components' tab to see all included VIBs. Cross-reference with the vSAN Compatibility Guide to ensure all are supported.

Task 2 Create and Publish a Cluster Image

Manageability

Image creation is where you assemble the base ESXi, drivers, and firmware into a versioned, certified bundle. Think of it as packaging: you're saying 'this image is the canonical configuration for my management cluster, and I guarantee it's been tested and validated.' This is the 'build and release' discipline in operations.

Step 1

Navigate to SDDC Manager > Lifecycle > Images > Create Image. Click 'Create New Image'.

Image creation wizard opens
Step 2

Step 1 - Select Base ESXi:

  • Base ESXi: Select 'ESXi 8.0 Update 3' (or the latest available 8.0.x version)
  • Release Date: Select the most recent 8.0.3 build
  • Click 'Next'
Base ESXi is selected and wizard moves to next step
Ensure you select Update 3 (or later), not Update 2. ESXi 8.0.2 has known issues with vLCM and NSX integration; 8.0.3 is the recommended baseline for VCF 9.0.
Step 3

Step 2 - Add Custom Components (VIBs and Drivers):

  • vSAN Components: vLCM auto-suggests vsan-vsanmgmt and vsan-esx VIBs for your ESXi version. Accept the defaults (these are mandatory for vSAN clusters).
  • NIC Drivers: If your cluster has Intel NICs, search for 'ixgbe' or 'i40e' driver VIBs. Select the version matching your NIC hardware (e.g., 'ixgbe-driver-8.0.3-X.X.x'). If you don't know the exact version, select the most recent one — vLCM will validate compatibility during precheck.
  • RAID Controllers: If your hosts have external RAID controllers (e.g., LSI MegaRAID), search for and add the corresponding driver VIB. If hosts are all-flash or software RAID, skip this.
  • Optional: Add any other certified VIBs (e.g., hyperconverged system VIBs, management agent VIBs).
  • Click 'Next'
All VIBs selected and displayed in a summary table
vLCM auto-validates VIB compatibility. If you select two conflicting VIBs, it will warn you before image creation. Pay attention to these warnings — they prevent remediation failures.
Step 4

Step 3 - Add Firmware (Optional but Recommended):

  • NIC Firmware: If you know your NIC firmware version (e.g., Intel X550 Ethernet Adapter firmware), upload or select it. Firmware updates are optional; if not specified, hosts keep their current firmware.
  • RAID Controller Firmware: Same as above — only if your hosts have external RAID controllers.
  • Baseboard Management Controller (BMC) Firmware: Only if you have standardized BMC firmware in your environment (e.g., iLO for HPE, iDRAC for Dell).
  • For this lab, you can skip firmware updates if they're not available. Click 'Next'
Firmware components are optional; you can skip to the next step
Firmware updates require host reboots and can cause longer remediation times. Only include firmware in the image if you have a specific compliance requirement (e.g., security patch for NIC firmware).
Step 5

Step 4 - Name and Publish:

  • Image Name: Enter 'ESXi-8.0-U3-vSAN-Ready-20240501' (use a descriptive name with date)
  • Description: 'Certified management domain image for VCF 9.0. Base ESXi 8.0.3, vSAN-ready, NSX-compatible. Validated for 4-node management cluster.'
  • Publish: Select 'Yes, publish this image' so it's available for assignment to clusters
  • Click 'Create'
Image is created and published; you're redirected to the Images list view
Image naming convention matters. Include the base version (8.0-U3), purpose (vSAN-Ready), and date (20240501). This makes it easy to track and audit image versions over time.
Step 6

Verify image creation: Back in SDDC Manager > Lifecycle > Images, you should see 'ESXi-8.0-U3-vSAN-Ready-20240501' in the list. Click on it to view the image details:

  • Components tab: Shows all included VIBs and firmware
  • Clusters tab: Shows which clusters this image is assigned to (should be empty at this point)
  • Click on 'Components' and take a screenshot showing all VIBs included
Image exists with all selected components visible in the Components tab
The image creation can take 5-10 minutes as SDDC Manager validates and indexes the image. If you see 'Creating...' status, wait and refresh after 5 minutes.

Validation Gate

Check: Image 'ESXi-8.0-U3-vSAN-Ready-20240501' exists in SDDC Manager > Lifecycle > Images, is Published, and includes the expected VIBs (vsan-vsanmgmt, vsan-esx, NIC driver, optional firmware)

Expected: You can view the image, see all components, and confirm it's ready for assignment to clusters

Common Errors

Image creation fails with 'Incompatible VIB combination' or 'VIB signature validation failed'
Cause: Two selected VIBs have dependency conflicts, or a VIB is not signed/certified for your ESXi version
Fix: Remove the conflicting VIB and select an alternative. Consult the vSAN Compatibility Guide (https://compatibilityguide.broadcom.com/) to verify all VIBs are certified for ESXi 8.0.3.
Image creation succeeds but components are not visible in the Components tab
Cause: SDDC Manager is still indexing the image (background process)
Fix: Wait 5 minutes and refresh the page. If still missing, try logging out and back into SDDC Manager.
You can't find the NIC driver VIB or RAID driver VIB in the component library
Cause: Driver VIB has not been uploaded to SDDC Manager, or it's named differently than expected
Fix: Manually upload the VIB to SDDC Manager: Lifecycle > Images > Upload Component. Select the .vib file and wait for indexing (5-10 minutes). Then, try creating the image again.

Task 3 Assign Image to Cluster and Interpret Precheck Results

Manageability

Assigning an image triggers a precheck validation — vLCM verifies that each host in the cluster can be remediated with the image. Precheck catches show-stoppers (e.g., incompatible CPU, missing RAID controller driver) before you stage/remediate. Understanding precheck results is critical: if precheck fails, you must fix the root cause before proceeding. A panelist will ask: 'Tell me about a precheck failure you hit and how you resolved it.'

Step 1

Navigate to SDDC Manager > Inventory > Clusters > management-cluster. Click the 'Image Compliance' tab.

You see the current image compliance state (likely showing Desired Image = None, Actual = [current ESXi version])
Step 2

Click 'Set Desired Image'. A dialog opens showing published images. Select 'ESXi-8.0-U3-vSAN-Ready-20240501'. Click 'Next'.

Image is selected and you're prompted to proceed with precheck
Step 3

Click 'Run Precheck' or 'Continue' (depending on SDDC Manager version). vLCM will validate each host against the image. This takes 2-5 minutes. Monitor the precheck progress in the UI.

Precheck completes with results: either all Passed, or some Failed/Warnings
Precheck is non-destructive; it doesn't modify hosts. It's purely a validation step. Pay close attention to any FAILED items — these MUST be resolved before staging/remediation.
Step 4

Review precheck results in detail. The UI should show a per-host summary:

  • ESXi-01: Precheck Passed (all checks green)
  • ESXi-02: Precheck Passed
  • ESXi-03: Precheck Passed (or Failed if there's an issue)
  • ESXi-04: Precheck Passed

For each host, expand the details and document any warnings or failures.

Precheck results table shows status for each host
Common precheck warnings (not failures): 'Host firmware is older than image specifies' (warning only, host will boot OK), 'Custom VIBs detected' (warning, host may have third-party drivers). These are safe to proceed. Failures (e.g., 'CPU not supported', 'RAID controller driver not found') require investigation.
Step 5

If any host shows a Precheck Failure:

Case 1: 'CPU not supported by ESXi 8.0.3'

  • Cause: Host has a legacy CPU that doesn't support AVX (Advanced Vector Extensions)
  • Fix: Add 'allowLegacyCPU=TRUE' to the boot.cfg of that host (see common errors below). Re-run precheck.

Case 2: 'RAID controller driver not found'

  • Cause: The image doesn't include a driver for the host's RAID controller
  • Fix: Go back to Task 2, add the missing driver VIB to the image, and publish a new image version (e.g., 'ESXi-8.0-U3-vSAN-Ready-with-LSI-20240501'). Then re-assign the new image.

Case 3: 'VIB conflict detected' (e.g., two versions of the same driver package)

  • Cause: The image includes two incompatible VIBs
  • Fix: Edit the image, remove the conflicting VIB, and re-publish. Then re-run precheck.

If no failures, skip this step and continue to Step 6.

All precheck failures are documented and resolved, or none exist
Step 6

Once precheck passes for all hosts, take a screenshot of the precheck summary and save it in your lab notes. Document:

  • Image name: 'ESXi-8.0-U3-vSAN-Ready-20240501'
  • Cluster: 'management-cluster'
  • Precheck status: 'Passed for all 4 hosts'
  • Date: [today's date]
Screenshot and documentation of successful precheck

Validation Gate

Check: Precheck results show 'Passed' for all 4 hosts (ESXi-01 through ESXi-04). No Failed checks. Any Warnings are documented and understood.

Expected: You can explain: (1) what each precheck validates, (2) any warnings/failures found, (3) why they don't block remediation (or how you fixed them)

Common Errors

Precheck fails with 'CPU not supported by ESXi 8.0.3' on one or more hosts
Cause: Host has a legacy CPU that doesn't support advanced CPU features required by ESXi 8.0.3 (e.g., RDRAND instruction set)
Fix: SSH to the affected host, edit /etc/vmware/esx.conf, and add the line: /adv/Misc/allowLegacyCPU=true. Reboot the host. Re-run precheck — it should pass now. (Note: this is a lab workaround; in production, upgrade the host CPU.)
Precheck hangs or times out (stuck at 'Validating hosts...' for >10 minutes)
Cause: One or more hosts are not responding to precheck validation (network issue, SDDC Manager→Host connectivity problem)
Fix: SSH to the slow host and verify it can reach SDDC Manager: ssh root@esxi-01.vcf.local, then 'ping sddc-manager.vcf.local'. If no response, check the management network configuration. Restart the management agents on the host: /etc/init.d/vpxa restart. Re-run precheck.
Precheck passes, but you assign the image and it immediately shows 'Desired: Image-A, Actual: Old-ESXi-8.0.0' (not remediating)
Cause: Image assignment succeeded but staging/remediation hasn't started. This is normal — you must explicitly click 'Stage Image' in the next step.
Fix: Proceed to Task 4 to stage the image to hosts. Assignment is just declaring the desired state; staging/remediation are the actual changes.

Task 4 Stage Image to Hosts (Prepare for Remediation)

Availability

Staging copies the cluster image to each host's /boot partition and validates that all image bits are available before committing to the reboot-based remediation. Staging is non-disruptive; hosts stay online and VMs keep running. This is a critical checkpoint: if staging fails, you identify issues before triggering reboots.

Step 1

In SDDC Manager > Inventory > Clusters > management-cluster > Image Compliance, click 'Stage Image'. This initiates staging of the image to all 4 hosts.

Staging begins; UI shows 'Staging in progress' with a progress bar
Step 2

Monitor staging progress. This takes 10-30 minutes depending on image size and host I/O capacity. While staging is in progress, open a new terminal and SSH into one of the hosts to observe the staging process in real-time:
ssh root@esxi-01.vcf.local
tail -f /var/log/vLCM.log

You should see log lines like:
'[INFO] Staging image 'ESXi-8.0-U3-vSAN-Ready-20240501' to /boot'
'[INFO] Verifying image integrity...'
'[INFO] Staging complete. Ready for remediation.'

vLCM logs show staging progress; /boot is being populated with image bits
Staging logs are invaluable for troubleshooting. If staging fails, the logs will show exactly where the failure occurred (e.g., 'Failed to extract VIB: ixgbe-driver' or 'Insufficient space in /boot').
Step 3
Check /boot free space during staging: In the same SSH session, run 'df -h /boot' periodically. You should see /boot free space decreasing as the image is staged (e.g., 500 GB free → 250 GB free as 250 GB image is copied).
/boot free space decreases as staging progresses
Step 4

Wait for staging to complete on all 4 hosts. SDDC Manager UI should show 'Staging Status: Complete' or 'Ready for Remediation'.

Staging completes for all 4 hosts; UI shows completion status
Step 5

Verify staging on all hosts. SSH into each host and run:
grep 'Staging complete' /var/log/vLCM.log | tail -1

This should show a timestamp confirming staging completed. Repeat for ESXi-02, ESXi-03, ESXi-04.

All 4 hosts show 'Staging complete' in their vLCM logs
This is a key validation checkpoint. If any host shows 'Staging incomplete' or 'Staging failed', investigate before proceeding to remediation.

Validation Gate

Check: Staging status is 'Complete' for all 4 hosts. vLCM logs on each host confirm staging success. /boot contains the staged image.

Expected: You can SSH into any host and verify /boot contains the new image: 'ls -la /boot/vmlinuz*' should show the new ESXi 8.0.3 kernel.

Common Errors

Staging fails with 'Insufficient space in /boot'
Cause: /boot partition is too small or too full. ESXi 8.0.3 image requires ~150-250 GB in /boot, depending on components.
Fix: SSH to the host and clean up /boot: rm -rf /boot/vmlinuz.old /boot/vmcore.old (careful with paths!). If space is still insufficient, you may need to skip staging and use a different image (one with fewer components). In production, ensure /boot has >250 GB free before any ESXi upgrade.
Staging times out or hangs after 30 minutes
Cause: Slow datastore I/O or network bandwidth issue. The image copy is bandwidth-limited.
Fix: Wait longer (staging can take up to 1 hour for large images). If it's stuck for >1 hour, cancel staging and investigate host I/O: 'esxtop' and check disk read/write throughput. Ensure the management network has sufficient bandwidth (no congestion).
Staging succeeds for ESXi-01 and ESXi-02, but fails on ESXi-03 with 'Network timeout'
Cause: One host has a network connectivity issue (e.g., management port is flapping, or host is becoming unresponsive)
Fix: SSH into ESXi-03 and check network health: esxcfg-vswitch -l (verify management vSwitch is up). Restart the management agents: /etc/init.d/vpxa restart. Then retry staging from SDDC Manager.

Task 5 Execute Staged Rolling Remediation (Zero-Downtime Patching)

Availability

Remediation is where vLCM shines: it reboots each host one at a time, with HA handling VM migration. The goal is zero workload downtime despite taking 4 hosts offline sequentially. This is the 'orchestrated rolling update' pattern. A panelist will probe: 'Walk me through a remediation of 4 hosts. How do you ensure no VM downtime? What if a host gets stuck in maintenance mode?'

Step 1

Pre-remediation checklist:
(a) Verify vSAN health: In vCenter > Cluster A > vSAN, status should be green (no rebuilding in progress from a prior failure)
(b) Verify DRS is enabled: vCenter > Cluster A > Configure > Services > vSphere DRS > Status should be 'Enabled'
(c) Verify HA is enabled: vCenter > Cluster A > Configure > Services > vSphere HA > Status should be 'Enabled', and 'Admission Control' should be 'Configured with 1 host failure tolerated'
(d) Check vMotion readiness: All hosts should have vMotion enabled on their VMkernel adapters
(e) Record current state: Take a screenshot of cluster summary (4 hosts online, VMs running, vSAN green)

If any of the above are not met, fix them before proceeding.

All pre-remediation checks pass. You have baseline screenshots.
Step 2

Initiate remediation on Host 1 (ESXi-01): In SDDC Manager > Inventory > Clusters > management-cluster > Image Compliance, click 'Remediate'. A dialog appears showing all 4 hosts and their remediation order.

Remediation dialog opens showing the remediation sequence (likely Host 1 → Host 2 → Host 3 → Host 4)
Step 3

By default, vLCM remediates hosts sequentially (one at a time). Confirm the remediation settings:

  • Mode: 'Sequential' (do not use Parallel unless you have >8 hosts and explicit approval)
  • Order: Host 1, Host 2, Host 3, Host 4
  • Pre-remediation: 'Enable Maintenance Mode' (checked)
  • Post-remediation: 'Exit Maintenance Mode' (checked)
  • Disable HA: (unchecked — we want HA to manage VM migration)

Click 'Start Remediation' to begin.

Remediation starts. UI shows 'Host 1: Remediation in progress — Entering maintenance mode'
Remediation is disruptive. Each host reboot takes 15-30 minutes. The entire 4-host rollout will take 60-120 minutes. Do not interrupt or close the SDDC Manager console during remediation.
Step 4

Monitor Host 1 remediation in parallel:

(a) In vCenter, watch Host 1 state: Cluster A > Hosts > ESXi-01. You should see it enter 'Maintenance Mode' (not accepting new VMs). VMs on ESXi-01 will be vMotioned to ESXi-02, ESXi-03, or ESXi-04 by HA/DRS. This takes 5-10 minutes depending on number of VMs and vMotion bandwidth.

(b) In SDDC Manager, check the current step: 'ESXi-01: Maintenance mode complete. Rebooting...'

(c) SSH to ESXi-01 and monitor the boot: tail -f /var/log/vmkernel.log. You should see kernel logs showing the system is shutting down and rebooting with the new image.

Host 1 enters maintenance mode, VMs are evacuated, host reboots, new kernel boots
Step 5

Post-reboot validation for Host 1 (takes 5-10 minutes after reboot):

(a) Host becomes available again: vCenter shows ESXi-01 as 'Connected' (no longer grayed out)
(b) Verify ESXi version: SSH to ESXi-01, run 'vmware -v'. You should see 'ESXi 8.0.3 build XXXXX' (or whatever your target version is).
(c) Verify NSX transport node health: NSX Manager > System > Fabric > Host Transport Nodes. ESXi-01 should show 'Configuration State: Success, Status: Up'
(d) vSAN recovery: vCenter > vSAN > Disks. You should see vSAN rebuilding data (if FTT=1, there's 1x replication overhead being restored after the host went down).

Do NOT proceed to Host 2 until Host 1 is fully operational (vSAN rebuild is in progress or complete, all services are healthy).

Host 1 is back online, running ESXi 8.0.3, NSX is healthy, vSAN is rebuilding
Step 6

Monitor vSAN rebuild on Host 1. Run: vsan.py -c localhost -u root -p [password] cluster get resilience

You should see output like:
'Capacity in compliance: 95%, Rebuilding: 5%, Time to complete: 10 minutes'

Wait for rebuild to reach 100% (Capacity in compliance = 100%, Rebuilding = 0%) before proceeding to Host 2. This prevents cascading failures if a second host goes down during rebuild.

vSAN rebuild completes on Host 1; cluster resilience returns to nominal
Step 7

Repeat steps 3-6 for Host 2 (ESXi-02), Host 3 (ESXi-03), and Host 4 (ESXi-04). The full remediation of 4 hosts takes approximately:

  • Host 1: 30 min (reboot + vSAN rebuild)
  • Host 2: 30 min (reboot + vSAN rebuild)
  • Host 3: 30 min (reboot + vSAN rebuild)
  • Host 4: 30 min (reboot + vSAN rebuild)
  • Total: ~120 minutes (2 hours)

During each host's remediation, monitor vCenter and vSAN to ensure no anomalies.

All 4 hosts complete remediation sequentially; each shows new ESXi 8.0.3 version

Validation Gate

Check: All 4 hosts complete remediation and are running ESXi 8.0.3. vCenter shows all 4 hosts as 'Connected' with no alarms. vSAN shows 'Healthy' status. No VM live migration errors in vCenter events.

Expected: vmware -v on each host shows ESXi 8.0.3. vCenter cluster summary shows 4 hosts with green health indicators. Total remediation time is documented.

Common Errors

Host gets stuck in 'Maintenance Mode' and won't exit
Cause: A VM cannot be vMotioned off the host (e.g., VM is powered off, or vMotion is blocked by affinity rules)
Fix: Navigate to vCenter, select the stuck host, and manually click 'Exit Maintenance Mode'. The host will exit immediately. Then, in SDDC Manager, retry remediation. If it fails again, manually vMotion any blocked VMs to another host, then exit maintenance mode.
Remediation hangs after reboot ('Host is rebooting...' for >20 minutes)
Cause: New ESXi image is booting very slowly (e.g., slow storage I/O during boot, or drivers are being loaded)
Fix: Wait up to 30 minutes. If still stuck, SSH to the host (if accessible) and check /var/log/vmkernel.log for boot errors. If host doesn't respond to SSH, it may be in bootloop — revert the host to the previous image by editing its boot.cfg to remove the new image selection.
Remediation completes but one host shows 'ESXi 8.0.0' (old version) instead of '8.0.3'
Cause: Host didn't actually boot the new image — it fell back to the old image (e.g., due to boot.cfg corruption or staging failure)
Fix: SSH to the host and verify the boot.cfg is correct: cat /etc/vmware/esx.conf | grep -i tboot. If it's not pointing to the new image, this is a boot configuration issue. Manually edit the file to boot from the new image, then reboot.
vSAN goes unhealthy during Host 2 remediation (shows 'Degraded' status)
Cause: Two hosts are offline simultaneously, but vSAN only tolerates one host offline (FTT=1). This is a race condition if remediation timing is tight.
Fix: This is critical: immediately STOP the remediation. Ensure Host 2 has fully booted and vSAN rebuild from Host 1 is complete before proceeding. Verify vSAN health: vsan.py -c localhost cluster get resilience. Rebuild must show 100% compliance before Host 3 starts remediation.
NSX transport node shows 'Install Failed' after host reboot
Cause: NSX VIBs were not properly staged or there's a VIB conflict with the new ESXi image
Fix: In NSX Manager, select the failed transport node and click 'Resolve' or 'Reinstall'. NSX will re-install the transport node VIBs. Wait 5-10 minutes for completion.

Task 6 Validate Compliance and Document Remediation Results

Manageability

Post-remediation validation is the proof that the cluster is in the desired state. You verify: (1) all hosts running the target image version, (2) cluster services (HA, DRS, vSAN) are healthy, (3) no VMs were lost or corrupted, (4) performance is acceptable. This is the 'acceptance testing' phase. The documentation becomes a VCDX artifact: 'I deployed this image, validated it, and here's proof it works.'

Step 1

Verify all hosts are running the target ESXi version. SSH into each host and run:
vmware -v

All 4 hosts should show 'ESXi 8.0.3' (or your target version). Create a table:

Host | ESXi Version | Build | Status
ESXi-01 | 8.0.3 | XXXXX | ✓ Correct
ESXi-02 | 8.0.3 | XXXXX | ✓ Correct
ESXi-03 | 8.0.3 | XXXXX | ✓ Correct
ESXi-04 | 8.0.3 | XXXXX | ✓ Correct

All 4 hosts report ESXi 8.0.3 (or target version)
Step 2

Verify cluster services are healthy:

(a) vCenter > Cluster A > Summary. You should see:

  • 'Connected Hosts: 4'
  • 'vSAN Status: Healthy'
  • 'DRS Status: Fully Automated'
  • 'HA Status: Enabled (Admission Control Configured)'

(b) Navigate to vCenter > Cluster A > Configure > Services > vSphere HA.

  • Status: 'Enabled'
  • 'Failures and Isolation Response: Power Off VMs'
  • 'Admission Control': 'Host Failures Tolerated: 1'

(c) Navigate to vCenter > Cluster A > Configure > Services > vSphere DRS.

  • Status: 'Fully Automated'
  • 'DRS Recommendations' should show ≤5 outstanding recommendations (a fresh cluster after remediation may have minor imbalances)
All cluster services show healthy status
Step 3

Verify vSAN health in detail:

(a) vCenter > vSAN > Disks. All disks should show 'Healthy' (green).

(b) vCenter > vSAN > Capacity. You should see:

  • Used capacity: X GB
  • Free capacity: Y GB
  • (Used + Free = Total vSAN datastore capacity)

(c) vCenter > vSAN > Performance. No alarms or warnings should be present.

vSAN disks are healthy, capacity is normal, no performance alerts
Step 4

Verify NSX transport nodes are healthy:

(a) NSX Manager > System > Fabric > Nodes > Host Transport Nodes.

  • All 4 hosts should show 'Configuration State: Success, Status: Up'
  • No 'Install Failed' or 'Not Ready' statuses

(b) If any host shows 'Install Failed', click on it and select 'Resolve' to reinstall transport node VIBs.

All 4 transport nodes show Success / Up
Step 5

Verify no VMs were lost or corrupted:

(a) vCenter > Cluster A > VMs. Count the total number of VMs and compare to your baseline (from Task 5, Step 1). The count should be identical.

(b) Select each VM and verify:

  • 'Power State: Powered On' (or Powered Off if that was the pre-remediation state)
  • 'vSAN SPBM Compliance Status: Compliant' (if vSAN policies are configured)

(c) Optionally, log into 1-2 critical VMs and verify they're responsive (ping, SSH, application health check).

All VMs are present, powered on (if expected), and compliant
Step 6

Check cluster events for remediation-related errors:

(a) vCenter > Cluster A > Events. Filter for events in the last 2-3 hours (during remediation window).

(b) Look for:

  • 'HA VM Failover' events: None expected (VMs should have migrated gracefully via vMotion, not via HA)
  • 'DRS Recommendations' events: A few expected as hosts were rebooted
  • 'vMotion' events: Many expected (VM migration off rebooting hosts)
  • Errors or Warnings: None expected; any errors should be investigated

Take a screenshot of the events log.

Events log shows normal vMotion migration activity, no unexpected HA failovers or errors
Step 7

Create a comprehensive remediation report. Document:

'vLCM Remediation Report: ESXi-8.0-U3-vSAN-Ready-20240501
Cluster: management-cluster
Remediation Date: [date]
Remediation Duration: [total time in hours:minutes]

Pre-Remediation State:

  • 4 hosts running ESXi 8.0.0
  • All hosts connected, vSAN healthy, HA enabled
  • [X] VMs running across cluster

Post-Remediation State:

  • All 4 hosts running ESXi 8.0.3
  • All hosts connected, vSAN healthy, HA enabled
  • [X] VMs running (same count as pre-remediation)
  • No VM downtime events observed
  • No vSAN rebuild failures
  • NSX transport nodes all healthy

Compliance Status:

  • Cluster Image Desired: ESXi-8.0-U3-vSAN-Ready-20240501 ✓
  • Cluster Image Actual: ESXi-8.0-U3-vSAN-Ready-20240501 ✓
  • Cluster is IN COMPLIANCE

Issues Encountered: [None, or list any]

Key Metrics:

  • Remediation time per host: ~30 minutes (including vSAN rebuild)
  • Total remediation time: ~120 minutes (4 hosts)
  • VMs migrated via vMotion: [total count]
  • vMotion failures: 0
  • HA failover events: 0
  • vSAN rebuild time: ~10 minutes per host

Conclusion: Cluster remediation completed successfully with zero workload downtime. All hosts are now running the certified, validated image.'

Comprehensive remediation report is documented

Validation Gate

Check: All 4 hosts running ESXi 8.0.3. Cluster services (HA, DRS, vSAN) are healthy. All VMs accounted for and operational. Remediation report is documented. Image Compliance in SDDC Manager shows 'Desired = Actual = ESXi-8.0-U3-vSAN-Ready-20240501'.

Expected: SDDC Manager > Inventory > Clusters > management-cluster > Image Compliance shows green checkmark: 'Cluster is compliant with desired image version'.

Common Errors

Image Compliance in SDDC Manager still shows 'Desired = Image-A, Actual = Old-ESXi-8.0.0' even though all hosts show ESXi 8.0.3
Cause: SDDC Manager UI cache is stale, or the compliance check hasn't refreshed yet
Fix: Refresh the SDDC Manager page (Ctrl+F5). Wait 5 minutes. If still not updating, log out of SDDC Manager and log back in.
One host shows 'ESXi 8.0.3' in 'vmware -v', but SDDC Manager shows it as 'Not Compliant'
Cause: The host's boot.cfg or boot loader is pointing to a different image location, or SDDC Manager hasn't discovered the new version yet
Fix: In SDDC Manager, click 'Run Compliance Check' or 'Refresh Compliance' to force a rescan of all hosts. Wait 5 minutes.
DRS is showing >10 recommendations post-remediation
Cause: After rebooting hosts, VMs may not be perfectly balanced. DRS is making recommendations to rebalance.
Fix: This is normal post-remediation. DRS will automatically apply recommendations within 5-10 minutes. If you want to accelerate balancing, click 'Apply All Recommendations' in vCenter > Cluster A > Recommendations.

Final Validation

vLCM remediation is complete. The management cluster is now running a certified, validated cluster image (ESXi 8.0.3 + components). Zero workload downtime was achieved through HA-aware, sequential remediation. The cluster remains compliant: actual host versions match the desired image. This is production-grade patch management.

✓ All 4 hosts running target ESXi version (8.0.3) → vmware -v on each host confirms ESXi 8.0.3

✓ Cluster image compliance is GREEN (desired = actual) → SDDC Manager > Image Compliance shows checkmark and 'Compliant'

✓ All cluster services healthy (HA, DRS, vSAN) → vCenter shows 4 connected hosts, vSAN healthy, HA enabled with 1 failure tolerated, DRS enabled

✓ All VMs present and operational → VM count matches pre-remediation baseline; all VMs are connected/powered on

✓ NSX transport nodes all healthy → NSX Manager shows 4 host transport nodes with Configuration State = Success, Status = Up

✓ No vSAN rebuild failures or data loss → vSAN shows healthy status; no missing or degraded objects

✓ Remediation report is documented → Report includes pre/post state, timeline, metrics, and proof of zero workload downtime

Cleanup / Restore

Snapshot: vcp-admin-02-complete

• Take a final snapshot of all 4 hosts in the cluster: 'vcp-admin-02-complete' (for use in future labs)

• Export the remediation report as PDF for your VCDX design artifacts

• Clean up any test VIBs or staging artifacts from /boot on ESXi hosts (optional, but good hygiene)

• Document any lessons learned (e.g., 'vSAN rebuild took longer than expected due to network bandwidth', 'one host took 35 min to boot instead of 15')

Design Reflection (VCDX)

VCDX panelists probe your operational maturity by asking about patch/update strategy. vLCM demonstrates that you've moved beyond 'patch individual hosts' to 'orchestrated cluster remediation'. Key talking points: (1) Why cluster images? They eliminate version skew and configuration drift. (2) How do you handle HA during remediation? Sequential remediation with HA-aware vMotion ensures zero workload downtime. (3) What if staging fails on a host? Precheck catches most issues; if staging fails, you can retry or revert to a previous image. (4) How often do you remediate?
In production, apply security patches monthly, driver updates quarterly, major version upgrades annually. (5) Compliance: How do you prove your cluster stays at a known version? Image Compliance in SDDC Manager is your proof — actual versions match desired image. Be prepared to articulate the trade-off between 'strict change management' (only approved images) and 'rapid patching' (need for emergency security patches).

Requirements

  • All hosts in a cluster must run the same ESXi version and certified driver/firmware combination for operational consistency
  • Cluster image updates must be staged before remediation to validate image integrity and detect incompatibilities
  • Remediation must be coordinated with HA/DRS to ensure zero workload downtime during host reboots
  • Post-remediation validation must verify cluster services are healthy and no data loss has occurred
  • Image compliance status must be queryable (SDDC Manager > Image Compliance) to demonstrate adherence to desired state

Constraints

  • vLCM requires at least 100 GB free space in /boot on each host for image staging
  • Sequential remediation takes 15-30 min per host × number of hosts = 60-120 min for a 4-node cluster
  • HA Admission Control must be sized to tolerate 1 host offline during remediation; if configured for 0 failures tolerated, VMs cannot be migrated safely
  • vSAN requires rebuild time (10-30 min per host) after each host reboot; parallel remediation risks vSAN degradation if 2+ hosts reboot before rebuild completes
  • Precheck validation depends on data source connectivity; if SDDC Manager cannot reach a host, precheck hangs or times out

Assumptions

  • All 4 hosts have identical hardware (CPU, NIC, RAID controller) so one image works for all
  • ESXi ISO and all component VIBs (vSAN, drivers) are available in SDDC Manager's content library
  • Network connectivity and vMotion bandwidth are sufficient to migrate all VMs off a host within 10 minutes
  • vSAN rebuild bandwidth is available (vSAN network not saturated) so rebuild completes within 30 min per host
  • No affinity rules or other constraints block VM migration during host reboot

Risks

  • Precheck passes but staging fails on a host due to late-discovered incompatibility (e.g., custom VIB on host not in image) — IMPACT: host cannot boot new image — MITIGATION: use precheck to catch most issues; have rollback image ready if needed
  • vSAN goes unhealthy if 2 hosts reboot within rebuild window (e.g., remediation timing is too tight) — IMPACT: data loss if FTT<2 — MITIGATION: ensure FTT=1 minimum and allow full vSAN rebuild before next host remediation
  • HA evacuation fails if vMotion network is congested, leaving VMs on rebooting host — IMPACT: unplanned VM downtime — MITIGATION: monitor vMotion network saturation; use bandwidth reservations if needed
  • Firmware update in image causes extended boot time, delaying remediation timeline — IMPACT: remediation takes 2x longer than planned — MITIGATION: skip firmware updates unless security-critical; if included, test boot time in lab first

Self-Assessment Discussion Prompts

  1. You designed a cluster image (ESXi 8.0.3 + vSAN VIBs + NIC drivers). Walk me through your rationale for each component. Why that specific vSAN driver version? Why that NIC driver version?
  2. Precheck found a warning: 'NIC firmware is older than image specifies'. Do you block remediation or proceed? Justify your answer.
  3. Your remediation of 4 hosts took 2 hours total. The estimate was 1.5 hours. What caused the delay? How would you optimize for the next remediation?
  4. During Host 2 remediation, vSAN shows 'Degraded' status (warning). How do you respond? Do you pause the rollout? Why or why not?
  5. Your cluster is now at ESXi 8.0.3. In 3 months, ESXi 8.0.4 is released with a critical security patch. How do you approach the upgrade? Do you update the cluster image immediately, or wait for more testing?
  6. You're remedying a production cluster (500+ VMs, not just 4 test VMs). What additional safeguards would you implement? Longer precheck window? Parallel remediation? Rollback plan?
  7. Compliance in SDDC Manager shows 'Desired: Image-A (8.0.3), Actual: Image-A on 3 hosts, Image-B on 1 host'. What happened and how do you fix it?

Extensions

Multi-Cluster Image Rollout

Extend the lab to 2 workload domain clusters. Design a coordinated rollout plan: (1) Which cluster gets patched first? (2) How long is the window between cluster remediations? (3) How do you monitor inter-cluster dependencies (e.g., stretched vSAN, NSX cross-cluster services)? This mirrors production scenarios where you manage multiple clusters and must coordinate updates.

harder

vLCM 5.2.x vs. 9.0.x Comparison

If you have access to a VCF 5.2.x environment, document the operational differences: VCF 5.2.x uses separate vSAN Image Builder + vSphere ESXi Image Builder vs. unified vLCM in 9.0.x. Compare the UI, workflow, and time-to-remediate. Document 5 operational improvements in 9.0.x vLCM.

same

Rollback Scenario: Staged Revert

After completing remediation, intentionally introduce a failure (e.g., disable a NIC driver in the image). Create a second 'rollback' image (previous ESXi 8.0.2 + original drivers) and execute a remediation back to that image. This tests your ability to recover from a bad image. Document the rollback procedure and MTTR (time to revert all 4 hosts).

harder

Custom VIB Integration

Integrate a custom or vendor-specific VIB into the cluster image (e.g., monitoring agent, firmware management tool). Go through the full cycle: (a) add VIB to image, (b) run precheck to validate compatibility, (c) stage and remediate with the custom VIB, (d) verify the custom VIB is active post-remediation. This exercises real-world scenarios where vendors ship custom VIBs for their appliances.

moderate

References

  • VMware Cloud Foundation 9.0 Cluster Image Lifecycle ManagementTier 1 — Official
    Official Broadcom documentation for vLCM architecture, image design, staging, remediation, and compliance checking. The authoritative reference for all vLCM concepts and workflows.
  • VCF 9.0 Lifecycle Operations and Management GuideTier 1 — Official
    Deep dive on SDDC Manager's Lifecycle > Images section, precheck validation, and compliance policies. Critical for understanding the SDDC Manager UI workflow.
  • vSAN Compatibility GuideTier 1 — Official
    Broadcom's official compatibility matrix for vSAN drivers, firmware, and supported ESXi versions. Essential for validating VIB choices and firmware versions in your cluster image.
  • ESXi Patching and Updates Best PracticesTier 1 — Official
    VMware best practices for ESXi lifecycle management, including vLCM, patching windows, HA coordination, and compliance tracking.
  • William Lam — VCF vLCM Deep DiveTier 3 — Expert Blog
    Expert blog exploring vLCM implementation, troubleshooting, and production lessons learned. Often covers edge cases and best practices not found in official docs.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.