Upgrade Troubleshooting Scenario
Objectives
- Diagnose a hung BIOS/firmware flash during ESXi host upgrade orchestrated by SDDC Manager and vLCM
- Evaluate safe interrupt strategies that avoid data corruption on vSAN and host boot devices
- Execute a roll-forward approach to resume an ESXi upgrade from its fail-point using vLCM remediation
- Perform a rollback procedure when roll-forward is not viable, restoring the host to its pre-upgrade image
- Validate post-recovery cluster health across vSAN, NSX transport nodes, and SDDC Manager inventory
- Apply the 19-step VCF upgrade sequence (Broadcom KB 390634) to plan upgrade order and dependency awareness
- Document upgrade failure recovery using RCAR methodology for VCDX defense readiness
Prerequisites
VCF 9.0 management domain fully deployed with SDDC Manager, vCenter, NSX Manager, and vSAN cluster operational. At least one pending ESXi host image update configured in vLCM.
Prior labs: vcp-support-01, vcp-support-05
Required skills:
- SSH access to ESXi hosts and VCSA
- Familiarity with vSphere Lifecycle Manager (vLCM) image-based management
- Understanding of vSAN maintenance mode and data evacuation policies
- Basic knowledge of SDDC Manager upgrade workflows and bundle management
Lab Environment
VCF management domain with SDDC Manager, vCenter, NSX Manager, and a minimum 4-node vSAN ESA cluster. One host designated as the upgrade target with a pending vLCM image containing a newer ESXi build and firmware/driver updates.
Tasks
Task 1 Diagnose Firmware Flash Hang During ESXi Upgrade
Build the diagnostic skills to identify why a BIOS/firmware flash has hung during a vLCM-orchestrated ESXi upgrade, distinguish between a true hang and a long-running but healthy flash, and assess the risk before taking any corrective action.
In the vSphere Client, navigate to the vLCM cluster view (Cluster > Updates > Image). Check the remediation status for the target host. Document the current remediation state: is the host shown as 'Remediating', 'Failed', or 'Non-Compliant'? Note the stage at which the process stopped (e.g., 'Installing host image', 'Applying firmware updates', 'Rebooting').
SSH into the VCSA and examine the vLCM logs at /var/log/vmware/vmware-updatemgr/vum-server/vmware-vum-server.log. Search for the target host FQDN or IP to find the remediation timeline. Also check /var/log/vmware/lifecycle/lifecycle.log for SDDC Manager-initiated upgrade workflow entries. Identify the last successful substage and the first error or timeout message.
Check the Hardware Support Manager (HSM) integration status. In vSphere Client, navigate to Lifecycle Manager > Settings > Hardware Support. Verify the HSM provider (e.g., Dell OpenManage, HPE OneView) is connected and its depot is accessible. Review the HSM logs for firmware compatibility warnings. On the ESXi host (if accessible via SSH or iLO/iDRAC), check /var/log/esxupdate.log and /var/log/sfcb/sfcbd.log for firmware update progress.
Determine whether the firmware flash is truly hung or merely slow. Check the host out-of-band management console (iLO, iDRAC, or CIMC) for BIOS POST status. If the host shows a firmware update progress screen, document the percentage and wait 10 minutes to confirm no progress. If the host shows a normal BIOS POST screen or is at an ESXi boot prompt, the firmware flash may have completed but vLCM lost track. Document your assessment: (a) truly hung, (b) slow but progressing, or (c) completed but unreported.
Validation Gate
Check: Document the root cause diagnosis with evidence from vLCM logs, HSM status, and OOB console state
Expected: Root cause identified as one of: HSM connectivity loss, firmware package incompatibility, vLCM timeout on long-running flash, or actual firmware tool crash. Evidence collected from at least three sources (vLCM task, vum-server.log, OOB console).
Common Errors
Task 2 Design and Execute Safe Interrupt Strategy
When a firmware flash is confirmed hung, design an interrupt strategy that minimizes risk of data corruption on vSAN disks and host boot devices, preserving the ability to recover the host without hardware replacement.
Before any interrupt action, verify the vSAN data protection state. In vSphere Client, navigate to the vSAN cluster and check: (a) the host was placed in maintenance mode with 'Ensure Data Accessibility' before remediation started, (b) vSAN health shows no objects with reduced availability, (c) resync operations are not in progress. Run esxcli vsan health cluster list from another healthy host to confirm cluster health.
Assess the interrupt options in order of increasing risk: (1) Cancel the vLCM remediation task from vSphere Client (least disruptive -- sends a cancel signal to the remediation engine), (2) Reset the host BMC/IPMI via OOB management (moderate risk -- forces BMC restart without affecting main CPU/firmware), (3) Graceful shutdown via OOB management ACPI power button (moderate-high risk -- OS-level shutdown if ESXi is responsive), (4) Hard power cycle via OOB management (highest risk -- immediate power loss, potential firmware corruption). Document which option you will use and why.
Execute the chosen interrupt strategy. If using vLCM cancel: right-click the remediation task in Recent Tasks and select Cancel. If using BMC reset: access the OOB console and use the BMC cold reset option (ipmitool mc reset cold from a management station). Monitor the host state after the interrupt. Document the time of interrupt and the host response.
Verify host integrity after the interrupt. SSH into the ESXi host (if accessible) and run: (a) esxcli system version get -- confirm ESXi version, (b) esxcli hardware platform get -- confirm BIOS version, (c) esxcli software vib list | grep -i firmware -- check firmware VIB state, (d) vsish -e get /hardware/bios/biosInfo -- BIOS details. Compare the reported firmware versions against the pre-upgrade baseline to determine if the firmware was partially updated.
If the host booted successfully, verify vSAN disk group health from the host perspective. Run: esxcli vsan storage list to confirm all disk groups are present. Check for any disk errors in /var/log/vmkernel.log related to vSAN or storage controller. Do NOT exit maintenance mode yet -- the host needs further evaluation in Tasks 3 and 4.
Validation Gate
Check: Host successfully interrupted and stabilized with verified disk and firmware integrity
Expected: Interrupt executed using the least-disruptive viable method. Host reached a stable boot state. Firmware versions documented. vSAN storage pools confirmed healthy. Host remains in maintenance mode for further recovery.
Common Errors
Task 3 Roll-Forward and Rollback Recovery Procedures
Execute the appropriate recovery path: roll-forward (resume the upgrade from the fail-point) when the host is in a recoverable state, or rollback (restore the pre-upgrade image) when roll-forward is not viable.
Evaluate roll-forward viability. Check the following criteria: (a) ESXi host booted successfully to the DCUI, (b) firmware versions are either fully old or fully new (not mixed), (c) vLCM can reach the host and reports it as 'Non-Compliant' (not 'Unknown'), (d) the Hardware Support Manager shows the host firmware baseline is assessable. If all criteria are met, roll-forward is viable. If any criterion fails, proceed to rollback evaluation in Step 3.
Execute roll-forward via vLCM remediation. In vSphere Client, navigate to the cluster > Updates > Image. Select the target host and click 'Remediate'. Confirm the host is already in maintenance mode. Monitor the remediation progress in Recent Tasks. This time, before starting, verify HSM connectivity and set a longer remediation timeout if the failure was caused by a timeout: vSphere Client > Lifecycle Manager > Settings > Remediation Timeout.
If roll-forward is not viable (mixed firmware state, host not booting cleanly, or vLCM cannot assess compliance), execute rollback. For ESXi image rollback: (a) if the host boots, use esxcli software profile rollback to revert to the pre-upgrade ESXi image (the altbootbank contains the previous version), (b) for firmware rollback, access the OOB management console and use the vendor firmware recovery tool (HPE iLO Recovery Set, Dell iDRAC firmware rollback, or Cisco CIMC recovery). Document each rollback step executed.
After either roll-forward or rollback completes, update the SDDC Manager upgrade workflow status. If the upgrade was SDDC Manager-orchestrated: (a) check the SDDC Manager UI > Lifecycle > Upgrade History for the failed host upgrade task, (b) use the SDDC Manager API to retry or skip the host: GET /v1/upgrades to find the upgrade ID, then POST /v1/upgrades/{id}/retry or use the skip option for the specific host. Reference the VCF 19-step upgrade sequence (KB 390634) to understand where this host upgrade falls in the overall workflow.
Document the recovery decision and outcome for the VCDX design journal. Record: (a) the failure symptom, (b) root cause from Task 1, (c) interrupt method from Task 2, (d) recovery path chosen (roll-forward or rollback), (e) recovery outcome, (f) time to recovery (TTR), (g) lessons learned. Frame the documentation using RCAR methodology: what requirement was at risk, what constraint influenced the recovery decision, what assumptions were validated or invalidated, and what risk materialized.
Validation Gate
Check: Host successfully recovered via roll-forward or rollback with SDDC Manager workflow updated
Expected: Host is either upgraded to the target image (roll-forward) or restored to the pre-upgrade image (rollback). vLCM compliance status is accurate. SDDC Manager upgrade workflow reflects the current state. Recovery documented with RCAR analysis.
Common Errors
Task 4 Post-Recovery Cluster Validation and Health Checks
Systematically validate that the cluster, vSAN, NSX transport nodes, and SDDC Manager inventory are all healthy after the upgrade interruption and recovery, ensuring no latent issues remain.
Exit the host from maintenance mode and verify vSAN resynchronization. In vSphere Client, right-click the host > Exit Maintenance Mode. Monitor vSAN resync progress: navigate to Cluster > Monitor > vSAN > Resyncing Components. Wait for all resync operations to complete before proceeding. Run esxcli vsan health cluster list from any host in the cluster to get a full health report.
Validate NSX transport node status. In NSX Manager UI, navigate to System > Fabric > Nodes > Host Transport Nodes. Verify the recovered host shows 'Success' for Configuration State and 'Up' for Node Status. If the host was rolled back, NSX transport node configuration may need to be re-applied. Check the NSX Manager API: GET /api/v1/transport-nodes/<node-id>/state for detailed status.
Verify SDDC Manager inventory consistency. In SDDC Manager UI, navigate to Inventory > Hosts. Confirm the recovered host shows correct ESXi version, firmware baseline, and lifecycle status. Run the SDDC Manager SoS (Suite of Services) diagnostic utility: /opt/vmware/sddc-support/sos --health-check on the SDDC Manager appliance to validate all management domain components.
Run a comprehensive cluster health validation checklist: (a) vSAN health -- all tests green, no objects with reduced availability, (b) DRS -- host participating in load balancing, resource pools balanced, (c) HA -- host admission control satisfied, host is HA-capable, (d) networking -- all vmknics up, vMotion test between recovered host and another host succeeds, (e) licensing -- per-core license consumed correctly for VCF 9.0 (verify in vCenter > Administration > Licensing). Document any findings.
Create a post-incident report summarizing the full upgrade failure and recovery lifecycle. Include: (a) timeline of events from upgrade start to recovery completion, (b) root cause analysis with evidence, (c) recovery method used and justification, (d) validation results confirming full recovery, (e) preventive recommendations to avoid recurrence (e.g., pre-check HSM connectivity, validate firmware compatibility against HCL, extend remediation timeouts for known slow firmware updates). Present this as a VCDX-quality operational runbook entry.
Validation Gate
Check: All cluster health checks pass and post-incident documentation is complete
Expected: Host fully operational in the cluster. vSAN resync complete with all objects healthy. NSX transport node active. SDDC Manager inventory accurate. SoS health check all-green. Post-incident report documented.
Common Errors
Final Validation
Complete upgrade failure troubleshooting lifecycle from diagnosis through safe interrupt, recovery (roll-forward or rollback), and comprehensive post-recovery validation across the entire VCF stack
✓ Firmware flash hang correctly diagnosed → Root cause identified with evidence from vLCM logs, HSM status, and OOB console -- true hang vs. slow flash vs. completed-but-unreported distinguished
✓ Safe interrupt executed without data corruption → Least-disruptive interrupt method chosen and executed. vSAN data integrity confirmed. No firmware corruption from the interrupt.
✓ Recovery path successfully completed → Host either upgraded to target image (roll-forward) or restored to pre-upgrade image (rollback). vLCM compliance status accurate.
✓ SDDC Manager upgrade workflow updated → SDDC Manager reflects current host state. Upgrade workflow either retried or host skipped with a documented remediation plan.
✓ Full cluster health validated post-recovery → vSAN resync complete, NSX transport node active, SoS health check all-green, licensing verified, post-incident report documented.
Cleanup / Restore
• Verify all hosts are out of maintenance mode and fully participating in the vSAN cluster
• Confirm SDDC Manager SoS health check shows all-green for management domain components
• Remove any temporary firmware packages or staged VIBs that were used during troubleshooting
• Revert to lab snapshot if configuration changes were made that should not persist
Design Reflection (VCDX)
Upgrade failure recovery demonstrates mastery of VCF lifecycle management, the ability to make risk-based decisions under pressure, and understanding of the blast radius concept (management plane vs. data plane, single host vs. cluster vs. VCF stack). The structured approach -- diagnose, assess risk, choose least-disruptive recovery, validate comprehensively -- is the hallmark of VCDX-level operational design thinking.
Requirements
- R-001: Complete VCF 9.0 upgrade within the approved maintenance window with zero data loss
- R-002: Maintain vSAN data accessibility during rolling host upgrades (no object availability reduction)
- R-003: Ensure all hosts reach target ESXi image and firmware baseline per the vLCM cluster image
- R-004: Preserve SDDC Manager lifecycle workflow integrity throughout the upgrade process
Constraints
- Cannot power cycle a host during active firmware flash -- risk of bricking BMC/BIOS
- vSAN FTT=1 policy allows only one host in maintenance mode at a time in a 4-node cluster
- SDDC Manager orchestrated upgrades follow a fixed 19-step sequence -- cannot skip dependency steps
- Firmware updates require Hardware Support Manager connectivity to the vendor depot
Assumptions
- ESXi altbootbank contains the pre-upgrade image for rollback capability
- Out-of-band management (iLO/iDRAC/CIMC) is accessible for host console and power control
- Hardware Support Manager firmware packages are validated against the vendor HCL before deployment
- Operations team has physical or remote access to host OOB management interfaces
Risks
- Firmware flash hang during BIOS update can brick the host if interrupted unsafely -- mitigated by OOB console verification before any interrupt
- Partial firmware update leaves host in mixed firmware state -- mitigated by vendor-specific firmware recovery tools and HCL validation
- SDDC Manager upgrade workflow stuck in failed state blocks all subsequent lifecycle operations -- mitigated by API-based retry/skip capability
- vSAN resync after host recovery consumes cluster bandwidth -- mitigated by scheduling upgrades during low-I/O periods
Self-Assessment Discussion Prompts
- How would you modify the upgrade strategy for a 32-host production cluster where the maintenance window is only 4 hours?
- What pre-upgrade checks would you add to prevent firmware flash hangs from occurring in the first place?
- How does the VCF 19-step upgrade sequence change your rollback strategy if the failure occurs at step 12 (ESXi hosts) vs. step 4 (vCenter)?
- If two hosts in a 4-node vSAN cluster both fail firmware updates simultaneously, what is your recovery priority and why?
Extensions
Automated Pre-Upgrade Firmware Compatibility Validation
Build a PowerCLI or Python script that queries each host's current BIOS, BMC, and storage controller firmware versions, compares them against the vLCM desired image firmware baseline, and validates compatibility against the vendor HCL. Run this before every upgrade cycle to catch incompatibilities before they cause flash hangs.
Multi-Host Parallel Upgrade Failure Simulation
In a Holodeck lab with 8+ hosts, simulate concurrent firmware failures on two hosts in the same vSAN fault domain. Practice the recovery decision: which host to recover first, how to maintain vSAN object availability with two hosts down, and how to coordinate SDDC Manager workflow retry for multiple failed hosts.
VCF Upgrade Rollback Decision Framework
Create a decision tree document that maps every step in the VCF 19-step upgrade sequence to its rollback procedure, blast radius, and estimated rollback time. Include decision criteria for when to rollback a single component vs. rollback the entire upgrade. Test the framework against three historical upgrade failure scenarios.
⚠ Known Pitfalls (from Community KB)
References
- Broadcom KB 390634 -- VCF Lifecycle Upgrade Sequence: 19-step upgrade order and dependencies for VCF 9.0
- vSphere Lifecycle Manager 9.0 Documentation -- Image-based host management, firmware updates, and Hardware Support Manager integration
- vSAN 9.0 Troubleshooting Guide -- Maintenance mode considerations, data evacuation policies, and resync operations
- ESXi 9.0 Upgrade Guide -- Boot bank architecture, altbootbank rollback, and esxcli software profile commands
- VMware VCF 9.0 SDDC Manager API Reference -- Upgrade workflow management, retry/skip operations, and SoS diagnostic utility