Upgrade Failure & Remediation
Objectives
- Understand VCF upgrade sequence and component dependencies (19-step order per KB 390634)
- Diagnose upgrade failures using SDDC Manager logs, vCenter upgrade logs, and NSX upgrade status
- Execute vCenter upgrade rollback via VM snapshot restoration
- Recover failed SDDC Manager upgrade workflows using API retry and manual intervention
- Design pre-upgrade validation checklists to prevent common upgrade failures
- Document upgrade failure recovery using RCAR methodology
Prerequisites
VCF lab with management domain running pre-upgrade version
Prior labs: vcp-support-01
Required skills:
- SDDC Manager UI and API navigation
- vCenter snapshot management
- SSH access to VCF components
Lab Environment
VCF management domain with SDDC Manager, vCenter, NSX Manager cluster, and 4 ESXi hosts
Tasks
Task 1 VCF Upgrade Architecture & Sequence Planning
Master the 19-step VCF upgrade sequence, understand component interdependencies, and identify which steps are most likely to fail and why.
Document the VCF 9.0 upgrade sequence per Broadcom KB 390634. The 19 steps in order: (1) SDDC Manager, (2) vRealize Suite Lifecycle Manager, (3) Workspace ONE Access, (4) vRealize/Aria Automation, (5) vRealize/Aria Operations, (6) vRealize/Aria Operations for Logs, (7) vRealize/Aria Operations for Networks, (8) vCenter Server (management domain), (9) NSX Manager cluster (management domain), (10) ESXi hosts (management domain), (11) vSAN on-disk format (management domain), (12-15) Repeat steps 8-11 for each workload domain, (16) HCX, (17) vSAN Witness, (18) SDDC Manager drift reconciliation, (19) Post-upgrade validation.
Identify the most failure-prone upgrade steps and common causes. Document: (a) SDDC Manager upgrade — database migration issues, insufficient disk space; (b) vCenter upgrade — VCSA migration assistant failures, customization conflicts; (c) NSX Manager cluster upgrade — cluster stability during rolling upgrade, backup controller issues; (d) ESXi host upgrade — vSAN maintenance mode failures, incompatible hardware/drivers.
Map the rollback options for each component. vCenter: revert VM snapshot (fastest, most reliable). SDDC Manager: restore from backup (more complex, requires database consistency). NSX Manager: restore from backup or rebuild (most complex — 3-node cluster coordination). ESXi: cannot easily downgrade — revert at vSAN cluster level or reinstall. Document the rollback time estimate for each.
Review the pre-upgrade validation checklist. Items: (a) verify bundle compatibility matrix in SDDC Manager, (b) check all VCF component health (SDDC Manager dashboard), (c) verify vSAN health and capacity (must be below 80% for host upgrades), (d) confirm DNS forward/reverse resolution for all components, (e) verify NTP synchronization across all components (<5 second drift), (f) take VM snapshots of vCenter, SDDC Manager, NSX Managers, (g) export NSX Manager backup, (h) document current versions of all components.
Validation Gate
Check: Complete VCF upgrade sequence and pre-upgrade checklist documentation
Expected: 19-step sequence documented with dependency justification. Failure matrix and rollback matrix created. Pre-upgrade checklist with 15+ items.
Common Errors
Task 2 Upgrade Failure Diagnosis & Log Analysis
Develop systematic log analysis skills for diagnosing VCF upgrade failures — from SDDC Manager workflow failures to vCenter migration issues to ESXi remediation errors.
Map the critical log locations for VCF upgrade troubleshooting. SDDC Manager: /var/log/vmware/vcf/sddc-manager-ui-app/sddc-manager-ui-app.log (UI workflow), /var/log/vmware/vcf/operationsmanager/operations-manager.log (orchestration). vCenter upgrade: /var/log/vmware/upgrade/upgrade-runner.log. NSX Manager: /var/log/proton/nsxapi.log. ESXi remediation: /var/log/vmware/vcf/lcm/lcm-debug.log.
Analyze a simulated vCenter upgrade failure. Common scenario: VCSA upgrade fails at database migration step. Log analysis: search upgrade-runner.log for 'ERROR' and 'FAILED'. The error typically shows: database schema migration failure, disk space exhaustion, or service dependency conflict. Document the error pattern and correlate with pre-upgrade checklist items.
Analyze SDDC Manager workflow failure. Navigate to SDDC Manager UI > Lifecycle Management > view failed workflow. Document: workflow ID, failed step, error message, and retry eligibility. Via API: GET /v1/tasks/{taskId} to retrieve detailed error information. Check if the workflow can be retried (some failures are retryable, others require manual intervention).
Practice ESXi host upgrade failure diagnosis. Common scenario: host enters maintenance mode but vSAN data evacuation stalls. Diagnose: (a) check vSAN resync status (objects still syncing from previous operation), (b) verify cluster capacity allows data migration, (c) check for VM anti-affinity rules preventing evacuation, (d) verify no VMs are pinned to the host. Document remediation for each cause.
Build an upgrade failure escalation decision tree. Level 1: retry the failed workflow (if retryable). Level 2: analyze logs, fix the identified issue, retry. Level 3: rollback the failed component (snapshot/backup restore), fix root cause, re-attempt. Level 4: escalate to VMware support with: support bundle, upgrade logs, and detailed failure description.
Validation Gate
Check: Complete upgrade failure diagnosis with log analysis and escalation procedures
Expected: Log locations mapped for all VCF components. Failure patterns documented. Escalation decision tree created with time-boxing.
Common Errors
Task 3 Rollback Procedures & Recovery Operations
Execute component-level rollback procedures for failed VCF upgrades — from vCenter snapshot revert to SDDC Manager backup restore to NSX cluster recovery.
Practice vCenter rollback via snapshot revert. Steps: (1) Power off the VCSA VM, (2) Revert to pre-upgrade snapshot in ESXi host client (not vSphere client — vCenter is down), (3) Power on the VCSA VM, (4) Wait for all services to start (5-15 minutes), (5) Verify vCenter is responsive and inventory is intact, (6) Notify SDDC Manager of rollback (may need to reconcile state).
Practice SDDC Manager rollback from backup. Steps: (1) Power off the SDDC Manager VM, (2) Deploy a fresh SDDC Manager OVA from the original version media, (3) Restore from backup using the SDDC Manager restore wizard, (4) Verify SDDC Manager UI is accessible, (5) Verify all domains and hosts appear in inventory, (6) Reconcile any state differences.
Practice NSX Manager cluster recovery after failed upgrade. Scenario: one NSX Manager node in the 3-node cluster fails during upgrade. Steps: (1) Verify remaining 2 nodes are healthy (nsxcli -c 'get cluster status'), (2) Remove the failed node from the cluster, (3) Deploy a fresh NSX Manager node at the target version, (4) Join the new node to the cluster, (5) Verify cluster stability (all 3 nodes STABLE). Document the timeline.
Handle the partial upgrade scenario: vCenter upgraded successfully but NSX upgrade failed. Now the environment has mixed versions (vCenter 9.0, NSX 4.x). Document: (a) is this a supported mixed-version state? Check VMware Interoperability Matrix. (b) Can you proceed with NSX upgrade retry? (c) Should you rollback vCenter to match NSX? Analyze the decision tree.
Design post-rollback validation procedure. After any rollback: (1) Verify all VCF component versions match expected state, (2) Run SDDC Manager health check, (3) Verify vSAN cluster health, (4) Verify NSX controller cluster stability, (5) Run DFW Traceflow tests for critical paths, (6) Verify all VMs are running and accessible, (7) Test VCF operations (VM provisioning, host add). Document as a runbook.
Validation Gate
Check: Execute rollback procedures and validate environment recovery
Expected: vCenter snapshot revert practiced, SDDC Manager restore documented, NSX cluster recovery understood, post-rollback validation checklist created.
Common Errors
Task 4 Upgrade Prevention Strategy & VCDX Defense
Design a comprehensive upgrade strategy that minimizes failure risk and prepare to defend VCF lifecycle management design decisions in VCDX context.
Design a VCF upgrade strategy document. Include: (a) upgrade policy (frequency: quarterly patches, semi-annual minor, annual major), (b) maintenance window schedule (management domain: weekend 1, workload domain 1: weekend 2, workload domain 2: weekend 3), (c) pre-upgrade checklist reference, (d) rollback decision criteria (time-boxed: if not resolved in 2 hours, rollback), (e) communication plan (stakeholders notified 2 weeks before, day-of status updates).
Calculate maintenance window requirements. For each upgrade step, estimate: preparation time, execution time, verification time, and rollback buffer. Total management domain upgrade: pre-checks (2h) + SDDC Manager (1h) + vCenter (2h) + NSX (3h) + ESXi hosts (4h for rolling upgrade of 4 hosts) + validation (1h) = ~13 hours + 2h rollback buffer = 15-hour maintenance window.
Design upgrade testing strategy using Holodeck. Steps: (1) Deploy Holodeck VCF environment matching production version, (2) Run the complete upgrade in Holodeck first, (3) Document any issues encountered and resolution, (4) Validate upgrade success in Holodeck, (5) Use Holodeck results to refine production upgrade plan. Document the value of pre-production upgrade testing.
Prepare four VCDX defense responses: (1) 'Why do you upgrade management domain before workload domains?' — SDDC Manager orchestrates all upgrades; it must be at the latest version to manage the upgrade workflow for workload domains. vCenter and NSX in management domain are upgraded next because workload domain upgrades depend on management domain health. (2) 'What is your rollback strategy for a failed ESXi host upgrade?' — ESXi host rollback via Alt-F11 boot bank switch (if dual-bank supported). If not, rebuild host from image profile and rejoin cluster. vSAN rebuild handles data redistribution. (3) 'How do you handle a upgrade that spans multiple maintenance windows?' — Document the partial upgrade state, verify component interoperability using VMware matrix, schedule remaining steps in next window. Never leave an environment in untested partial upgrade state indefinitely. (4) 'What is the most critical pre-upgrade step?' — VM snapshot of ALL VCF management components. Without snapshots, rollback requires full rebuild. This one step reduces rollback RTO from days to minutes.
Validation Gate
Check: Complete upgrade strategy document and VCDX defense preparation
Expected: Upgrade strategy with policy, scheduling, and communication plan. Maintenance window calculations documented. Holodeck testing strategy defined. Four VCDX defense responses prepared.
Common Errors
Final Validation
Complete VCF upgrade failure remediation with rollback procedures, prevention strategy, and VCDX defense
✓ Upgrade sequence mastered → 19-step VCF upgrade order documented with dependency justification
✓ Failure diagnosis capability → Log analysis for SDDC Manager, vCenter, NSX, and ESXi upgrade failures
✓ Rollback procedures tested → vCenter snapshot revert, SDDC Manager restore, NSX cluster recovery documented
✓ Prevention strategy designed → Pre-upgrade checklist, Holodeck testing, maintenance window planning
✓ VCDX defense prepared → Four defense responses with lifecycle management expertise
Cleanup / Restore
• Verify all VCF components are running at consistent versions
• Remove any temporary snapshots taken during practice
• Confirm SDDC Manager health dashboard is green
• Revert to snapshot if needed
Design Reflection (VCDX)
VCF upgrade lifecycle management demonstrates operational maturity and risk management. The design connects upgrade strategy to business continuity requirements with structured maintenance windows, Holodeck pre-validation, and graduated rollback procedures. The 19-step upgrade sequence is a core VCF operational competency.
Requirements
- R-001: VCF upgrades must complete within defined maintenance windows
- R-002: Upgrade rollback must be possible within 2 hours of failure detection
- R-003: Zero data loss during upgrade operations
Constraints
- Upgrade sequence must follow KB 390634 order
- Maintenance windows limited to weekends (business constraint)
- Holodeck test environment must match production version baseline
Assumptions
- VM snapshots taken before every component upgrade
- SDDC Manager and NSX Manager backups available and tested
- Operations team trained on rollback procedures
Risks
- Product bug in upgrade path causes non-retryable failure — mitigate with Holodeck pre-testing and VMware SR readiness
- Maintenance window exceeded due to unexpected failure — time-boxed rollback decision prevents extended outage
- Partial upgrade state left across multiple windows — verify interoperability matrix and schedule remaining steps promptly
Self-Assessment Discussion Prompts
- How would you handle a VCF upgrade in a stretched cluster environment where both sites must be upgraded?
- What is your strategy for upgrading 10+ workload domains efficiently?
- How do you decide between patch-level updates and full version upgrades?
- What metrics would you track to measure upgrade program maturity?
Extensions
Automated Pre-Upgrade Validation
Build a PowerCLI or Python script that runs the complete pre-upgrade checklist automatically: DNS/NTP validation, component health, vSAN capacity, snapshot verification, backup verification. Output a go/no-go report.
VCF Upgrade with Async Patch Tool
Practice using the VCF Async Patch Tool for applying emergency patches outside the normal upgrade cycle. Document: when to use Async Patch vs. full upgrade, how to validate async patch compatibility, and rollback procedures.
Multi-Domain Upgrade Orchestration
Design an upgrade plan for a VCF environment with 1 management domain and 5 workload domains. Calculate total maintenance window requirements, stagger domain upgrades, and design the verification matrix for cross-domain compatibility.
⚠ Known Pitfalls (from Community KB)
References
- Broadcom KB 390634 — VCF Upgrade Sequence and Component Order
- VMware VCF 9.0 Lifecycle Management Guide — Upgrade Planning and Execution
- vCenter Server 9.0 Upgrade Guide — VCSA Migration and Rollback
- NSX 9.0 Upgrade Guide — Manager Cluster Upgrade Procedures
- VMware VCF Async Patch Tool Documentation