Academy/VCF 9.0 Support (2V0-15.25)/Upgrade Failure & Remediation
This lab targets VCF 9.0

Upgrade Failure & Remediation

VCF 9.0Advancedvcp-foundationvcp-support⏱ 150 min

VCF lifecycle upgrade failure scenarios with rollback, log analysis, and SDDC Manager workflow recovery.

Objectives

  • Understand VCF upgrade sequence and component dependencies (19-step order per KB 390634)
  • Diagnose upgrade failures using SDDC Manager logs, vCenter upgrade logs, and NSX upgrade status
  • Execute vCenter upgrade rollback via VM snapshot restoration
  • Recover failed SDDC Manager upgrade workflows using API retry and manual intervention
  • Design pre-upgrade validation checklists to prevent common upgrade failures
  • Document upgrade failure recovery using RCAR methodology

Prerequisites

VCF lab with management domain running pre-upgrade version

Prior labs: vcp-support-01

Required skills:

  • SDDC Manager UI and API navigation
  • vCenter snapshot management
  • SSH access to VCF components

Lab Environment

VCF management domain with SDDC Manager, vCenter, NSX Manager cluster, and 4 ESXi hosts

Tasks

Task 1 VCF Upgrade Architecture & Sequence Planning

Foundation for upgrade troubleshooting. Panelists probe: Why this specific order? What happens if you skip a step? Which components can be rolled back independently?

Master the 19-step VCF upgrade sequence, understand component interdependencies, and identify which steps are most likely to fail and why.

Step 1

Document the VCF 9.0 upgrade sequence per Broadcom KB 390634. The 19 steps in order: (1) SDDC Manager, (2) vRealize Suite Lifecycle Manager, (3) Workspace ONE Access, (4) vRealize/Aria Automation, (5) vRealize/Aria Operations, (6) vRealize/Aria Operations for Logs, (7) vRealize/Aria Operations for Networks, (8) vCenter Server (management domain), (9) NSX Manager cluster (management domain), (10) ESXi hosts (management domain), (11) vSAN on-disk format (management domain), (12-15) Repeat steps 8-11 for each workload domain, (16) HCX, (17) vSAN Witness, (18) SDDC Manager drift reconciliation, (19) Post-upgrade validation.

19-step sequence documented. Key dependency: SDDC Manager must upgrade first because it orchestrates all subsequent upgrades. vCenter before NSX because NSX depends on vCenter inventory. ESXi before vSAN ODF because new on-disk format requires new ESXi version.
The upgrade sequence is the most frequently tested topic in VCF support exams. Memorize the order and understand WHY each step must precede the next.
Step 2

Identify the most failure-prone upgrade steps and common causes. Document: (a) SDDC Manager upgrade — database migration issues, insufficient disk space; (b) vCenter upgrade — VCSA migration assistant failures, customization conflicts; (c) NSX Manager cluster upgrade — cluster stability during rolling upgrade, backup controller issues; (d) ESXi host upgrade — vSAN maintenance mode failures, incompatible hardware/drivers.

Failure matrix: Step | Common Failure | Root Cause | Prevention. SDDC Manager: DB migration = clean up stale data pre-upgrade. vCenter: migration assistant = verify DNS and NTP pre-upgrade. NSX: cluster instability = ensure 3-node cluster healthy pre-upgrade. ESXi: maintenance mode = verify vSAN capacity allows data evacuation.
80% of upgrade failures are preventable with proper pre-upgrade validation. The remaining 20% are product bugs requiring VMware SR escalation.
Step 3

Map the rollback options for each component. vCenter: revert VM snapshot (fastest, most reliable). SDDC Manager: restore from backup (more complex, requires database consistency). NSX Manager: restore from backup or rebuild (most complex — 3-node cluster coordination). ESXi: cannot easily downgrade — revert at vSAN cluster level or reinstall. Document the rollback time estimate for each.

Rollback matrix: vCenter snapshot revert (15-30 min), SDDC Manager backup restore (30-60 min), NSX Manager backup restore (60-120 min per node), ESXi downgrade (not supported — plan forward or rebuild). Key insight: always take VM snapshots before every component upgrade.
VCDX defense: the rollback strategy must be planned BEFORE the upgrade starts. Panelists ask: What is your rollback plan for step 8 failure? Answer with specific RTO and procedure.
Step 4

Review the pre-upgrade validation checklist. Items: (a) verify bundle compatibility matrix in SDDC Manager, (b) check all VCF component health (SDDC Manager dashboard), (c) verify vSAN health and capacity (must be below 80% for host upgrades), (d) confirm DNS forward/reverse resolution for all components, (e) verify NTP synchronization across all components (<5 second drift), (f) take VM snapshots of vCenter, SDDC Manager, NSX Managers, (g) export NSX Manager backup, (h) document current versions of all components.

Pre-upgrade checklist with 15+ items, each mapped to the upgrade failure it prevents. DNS/NTP issues cause 30% of upgrade failures. Snapshot/backup omissions make rollback impossible.
Create a pre-upgrade checklist template. Run it every time — even for minor patches. The checklist catches issues that are obvious in hindsight but easy to miss under upgrade pressure.

Validation Gate

Check: Complete VCF upgrade sequence and pre-upgrade checklist documentation

Expected: 19-step sequence documented with dependency justification. Failure matrix and rollback matrix created. Pre-upgrade checklist with 15+ items.

Common Errors

Attempting to upgrade vCenter before SDDC Manager — SDDC Manager must be upgraded first to orchestrate subsequent steps
Skipping VM snapshots before component upgrades — makes rollback impossible if upgrade fails
Not verifying vSAN capacity before host upgrades — maintenance mode fails if capacity is insufficient for data evacuation
Ignoring DNS/NTP pre-checks — account for 30% of upgrade failures

Task 2 Upgrade Failure Diagnosis & Log Analysis

Core troubleshooting skill. Panelists ask: Where do you look first when an upgrade fails? How do you distinguish a transient error from a blocking issue? When do you escalate to VMware support?

Develop systematic log analysis skills for diagnosing VCF upgrade failures — from SDDC Manager workflow failures to vCenter migration issues to ESXi remediation errors.

Step 1

Map the critical log locations for VCF upgrade troubleshooting. SDDC Manager: /var/log/vmware/vcf/sddc-manager-ui-app/sddc-manager-ui-app.log (UI workflow), /var/log/vmware/vcf/operationsmanager/operations-manager.log (orchestration). vCenter upgrade: /var/log/vmware/upgrade/upgrade-runner.log. NSX Manager: /var/log/proton/nsxapi.log. ESXi remediation: /var/log/vmware/vcf/lcm/lcm-debug.log.

Each component has specific log files for upgrade operations. SDDC Manager operations-manager.log is the primary orchestration log — it shows the workflow state machine and identifies which step failed.
Start with SDDC Manager operations-manager.log for ANY VCF upgrade failure — even if the failure appears to be in vCenter or NSX, SDDC Manager orchestrates the workflow and captures the triggering error.
Step 2

Analyze a simulated vCenter upgrade failure. Common scenario: VCSA upgrade fails at database migration step. Log analysis: search upgrade-runner.log for 'ERROR' and 'FAILED'. The error typically shows: database schema migration failure, disk space exhaustion, or service dependency conflict. Document the error pattern and correlate with pre-upgrade checklist items.

Error patterns: 'Database migration failed' = insufficient /storage/db space or corrupted VCDB entries. 'Service start failed' = dependency service (SSO/STS) not healthy. 'Migration assistant failed' = DNS resolution failure between source and target VCSA.
For vCenter upgrade failures, the error in upgrade-runner.log points to the root cause 90% of the time. The remaining 10% require SDDC Manager log correlation.
Step 3

Analyze SDDC Manager workflow failure. Navigate to SDDC Manager UI > Lifecycle Management > view failed workflow. Document: workflow ID, failed step, error message, and retry eligibility. Via API: GET /v1/tasks/{taskId} to retrieve detailed error information. Check if the workflow can be retried (some failures are retryable, others require manual intervention).

SDDC Manager workflows have states: IN_PROGRESS, SUCCESSFUL, FAILED, CANCELLED. Failed workflows show the specific sub-task that failed. Retryable failures: transient network issues, temporary service unavailability. Non-retryable: version incompatibility, missing prerequisites.
SDDC Manager API is your best friend for upgrade troubleshooting. The UI shows summary information; the API provides detailed error context including stack traces.
Step 4

Practice ESXi host upgrade failure diagnosis. Common scenario: host enters maintenance mode but vSAN data evacuation stalls. Diagnose: (a) check vSAN resync status (objects still syncing from previous operation), (b) verify cluster capacity allows data migration, (c) check for VM anti-affinity rules preventing evacuation, (d) verify no VMs are pinned to the host. Document remediation for each cause.

ESXi upgrade failure during maintenance mode entry: vSAN evacuation stall is the most common cause. Remediation: wait for pending resyncs to complete, verify capacity, check DRS constraints. If a specific VM cannot migrate, check affinity rules and resource reservations.
Never force maintenance mode on a vSAN host — this bypasses data evacuation and risks data loss. Always resolve the evacuation blocker before proceeding.
Step 5

Build an upgrade failure escalation decision tree. Level 1: retry the failed workflow (if retryable). Level 2: analyze logs, fix the identified issue, retry. Level 3: rollback the failed component (snapshot/backup restore), fix root cause, re-attempt. Level 4: escalate to VMware support with: support bundle, upgrade logs, and detailed failure description.

Escalation decision tree with 4 levels, each with specific criteria for escalation to the next level. Include time-boxing: Level 1 (5 min), Level 2 (30 min), Level 3 (60 min), Level 4 (open SR). Collect support bundle at Level 2 for VMware SR preparation.
Time-boxing escalation levels prevents excessive troubleshooting during maintenance windows. If Level 2 analysis doesn't resolve in 30 minutes, proceed to rollback rather than extending the outage.

Validation Gate

Check: Complete upgrade failure diagnosis with log analysis and escalation procedures

Expected: Log locations mapped for all VCF components. Failure patterns documented. Escalation decision tree created with time-boxing.

Common Errors

Not starting with SDDC Manager operations-manager.log — it is the orchestration log that captures all workflow failures
Retrying failed workflows without fixing the root cause — same failure recurs
Forcing ESXi maintenance mode when vSAN evacuation stalls — risks data loss
Not collecting support bundle before attempting fixes — evidence lost for VMware SR

Task 3 Rollback Procedures & Recovery Operations

Critical operational skill. Panelists ask: What is your rollback procedure for each component? How do you ensure consistency after partial upgrade? What is the maximum rollback window?

Execute component-level rollback procedures for failed VCF upgrades — from vCenter snapshot revert to SDDC Manager backup restore to NSX cluster recovery.

Step 1

Practice vCenter rollback via snapshot revert. Steps: (1) Power off the VCSA VM, (2) Revert to pre-upgrade snapshot in ESXi host client (not vSphere client — vCenter is down), (3) Power on the VCSA VM, (4) Wait for all services to start (5-15 minutes), (5) Verify vCenter is responsive and inventory is intact, (6) Notify SDDC Manager of rollback (may need to reconcile state).

vCenter snapshot revert takes 15-30 minutes total. After revert, vCenter returns to pre-upgrade version. SDDC Manager may show version mismatch — use SDDC Manager API to update component version tracking. ESXi hosts reconnect automatically to the restored vCenter.
vCenter snapshot revert is the fastest and most reliable rollback option. Always take the snapshot AFTER stopping vCenter services and BEFORE starting the upgrade — this ensures a clean snapshot state.
Step 2

Practice SDDC Manager rollback from backup. Steps: (1) Power off the SDDC Manager VM, (2) Deploy a fresh SDDC Manager OVA from the original version media, (3) Restore from backup using the SDDC Manager restore wizard, (4) Verify SDDC Manager UI is accessible, (5) Verify all domains and hosts appear in inventory, (6) Reconcile any state differences.

SDDC Manager restore takes 30-60 minutes. Backup includes: database (PostgreSQL), configuration files, and certificate state. After restore, SDDC Manager must re-establish communication with all managed components (vCenter, NSX, ESXi hosts).
SDDC Manager rollback is more complex than vCenter because it is the orchestration layer. After restore, verify all managed component versions match SDDC Manager's expected state.
Step 3

Practice NSX Manager cluster recovery after failed upgrade. Scenario: one NSX Manager node in the 3-node cluster fails during upgrade. Steps: (1) Verify remaining 2 nodes are healthy (nsxcli -c 'get cluster status'), (2) Remove the failed node from the cluster, (3) Deploy a fresh NSX Manager node at the target version, (4) Join the new node to the cluster, (5) Verify cluster stability (all 3 nodes STABLE). Document the timeline.

NSX Manager cluster recovery: 60-120 minutes. During recovery with 2 of 3 nodes, NSX management plane operates in degraded mode (no HA). DFW rules continue enforcing on ESXi hosts regardless of NSX Manager state. New node join triggers data sync from surviving nodes.
NSX Manager cluster is designed for rolling upgrade — only one node upgrades at a time. If a node fails, the remaining 2 nodes maintain management plane availability. Never attempt to upgrade a second node while the first is in failed state.
Step 4

Handle the partial upgrade scenario: vCenter upgraded successfully but NSX upgrade failed. Now the environment has mixed versions (vCenter 9.0, NSX 4.x). Document: (a) is this a supported mixed-version state? Check VMware Interoperability Matrix. (b) Can you proceed with NSX upgrade retry? (c) Should you rollback vCenter to match NSX? Analyze the decision tree.

Partial upgrade state assessment: VMware Interoperability Matrix defines which version combinations are supported. vCenter N+1 with NSX N is typically supported for a limited upgrade window. Decision: retry NSX upgrade if failure was transient; rollback vCenter if NSX version is fundamentally incompatible.
VCDX design point: upgrade sequence design must account for partial failure states. Document the supported mixed-version windows and the decision criteria for retry vs. rollback.
Step 5

Design post-rollback validation procedure. After any rollback: (1) Verify all VCF component versions match expected state, (2) Run SDDC Manager health check, (3) Verify vSAN cluster health, (4) Verify NSX controller cluster stability, (5) Run DFW Traceflow tests for critical paths, (6) Verify all VMs are running and accessible, (7) Test VCF operations (VM provisioning, host add). Document as a runbook.

Post-rollback validation checklist with 7+ items ensuring environment integrity. Each item has specific commands/tests and expected results. All checks must pass before declaring rollback successful.
Post-rollback validation is as important as the rollback itself. A 'successful' rollback that leaves orphaned objects or broken state is worse than the original failure.

Validation Gate

Check: Execute rollback procedures and validate environment recovery

Expected: vCenter snapshot revert practiced, SDDC Manager restore documented, NSX cluster recovery understood, post-rollback validation checklist created.

Common Errors

Reverting vCenter snapshot while VMs are still referencing new-version features — can cause VM configuration issues
Not reconciling SDDC Manager state after vCenter rollback — SDDC Manager shows version mismatch
Attempting to upgrade second NSX Manager node when first is in failed state — risks losing cluster quorum
Skipping post-rollback validation — orphaned objects and broken state go undetected

Task 4 Upgrade Prevention Strategy & VCDX Defense

Synthesis task connecting upgrade operations to architectural design. Panelists expect: upgrade strategy document, maintenance window planning, and risk mitigation.

Design a comprehensive upgrade strategy that minimizes failure risk and prepare to defend VCF lifecycle management design decisions in VCDX context.

Step 1

Design a VCF upgrade strategy document. Include: (a) upgrade policy (frequency: quarterly patches, semi-annual minor, annual major), (b) maintenance window schedule (management domain: weekend 1, workload domain 1: weekend 2, workload domain 2: weekend 3), (c) pre-upgrade checklist reference, (d) rollback decision criteria (time-boxed: if not resolved in 2 hours, rollback), (e) communication plan (stakeholders notified 2 weeks before, day-of status updates).

Upgrade strategy document with 5 sections covering policy, scheduling, preparation, execution, and communication. Each section has specific templates and decision criteria.
A well-documented upgrade strategy demonstrates operational maturity. VCDX panelists specifically ask about lifecycle management — show that your design includes day-2 operations, not just day-0 deployment.
Step 2

Calculate maintenance window requirements. For each upgrade step, estimate: preparation time, execution time, verification time, and rollback buffer. Total management domain upgrade: pre-checks (2h) + SDDC Manager (1h) + vCenter (2h) + NSX (3h) + ESXi hosts (4h for rolling upgrade of 4 hosts) + validation (1h) = ~13 hours + 2h rollback buffer = 15-hour maintenance window.

Maintenance window calculation showing each step with time estimate. Total management domain upgrade: 13-15 hours. Each workload domain: 6-8 hours (vCenter + NSX + ESXi only). Stagger across weekends to limit blast radius.
Maintenance window planning is a VCDX design deliverable. Show the math: hours per component, rollback buffer, verification time. Under-estimating maintenance windows causes rushed upgrades and increased failure risk.
Step 3

Design upgrade testing strategy using Holodeck. Steps: (1) Deploy Holodeck VCF environment matching production version, (2) Run the complete upgrade in Holodeck first, (3) Document any issues encountered and resolution, (4) Validate upgrade success in Holodeck, (5) Use Holodeck results to refine production upgrade plan. Document the value of pre-production upgrade testing.

Holodeck upgrade testing catches product bugs, validates procedures, and trains operations team before production upgrade. Investment: 4-8 hours of testing saves potential multi-day outage from untested upgrade.
Holodeck pre-production testing is your best upgrade risk mitigation tool. VCDX panelists ask: How do you validate your upgrade procedure? 'We test in Holodeck first' is a strong answer.
Step 4

Prepare four VCDX defense responses: (1) 'Why do you upgrade management domain before workload domains?' — SDDC Manager orchestrates all upgrades; it must be at the latest version to manage the upgrade workflow for workload domains. vCenter and NSX in management domain are upgraded next because workload domain upgrades depend on management domain health. (2) 'What is your rollback strategy for a failed ESXi host upgrade?' — ESXi host rollback via Alt-F11 boot bank switch (if dual-bank supported). If not, rebuild host from image profile and rejoin cluster. vSAN rebuild handles data redistribution. (3) 'How do you handle a upgrade that spans multiple maintenance windows?' — Document the partial upgrade state, verify component interoperability using VMware matrix, schedule remaining steps in next window. Never leave an environment in untested partial upgrade state indefinitely. (4) 'What is the most critical pre-upgrade step?' — VM snapshot of ALL VCF management components. Without snapshots, rollback requires full rebuild. This one step reduces rollback RTO from days to minutes.

Four defense responses demonstrating lifecycle management expertise and operational maturity.
VCDX defense for upgrades: demonstrate that you plan for failure, not just success. The rollback strategy is as important as the upgrade strategy.

Validation Gate

Check: Complete upgrade strategy document and VCDX defense preparation

Expected: Upgrade strategy with policy, scheduling, and communication plan. Maintenance window calculations documented. Holodeck testing strategy defined. Four VCDX defense responses prepared.

Common Errors

Underestimating maintenance window duration — rushed upgrades increase failure risk
Not testing upgrades in Holodeck first — production becomes the test environment
Leaving environments in partial upgrade state across multiple weeks — untested mixed-version states accumulate risk
Not including rollback buffer in maintenance window — no time for recovery if upgrade fails

Final Validation

Complete VCF upgrade failure remediation with rollback procedures, prevention strategy, and VCDX defense

✓ Upgrade sequence mastered → 19-step VCF upgrade order documented with dependency justification

✓ Failure diagnosis capability → Log analysis for SDDC Manager, vCenter, NSX, and ESXi upgrade failures

✓ Rollback procedures tested → vCenter snapshot revert, SDDC Manager restore, NSX cluster recovery documented

✓ Prevention strategy designed → Pre-upgrade checklist, Holodeck testing, maintenance window planning

✓ VCDX defense prepared → Four defense responses with lifecycle management expertise

Cleanup / Restore

• Verify all VCF components are running at consistent versions

• Remove any temporary snapshots taken during practice

• Confirm SDDC Manager health dashboard is green

• Revert to snapshot if needed

Design Reflection (VCDX)

VCF upgrade lifecycle management demonstrates operational maturity and risk management. The design connects upgrade strategy to business continuity requirements with structured maintenance windows, Holodeck pre-validation, and graduated rollback procedures. The 19-step upgrade sequence is a core VCF operational competency.

Requirements

  • R-001: VCF upgrades must complete within defined maintenance windows
  • R-002: Upgrade rollback must be possible within 2 hours of failure detection
  • R-003: Zero data loss during upgrade operations

Constraints

  • Upgrade sequence must follow KB 390634 order
  • Maintenance windows limited to weekends (business constraint)
  • Holodeck test environment must match production version baseline

Assumptions

  • VM snapshots taken before every component upgrade
  • SDDC Manager and NSX Manager backups available and tested
  • Operations team trained on rollback procedures

Risks

  • Product bug in upgrade path causes non-retryable failure — mitigate with Holodeck pre-testing and VMware SR readiness
  • Maintenance window exceeded due to unexpected failure — time-boxed rollback decision prevents extended outage
  • Partial upgrade state left across multiple windows — verify interoperability matrix and schedule remaining steps promptly

Self-Assessment Discussion Prompts

  1. How would you handle a VCF upgrade in a stretched cluster environment where both sites must be upgraded?
  2. What is your strategy for upgrading 10+ workload domains efficiently?
  3. How do you decide between patch-level updates and full version upgrades?
  4. What metrics would you track to measure upgrade program maturity?

Extensions

Automated Pre-Upgrade Validation

Build a PowerCLI or Python script that runs the complete pre-upgrade checklist automatically: DNS/NTP validation, component health, vSAN capacity, snapshot verification, backup verification. Output a go/no-go report.

VCF Upgrade with Async Patch Tool

Practice using the VCF Async Patch Tool for applying emergency patches outside the normal upgrade cycle. Document: when to use Async Patch vs. full upgrade, how to validate async patch compatibility, and rollback procedures.

Multi-Domain Upgrade Orchestration

Design an upgrade plan for a VCF environment with 1 management domain and 5 workload domains. Calculate total maintenance window requirements, stagger domain upgrades, and design the verification matrix for cross-domain compatibility.

⚠ Known Pitfalls (from Community KB)

Upgrading vCenter before SDDC Manager — SDDC Manager must upgrade first to orchestrate subsequent components; out-of-order upgrade breaks the workflow engine
Not taking VM snapshots before every component upgrade — makes rollback impossible; snapshot revert is the fastest and most reliable rollback method
Forcing ESXi host into maintenance mode when vSAN evacuation stalls — risks data loss; diagnose and resolve the evacuation blocker first
Not time-boxing upgrade troubleshooting — spending 4+ hours diagnosing a failure during a maintenance window when rollback takes 30 minutes; define rollback decision criteria upfront

References

  • Broadcom KB 390634 — VCF Upgrade Sequence and Component Order
  • VMware VCF 9.0 Lifecycle Management Guide — Upgrade Planning and Execution
  • vCenter Server 9.0 Upgrade Guide — VCSA Migration and Rollback
  • NSX 9.0 Upgrade Guide — Manager Cluster Upgrade Procedures
  • VMware VCF Async Patch Tool Documentation
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.