VCF 5.2 to 9.0 Upgrade Planning Workshop
Objectives
- Plan a production VCF 5.2 to 9.0 upgrade with pre-flight assessment, compatibility validation, and component sequencing
- Execute upgrade steps in Holodeck lab environment following the rigid upgrade order (SDDC Manager → vCenter → NSX → ESXi)
- Handle a simulated failure (NSX upgrade stalls) and demonstrate recovery/rollback procedures
- Validate post-upgrade services, workload connectivity, and new VCF 9.0 features (VCF Operations UI)
Prerequisites
VCF 5.2 lab environment deployed and operational (ideally holodeck-02 with management domain; or holodeck-03+ with workload domains). Alternatively, design-exercise mode using VCF Upgrade Planner documentation and upgrade runbooks.
Prior labs: holodeck-02 (management domain deployment), vcf-evolution-01 (architectural understanding of 5.x vs 9.0)
Required skills:
- VCF operational management (SDDC Manager, vCenter, NSX administration)
- Change management (maintenance windows, communication, rollback planning)
- Certificate and DNS management (renewal, validation, drift detection)
- Storage and networking troubleshooting (vSAN health, NSX connectivity, VLAN isolation)
- Failure diagnosis and root cause analysis (logs, monitoring, escalation paths)
Lab Environment
Holodeck lab environment: 1 management domain (vCenter, SDDC Manager, NSX, vSAN 3-node cluster minimum). Alternatively, documentation-based planning exercise using VCF Upgrade Planner workflow.
graph TB Client[SDDC Client/Upgrade Tool] -->|HTTPS| SDDCMgr[SDDC Manager 5.2] SDDCMgr -->|HTTPS| vCenter[vCenter 5.x] SDDCMgr -->|HTTPS| NSXMgr[NSX Manager 5.x] vCenter -->|mgmt| ESXi1[ESXi 1] vCenter -->|mgmt| ESXi2[ESXi 2] vCenter -->|mgmt| ESXi3[ESXi 3] NSXMgr -->|overlay| ESXi1 NSXMgr -->|overlay| ESXi2 NSXMgr -->|overlay| ESXi3 ESXi1 -->|heartbeat| vSAN[(vSAN Cluster)] ESXi2 -->|heartbeat| vSAN ESXi3 -->|heartbeat| vSAN
Tasks
Task 1 Pre-Upgrade Assessment: Validate HCL, Certificates, NTP/DNS, and Component Versions
reliabilityUpgrade failures are rarely due to the upgrade process itself — they're due to pre-existing environment issues (clock skew, certificate expiry, incompatible hardware, outdated DNS records). VCDX candidates must be methodical about pre-flight validation. This task forces you to think like a change manager: 'What can go wrong, and how do I prevent it?'
Run VCF Upgrade Planner (if available) or manually download and review the HCL (Hardware Compatibility List) for VCF 9.0.2. Verify each ESXi host in the cluster is on the Supported HCL. Document: [Host Model | ESXi Version | CPU SKU | RAM | Supported in 9.0?]
Check SSL certificate expiry dates on all VCF components: SDDC Manager, vCenter, NSX Manager, each ESXi host. Command: 'echo | openssl s_client -connect <ip>:443 2>/dev/null | openssl x509 -noout -dates'. Document: [Component | IP | Issuer | Expiry Date | Days Until Expiry | Action Needed?]
Verify certificate SANs (Subject Alternate Names) match FQDN and IP addresses used in upgrade. Example: SDDC Manager cert should have CN=sddc-manager.lab.local AND SAN=sddc-manager.lab.local, 10.0.1.10. Document discrepancies.
Verify NTP synchronization on all components (SDDC Manager, vCenter, NSX Manager, all ESXi hosts). Command: 'ntpstat' or 'timedatectl' (Linux) / 'w32tm /query /status' (Windows). Document: [Component | NTP Server | Offset (ms) | Sync Status]
Verify DNS resolution in both directions. Test: 'nslookup sddc-manager.lab.local' (forward), 'nslookup 10.0.1.10' (reverse). Verify DNS used by upgrade process is consistent. Document: [FQDN | IP | Forward Resolution | Reverse Resolution | DNS Server Used]
Document current component versions: SDDC Manager 5.x.x, vCenter 5.x.x, NSX Manager 5.x.x, vSAN driver version, ESXi versions (all should be aligned). Use SDDC Manager API: 'curl -k https://sddc-manager/api/v1/system/version' or UI. Document: [Component | Current Version | Supported in 9.0? | Pre-req for 9.0?]
Validate current vSAN cluster health: 'Cluster > Configuration > vSAN > Cluster Information'. Check: Cluster health = Healthy, no unsynced objects, no degraded disks, all hosts online. Document: [Metric | Current State | Healthy?] Example: Sync Status (Synced/Resyncing), Physical Memory Allocated (%), Disk Errors (0?)
Create a pre-upgrade snapshot/baseline: Holodeck: 'New-Snapshot -VM <sddc>, <vcenter>, <nsx> -Name pre-upgrade-baseline'. Production: Create backups of SDDC Manager and vCenter VMs. Document snapshot/backup timestamps and verify restoration is possible.
Synthesize pre-flight findings into a checklist: [Item | Status | Action If Failed | Owner | Target Completion]. Include HCL, certificates, NTP, DNS, versions, vSAN health, backups. Mark as 'Ready for Upgrade' only when all items are GREEN.
Validation Gate
Check: Pre-flight checklist completed and signed off; no blockers remaining
Expected: All 9 substeps validated; environment is confirmed ready for VCF 9.0 upgrade
Common Errors
Task 2 Plan Upgrade Sequence, Maintenance Window, and Rollback Strategy
manageabilityVCF upgrades have a rigid component order: SDDC Manager first (it orchestrates the rest), then vCenter, then NSX, then ESXi via LCM. Deviating from this order causes cascading failures. VCDX candidates must understand this sequencing deeply and plan for failure scenarios — 'If NSX upgrade fails, can we rollback without losing the SDDC Manager and vCenter upgrades?'
Document the upgrade sequence in tabular form: [Step | Component | Version | Task | Estimated Time | Rollback Point?] Example:
Step 1: SDDC Manager 5.2.1 → 9.0.2 (in-place upgrade, no rollback to 5.2 once complete) Step 2: vCenter 5.x → 9.0.2 (in-place upgrade, no rollback) Step 3: NSX Manager 5.x → 9.0.2 (rolling upgrade, edge nodes upgraded sequentially) Step 4: ESXi hosts 5.x → 9.0.2 (via LCM, one host at a time with vMotion)
Calculate total upgrade window duration: Pre-flight (1 hour) + SDDC Manager (1.5 hours) + vCenter (1 hour) + NSX (2 hours) + ESXi LCM (1 hour per host × 3 hosts = 3 hours) + Validation (1 hour) = ~9.5 hours total. Add 30% contingency buffer = ~12 hours. Document: [Phase | Duration | Rollback Possible?] and overall maintenance window time.
Plan stakeholder communication: Email to business owners at T-7 days (announcement), T-1 day (reminder with start/end times, contact info), T0 (go-live notification), post-upgrade (validation results). Document each communication: [Audience | Message | Timing | Owner]
Define rollback strategy for each component: SDDC Manager (restore from snapshot/backup, requires ~2 hours), vCenter (similar, ~1.5 hours), NSX (rolling rollback of edge nodes, ~2 hours), ESXi (revert LCM update or use previous snapshot, ~30 min per host). Document: [Component | Rollback Method | Estimated Time | Blast Radius | Data Loss Risk]
Identify failure scenarios and decision points: 'If SDDC Manager upgrade fails at 30%, do we proceed or rollback?' (Answer: Rollback, because SDDC Manager must be stable before vCenter upgrade). 'If vCenter upgrade hangs at 50%, how long do we wait before calling rollback?' (Answer: 1 hour; if not progressing, assume stuck). Document: [Scenario | Decision Point | Condition | Action | Responsible]
Plan resource allocation: Who is on call? Who is the primary upgrade engineer, secondary, escalation contact? What's the communication channel (Slack, Zoom, phone)? Document: [Role | Name/Team | Contact | Responsibilities During Upgrade | Availability (hours)]
Plan validation checkpoints: After SDDC Manager upgrade completes (check API endpoints), after vCenter (verify vSphere Client login), after NSX (verify logical networks), after ESXi (verify workload connectivity). Document: [Checkpoint | Component | Validation Test | Expected Result | Time Estimate]
Create a 1-page summary for executives: 'VCF Upgrade Plan: 5.2 → 9.0 on [DATE]. 12-hour maintenance window. Rollback available if upgrade fails before 50% complete; beyond that point, proceed or accept downtime. Risk: Low (multi-component redundancy ensures no data loss, but services will be down 12 hours).' Include go/no-go decision criteria.
Validation Gate
Check: Upgrade plan (sequencing, timeline, communication, rollback) is complete and approved
Expected: Ready to proceed to Task 3 (execution); all stakeholders informed and decision rights established
Common Errors
Task 3 Execute Upgrade in Holodeck Lab: SDDC Manager, vCenter, NSX, ESXi with Simulated Failure Handling
reliabilityTheory is useless without practice. This task walks you through each upgrade step, validating at each stage. Critically, you'll encounter a simulated failure (NSX upgrade stalls) and demonstrate how to diagnose and recover — a real skill VCDX panelists value.
Pre-upgrade validation (from Task 1): Verify pre-flight checklist is GREEN. Take baseline snapshot: 'New-Snapshot -VM vcf-sd-<mgmt-domain> -Name pre-upgrade-baseline'. Document snapshot times. Declare maintenance window open.
SDDC Manager upgrade: Download VCF 9.0.2 SDDC Manager OVA. In SDDC Manager UI, navigate to Upgrade > Available Patches. (Or via PowerShell: Update-SddcManager -Version 9.0.2). Monitor progress. Expected: 30-45 min. Validate: SDDC Manager API returns 9.0.2, UI is responsive, no errors in logs (tail /var/log/upgrade.log).
vCenter upgrade: Trigger vCenter 9.0.2 update via SDDC Manager Upgrade UI or vCenter Update Manager. Monitor progress (vCenter will briefly reboot). Expected: 1-1.5 hours. Validate: vSphere Client logs in, vCenter cluster shows all hosts connected, no errors in vCenter logs.
NSX Manager upgrade: Trigger NSX Manager upgrade via SDDC Manager or NSX UI. This step updates NSX Manager, then rolling upgrade of Edge nodes. Expected: 1.5-2 hours. Validate after NSX Mgr completes: NSX API 'GET /api/v1/infra/domains' works, logical networks are intact.
SIMULATED FAILURE: NSX edge node upgrade stalls. Symptom: LCM shows 'Edge-01 upgrade: 45% complete. No progress for >10 min.' Troubleshoot: SSH into edge node (ssh admin@edge-01), check: 'systemctl status nsxd' (running?), 'journalctl -xe' (errors?). Try: 'systemctl restart nsxd'. If still stuck after 2 min, escalate.
If edge node is unrecoverable, demonstrate rollback: Rollback that specific edge from 9.0 to 5.x (via LCM: Rollback Node). Verify: That edge continues to serve traffic at 5.x version (mixed version cluster is unsupported long-term but acceptable for recovery). Document decision and timeline (this added 30 min to upgrade).
ESXi upgrade via LCM: After NSX completes (or is recovered), trigger ESXi upgrade via LCM. Upgrade one host at a time (cluster size 3, so 3 rounds). Each host: vMotion workloads to other hosts → upgrade ESXi → reboot → rejoin cluster → validate. Expected: 30 min per host × 3 = 1.5 hours total. Validate: Each host shows 9.0.2 version after reboot, cluster shows all hosts connected.
Post-upgrade immediate validation: SDDC Manager UI shows all components healthy (vCenter, NSX, ESXi all green). Check vSAN cluster: all hosts online, no resyncing objects. Verify workload connectivity: ping a test VM to external network (should work). Document: [Component | Status | Errors | Timestamp]
Record upgrade metrics: [Component | Start Time | End Time | Duration | Rollback Used? | Issues Encountered]. Compare to plan: Did actual times match estimates? What took longer? Document lessons learned.
Validation Gate
Check: All upgrade steps completed; simulated failure diagnosed and resolved; post-upgrade validation passed
Expected: VCF 5.2 → 9.0 upgrade successful with documented failure recovery; ready for Task 4 (comprehensive validation)
Common Errors
Task 4 Post-Upgrade Validation: Services, Workload Connectivity, and New VCF 9.0 Features
manageabilityThe upgrade isn't done until you verify everything works and new features are functional. This task ensures you know what to test (not just 'does the UI load?') and demonstrates mastery of end-to-end VCF operations.
Comprehensive service health check: vCenter Cluster Health (all hosts green), SDDC Manager System Status (all health checks pass), NSX System Status (all nodes healthy, controllers consensus reached), vSAN Cluster Health (all hosts operational, capacity balanced). Document: [Service | Status | Any Alarms?] Expected: All GREEN, 0 alarms.
Workload connectivity validation: Test VM-to-VM traffic on same segment (local traffic), VM-to-VM on different segments (L3 routing via T0/T1), VM-to-external (north-south via edge). Command: 'ping' from one VM to another, or use iperf for throughput. Document: [Traffic Flow | Source | Destination | Result | Latency/Throughput]
Verify new VCF 9.0 features are enabled: VCF Operations (Unified UI). Log into 'https://sddc-manager/vcf' and verify 'Cluster Operations' and 'Infrastructure' dashboards load. Check: [Feature | Accessible? | Data Populated?] Examples: Cluster Overview, Capacity, Compliance.
Verify NSX 4.1+ features (if applicable): Federation capability status, advanced routing (if configured), stretched segments (if multi-site). Document: [Feature | Availability | Configured? | Working?]
Certificate validation: Verify all SSL certificates are renewed/valid post-upgrade. Check expiry dates on SDDC Manager, vCenter, NSX Manager, each Edge. Command: 'openssl s_client -connect <ip>:443 2>/dev/null | openssl x509 -noout -dates'. Document: [Component | Issuer | Expiry | Days Until | Action Needed?]
Verify DNS and NTP post-upgrade: Ensure DNS still resolves correctly for all components (FQDNs, IPs). Ensure NTP offset is <100ms on all nodes. Command: 'nslookup <fqdn>', 'ntpstat'. Document: [Component | DNS Status | NTP Offset | Healthy?]
Run upgrade runbook to document: [Item | Pre-Upgrade State | Post-Upgrade State | Status | Notes]. Example: SDDC Manager version, vCenter version, NSX version, ESXi versions, vSAN cluster health, workload count, network segments count, policies, etc.
Close maintenance window: Send post-upgrade notification to stakeholders. Include: Upgrade completed successfully, services healthy, workloads accessible, any issues encountered and resolutions. Thank you message. Document: [Date/Time | Audience | Message Sent By]
Create post-upgrade documentation: Upgrade timeline, issues encountered (NSX edge stall) and resolution (restarted nsxd), lessons learned, any follow-up actions (e.g., 'Review NSX edge deployment for stability', 'Update runbooks for 9.0'). Suitable for knowledge base or team wiki.
Validation Gate
Check: All validation steps completed; VCF 9.0 upgrade fully validated and documented; services operational; stakeholders informed
Expected: Successful upgrade with comprehensive post-upgrade documentation; confidence in 9.0 environment; ready for production use
Common Errors
Final Validation
You have planned and executed a production-grade VCF 5.2 to 9.0 upgrade, including pre-flight validation, sequencing, failure handling, and comprehensive post-upgrade validation. You demonstrated the ability to diagnose a simulated failure (NSX edge stall) and recover without data loss. This is a VCDX-level operational skill — panelists will expect you to walk through an upgrade scenario and articulate contingency plans.
✓ Pre-upgrade assessment complete → HCL validation, certificate dates, NTP sync, DNS resolution, component versions, vSAN health verified; checklist signed off
✓ Upgrade plan (sequencing, timeline, communication, rollback) → Documented component order (SDDC Mgr → vCenter → NSX → ESXi), 12-hour maintenance window, stakeholder communication, rollback procedures, failure decision tree
✓ Upgrade execution with failure handling → All components upgraded successfully; simulated NSX edge failure diagnosed and recovered; post-upgrade services healthy
✓ Post-upgrade validation → All services green, workloads connected, new VCF 9.0 features accessible, certificates valid, DNS/NTP synchronized, documentation complete
✓ Runbook and lessons learned → Post-upgrade documentation recorded; timeline, issues, resolutions, and recommendations for future upgrades documented
Cleanup / Restore
• Archive upgrade logs and runbooks for audit trail and knowledge base
• Snapshot the VCF 9.0 environment as 'vcf-90-post-upgrade' for future labs
• Decommission pre-upgrade snapshots after 30-day retention (in production, follow ORO retention policy)
• Update runbooks for 9.0 (if any changes to procedures)
• Conduct post-upgrade retrospective (team meeting): What went well? What was harder than expected? How do we improve next time?
Design Reflection (VCDX)
A VCDX panelist will ask: 'Walk us through a VCF upgrade you planned or executed. What went wrong, and how did you recover?' Your answer should include: (1) Specific pre-flight validation you did (not generic 'checked health'). (2) The exact sequencing of components (SDDC Mgr first, then vCenter, then NSX, then ESXi) and why that order matters. (3) A concrete failure scenario and how you diagnosed it (logs, API calls, not just rebooting). (4) Decision-making under pressure (NSX edge failed; do you rollback the entire upgrade or just that component?). (5) Post-upgrade validation that goes beyond 'UI loads' (workload connectivity, new features, certificates, DNS/NTP).
Requirements
- Pre-flight validation complete: HCL compatibility, certificate expiry, NTP/DNS sync, component versions, vSAN health, backups
- Upgrade sequence documented: SDDC Manager → vCenter → NSX → ESXi via LCM, in strict order, with time estimates and rollback points
- Maintenance window scheduled: 12-hour block, communicated to stakeholders 7 days in advance, with go/no-go criteria
- Failure handling procedures: Decision tree for rollback vs. retry, thresholds for escalation, staffing plan with on-call engineers
- Simulated failure resolution: NSX edge upgrade stall diagnosed, recovered, and impact assessed
- Post-upgrade validation: Services health, workload connectivity, new features, certificates, DNS/NTP, documentation
- Upgrade runbook and lessons learned: Timeline, issues, resolutions, recommendations for next upgrade
Constraints
- VCF upgrade sequence is rigid — cannot parallelize components or deviate from order without cascading failures
- NSX edge upgrade is serial (one at a time); cannot upgrade all edges in parallel, which extends upgrade window
- Holodeck lab environment has single failure domain (one physical host); production environments have HA for SDDC Manager and vCenter, reducing downtime risk
- Rollback after SDDC Manager upgrade completes requires restore from backup; post-upgrade changes are lost during rollback
- Mixed-version clusters (some hosts on 5.x, some on 9.0) are unsupported for extended periods; ESXi hosts must be upgraded in quick succession
- No in-place upgrade path from VCF 3.x to 9.0; must upgrade to intermediate version (e.g., 5.2) first, then to 9.0
Assumptions
- All pre-flight checks pass (HCL, certificates, NTP, DNS, versions, backups); no blockers remain at upgrade start
- Maintenance window of 12 hours is acceptable to business; no critical workloads require 24/7 uptime during upgrade
- Network and storage are stable during upgrade; no planned maintenance on upstream switches, routers, or SAN during VCF upgrade window
- Upgrade operators are trained and experienced with VCF; they can follow runbook sequencing and troubleshoot failures without vendor support
- Rollback is an acceptable failure mode; business accepts brief downtime (2-4 hours) if upgrade fails and must be rolled back
- Post-upgrade, all workloads return to running state; no application data corruption or permanent loss of configuration
Risks
- Certificate expiry during upgrade → SSL connections fail, components can't communicate — MITIGATION: Validate certificate expiry 90 days in advance; renew if needed before upgrade start
- NTP drift during upgrade → Time-dependent services (Kerberos, SSL handshakes) fail — MITIGATION: Verify NTP sync on all nodes; fix drift >100ms before upgrade
- NSX edge upgrade hangs indefinitely → NSX services unavailable, workloads lose connectivity — MITIGATION: Monitor edge upgrade progress; set timeout threshold (10-15 min); if no progress, restart nsxd and retry
- ESXi vMotion stalls → VMs cannot move to other hosts during host upgrade; upgrade blocks indefinitely — MITIGATION: Verify vMotion network health before starting ESXi upgrades; ensure NSX logical switches are operational
- SDDC Manager upgrade fails partway through → Cannot rollback SDDC Manager without restore from backup (data loss); subsequent component upgrades blocked — MITIGATION: Ensure backup is recent and restorable; test backup restore before upgrade window
- Post-upgrade, workload VMs have lost network connectivity → NSX rules not applied correctly during upgrade; VMs show IP but can't communicate — MITIGATION: Validate NSX firewall rules post-upgrade; compare pre/post configurations
Self-Assessment Discussion Prompts
- You're planning a VCF upgrade for a 16-host cluster with 500+ workload VMs. The maintenance window can only be 8 hours (compressed from typical 12 hours). How do you optimize the upgrade to fit?
- NSX edge upgrade fails; you have 2 options: (1) Rollback the entire upgrade (lose SDDC Manager and vCenter upgrades too, restart tomorrow), (2) Rollback just the edge to 5.x, continue with ESXi upgrades. Which do you choose, and why?
- Post-upgrade, you discover a workload VM lost network connectivity. You check NSX firewall rules and see a policy rule was added post-upgrade that blocks the VM's traffic. The rule wasn't there pre-upgrade. Explain how this could happen and how you'd fix it.
- Your pre-flight check missed an expired SSL certificate on NSX Manager (cert expired yesterday). Upgrade starts; NSX Manager can't communicate with SDDC Manager (SSL verification fails). Can you proceed, or must you rollback? How do you recover?
- A 3-host vSAN cluster is being upgraded. Host 1 upgrade succeeds, host 2 upgrade is in progress (vMotion is halfway done), then host 2's network is interrupted (physical switch reboot). How does this affect the upgrade? What's your recovery?
Extensions
Extend Upgrade to Multi-Workload Domain Scenario
Plan and execute an upgrade of VCF 5.2 to 9.0 with multiple workload domains (VI 2, VI 3, etc. in addition to management domain). Document the upgrade sequence for multi-domain, including dependencies between management and workload domain upgrades. Update runbook accordingly.
harderUpgrade vSAN from 6.x to 7.0+ During VCF Upgrade
Plan a concurrent vSAN upgrade (6.x → 7.0) during VCF upgrade (5.2 → 9.0). Document interaction: Can vSAN upgrade happen in parallel with ESXi upgrade, or must it be sequential? Handle the scenario where vSAN upgrade and ESXi upgrade conflict (both require node reboot).
harderBuild an Automated Upgrade Orchestration Script
Write a PowerShell or Python script that automates the VCF upgrade sequencing (SDDC Manager → vCenter → NSX → ESXi via LCM). Include pre-flight validation, progress monitoring, failure detection, and automatic rollback decision logic. This mirrors production automation.
harderPlan Cross-Version Upgrade Path from VCF 3.x
Design an upgrade path from VCF 3.x (socket-based licensing, NSX-V, no Aria) to VCF 9.0 (subscription licensing, NSX 4.1, Aria integrated). Document intermediate stops (e.g., 3.x → 5.2 → 9.0), licensing model transition, and NSX-V to NSX-T migration steps.
harder⚠ Known Pitfalls (from Community KB)
References
- VMware Cloud Foundation Upgrade Guide (VCF 9.0)Tier 1 — Official
Official Broadcom documentation for VCF upgrade procedures. Covers pre-flight validation, sequencing, rollback procedures, and known issues. Essential reference. - VCF Upgrade Planner ToolTier 1 — Official
Interactive tool that validates HCL compatibility, component versions, and generates upgrade sequence recommendations. Saves significant time on planning. - VCF Release Notes and Known IssuesTier 1 — Official
Critical known issues, workarounds, and supported upgrade paths. Always review before starting upgrade. - NSX Upgrade Guide (NSX-T to NSX 4.1)Tier 1 — Official
NSX-specific upgrade procedures, edge node upgrade sequencing, and federation rollback strategies. - VMware Support Lifecycle and EOL DatesTier 1 — Official
Dates when VCF versions go end-of-support (EOS) and end-of-general-support (EOGS). Critical for planning upgrade timeline. - William Lam — VCF Upgrade Experiences and TroubleshootingTier 3 — Expert Blog
Community-authored blog covering real-world upgrade experiences, common pitfalls, and troubleshooting techniques.