This lab targets VCF 5.2 to 9.0

VCF 5.2 to 9.0 Upgrade Planning Workshop

VCF 5.2 to 9.0Advancedvcdx-distinguished⏱ 90 min

This lab simulates a production VCF 5.2 upgrade to 9.0. Works with Holodeck lab environment (holodeck-02 or later) with a functional VCF domain, or can be adapted for design-exercise mode using documentation and runbooks.

Objectives

  • Plan a production VCF 5.2 to 9.0 upgrade with pre-flight assessment, compatibility validation, and component sequencing
  • Execute upgrade steps in Holodeck lab environment following the rigid upgrade order (SDDC Manager → vCenter → NSX → ESXi)
  • Handle a simulated failure (NSX upgrade stalls) and demonstrate recovery/rollback procedures
  • Validate post-upgrade services, workload connectivity, and new VCF 9.0 features (VCF Operations UI)

Prerequisites

VCF 5.2 lab environment deployed and operational (ideally holodeck-02 with management domain; or holodeck-03+ with workload domains). Alternatively, design-exercise mode using VCF Upgrade Planner documentation and upgrade runbooks.

Prior labs: holodeck-02 (management domain deployment), vcf-evolution-01 (architectural understanding of 5.x vs 9.0)

Required skills:

  • VCF operational management (SDDC Manager, vCenter, NSX administration)
  • Change management (maintenance windows, communication, rollback planning)
  • Certificate and DNS management (renewal, validation, drift detection)
  • Storage and networking troubleshooting (vSAN health, NSX connectivity, VLAN isolation)
  • Failure diagnosis and root cause analysis (logs, monitoring, escalation paths)

Lab Environment

Holodeck lab environment: 1 management domain (vCenter, SDDC Manager, NSX, vSAN 3-node cluster minimum). Alternatively, documentation-based planning exercise using VCF Upgrade Planner workflow.

graph TB
  Client[SDDC Client/Upgrade Tool] -->|HTTPS| SDDCMgr[SDDC Manager 5.2]
  SDDCMgr -->|HTTPS| vCenter[vCenter 5.x]
  SDDCMgr -->|HTTPS| NSXMgr[NSX Manager 5.x]
  vCenter -->|mgmt| ESXi1[ESXi 1]
  vCenter -->|mgmt| ESXi2[ESXi 2]
  vCenter -->|mgmt| ESXi3[ESXi 3]
  NSXMgr -->|overlay| ESXi1
  NSXMgr -->|overlay| ESXi2
  NSXMgr -->|overlay| ESXi3
  ESXi1 -->|heartbeat| vSAN[(vSAN Cluster)]
  ESXi2 -->|heartbeat| vSAN
  ESXi3 -->|heartbeat| vSAN

Tasks

Task 1 Pre-Upgrade Assessment: Validate HCL, Certificates, NTP/DNS, and Component Versions

reliability

Upgrade failures are rarely due to the upgrade process itself — they're due to pre-existing environment issues (clock skew, certificate expiry, incompatible hardware, outdated DNS records). VCDX candidates must be methodical about pre-flight validation. This task forces you to think like a change manager: 'What can go wrong, and how do I prevent it?'

Step 1

Run VCF Upgrade Planner (if available) or manually download and review the HCL (Hardware Compatibility List) for VCF 9.0.2. Verify each ESXi host in the cluster is on the Supported HCL. Document: [Host Model | ESXi Version | CPU SKU | RAM | Supported in 9.0?]

HCL validation table showing all hosts are supported; note any hosts at risk of EOL or not on HCL (which blocks upgrade)
Step 2

Check SSL certificate expiry dates on all VCF components: SDDC Manager, vCenter, NSX Manager, each ESXi host. Command: 'echo | openssl s_client -connect <ip>:443 2>/dev/null | openssl x509 -noout -dates'. Document: [Component | IP | Issuer | Expiry Date | Days Until Expiry | Action Needed?]

Certificate inventory with expiry >90 days for all components; any certificates expiring <90 days flagged for renewal before upgrade
Step 3

Verify certificate SANs (Subject Alternate Names) match FQDN and IP addresses used in upgrade. Example: SDDC Manager cert should have CN=sddc-manager.lab.local AND SAN=sddc-manager.lab.local, 10.0.1.10. Document discrepancies.

Certificate SAN validation: all critical FQDNs and IPs are in the certificate; no mismatches that would cause SSL errors during upgrade
Step 4

Verify NTP synchronization on all components (SDDC Manager, vCenter, NSX Manager, all ESXi hosts). Command: 'ntpstat' or 'timedatectl' (Linux) / 'w32tm /query /status' (Windows). Document: [Component | NTP Server | Offset (ms) | Sync Status]

NTP status table showing all components synchronized to the same stratum, offset <100ms, no outliers
Step 5

Verify DNS resolution in both directions. Test: 'nslookup sddc-manager.lab.local' (forward), 'nslookup 10.0.1.10' (reverse). Verify DNS used by upgrade process is consistent. Document: [FQDN | IP | Forward Resolution | Reverse Resolution | DNS Server Used]

DNS validation showing all FQDNs and IPs resolve consistently; reverse DNS working; no NXDOMAIN or timeout issues
Step 6

Document current component versions: SDDC Manager 5.x.x, vCenter 5.x.x, NSX Manager 5.x.x, vSAN driver version, ESXi versions (all should be aligned). Use SDDC Manager API: 'curl -k https://sddc-manager/api/v1/system/version' or UI. Document: [Component | Current Version | Supported in 9.0? | Pre-req for 9.0?]

Version inventory showing all components are supported upgrade paths (e.g., SDDC Manager 5.2.1 → 9.0.2 is supported; 5.0.0 → 9.0.2 is not; must upgrade to 5.2.x first)
Step 7

Validate current vSAN cluster health: 'Cluster > Configuration > vSAN > Cluster Information'. Check: Cluster health = Healthy, no unsynced objects, no degraded disks, all hosts online. Document: [Metric | Current State | Healthy?] Example: Sync Status (Synced/Resyncing), Physical Memory Allocated (%), Disk Errors (0?)

vSAN health report showing cluster is ready for upgrade (no resyncing, all disks healthy, no capacity issues)
Step 8

Create a pre-upgrade snapshot/baseline: Holodeck: 'New-Snapshot -VM <sddc>, <vcenter>, <nsx> -Name pre-upgrade-baseline'. Production: Create backups of SDDC Manager and vCenter VMs. Document snapshot/backup timestamps and verify restoration is possible.

Baseline snapshots or backups documented with timestamps; snapshot consistency verified (all VMs created at same time)
Step 9

Synthesize pre-flight findings into a checklist: [Item | Status | Action If Failed | Owner | Target Completion]. Include HCL, certificates, NTP, DNS, versions, vSAN health, backups. Mark as 'Ready for Upgrade' only when all items are GREEN.

Signed-off pre-flight checklist; green light to proceed to upgrade planning phase (Task 2)

Validation Gate

Check: Pre-flight checklist completed and signed off; no blockers remaining

Expected: All 9 substeps validated; environment is confirmed ready for VCF 9.0 upgrade

Common Errors

Certificate expiry validation is skipped; discovers expired cert during upgrade (blocks SSL connections)
Cause: Assuming certificates are good; not validating before maintenance window
Fix: Always run 'openssl s_client' on every component before upgrade. Check 'Expiry Date' field. If <90 days, renew now (not after upgrade starts).
NTP offset is documented but not acted on; NTP is 'almost synchronized' (offset 500ms)
Cause: Confusing 'synchronized' with 'slightly out of sync'
Fix: NTP offset must be <100ms (ideally <50ms). If offset is >100ms, fix NTP before upgrade. Stale time causes SSL cert validation failures, component communication timeouts, and silent data corruption.
DNS reverse lookup fails for one host; not discovered until upgrade attempts to contact it
Cause: Validating forward resolution but not reverse; reverse DNS is often forgotten
Fix: Test both directions: 'nslookup sddc-manager' and 'nslookup 10.0.1.10'. If reverse fails, update DNS admin's PTR records before upgrade.
vSAN is healthy overall but one disk is failed/degraded; not noticed until upgrade stalls waiting for vSAN rebalance
Cause: Skimming vSAN health; not checking individual disk status
Fix: Detailed check: vCenter > Configuration > vSAN > Cluster > Physical Disks. All must show 'Healthy'. If any failed, replace disk or remove host before upgrade.

Task 2 Plan Upgrade Sequence, Maintenance Window, and Rollback Strategy

manageability

VCF upgrades have a rigid component order: SDDC Manager first (it orchestrates the rest), then vCenter, then NSX, then ESXi via LCM. Deviating from this order causes cascading failures. VCDX candidates must understand this sequencing deeply and plan for failure scenarios — 'If NSX upgrade fails, can we rollback without losing the SDDC Manager and vCenter upgrades?'

Step 1

Document the upgrade sequence in tabular form: [Step | Component | Version | Task | Estimated Time | Rollback Point?] Example:

  Step 1: SDDC Manager 5.2.1 → 9.0.2 (in-place upgrade, no rollback to 5.2 once complete)
  Step 2: vCenter 5.x → 9.0.2 (in-place upgrade, no rollback)
  Step 3: NSX Manager 5.x → 9.0.2 (rolling upgrade, edge nodes upgraded sequentially)
  Step 4: ESXi hosts 5.x → 9.0.2 (via LCM, one host at a time with vMotion)
Upgrade sequencing table showing component order, estimated time per component, and rollback points identified (e.g., 'Can rollback after pre-flight but before SDDC Manager upgrade starts'; 'No rollback after SDDC Manager completes')
Step 2

Calculate total upgrade window duration: Pre-flight (1 hour) + SDDC Manager (1.5 hours) + vCenter (1 hour) + NSX (2 hours) + ESXi LCM (1 hour per host × 3 hosts = 3 hours) + Validation (1 hour) = ~9.5 hours total. Add 30% contingency buffer = ~12 hours. Document: [Phase | Duration | Rollback Possible?] and overall maintenance window time.

Maintenance window schedule: Start time, component-by-component timeline, end time, with buffer; e.g., 'Saturday 6 PM to Sunday 6 AM (12 hours)'
Step 3

Plan stakeholder communication: Email to business owners at T-7 days (announcement), T-1 day (reminder with start/end times, contact info), T0 (go-live notification), post-upgrade (validation results). Document each communication: [Audience | Message | Timing | Owner]

Communication plan with 4-5 messages scheduled; each includes who sends it, when, what stakeholders are notified, and success criteria (e.g., 'Workloads can access external networks post-upgrade')
Step 4

Define rollback strategy for each component: SDDC Manager (restore from snapshot/backup, requires ~2 hours), vCenter (similar, ~1.5 hours), NSX (rolling rollback of edge nodes, ~2 hours), ESXi (revert LCM update or use previous snapshot, ~30 min per host). Document: [Component | Rollback Method | Estimated Time | Blast Radius | Data Loss Risk]

Rollback playbook showing how to recover each component; key insight: SDDC Manager and vCenter rollback require restore from backup (data loss possible if changes made post-upgrade); NSX and ESXi can be rolled back more easily
Step 5

Identify failure scenarios and decision points: 'If SDDC Manager upgrade fails at 30%, do we proceed or rollback?' (Answer: Rollback, because SDDC Manager must be stable before vCenter upgrade). 'If vCenter upgrade hangs at 50%, how long do we wait before calling rollback?' (Answer: 1 hour; if not progressing, assume stuck). Document: [Scenario | Decision Point | Condition | Action | Responsible]

Failure decision tree showing criteria for rollback vs. retry vs. escalation; includes thresholds (time, error count) for decision making
Step 6

Plan resource allocation: Who is on call? Who is the primary upgrade engineer, secondary, escalation contact? What's the communication channel (Slack, Zoom, phone)? Document: [Role | Name/Team | Contact | Responsibilities During Upgrade | Availability (hours)]

Staffing plan showing at least 2 capable engineers throughout the upgrade window; escalation path to management defined
Step 7

Plan validation checkpoints: After SDDC Manager upgrade completes (check API endpoints), after vCenter (verify vSphere Client login), after NSX (verify logical networks), after ESXi (verify workload connectivity). Document: [Checkpoint | Component | Validation Test | Expected Result | Time Estimate]

Validation checkpoint list with concrete tests (API calls, login attempts, ping tests) that confirm each phase completed successfully
Step 8
Create a 1-page summary for executives: 'VCF Upgrade Plan: 5.2 → 9.0 on [DATE]. 12-hour maintenance window. Rollback available if upgrade fails before 50% complete; beyond that point, proceed or accept downtime. Risk: Low (multi-component redundancy ensures no data loss, but services will be down 12 hours).' Include go/no-go decision criteria.
Executive summary suitable for sending to CIO/COO with approval sign-off line

Validation Gate

Check: Upgrade plan (sequencing, timeline, communication, rollback) is complete and approved

Expected: Ready to proceed to Task 3 (execution); all stakeholders informed and decision rights established

Common Errors

Upgrade sequence deviates from VCF order (e.g., trying to upgrade NSX before SDDC Manager completes)
Cause: Misunderstanding dependencies; assuming components can upgrade in parallel
Fix: Emphasize: SDDC Manager orchestrates all subsequent upgrades. It must complete first. vCenter and NSX have dependencies on SDDC Manager 9.0 APIs. ESXi upgrade is coordinated by LCM (part of vCenter). Strict order: SDDC Mgr → vCenter → NSX → ESXi.
Timeline is too optimistic (e.g., 4 hours total for a 3-host cluster); upgrade runs over and hits business hours
Cause: Not accounting for validation time, unexpected delays, or safety margins
Fix: Real-world upgrade times: SDDC Manager 1.5-2 hrs, vCenter 1.5 hrs, NSX 2-3 hrs (edge upgrades serial), ESXi 30 min-1 hr per host. Add 30% buffer (delays happen). For 3-host cluster: 10-12 hours is realistic. Schedule in maintenance window; avoid business hours.
Rollback strategy assumes 'just restore from backup' without testing restore procedure beforehand
Cause: Not validating backup/snapshot functionality before upgrade
Fix: Before upgrade day: Test snapshot restore on a non-prod environment. Know exactly how long restore takes. For SDDC Manager: Test that configuration is recoverable. For vCenter: Test that vSphere objects are recoverable.
Stakeholder communication happens only after upgrade starts (too late for go/no-go decision)
Cause: Underestimating communication lead time
Fix: Send announcements 7 days before, reminders 1 day before, start notification at T0. Collect feedback early (any last-minute blockers from business?). Only proceed with upgrade if stakeholders confirm readiness.

Task 3 Execute Upgrade in Holodeck Lab: SDDC Manager, vCenter, NSX, ESXi with Simulated Failure Handling

reliability

Theory is useless without practice. This task walks you through each upgrade step, validating at each stage. Critically, you'll encounter a simulated failure (NSX upgrade stalls) and demonstrate how to diagnose and recover — a real skill VCDX panelists value.

Step 1

Pre-upgrade validation (from Task 1): Verify pre-flight checklist is GREEN. Take baseline snapshot: 'New-Snapshot -VM vcf-sd-<mgmt-domain> -Name pre-upgrade-baseline'. Document snapshot times. Declare maintenance window open.

Baseline snapshot created; checklist signed off; maintenance window officially open (notifications sent to stakeholders)
Step 2

SDDC Manager upgrade: Download VCF 9.0.2 SDDC Manager OVA. In SDDC Manager UI, navigate to Upgrade > Available Patches. (Or via PowerShell: Update-SddcManager -Version 9.0.2). Monitor progress. Expected: 30-45 min. Validate: SDDC Manager API returns 9.0.2, UI is responsive, no errors in logs (tail /var/log/upgrade.log).

SDDC Manager upgraded to 9.0.2; API endpoint 'GET /api/v1/system/version' returns 9.0.2; UI login successful
Step 3

vCenter upgrade: Trigger vCenter 9.0.2 update via SDDC Manager Upgrade UI or vCenter Update Manager. Monitor progress (vCenter will briefly reboot). Expected: 1-1.5 hours. Validate: vSphere Client logs in, vCenter cluster shows all hosts connected, no errors in vCenter logs.

vCenter upgraded; all ESXi hosts still connected to vCenter; no vCenter alarms
Step 4

NSX Manager upgrade: Trigger NSX Manager upgrade via SDDC Manager or NSX UI. This step updates NSX Manager, then rolling upgrade of Edge nodes. Expected: 1.5-2 hours. Validate after NSX Mgr completes: NSX API 'GET /api/v1/infra/domains' works, logical networks are intact.

NSX Manager upgraded; logical routers/segments still present; NSX API responding normally
Step 5

SIMULATED FAILURE: NSX edge node upgrade stalls. Symptom: LCM shows 'Edge-01 upgrade: 45% complete. No progress for >10 min.' Troubleshoot: SSH into edge node (ssh admin@edge-01), check: 'systemctl status nsxd' (running?), 'journalctl -xe' (errors?). Try: 'systemctl restart nsxd'. If still stuck after 2 min, escalate.

Diagnosed failure: nsxd service crashed during upgrade. Attempted recovery by restarting service. Service recovered, LCM resumed upgrade.
Step 6

If edge node is unrecoverable, demonstrate rollback: Rollback that specific edge from 9.0 to 5.x (via LCM: Rollback Node). Verify: That edge continues to serve traffic at 5.x version (mixed version cluster is unsupported long-term but acceptable for recovery). Document decision and timeline (this added 30 min to upgrade).

Edge node rolled back to 5.x; traffic is still flowing (mixed versions are tolerated for short time); plan to retry edge upgrade post-upgrade window
Step 7
ESXi upgrade via LCM: After NSX completes (or is recovered), trigger ESXi upgrade via LCM. Upgrade one host at a time (cluster size 3, so 3 rounds). Each host: vMotion workloads to other hosts → upgrade ESXi → reboot → rejoin cluster → validate. Expected: 30 min per host × 3 = 1.5 hours total. Validate: Each host shows 9.0.2 version after reboot, cluster shows all hosts connected.
All 3 ESXi hosts upgraded to 9.0.2; cluster membership intact; vSAN quorum maintained throughout (no data loss)
Step 8

Post-upgrade immediate validation: SDDC Manager UI shows all components healthy (vCenter, NSX, ESXi all green). Check vSAN cluster: all hosts online, no resyncing objects. Verify workload connectivity: ping a test VM to external network (should work). Document: [Component | Status | Errors | Timestamp]

All services healthy; no alarms; workloads can access external networks; VCF Operations UI (new in 9.0) is accessible
Step 9

Record upgrade metrics: [Component | Start Time | End Time | Duration | Rollback Used? | Issues Encountered]. Compare to plan: Did actual times match estimates? What took longer? Document lessons learned.

Upgrade timeline showing actual durations and deviations from plan; issues and resolutions documented (e.g., 'NSX edge stalled for 15 min, restarted nsxd, succeeded on retry')

Validation Gate

Check: All upgrade steps completed; simulated failure diagnosed and resolved; post-upgrade validation passed

Expected: VCF 5.2 → 9.0 upgrade successful with documented failure recovery; ready for Task 4 (comprehensive validation)

Common Errors

NSX edge fails and is marked unhealthy; all NSX services are down (no traffic can flow)
Cause: Attempting to continue upgrade without rolling back failed edge; trying to upgrade remaining components while one edge is broken
Fix: If edge upgrade fails: (1) Do NOT proceed to next edge until current edge is fixed or rolled back. (2) If unfixable, rollback to 5.x, document the failure, and retry in controlled environment. (3) A broken edge blocks the entire NSX service.
ESXi upgrade starts, vMotion stalls (workload VMs can't move to other hosts)
Cause: Network connectivity issues during vMotion (vMotion VLAN down or NSX issues from previous step)
Fix: Verify vMotion network is up before starting ESXi upgrade. If vMotion fails: Investigate NSX logical network health, verify vMotion VLAN on physical switches, ensure no certificate issues blocking vMotion.
Post-upgrade, SDDC Manager shows vCenter as 'Offline' even though vCenter is running and upgraded
Cause: Certificate issue from vCenter upgrade; old certificate cached in SDDC Manager
Fix: SDDC Manager may need to re-validate certificate for vCenter. Force certificate refresh: SDDC Manager UI > Infrastructure > vCenter Servers > [vCenter] > Edit > Save. Or wait 10 min for automatic re-validation.

Task 4 Post-Upgrade Validation: Services, Workload Connectivity, and New VCF 9.0 Features

manageability

The upgrade isn't done until you verify everything works and new features are functional. This task ensures you know what to test (not just 'does the UI load?') and demonstrates mastery of end-to-end VCF operations.

Step 1

Comprehensive service health check: vCenter Cluster Health (all hosts green), SDDC Manager System Status (all health checks pass), NSX System Status (all nodes healthy, controllers consensus reached), vSAN Cluster Health (all hosts operational, capacity balanced). Document: [Service | Status | Any Alarms?] Expected: All GREEN, 0 alarms.

Service health report showing all VCF components operational post-upgrade; no critical alarms
Step 2

Workload connectivity validation: Test VM-to-VM traffic on same segment (local traffic), VM-to-VM on different segments (L3 routing via T0/T1), VM-to-external (north-south via edge). Command: 'ping' from one VM to another, or use iperf for throughput. Document: [Traffic Flow | Source | Destination | Result | Latency/Throughput]

All traffic flows working; no routing errors; latency within expected range (<1ms for local, <10ms for external)
Step 3

Verify new VCF 9.0 features are enabled: VCF Operations (Unified UI). Log into 'https://sddc-manager/vcf' and verify 'Cluster Operations' and 'Infrastructure' dashboards load. Check: [Feature | Accessible? | Data Populated?] Examples: Cluster Overview, Capacity, Compliance.

VCF Operations UI accessible; dashboards show cluster data; no 'loading' spinners indefinitely
Step 4

Verify NSX 4.1+ features (if applicable): Federation capability status, advanced routing (if configured), stretched segments (if multi-site). Document: [Feature | Availability | Configured? | Working?]

NSX 4.1 features discoverable in UI; federation UI shows status (if multi-site); advanced routing works if configured
Step 5

Certificate validation: Verify all SSL certificates are renewed/valid post-upgrade. Check expiry dates on SDDC Manager, vCenter, NSX Manager, each Edge. Command: 'openssl s_client -connect <ip>:443 2>/dev/null | openssl x509 -noout -dates'. Document: [Component | Issuer | Expiry | Days Until | Action Needed?]

All certificates valid (not expired), expiry >90 days; no cert renewal needed immediately post-upgrade
Step 6

Verify DNS and NTP post-upgrade: Ensure DNS still resolves correctly for all components (FQDNs, IPs). Ensure NTP offset is <100ms on all nodes. Command: 'nslookup <fqdn>', 'ntpstat'. Document: [Component | DNS Status | NTP Offset | Healthy?]

DNS resolution correct; NTP synchronized; no drift detected
Step 7

Run upgrade runbook to document: [Item | Pre-Upgrade State | Post-Upgrade State | Status | Notes]. Example: SDDC Manager version, vCenter version, NSX version, ESXi versions, vSAN cluster health, workload count, network segments count, policies, etc.

Before/after snapshot comparing configurations; verifies nothing was inadvertently changed or lost during upgrade
Step 8

Close maintenance window: Send post-upgrade notification to stakeholders. Include: Upgrade completed successfully, services healthy, workloads accessible, any issues encountered and resolutions. Thank you message. Document: [Date/Time | Audience | Message Sent By]

Stakeholder notification sent; upgrade officially declared complete
Step 9

Create post-upgrade documentation: Upgrade timeline, issues encountered (NSX edge stall) and resolution (restarted nsxd), lessons learned, any follow-up actions (e.g., 'Review NSX edge deployment for stability', 'Update runbooks for 9.0'). Suitable for knowledge base or team wiki.

Post-upgrade runbook documenting what happened, why, and recommendations for next time

Validation Gate

Check: All validation steps completed; VCF 9.0 upgrade fully validated and documented; services operational; stakeholders informed

Expected: Successful upgrade with comprehensive post-upgrade documentation; confidence in 9.0 environment; ready for production use

Common Errors

Post-upgrade validation shows 'vCenter Alarms: 100+ critical', but environment seems functional
Cause: Alarms not cleared post-upgrade; old 5.x alarms lingering in vCenter
Fix: Review alarm list. Transient alarms during upgrade (host connectivity, storage latency) will auto-resolve within 5-10 min. If alarms persist >15 min, investigate: vCenter logs, host connectivity, network reachability.
NSX logical segments not showing up in vSphere Client after upgrade (segments are present in NSX Manager UI)
Cause: vSphere plugin not refreshed; cached data from 5.x NSX
Fix: Force vSphere Client refresh: Logout and log back in. Or restart vCenter service. NSX segments should appear in 'Networking > Segments' view post-upgrade.
Workload VMs lose network connectivity post-upgrade; VMs show IP addresses but can't ping
Cause: NSX distributed firewall rules not applied correctly after upgrade; old NSX-V rules didn't translate to NSX-T format
Fix: Check NSX Distributed Firewall rules post-upgrade. Compare pre-upgrade (NSX 5.x) to post (NSX 9.0). If rules are missing, apply them from backup or reconfigure manually.

Final Validation

You have planned and executed a production-grade VCF 5.2 to 9.0 upgrade, including pre-flight validation, sequencing, failure handling, and comprehensive post-upgrade validation. You demonstrated the ability to diagnose a simulated failure (NSX edge stall) and recover without data loss. This is a VCDX-level operational skill — panelists will expect you to walk through an upgrade scenario and articulate contingency plans.

✓ Pre-upgrade assessment complete → HCL validation, certificate dates, NTP sync, DNS resolution, component versions, vSAN health verified; checklist signed off

✓ Upgrade plan (sequencing, timeline, communication, rollback) → Documented component order (SDDC Mgr → vCenter → NSX → ESXi), 12-hour maintenance window, stakeholder communication, rollback procedures, failure decision tree

✓ Upgrade execution with failure handling → All components upgraded successfully; simulated NSX edge failure diagnosed and recovered; post-upgrade services healthy

✓ Post-upgrade validation → All services green, workloads connected, new VCF 9.0 features accessible, certificates valid, DNS/NTP synchronized, documentation complete

✓ Runbook and lessons learned → Post-upgrade documentation recorded; timeline, issues, resolutions, and recommendations for future upgrades documented

Cleanup / Restore

• Archive upgrade logs and runbooks for audit trail and knowledge base

• Snapshot the VCF 9.0 environment as 'vcf-90-post-upgrade' for future labs

• Decommission pre-upgrade snapshots after 30-day retention (in production, follow ORO retention policy)

• Update runbooks for 9.0 (if any changes to procedures)

• Conduct post-upgrade retrospective (team meeting): What went well? What was harder than expected? How do we improve next time?

Design Reflection (VCDX)

A VCDX panelist will ask: 'Walk us through a VCF upgrade you planned or executed. What went wrong, and how did you recover?' Your answer should include: (1) Specific pre-flight validation you did (not generic 'checked health'). (2) The exact sequencing of components (SDDC Mgr first, then vCenter, then NSX, then ESXi) and why that order matters. (3) A concrete failure scenario and how you diagnosed it (logs, API calls, not just rebooting). (4) Decision-making under pressure (NSX edge failed; do you rollback the entire upgrade or just that component?). (5) Post-upgrade validation that goes beyond 'UI loads' (workload connectivity, new features, certificates, DNS/NTP).

Requirements

  • Pre-flight validation complete: HCL compatibility, certificate expiry, NTP/DNS sync, component versions, vSAN health, backups
  • Upgrade sequence documented: SDDC Manager → vCenter → NSX → ESXi via LCM, in strict order, with time estimates and rollback points
  • Maintenance window scheduled: 12-hour block, communicated to stakeholders 7 days in advance, with go/no-go criteria
  • Failure handling procedures: Decision tree for rollback vs. retry, thresholds for escalation, staffing plan with on-call engineers
  • Simulated failure resolution: NSX edge upgrade stall diagnosed, recovered, and impact assessed
  • Post-upgrade validation: Services health, workload connectivity, new features, certificates, DNS/NTP, documentation
  • Upgrade runbook and lessons learned: Timeline, issues, resolutions, recommendations for next upgrade

Constraints

  • VCF upgrade sequence is rigid — cannot parallelize components or deviate from order without cascading failures
  • NSX edge upgrade is serial (one at a time); cannot upgrade all edges in parallel, which extends upgrade window
  • Holodeck lab environment has single failure domain (one physical host); production environments have HA for SDDC Manager and vCenter, reducing downtime risk
  • Rollback after SDDC Manager upgrade completes requires restore from backup; post-upgrade changes are lost during rollback
  • Mixed-version clusters (some hosts on 5.x, some on 9.0) are unsupported for extended periods; ESXi hosts must be upgraded in quick succession
  • No in-place upgrade path from VCF 3.x to 9.0; must upgrade to intermediate version (e.g., 5.2) first, then to 9.0

Assumptions

  • All pre-flight checks pass (HCL, certificates, NTP, DNS, versions, backups); no blockers remain at upgrade start
  • Maintenance window of 12 hours is acceptable to business; no critical workloads require 24/7 uptime during upgrade
  • Network and storage are stable during upgrade; no planned maintenance on upstream switches, routers, or SAN during VCF upgrade window
  • Upgrade operators are trained and experienced with VCF; they can follow runbook sequencing and troubleshoot failures without vendor support
  • Rollback is an acceptable failure mode; business accepts brief downtime (2-4 hours) if upgrade fails and must be rolled back
  • Post-upgrade, all workloads return to running state; no application data corruption or permanent loss of configuration

Risks

  • Certificate expiry during upgrade → SSL connections fail, components can't communicate — MITIGATION: Validate certificate expiry 90 days in advance; renew if needed before upgrade start
  • NTP drift during upgrade → Time-dependent services (Kerberos, SSL handshakes) fail — MITIGATION: Verify NTP sync on all nodes; fix drift >100ms before upgrade
  • NSX edge upgrade hangs indefinitely → NSX services unavailable, workloads lose connectivity — MITIGATION: Monitor edge upgrade progress; set timeout threshold (10-15 min); if no progress, restart nsxd and retry
  • ESXi vMotion stalls → VMs cannot move to other hosts during host upgrade; upgrade blocks indefinitely — MITIGATION: Verify vMotion network health before starting ESXi upgrades; ensure NSX logical switches are operational
  • SDDC Manager upgrade fails partway through → Cannot rollback SDDC Manager without restore from backup (data loss); subsequent component upgrades blocked — MITIGATION: Ensure backup is recent and restorable; test backup restore before upgrade window
  • Post-upgrade, workload VMs have lost network connectivity → NSX rules not applied correctly during upgrade; VMs show IP but can't communicate — MITIGATION: Validate NSX firewall rules post-upgrade; compare pre/post configurations

Self-Assessment Discussion Prompts

  1. You're planning a VCF upgrade for a 16-host cluster with 500+ workload VMs. The maintenance window can only be 8 hours (compressed from typical 12 hours). How do you optimize the upgrade to fit?
  2. NSX edge upgrade fails; you have 2 options: (1) Rollback the entire upgrade (lose SDDC Manager and vCenter upgrades too, restart tomorrow), (2) Rollback just the edge to 5.x, continue with ESXi upgrades. Which do you choose, and why?
  3. Post-upgrade, you discover a workload VM lost network connectivity. You check NSX firewall rules and see a policy rule was added post-upgrade that blocks the VM's traffic. The rule wasn't there pre-upgrade. Explain how this could happen and how you'd fix it.
  4. Your pre-flight check missed an expired SSL certificate on NSX Manager (cert expired yesterday). Upgrade starts; NSX Manager can't communicate with SDDC Manager (SSL verification fails). Can you proceed, or must you rollback? How do you recover?
  5. A 3-host vSAN cluster is being upgraded. Host 1 upgrade succeeds, host 2 upgrade is in progress (vMotion is halfway done), then host 2's network is interrupted (physical switch reboot). How does this affect the upgrade? What's your recovery?

Extensions

Extend Upgrade to Multi-Workload Domain Scenario

Plan and execute an upgrade of VCF 5.2 to 9.0 with multiple workload domains (VI 2, VI 3, etc. in addition to management domain). Document the upgrade sequence for multi-domain, including dependencies between management and workload domain upgrades. Update runbook accordingly.

harder

Upgrade vSAN from 6.x to 7.0+ During VCF Upgrade

Plan a concurrent vSAN upgrade (6.x → 7.0) during VCF upgrade (5.2 → 9.0). Document interaction: Can vSAN upgrade happen in parallel with ESXi upgrade, or must it be sequential? Handle the scenario where vSAN upgrade and ESXi upgrade conflict (both require node reboot).

harder

Build an Automated Upgrade Orchestration Script

Write a PowerShell or Python script that automates the VCF upgrade sequencing (SDDC Manager → vCenter → NSX → ESXi via LCM). Include pre-flight validation, progress monitoring, failure detection, and automatic rollback decision logic. This mirrors production automation.

harder

Plan Cross-Version Upgrade Path from VCF 3.x

Design an upgrade path from VCF 3.x (socket-based licensing, NSX-V, no Aria) to VCF 9.0 (subscription licensing, NSX 4.1, Aria integrated). Document intermediate stops (e.g., 3.x → 5.2 → 9.0), licensing model transition, and NSX-V to NSX-T migration steps.

harder

⚠ Known Pitfalls (from Community KB)

Skipped pre-flight certificate validation; cert expired during upgrade COMMON
Problem: Operator assumed certificates were valid (not expiring anytime soon). During upgrade, NSX Manager cannot communicate with SDDC Manager due to SSL certificate expiry. Upgrade is partially complete (SDDC Manager upgraded, vCenter attempted but stalled).
Resolution: Always validate certificate expiry on all components 90 days in advance. Use 'openssl s_client' to check dates. If cert expires <90 days, renew before upgrade. Never assume certs are good.
Attempted to upgrade multiple components in parallel (NSX and ESXi together) COMMON
Problem: Trying to speed up upgrade, operator starts ESXi upgrade while NSX upgrade is still in progress. NSX upgrade gets stuck because ESXi nodes are rebooting and not responsive to NSX Manager.
Resolution: VCF upgrade sequence is STRICT: SDDC Mgr → vCenter → NSX → ESXi. No parallelization. Violating this order causes cascading failures and forces rollback of all components.
NSX edge upgrade hangs; operator didn't have a timeout/escalation procedure COMMON
Problem: NSX edge upgrade stalls at 45% for 45 minutes. Operator is unsure: Do I wait longer? Do I cancel? Do I restart the edge? Lacks decision criteria.
Resolution: Set clear timeout threshold: If edge upgrade shows 0% progress for >10 minutes, investigate (SSH into edge, check logs, restart nsxd). If still stuck after 2 more minutes, rollback that edge to 5.x and proceed with rest of upgrade. Document decision.
Rolled back mid-upgrade but didn't verify backup integrity beforehand CRITICAL
Problem: SDDC Manager upgrade failed. Operator initiated rollback to snapshot/backup. Snapshot is corrupted or outdated (6 months old). Configuration data is lost.
Resolution: Before upgrade: Test backup/snapshot restore on non-prod environment. Verify that SDDC Manager, vCenter data is recoverable and recent (<24 hrs old). Only use a tested backup for rollback.
Post-upgrade, workloads show no network connectivity; NSX rules were not applied COMMON
Problem: VMs can ping locally but cannot reach external networks or other segments. NSX firewall rules are present in NSX Manager but not enforced on hosts. Appears to be a policy delivery failure.
Resolution: Check NSX Manager: Policies > Realization State. Verify policies are realized on ESXi hosts. If not realized, restart nsx-node service on problematic ESXi hosts or on NSX Manager. Compare pre/post firewall rules for differences.

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.