Academy/VCAP — VCF VKS (3V0-24.25)/Kubernetes Version Upgrade (TKR v1.27 → v1.28)
This lab targets VCF 9.0

Kubernetes Version Upgrade (TKR v1.27 → v1.28)

VCF 9.0Intermediatevcap-advanced⏱ 120 min

Objectives

  • Perform zero-downtime Kubernetes version upgrade with Velero backup.

Prerequisites

VCF lab environment deployed and operational

Lab Environment

Standard VCF lab environment for Advanced VCF 9.0 VKS (vSphere Kubernetes Service)

Tasks

Task 1 Kubernetes Version Upgrade (TKR v1.27 → v1.28)

Perform zero-downtime Kubernetes version upgrade with Velero backup.

Step 1

Verify current TKR: kubectl version --short

Step 2

Create Velero backup of running cluster: velero backup create pre-upgrade-backup

Step 3

Update ClusterClass YAML: Change TKR reference to v1.28.5.

Step 4

Apply updated ClusterClass: kubectl patch clusterclass ... --patch '{"spec.version":"v1.28.5"}'

Step 5

Monitor rolling update: kubectl get machinedeployments -w

Step 6

Verify no pod disruption: Deploy test app; confirm requests succeed during upgrade.

Step 7

Post-upgrade: Verify etcd cluster health: kubectl -n etcd exec etcd-0 -- etcdctl endpoint health

Step 8

Cleanup: Prune old backups: velero backup delete --all-filters

Validation Gate

Check: Verify lab completion

Expected: Lab exercise completed successfully

Common Errors

Attempting in-place TKR upgrade without checking compatibility matrix
Fix: TKG supports upgrade from TKR N to N+1 only (e.g., 1.27 → 1.28, not 1.26 → 1.28). Check the TKR compatibility matrix before planning upgrades. Multi-version jumps require sequential upgrades with validation between each hop.
Not draining worker nodes before upgrade
Fix: TKG upgrade cordons and drains worker nodes automatically, but custom PodDisruptionBudgets (PDBs) may block draining. Review PDBs before upgrade: kubectl get pdb -A. A PDB with minAvailable=100% prevents any pod eviction, blocking the upgrade.
Upgrading control plane and workers simultaneously
Fix: Always upgrade control plane first, then workers. Control plane manages the cluster — upgrading workers first can create version skew (workers ahead of control plane), which is an unsupported configuration.

Final Validation

Lab completed successfully

✓ All steps completed → No errors observed

Cleanup / Restore

• Revert to snapshot if needed

Design Reflection (VCDX)

Kubernetes lifecycle management on VCF shows operational maturity. VCDX panelists test upgrade methodology, version compatibility awareness, and rollback planning.

⚠ Known Pitfalls (from Community KB)

Attempting multi-version TKR jumps — only N→N+1 is supported.
Not checking PodDisruptionBudgets before upgrade — PDBs can block node draining indefinitely.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.