Academy/VCAP — VCF VKS (3V0-24.25)/Provision TKG Cluster via ClusterClass
This lab targets VCF 9.0

Provision TKG Cluster via ClusterClass

VCF 9.0Intermediatevcap-advanced⏱ 120 min

Objectives

  • Deploy a 3-control-plane, 5-worker TKG cluster with autoscaling enabled.

Prerequisites

VCF lab environment deployed and operational

Lab Environment

Standard VCF lab environment for Advanced VCF 9.0 VKS (vSphere Kubernetes Service)

Tasks

Task 1 Provision TKG Cluster via ClusterClass

Deploy a 3-control-plane, 5-worker TKG cluster with autoscaling enabled.

Step 1

Create ClusterClass YAML referencing TKR v1.28.5 from content library.

Step 2

Define two MachineDeployments: zone-a (3 replicas) and zone-b (2 replicas).

Step 3

Enable ClusterAutoscaler; set min=2, max=8 per zone.

Step 4

Deploy cluster: kubectl apply -f cluster.yaml

Step 5

Monitor provisioning: kubectl get clusters -o wide; kubectl get nodes

Step 6

Verify PVC creation and CSI integration: kubectl get pvc -A; kubectl get storageclasses

Step 7

Trigger pod deployment; verify autoscaler scales zones independently.

Step 8

Drain a zone; confirm PDB-aware rescheduling and autoscaler rebalance.

Validation Gate

Check: Verify lab completion

Expected: Lab exercise completed successfully

Common Errors

ClusterClass configuration with incorrect infrastructure references
Fix: ClusterClass templates reference vSphere infrastructure (storage class, content library, network). A typo in storageClassName or incorrect TKR image name causes cluster creation to fail with an 'infrastructure provider' error. Validate all references against actual vSphere objects before applying the ClusterClass manifest.
Insufficient worker node sizing for workload requirements
Fix: Default TKG worker nodes (2 vCPU, 4GB RAM) are insufficient for most production workloads. Size worker nodes based on pod density and resource requests: total pod CPU requests / worker vCPU should not exceed 80% for headroom. Use ClusterClass variables to parameterize node sizing.
Not configuring storage class for persistent volumes
Fix: TKG clusters need a default storage class mapped to a vSAN storage policy. Without this, PVC creation fails with 'no persistent volumes available'. Verify: kubectl get storageclass shows a class with 'default' annotation.

Final Validation

Lab completed successfully

✓ All steps completed → No errors observed

Cleanup / Restore

• Revert to snapshot if needed

Design Reflection (VCDX)

Kubernetes cluster provisioning on VCF demonstrates container platform design. VCDX panelists test worker node sizing methodology and how you integrate K8s storage with vSAN.

⚠ Known Pitfalls (from Community KB)

Using default worker node sizes for production — 2 vCPU/4GB is a PoC size, not production.
Not mapping storage classes to vSAN policies — PVCs fail silently.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.