Academy/vSphere Foundation 9.0 Administrator (2V0-16.25)/Lab: Deploy TKG Cluster via VMware Kubernetes Services (VKS)
This lab targets VCF 9.0

Lab: Deploy TKG Cluster via VMware Kubernetes Services (VKS)

VCF 9.0Advancedadminkubernetes⏱ 120 min

VKS (formerly TKG on Supervisor). Uses Cluster API. Antrea CNI in VVF (no NSX NCP). Requires Supervisor enabled (Lab 10).

Objectives

  • Create a TKG cluster manifest using Cluster API topology
  • Deploy a managed Kubernetes cluster via kubectl apply
  • Authenticate to TKG cluster and deploy a sample workload
  • Understand VKS architecture: Cluster API, Antrea CNI, vSAN CSI
  • Perform cluster lifecycle operations: scale workers, upgrade K8s version

Prerequisites

Supervisor enabled on VVF cluster (Lab 10 completed). Namespace 'dev-team' with resource quotas. Minimum 48GB additional RAM available for TKG cluster (3 control + 3 worker nodes).

Prior labs: vvf-admin-10

Required skills:

  • Kubernetes administration (kubectl, deployments, services)
  • YAML manifest authoring
  • Container concepts (images, registries, pods)

Lab Environment

Holodeck VVF pod with Supervisor running. TKG cluster deployed as set of VMs managed by Cluster API.

Credentials

SystemUsernamePassword
Supervisoradministrator@vsphere.localSame as vCenter
TKG Clusteradministrator@vsphere.localSame as vCenter

Tasks

Task 1 Create and Deploy TKG Cluster

manageability

VKS uses Cluster API to provision managed Kubernetes clusters. Understanding the manifest structure, version selection, and deployment process is exam-critical for Obj 4.4.

Step 1

Authenticate to Supervisor:
kubectl vsphere login --server=<supervisor-api-ip> --vsphere-username=administrator@vsphere.local --insecure-skip-tls-verify
kubectl config use-context dev-team

Step 2

List available Kubernetes versions:
kubectl get tanzukubernetesreleases

Note the available versions (e.g., v1.27.5+vmware.1, v1.28.3+vmware.1).

List of available TKR versions with READY=True status.
Step 3

Create TKG cluster manifest (tkg-cluster.yaml):

apiVersion: cluster.x-k8s.io/v1beta1
kind: Cluster
metadata:
name: lab-cluster-01
namespace: dev-team
spec:
clusterNetwork:
pods:
cidrBlocks:

  • 10.244.0.0/16

services:
cidrBlocks:

  • 10.96.0.0/16

topology:
class: tanzukubernetescluster
version: v1.27.5+vmware.1
controlPlane:
replicas: 1
metadata: {}
workers:
machineDeployments:

  • class: node-pool

name: worker-pool-1
replicas: 2
variables:

  • name: vmClass

value: best-effort-small

  • name: storageClass

value: vsan-default-storage-policy

For lab, use 1 control plane + 2 workers to conserve resources. Production: 3 control plane + 3+ workers.
Step 4

Deploy the cluster:
kubectl apply -f tkg-cluster.yaml

cluster.cluster.x-k8s.io/lab-cluster-01 created
Step 5

Monitor deployment progress (takes 15-25 minutes):
kubectl get cluster lab-cluster-01 -n dev-team -w

Also monitor in vCenter: VMs named 'lab-cluster-01-*' will be created.

Cluster phases: Provisioning → Provisioned. VMs appear in vCenter inventory under dev-team resource pool.
Step 6

Verify cluster is ready:
kubectl get cluster lab-cluster-01 -n dev-team

PHASE=Provisioned, READY=True

Validation Gate

Check: kubectl get cluster lab-cluster-01 -n dev-team

Expected: PHASE=Provisioned, 1 control plane + 2 worker VMs running in vCenter

Common Errors

Cluster stuck in 'Provisioning' for >30 minutes
Cause: Insufficient resources in namespace quota or TKR image not downloaded
Fix: kubectl describe cluster lab-cluster-01 — check Events. Verify namespace quota has enough CPU/memory. Check TKR availability: kubectl get tkr.
'no matches for kind Cluster' error
Cause: Cluster API CRDs not installed on Supervisor
Fix: Verify Supervisor is healthy: vCenter → Workload Management. TKG Service may need to be enabled on the Supervisor.
Worker nodes not joining cluster
Cause: Antrea CNI pod network issue or HAProxy not assigning VIPs
Fix: Check HAProxy logs: journalctl -u haproxy. Verify pod CIDR (10.244.0.0/16) doesn't conflict with existing networks.

Task 2 Authenticate to TKG Cluster and Deploy Workload

manageability

Authenticating to the TKG cluster and deploying a workload validates the full developer experience — from cluster provisioning to application deployment.

Step 1

Login to TKG cluster:
kubectl vsphere login --server=<supervisor-api-ip> --vsphere-username=administrator@vsphere.local --tanzu-kubernetes-cluster-name=lab-cluster-01 --tanzu-kubernetes-cluster-namespace=dev-team --insecure-skip-tls-verify

Step 2

Switch context:
kubectl config use-context lab-cluster-01

Step 3

Verify cluster nodes:
kubectl get nodes

1 control-plane node + 2 worker nodes, all in Ready status.
Step 4

Check cluster components:
kubectl get pods -A

System pods running: antrea-agent (on each node), antrea-controller, coredns, etcd, kube-apiserver, kube-controller-manager, kube-scheduler.
Step 5

Deploy sample application:
kubectl create deployment nginx --image=nginx:latest --replicas=2
kubectl expose deployment nginx --port=80 --type=LoadBalancer

deployment.apps/nginx created service/nginx exposed
Step 6

Wait for LoadBalancer IP (HAProxy assigns VIP):
kubectl get svc nginx -w

EXTERNAL-IP assigned from HAProxy VIP range (192.168.20.128-191).
If EXTERNAL-IP stays '<pending>', check HAProxy VIP pool availability and connectivity.
Step 7

Test application:
curl http://<EXTERNAL-IP>

nginx default welcome page HTML returned.

Validation Gate

Check: curl http://<EXTERNAL-IP> returns nginx welcome page

Expected: HTTP 200 with nginx default page content

Task 3 Scale Workers and Perform Cluster Lifecycle Operations

availability

Day-2 operations — scaling and upgrading — are key exam topics. Understanding how Cluster API handles rolling updates is critical.

Step 1

Scale worker pool from 2 to 3:
kubectl patch cluster lab-cluster-01 -n dev-team --type merge -p '{"spec":{"topology":{"workers":{"machineDeployments":[{"class":"node-pool","name":"worker-pool-1","replicas":3}]}}}}'

cluster.cluster.x-k8s.io/lab-cluster-01 patched
Step 2

Monitor scale-out (switch back to Supervisor context first):
kubectl config use-context dev-team
kubectl get machines -n dev-team -w

New Machine object created → VM provisioned → node joins cluster. Takes 3-5 minutes.
Step 3

Verify 3 workers:
kubectl config use-context lab-cluster-01
kubectl get nodes

1 control plane + 3 workers, all Ready.
Step 4

Scale back to 2 workers:
kubectl config use-context dev-team
kubectl patch cluster lab-cluster-01 -n dev-team --type merge -p '{"spec":{"topology":{"workers":{"machineDeployments":[{"class":"node-pool","name":"worker-pool-1","replicas":2}]}}}}'

One worker node cordoned, drained, and deleted. Workloads rescheduled to remaining nodes.
Step 5

Verify nginx pods redistributed:
kubectl config use-context lab-cluster-01
kubectl get pods -o wide

2 nginx pods running on remaining 2 worker nodes.

Validation Gate

Check: kubectl get nodes shows 1 control plane + 2 workers, kubectl get pods shows nginx pods running

Expected: Cluster scaled back to 2 workers, all pods healthy

Final Validation

TKG cluster deployed via VKS, workload running with LoadBalancer, scaling demonstrated.

✓ TKG cluster lab-cluster-01 in Provisioned state → kubectl get cluster shows READY=True

✓ nginx deployment accessible via LoadBalancer VIP → curl returns nginx welcome page

✓ Cluster scale-out and scale-in completed successfully → Worker count changed 2→3→2 without workload disruption

Cleanup / Restore

Snapshot: post-vks-lab

• Delete nginx deployment: kubectl delete deployment nginx && kubectl delete svc nginx

• Optionally delete TKG cluster: kubectl config use-context dev-team && kubectl delete cluster lab-cluster-01 (takes 10-15 min)

• Take snapshot 'post-vks-lab'

Design Reflection (VCDX)

Panelist: What's the blast radius if Supervisor fails — does the TKG cluster keep running? How do you handle K8s version upgrades across multiple TKG clusters? What's your persistent storage strategy for stateful workloads on VKS?

Requirements

  • Managed Kubernetes clusters for developer teams
  • Self-service cluster provisioning
  • Automated scaling for variable workloads
  • Persistent storage for stateful applications

Constraints

  • Antrea CNI in VVF (no NSX network policies beyond basic K8s NetworkPolicy)
  • HAProxy VIP pool size limits number of LoadBalancer services
  • TKG cluster resources consume namespace quota
  • K8s version tied to TKR availability (VMware release cadence)

Assumptions

  • Developers manage their own TKG clusters within namespace boundaries
  • HAProxy VIP range sufficient for all LoadBalancer services
  • Container images accessible from TKG nodes (registry connectivity)

Risks

  • TKG cluster upgrade failure can leave cluster in mixed-version state
  • Supervisor failure doesn't affect running TKG clusters but blocks management operations
  • Resource exhaustion if multiple large TKG clusters share namespace
  • Antrea CNI network policies less mature than NSX NCP

Self-Assessment Discussion Prompts

  1. How does VKS handle K8s version upgrades — in-place or blue-green?
  2. What monitoring should you add for TKG cluster health beyond Supervisor status?
  3. How would you design multi-tenancy with VKS — one cluster per team or shared cluster with namespace isolation?

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.