Master Supervisor cluster deployment, TKG cluster lifecycle, networking, storage, and security for Kubernetes workloads on vSphere. Covers Tanzu Kubernetes Grid, persistent volumes, NSX integration, and production Day-2 operations.
3V0-24.25
VCAP
60
Questions
135m
Duration
300/500
Pass Score
32
Objectives
Exam Blueprint Weights
Section titles, groupings and weights below are VCDX Academy study groupings, NOT the official Broadcom blueprint structure. Broadcom publishes no section weights. Always cross-check the official exam guide. Official exam guide ↗
Section 1 — Architecture
~15%
Section 2 — Cluster Lifecycle
~25%
High Weight
Section 3 — Networking
~20%
Section 4 — Storage
~15%
Section 5 — Security and Operations
~25%
High Weight
Version Evolution
vSphere Kubernetes Service (VKS, formerly Tanzu Kubernetes Grid) integrates container orchestration into VCF. Evolution: VCF 4.x introduced Workload Management (Supervisor) → VCF 5.x added TKG 2.0 with ClusterClass → VCF 9.0 rebranded to VKS with deeper NSX CNI integration. Key shift: from separate Tanzu products to integrated VKS within VCF Operations.
In VCDX defense, be ready to justify NSX vs. VDS+HAProxy choice. NSX provides superior multi-site support, GSLB, and identity-based firewall. VDS+HAProxy is cost-effective for single-site labs but lim
Design consideration: NSX+NCP provides tighter integration, GSLB, and federation support. Antrea is cost-effective for single-site, non-NSX deployments but limits multi-site HA design. In VCDX design
In VCDX design, emphasize multi-tenant isolation strategy: combine namespace quotas, network policies, PSA enforcement, and RBAC. Document assumptions about trust boundaries (are all tenants internal?
Supervisor Cluster Won't Enable:Check: vSAN health (all hosts healthy), DNS resolves supervisor-api FQDN, NSX LB configured and reachable, Content Library subscribed and synced within 24h, all 8 hosts
TKG Cluster Stuck in Creating:Check CAPI controllers: kubectl get events -A (look for Machine creation errors). Common: IP pool exhausted (increase IP range), VM class too small (min 2 CPU, 4 GB RAM),
VKS in VCF 9.0 introduces a simplified control plane based on Supervisor clusters—a native vSphere construct that transforms ESXi hosts into Kubernetes nodes. The Supervisor is not a virtual machine; it IS the vSphere cluster itself running the Kubernetes control plane as system VMs.
Supervisor Cluster Topology
A Supervisor requires:
1 control plane VM (Simple availability - the default when a Supervisor is activated during VCF workload-domain creation, scalable to 3 later) or 3 control plane VMs (High Availability, anti-affinity enforced, required for the Three Management Zone model): Each runs kubelet, API server, etcd shard, controller-manager, scheduler. These are 4-vCPU, 8GB RAM system pods on dedicated resource pool.
Worker ESXi nodes
: Regular cluster members host user workloads. Each gets kubelet agent injected.
vCenter registration
: Cluster must be managed by vCenter; standalone hosts not supported.
Networking stack
: NSX-T (preferred for network policies, GSLB) OR vSphere Distributed Switch + HAProxy/Avi (legacy path).
Persistent storage
: vSAN or external storage with vSphere CSI driver support.
Namespaces in Supervisor allow multi-tenancy with enforced quotas:
VM Classes
: Predefined CPU/RAM/storage combinations for pod requests (e.g., guaranteed-small: 2vCPU/4GB, guaranteed-xlarge: 8vCPU/32GB). Pod requests must match a defined class.
Storage Classes
: vSAN storage policies, vSAN file service for ReadWriteMany, external storage via CSI.
Content Library
: Required for TKR (Tanzu Kubernetes Release) image retrieval. Each Supervisor uses a shared library for base images.
RBAC + SSO
: Namespaces bound to vSphere SSO principals; kubectl access via vsphere login plugin.
TKG (Tanzu Kubernetes Grid) Cluster Provisioning
TKG clusters are provisioned on Supervisor using ClusterClass—a CRD-based declarative model:
ClusterClass
: Defines template for node image, control plane topology, worker node pools, CNI (Antrea), kubelet config.
Machine Deployments
: Scale-out worker nodes with autoscaling enabled. Topology labels for zone awareness (topology.kubernetes.io/zone).
Control Plane Scaling
: 1, 3, or 5 control plane VMs; etcd clustering supported at scale.
Kubernetes Version Management
: TKR images pinned (v1.27, v1.28, v1.29 available). Rolling cluster upgrades with node pool coordination.
: ALB Ingress Controller (Avi) for HTTP/HTTPS routing; NSX ALB required for production HA.
Service Mesh
: Istio/Anthos Service Mesh optional for advanced traffic management, mutual TLS.
Key Takeaways
In VCDX defense, be ready to justify NSX vs. VDS+HAProxy choice. NSX provides superior multi-site support, GSLB, and identity-based firewall. VDS+HAProxy is cost-effective for single-site labs but limits security design at scale.
Kubernetes RBAC bound to vSphere SSO users/groups:
kubectl vsphere login
: Single sign-on; kubeconfig auto-generates token bound to SSO session.
Role Bindings
: vSphere SSO group → Kubernetes ClusterRole/Role.
Audit Logging
: Audit Policy bound to vCenter event logging; track API calls, RBAC decisions.
Key Takeaways
For exam: Workload Management requires NSX-T Tier-0/Tier-1 (preferred) or VDS+external LB. Management network, control plane CIDR, node CIDR must not overlap.
For exam: Subscribed Content Library auto-syncs TKR versions. Capacity planning: 3-5GB per TKR version × number of versions. Always validate library sync before TKG cluster creation.
Design consideration: NSX+NCP provides tighter integration, GSLB, and federation support. Antrea is cost-effective for single-site, non-NSX deployments but limits multi-site HA design. In VCDX design doc, justify networking stack choice against SLA, multi-site requirements, and security posture.
In VCDX design, emphasize multi-tenant isolation strategy: combine namespace quotas, network policies, PSA enforcement, and RBAC. Document assumptions about trust boundaries (are all tenants internal? external SaaS?). Justify encryption-at-rest if regulated data present (PCI-DSS, HIPAA).
Three-Node Control Plane: HA & Leader Election via etcd
Supervisor clusters deploy one or three Kubernetes control-plane VMs on three separate ESXi hosts, providing N+1 fault tolerance. These VMs are not user-facing (they're system-managed) and run a minimal Linux OS with kubelet components.
Key Point:
In a High Availability Supervisor (three control-plane nodes), the three control-plane VMs form a single etcd cluster (port 2379-2380). Leader election ensures only one API server actively handles requests; the other two standby. If leader fails, etcd re-elects within 1-3 seconds.
vSphere Pods: Pod Isolation via Kata Containers
By default, pods on Supervisor run on TKG worker nodes (traditional Kubernetes VMs). Enable "vSphere Pods" — a pod class that runs each pod as its own minimal VM (via Kata containers), achieving near-bare-metal isolation. Trade-off: slower startup (1-2 sec vs milliseconds), higher resource footprint, but superior isolation.
TKG-on-Supervisor vs Supervisor-as-K8s
Two deployment patterns exist:
Pattern 1: TKG Clusters on Supervisor (Recommended for Multi-Tenancy)
Supervisor (3 control planes, system namespace for infra, no user pods directly)
+ TKG Cluster "prod-k8s-1" (user namespace ns-tenant-a, N worker nodes)
+ TKG Cluster "dev-k8s-1" (user namespace ns-tenant-b, M worker nodes)
Advantage: Namespace isolation, separate etcd per cluster, fault containment
Risk: Each TKG cluster = 3-node etcd + N workers = larger footprint
Supervisor (3 control planes + user-created TKG nodes via CAPI directly)
No TKG cluster CR; users deploy Kubernetes objects directly to Supervisor API
Supervisor manages TKG node VMs via CAPI MachineDeployment
Advantage: Single etcd, no nested cluster overhead, minimal infra
Risk: All workloads in shared namespace, lower isolation
Workload Management Enablement Prerequisites
Enabling Workload Management on a vSphere cluster requires these components present and healthy:
vSAN Cluster (Required)
All nodes in the vSphere cluster must have vSAN enabled (vSAN disk groups configured). Single-node vSAN is not supported for Supervisor. Minimum 3 nodes recommended for HA (Supervisor control planes spread across 3 hosts).
Distributed Virtual Switch (VDS)
All ESXi hosts must be connected to a VDS. VLAN or VXLAN uplink portgroups required for management network and pod network overlay. Standard vSwitch is not supported.
NSX or External Load Balancer
Supervisor API, Kubernetes API, and TKG service LoadBalancer endpoints need a load balancer. Options: NSX native (recommended, integrated DFW), or external LB (HAProxy, F5, Avi). LB must be layer 2 or 3 adjacent to management network.
Content Library (Subscribed)
Supervisor needs a subscribed content library to pull TKG OVA templates (Ubuntu, Photon OS). If offline, manual OVA upload. Check library sync status daily; out-of-date templates cause TKG creation failures.
DNS and NTP
All vSphere components and Supervisor must have synchronized clocks (within 1 sec). Forward/reverse DNS resolution for Supervisor API FQDN must work from all clusters. Forward lookup for vCenter FQDN required from ESXi.
Compute Resources
Minimum 3 ESXi hosts with 8 cores, 32 GB RAM each. Supervisor control-plane VMs consume 4 cores, 12 GB total. Plus TKG node capacity on same cluster (rule of thumb: reserve 20 percent headroom for Supervisor plus upgrades).
vSphere Network Setup
Supervisor Management Network (CIDR for control planes, assigned via network policy), TKG Service CIDR (pods), Service CIDR (ClusterIP ranges). No overlaps. Typical: 10.0.0.0/24 for mgmt, 10.1.0.0/16 for pod, 10.2.0.0/16 for service.
vSphere 8.0+ or vSphere 7 with U3+
Supervisor on vSphere 7 requires specific update levels; vSphere 8.0+ is standard. vCenter must be at same or newer version than ESXi.
Dependency Chain on Enablement Failure:
If content library sync fails, TKG creation blocked. If NSX LB down, new TKG service LoadBalancer assignments hang. If vSAN unhealthy, Supervisor control planes crash-loop. Always pre-flight all 8 requirements before attempting Workload Management enablement.
Supervisor Deployment Types and Control Plane Availability (VCF 9.0)
Simplified Supervisor Model — documented attributes: single control plane VM, single vNIC control plane VM, no load balancer. Benefit: supports running VM Service. Limitation, verbatim: 'Does not support running vSphere Pods and other Supervisor services.' It is real VCF 9.0 terminology, not a nickname.
Control Plane Availability is a separate axis:
├── Simple: ONE control-plane node. The default when a Supervisor is activated as part of VCF workload-domain creation. Scalable to three later. NOT available for the Three-Zone model.
└── High Availability: THREE control-plane nodes. Mandatory for the Three Management Zone model.
So 'a Supervisor always has three control-plane VMs' is FALSE for VCF 9.0 — three is the HA choice, not a universal rule. Related constraint: the Foundation Load Balancer (FLB) is VLAN-networking-only and is not supported with zonal HA. Source: techdocs VCF 9.0 Design — vSphere Supervisor Deployment Types.
Key Takeaways
Supervisor Cluster Won't Enable:Check: vSAN health (all hosts healthy), DNS resolves supervisor-api FQDN, NSX LB configured and reachable, Content Library subscribed and synced within 24h, all 8 hosts have 20 percent free disk space.
TKG Cluster Stuck in Creating:Check CAPI controllers: kubectl get events -A (look for Machine creation errors). Common: IP pool exhausted (increase IP range), VM class too small (min 2 CPU, 4 GB RAM), content library OVA stale (force resync).
kubectl vsphere login Fails:Check vSphere SSO health, user exists in AD, kubeconfig has correct server FQDN (ping supervisor-api). Proxy issues: if behind proxy, set HTTP_PROXY, HTTPS_PROXY env vars.
Pod Stuck in Pending:Check Pod events (kubectl describe pod), most common: CSI driver not ready (pvc pending), CNI plugin missing (no IP assigned), PSA blocking (image policy, securityContext). Verify StorageClass exists, CNI DaemonSet healthy.
Service LoadBalancer Stuck in Pending:NCP agent logs show NSX LB error (quota exceeded, wrong backend pool subnet). Check NSX LB capacity, verify service CIDR routable to NSX.
Exam: Supervisor control-plane count is ONE (Simple, the default via VCF workload-domain creation) or THREE (High Availability, required for Three Management Zones) — not 'always three'. The Simplified Supervisor Model is 1 CP VM, 1 vNIC, no load balancer, VM Service only (no vSphere Pods, no other Supervisor services).
Spherelet is a Kubernetes controller running on ESXi hosts that integrates Pod lifecycle with the hypervisor, enabling vSphere Pods. It's not a CNI plugin, monitoring agent, or NSX Edge service.
Production Supervisors deploy three Control Plane VMs in an HA set for etcd quorum. One provides no redundancy. Five is over-provisioned. Two with witness doesn't match Kubernetes HA patterns.
NSX-backed Supervisor provides NSX segments with a dedicated Tier-1 gateway per namespace, enabling overlay networking with full NSX security. VDS-backed uses distributed port groups. Antrea-only and host-only are not Supervisor backing options.
CAPV (Cluster API Provider vSphere) manages VKS cluster lifecycle via Cluster API CRDs, handling VM provisioning, scaling, and upgrades on vSphere. It doesn't scan images, run Harbor, or handle DNS resolution.
Tiny Supervisor sizing is appropriate for small dev/test labs with minimal resource requirements. XXL and custom XXL are for large production deployments. HA=7 is not a standard sizing option.
In a VCF 9.0 Supervisor, the Supervisor Control Plane VMs are deployed in an HA triplet. What happens if two of the three Control Plane VMs fail simultaneously?
The Supervisor continues fully operational because Kubernetes uses leader election
The Kubernetes API becomes unavailable because etcd loses quorum; Namespace operations and pod scheduling are interrupted until a CP VM is recovered
Spherelet takes over the control plane role on ESXi
vCenter becomes the Kubernetes control plane until recovery
If two of three CP VMs fail, etcd loses quorum (needs 2 of 3), making the Kubernetes API unavailable. Namespace operations and pod scheduling are interrupted. Spherelet doesn't take over CP. vCenter doesn't become the K8s control plane. Leader election needs quorum to function.
MachineDeployment replicas controls the number of worker nodes — editing this value scales the worker pool. StorageClass, ClusterIssuer, and HelmRelease are unrelated to node scaling.
cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size and -max-size annotations on MachineDeployments enable the Kubernetes Cluster Autoscaler. Other annotation patterns like kubernetes.io/role or metadata.autoscale are not valid.
MachineHealthCheck remediates unhealthy nodes by deleting and replacing the Machine object, triggering CAPI/CAPV to provision a new node. It doesn't just block scheduling, upgrade kernels, or cordon indefinitely.
In air-gapped environments, TKR images are imported locally into a private registry and referenced in cluster manifests. Public registry access isn't available. Not supported is incorrect. Building from source isn't required.
VM Service classes define compute specifications (CPU/memory) for VKS worker VMs, providing T-shirt sizing. Pod network policy, image credentials, and StorageClass provisioners serve different purposes.
An operator needs to perform a minor Kubernetes version upgrade on a VKS workload cluster (e.g., 1.29 -> 1.30). What is the correct upgrade sequencing?
Upgrade worker nodes first, then control plane
Upgrade the Supervisor to a version supporting the target TKR, subscribe the content library to the new TKR, then patch the Cluster/VKS cluster manifest
Directly edit the ClusterClass to bypass the Supervisor
Re-deploy the cluster from scratch; in-place upgrades are not supported
A MachineDeployment is annotated with cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size=3 and max-size=10. After a burst of pending pods, no new nodes are provisioned. Which is the most likely cause?
Cluster Autoscaler add-on is not installed on the workload cluster, so the annotations are ignored
If the Cluster Autoscaler add-on isn't installed on the workload cluster, the annotations are simply metadata with no effect — no component reads them. Anti-affinity, Supervisor locks, and TKR re-imports don't explain missing autoscaling when annotations are present but unused.
After the Ready=False condition persists past the 10m timeout, MHC deletes the Machine object, triggering CAPV to create a replacement node. It doesn't SSH to restart kubelet, just raise alerts, or drain-and-leave the node.
Air-gapped TKR ingestion: create a local subscribed content library, manually download TKR OVA files, import them, and associate the library with the Namespace. Internet isn't required. SMB mounts and DHCP option 66 are not valid methods.
Per KB 390634 the VCF component order is SDDC Manager > NSX > vCenter > Supervisor > ESXi, with TKRs and workload clusters last. NSX precedes vCenter — this is true in VCF 5.2 as well. In VCF 9.x the fleet/management layer (VCF Operations and related appliances) upgrades before SDDC Manager.
AntreaNetworkPolicy supports cluster-scoped policies and priorities across namespaces, unlike standard K8s NetworkPolicy which is namespace-scoped only. It supports both ingress and egress, isn't limited to L2 filtering, and uses label selectors.
Service type LoadBalancer on NSX-backed Supervisor is backed by Avi (NSX ALB) or native NSX LB for external IP allocation. HAProxy is for older TKG deployments. kube-proxy handles ClusterIP only. MetalLB is for bare-metal K8s.
Contour/Envoy provides Ingress controller functionality with path/host-based routing, TLS termination, and HTTP routing rules. It's not a CSI driver, Antrea replacement, or Cluster API provider.
Antrea Traceflow captures the packet walk between pods for troubleshooting, similar to NSX Traceflow but within the K8s overlay. It doesn't rotate certificates, replace kube-proxy, or upgrade Supervisor.
Each vSphere Namespace with NSX-backed networking gets a dedicated Tier-1 gateway and segments, providing network isolation per namespace. Shared segments, no isolation, and VLAN-only don't provide the NSX-backed isolation model.
Centralized Transit Gateway (CTGW) connects multiple VPCs through a common gateway for shared services and transit connectivity. It doesn't replace Antrea, isn't a StorageClass, and doesn't scan container images.
Creating a LoadBalancer Service dynamically provisions a virtual server on Avi or NSX native LB with the external IP allocated from the Ingress CIDR pool. It doesn't create DHCP segments, VRF gateways, or redeploy Edge clusters.
Two pods in different Namespaces of the same workload cluster cannot communicate. An AntreaClusterNetworkPolicy allows traffic. What is the next check?
Verify there is no Kubernetes NetworkPolicy denying the flow, because policies are additive and any deny by a policy that selects the pod takes effect
Kubernetes NetworkPolicy is additive — if any policy selects a pod and doesn't explicitly allow the flow, it's denied. Even if an AntreaClusterNetworkPolicy allows it, a namespace-scoped K8s NetworkPolicy deny takes effect. Restarting kube-proxy, rotating tokens, or changing CNI won't fix policy conflicts.
CTGW aggregates routing and north-south connectivity for multiple VPCs in a Project and connects upstream to the Tier-0 gateway. Each VPC doesn't have its own CTGW. CTGW doesn't replace NSX Manager or perform L2 bridging.
Envoy proxy managed by Contour terminates TLS using the TLS secret referenced in the Ingress/HTTPProxy resource. Application pods don't need to handle TLS when using Ingress. kube-proxy doesn't do TLS. NSX Edge TLS Passthrough is a different mechanism.
vSphere CSI driver references a StorageClass mapped to a vSAN storage policy for dynamic PV provisioning. NFS subdir provisioner is for external NFS. HostPath and local static PVs aren't dynamically provisioned by CSI.
RWX (ReadWriteMany) is supported by vSAN File Services-backed PVs. RWO is for block volumes. ROX is read-only many. vSAN File Services enables shared access via NFS v4.1.
Volume expansion requires allowVolumeExpansion: true on the StorageClass AND CSI driver support. reclaimPolicy controls what happens after PVC deletion. mountOptions define mount flags. Default settings don't enable expansion.
CNS (Cloud Native Storage) manages the PVC-to-VMDK lifecycle and provides visibility in vCenter for persistent volumes. It doesn't replace ESA, run Helm, or scan images.
A VolumeSnapshot requires a VolumeSnapshotClass referencing the CSI driver and snapshotter CRD. Just a PV, pod annotation, or StorageClass without provisioner is insufficient.
vSAN File Services-backed PVC with RWX supports shared access from multiple pods for stateful web uploads. Block RWO limits access to one pod. HostPath isn't shared across nodes. EmptyDir is ephemeral.
Online volume expansion requires allowVolumeExpansion=true, CSI driver support for online expansion, and vSAN policy allowing capacity growth. VMFS-6, PVC recreation, and vCenter upgrades are not the correct requirements.
Pod Security Admission 'restricted' enforces the strictest hardened baseline, prohibiting privileged containers, host namespaces, and dangerous capabilities. 'No policy' is baseline. Privileged defaults are the opposite. Network policies are separate.
Harbor provides a private container image registry with RBAC, vulnerability scanning (Trivy/Clair), image replication, and signing (Notary). It's not limited to replacing kube-proxy, providing CSI, or monitoring.
Velero backs up and restores cluster resources (manifests) and PV snapshots for disaster recovery. It doesn't handle autoscaling, ingress routing, or Antrea Traceflow.
Pending PVCs most often result from no matching StorageClass or unavailable storage policy/capacity. Incorrect pod security, missing ingress routes, and HPA misconfig don't affect PVC binding.
RoleBinding referencing a Role scoped to a namespace grants fine-grained access to that namespace only. ClusterRoleBinding to cluster-admin is overprivileged. No RBAC is insecure. Pinniped handles authentication, not authorization.
Velero backs up a namespace that contains PVCs. For vSAN-backed volumes using CSI snapshots, which requirement is needed for application-consistent backups of a MySQL StatefulSet?
Snapshots alone guarantee consistency
Use Velero backup hooks (pre/post freeze commands) or CSI snapshots combined with an application-quiescing tool such as fsfreeze or DB flush/lock scripts
A cluster admin applies the Pod Security Admission label pod-security.kubernetes.io/enforce=restricted to a namespace. Which pod spec is MOST likely to be rejected?
A pod with runAsNonRoot=true and no capabilities
A pod requesting hostNetwork=true and privileged=true
A pod with hostNetwork=true and privileged=true violates the 'restricted' PSA profile which prohibits host namespaces and privileged access. runAsNonRoot, readOnlyRootFilesystem, and resource limits are compliant with restricted.
If Harbor scanner (Trivy/Clair) is not installed or configured, scan capability is unavailable regardless of user roles. The developer role is sufficient to trigger scans. Image tag 'latest' and Notary signing are unrelated to scan availability.
First check the Supervisor LoadBalancer IP pool for available addresses and Ingress/Egress CIDR exhaustion — IP exhaustion is the most common cause of pending LoadBalancer Services. Pod deletion, kubelet restart, and Edge reboot won't fix IP allocation issues.
RoleBinding referencing a Role with verbs [get, list, watch] on pods in a specific namespace provides read-only pod access. ClusterRoleBinding with cluster-admin is over-privileged. RoleBinding to 'edit' grants write access. ServiceAccount without RBAC has no permissions.
Authentication succeeded (OIDC tokens are valid) but no RBAC binding grants permissions — the user/group needs a RoleBinding or ClusterRoleBinding to access cluster resources. Pinniped doesn't need NFS. The kubeconfig is valid if auth succeeded.
Prometheus scrapes metrics from kube-state-metrics but not from cAdvisor. Which configuration is most likely missing?
The ServiceMonitor (or scrape config) targeting the kubelet /metrics/cadvisor endpoint, and the correct RBAC permissions for the Prometheus service account to access nodes/metrics
The ServiceMonitor (or scrape config) targeting kubelet /metrics/cadvisor endpoint plus RBAC for the Prometheus service account to access nodes/metrics is missing. Grafana dashboards, Alertmanager webhooks, and Harbor certs don't affect metric scraping.
The vSphere cluster where VKS is enabled, with its Control Plane VMs running Kubernetes APIs and its ESXi hosts running Spherelet. It hosts vSphere Pods and manages downstream VKS guest clusters. It is the foundational layer of the three-layer VKS architecture.
Spherelet
A kubelet-equivalent process running on each ESXi host when the cluster is enabled as a Supervisor. It enables running pods natively on ESXi inside a CRX (Container Runtime Executive). Spherelet integrates Kubernetes scheduling with vSphere. Note: vSphere Pods are deprecated in VCF 9; VKS workloads primarily run in Tanzu Kubernetes guest clusters.
CRX (Container Runtime Executive)
A minimal, purpose-built VM that ESXi uses to isolate vSphere Pods at hardware-virtualization strength. CRX boots in under a second and carries only what a container needs. It provides VM-strong isolation for pods without a conventional guest OS.
Supervisor Service
A management-plane extension running on the Supervisor that provides platform capabilities (VKS, Contour, Harbor, External-DNS, LCI). Services are installed via the Supervisor Services catalog in vCenter. They are the pluggable ecosystem of VKS.
Cluster API Provider vSphere (CAPV)
The Kubernetes Cluster API implementation that provisions and lifecycles VKS guest clusters declaratively from the Supervisor. CAPV translates ClusterClass specs into VMs and kubeadm configuration. It is the engine behind 'create cluster' operations.
VDS-Backed Supervisor
A Supervisor networking mode that uses plain VDS port groups plus an external load balancer (Avi/HAProxy) instead of NSX. It is lighter-weight but lacks NSX-native isolation and service chaining. It is suitable for VVF-adjacent VKS-only scenarios.
Supervisor Control Plane VMs
Three control-plane VMs deployed by vCenter on the Supervisor cluster, running kube-apiserver, etcd, scheduler, controller-manager, plus VMware-specific controllers (VM-operator, CAPV, NCP/AKO).
vSphere Pod (vSphere Native Pod)
A legacy NSX-only Supervisor construct running pods directly on ESXi via CRX. Deprecated in favor of TKG clusters running on VM-operator VMs; still appears on exam objectives for older paths.
Namespace (Supervisor)
The primary multi-tenancy object in the Supervisor, binding user permissions, storage policies, VM classes, content libraries, and resource limits. Each namespace maps to a Kubernetes namespace on the Supervisor.
VM Class
A template describing the CPU/memory/GPU reservation of a VM provisioned by VM-operator or used as TKG node. Bound to namespaces; users pick a class when creating clusters. Custom classes allow right-sizing of TKG worker pools.
ClusterClass
A reusable Kubernetes Cluster API template defining worker/control-plane topology, machine specs, and add-ons, instantiated by Cluster objects. VKS ships curated ClusterClasses but customers can author their own. It is the blueprint of a VKS guest cluster.
MachineDeployment
A Cluster API resource that manages a set of worker-node VMs with desired replica count, similar to a Kubernetes Deployment but for VMs. Changing replicas triggers Spherelet to provision or drain nodes. It is the scaling primitive for VKS guest workers.
Tanzu Kubernetes Release (TKR)
A versioned bundle of Kubernetes, OS, and vSphere-supplied components used as the base for VKS guest cluster nodes. New TKRs bring upstream K8s versions and security fixes. Cluster upgrades consume newer TKRs via CAPV.
MachineHealthCheck
A Cluster API resource that monitors node conditions and automatically remediates unhealthy machines by replacement. It is the self-healing mechanism for VKS. It is configured with thresholds so transient issues do not trigger needless churn.
VM Operator
A Supervisor component that translates Kubernetes VM custom resources into vSphere VMs with full lifecycle management. It lets teams request VMs via kubectl, bridging VM and container workflows. It is closely related to VM Service classes.
Cluster Autoscaler (VKS)
An add-on that watches pending pods in guest clusters and scales MachineDeployment replicas up or down within configured min/max annotations. It translates pod demand into node capacity automatically. It is the standard mechanism for elastic worker pools.
TKG Control Plane Upgrade
A ClusterClass-driven rolling update that replaces control-plane VMs one at a time using a new TKR. etcd quorum is preserved as each node is cordoned, drained, and replaced.
Worker Rolling Update Strategy
MachineDeployments default to rolling update with configurable maxSurge and maxUnavailable; MachineHealthCheck can auto-remediate unhealthy nodes during the update.
MachineSet
The CAPI object that maintains a set of identical Machines (VMs); owned by a MachineDeployment. Scaling the MD creates/removes MachineSets during rolling updates.
VM Service (VM Operator)
Supervisor service that lets namespace users declaratively request standalone VMs via 'VirtualMachine' and 'VirtualMachineService' CRDs; VMs consume the same VM classes and images as TKG nodes.
Antrea CNI
The default container-network interface in VKS, providing pod networking, NetworkPolicies, and cluster-scoped policies using Open vSwitch. It is chosen for its NSX integration and visibility features. It powers both cluster and namespace-scoped network controls.
AntreaNetworkPolicy (ANP)
A namespace-scoped Kubernetes CRD that extends the standard NetworkPolicy with tiered rules, drop/reject actions, and richer selectors. It provides the expressiveness of enterprise firewalls inside Kubernetes. It is the recommended policy type for Antrea users.
ClusterNetworkPolicy (CNP)
A cluster-scoped Antrea CRD that enforces network rules across all namespaces, typically for platform-level security baselines. It is evaluated before namespace-scoped policies. It is administered by platform teams, not tenants.
LoadBalancer Service (VKS)
A Kubernetes Service type that VKS fulfills by requesting a VIP on the integrated load balancer (Avi or NSX LB). Each LB service consumes a VIP and backend pool. It is the typical ingress for non-HTTP workloads.
Contour Ingress
The default VKS Ingress controller (installed as a Supervisor Service) that translates Kubernetes Ingress and HTTPProxy resources into Envoy proxy configuration. It provides L7 HTTP/HTTPS routing at the cluster edge. It integrates with Harbor and External-DNS.
Antrea Traceflow
An Antrea-specific live packet-trace capability that shows each hop a packet takes through the CNI, including policy matches. It is invaluable when a pod can't reach another pod and NetworkPolicy may be involved. Its output is rendered both in YAML and in the Antrea UI.
NCP (NSX Container Plugin)
Controller that integrates Kubernetes with NSX, creating Tier-1 routers, segments, LBs, and DFW rules for namespaces, pods, and services. Required for NSX-backed Supervisors.
Egress IP Pool
Pool of NSX IPs from which namespace workloads SNAT when leaving the cluster. Allows firewall rules upstream to identify traffic by namespace.
Antrea Proxy Mode
Antrea can implement kube-proxy replacement in-kernel using OVS, improving Service throughput vs iptables. Proxy-all mode extends this to all Service types.
NetworkPolicy Object Ordering
Kubernetes NetworkPolicy is additive (union of allow rules). AntreaClusterNetworkPolicy adds tiers and priorities, enabling deny/allow ordering similar to NSX DFW categories.
vSphere CSI Driver
The Container Storage Interface implementation that provisions Kubernetes PersistentVolumes as vSAN or VMFS VMDKs on vSphere. It is installed automatically in VKS. It supports dynamic provisioning, snapshots, and expansion.
StorageClass (VKS)
A Kubernetes resource mapping a friendly name (for example, 'gold-ssd') to a vSAN storage policy plus access-mode defaults. StorageClasses decouple application manifests from storage-backend specifics. They are the tenant-visible storage catalog.
ReadWriteMany (RWX) Volume
A Kubernetes access mode that allows multiple pods on multiple nodes to mount the same volume concurrently, implemented in VKS via vSAN File Services-backed NFS. It is used for shared-config workloads and legacy apps. Not all StorageClasses support RWX.
VolumeSnapshot
A Kubernetes CRD that captures a point-in-time copy of a PersistentVolume, implemented in vSphere CSI as an underlying vSAN snapshot. It enables database backup-and-restore workflows without leaving kubectl. It requires a VolumeSnapshotClass bound to the right CSI driver.
Cloud Native Storage (CNS)
The vSphere subsystem that tracks Kubernetes PVCs as first-class vCenter objects, exposing relationship and metric data to Ops teams. It is what enables VCF Operations to show per-PVC capacity and performance. It is the backplane behind every VKS PVC.
PVC Binding Mode
A StorageClass attribute controlling when a PVC is bound to a PV: Immediate provisions at PVC creation, WaitForFirstConsumer delays until a pod is scheduled. WaitForFirstConsumer is recommended for topology-aware placement. It is a common VKS exam topic.
Storage Policy Mapping to StorageClass
vSphere storage policies exposed to a namespace appear as Kubernetes StorageClasses; PVCs reference them by name. Policy compliance is visible in vCenter and reported back to CNS.
CNS Volume Types
CNS backs PVCs with First-Class Disks (FCDs) on vSAN/VMFS/NFS and RWX via vSAN File Services. Type is chosen by StorageClass parameters (datastoreURL, fsType, CSI topology).
Volume Expansion
Online expansion of PVCs supported when StorageClass 'allowVolumeExpansion: true' and underlying datastore supports it. CSI driver orchestrates filesystem resize inside the pod.
VolumeSnapshotClass
CRD defining snapshot parameters (snapshotter, deletionPolicy). Binds CSI snapshotter to underlying FCD snapshots; used by VolumeSnapshot objects for per-PVC backups.
vSphere SSO Integration (VKS)
The mechanism that lets vCenter SSO users authenticate to VKS clusters via Kubernetes RBAC with Viewer/Editor/Owner role mapping. It avoids separate Kubernetes user stores for platform users. It is the default auth path for VKS in VCF.
Pinniped
An open-source authentication add-on that bridges external IdPs (OIDC, LDAP) into Kubernetes authentication for VKS. It issues short-lived credentials to kubectl. It is the recommended way to federate enterprise identities into tenant clusters.
Pod Security Admission (PSA)
The built-in Kubernetes admission controller replacing PodSecurityPolicy, evaluating pods against Privileged, Baseline, and Restricted profiles. Namespaces opt into a profile via labels. PSA is the default pod-security control in modern Kubernetes and VKS.
Harbor Registry (VKS)
The VKS-integrated private container registry Supervisor Service that stores and scans images for vulnerabilities with Trivy. It supports replication, signing, and RBAC. It is the preferred image source for VKS workloads.
Velero Backup
A Kubernetes-native backup tool that snapshots resources and PVs, commonly deployed for VKS with vSphere CSI snapshot support. It supports scheduled backups and cross-cluster restores. It is the standard Kubernetes BCDR utility on VKS.
ClusterRoleBinding
A Kubernetes RBAC resource that grants a ClusterRole (a cluster-wide permission set) to users, groups, or service accounts. It contrasts with RoleBinding, which is namespace-scoped. Correct use is essential for least-privilege VKS administration.
Workload Management Roles
vSphere SSO roles that expose VKS: 'Can Edit' (namespace owner), 'Can View' (read-only), plus custom roles built from Namespace-level privileges. Assigned via Namespace → Permissions.
Image Signing / Registry Scan
Harbor (Supervisor Service) supports Notary/Cosign signature validation and Trivy image scanning; pull policies can require signed images for a project.
Cluster Audit Logging
API-server audit log levels configured on TKG control planes via ClusterClass patches; shipped via a sidecar or Fluent Bit to external syslog/SIEM for compliance.
TKG Backup with Velero Plugin for vSphere
Plugin enables snapshot-based PV backup using CNS snapshots in parallel with Velero object backup. Required for application-consistent backups of stateful workloads.