Academy/vSphere Foundation 9.0 Support (2V0-18.25)/Lab: Troubleshoot K8s Pod Performance Issue
This lab targets VCF 9.0

Lab: Troubleshoot K8s Pod Performance Issue

VCF 9.0Intermediatevcp-foundation⏱ 135 min

Objectives

  • Diagnose pod performance problem: app is slow, determine if vSphere or K8s layer issue.

Prerequisites

VCF lab environment deployed and operational

Lab Environment

Standard VCF lab environment for vSphere Foundation 9.0 Support (Support Specialist)

Tasks

Task 1 Lab: Troubleshoot K8s Pod Performance Issue

Kubernetes pod performance troubleshooting on VCF Supervisor requires understanding the full stack: pod → container runtime → worker node VM → ESXi host → vSAN/network. Performance issues can originate at any layer.

Diagnose pod performance problem: app is slow, determine if vSphere or K8s layer issue.

Step 1

Deploy slow-app pod on TKG cluster (or use existing)

Step 2

User reports: "App is slow, taking 10s per request"

Step 3

Check K8s layer: kubectl logs pod-name (see if errors)

Step 4

Check pod resources: kubectl describe pod, see CPU/memory requests

Step 5

Check node resources: kubectl top nodes (CPU/memory available?)

Step 6

Check PVC/storage: kubectl describe pvc (latency issue?)

Step 7

Check vSphere layer: vCenter esxtop on worker VM host (CPU/disk latency?)

Step 8

Correlate: Is bottleneck vSAN latency, K8s resource contention, or app code?

Step 9

Make recommendation: increase RAID FTT, add pod replicas, or app code fix

Validation Gate

Check: Diagnose a slow pod: check pod resource usage (kubectl top), worker node VM performance (esxtop), vSAN datastore latency, and NSX connectivity. Identify root cause layer.

Expected: Root cause identified at specific layer: pod resource limits, worker node CPU/memory contention, vSAN I/O latency, or NSX network issue. Remediation applied at the correct layer.

Common Errors

Troubleshooting only inside the pod without checking the underlying infrastructure
Fix: K8s pod performance depends on the underlying VM (worker node) which depends on ESXi host resources. A pod with high CPU throttling may be caused by: pod resource limits too low, worker node VM CPU contention (check esxtop %RDY), or ESXi host over-commitment. Diagnose bottom-up: ESXi host → worker node VM → pod.
Not checking Kubernetes resource requests and limits
Fix: Pods without resource requests are best-effort — they get evicted first during resource pressure. Pods without limits can consume unlimited resources, starving other pods. Check: kubectl describe pod <name> for resource requests/limits. Best practice: always set both requests (guaranteed minimum) and limits (maximum allowed).
Ignoring persistent volume performance when diagnosing pod latency
Fix: Pods using persistent volumes (PVCs) on vSAN may experience I/O latency from: vSAN resync activity, storage policy mismatch (RAID-1 vs RAID-5 for the workload type), or datastore congestion. Check vSAN performance metrics for the specific datastore backing the PVC.
Not correlating pod events with infrastructure events
Fix: A pod restart may correlate with: ESXi host maintenance mode (VM migrated, pod restarted), vSAN resync (I/O latency spike caused pod health check failure), or NSX network interruption (pod lost connectivity). Check kubectl events alongside vCenter events and VCF Operations timeline for the same time window.

Final Validation

Lab completed successfully

✓ All steps completed → No errors observed

Cleanup / Restore

• Revert to snapshot if needed

Design Reflection (VCDX)

Container workload troubleshooting on VCF demonstrates full-stack operational knowledge. VCDX panelists may present a Kubernetes performance scenario to test whether you can diagnose across VM, storage, and network layers — not just within Kubernetes.

Requirements

  • Troubleshoot pod performance across the full stack
  • Correlate Kubernetes events with infrastructure events
  • Understand persistent volume performance on vSAN

Constraints

  • Pod resource requests/limits must be configured for predictable performance
  • Worker node VM sizing constrains pod density
  • vSAN performance affects all PVC-backed pods

Assumptions

  • kubectl access is available for Kubernetes diagnostics
  • VCF Operations and vCenter provide correlated infrastructure metrics

Risks

  • Misdiagnosing pod-level symptom when root cause is infrastructure
  • Pod density exceeding worker node VM capacity causing noisy neighbor issues

⚠ Known Pitfalls (from Community KB)

Diagnosing only inside Kubernetes when the root cause is ESXi host contention or vSAN latency — always check the full stack.
Running pods without resource requests and limits — unpredictable performance and eviction behavior during resource pressure.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.