Lab: Troubleshoot K8s Pod Performance Issue
Objectives
- Diagnose pod performance problem: app is slow, determine if vSphere or K8s layer issue.
Prerequisites
VCF lab environment deployed and operational
Lab Environment
Standard VCF lab environment for vSphere Foundation 9.0 Support (Support Specialist)
Tasks
Task 1 Lab: Troubleshoot K8s Pod Performance Issue
Diagnose pod performance problem: app is slow, determine if vSphere or K8s layer issue.
Deploy slow-app pod on TKG cluster (or use existing)
User reports: "App is slow, taking 10s per request"
Check K8s layer: kubectl logs pod-name (see if errors)
Check pod resources: kubectl describe pod, see CPU/memory requests
Check node resources: kubectl top nodes (CPU/memory available?)
Check PVC/storage: kubectl describe pvc (latency issue?)
Check vSphere layer: vCenter esxtop on worker VM host (CPU/disk latency?)
Correlate: Is bottleneck vSAN latency, K8s resource contention, or app code?
Make recommendation: increase RAID FTT, add pod replicas, or app code fix
Validation Gate
Check: Diagnose a slow pod: check pod resource usage (kubectl top), worker node VM performance (esxtop), vSAN datastore latency, and NSX connectivity. Identify root cause layer.
Expected: Root cause identified at specific layer: pod resource limits, worker node CPU/memory contention, vSAN I/O latency, or NSX network issue. Remediation applied at the correct layer.
Common Errors
Final Validation
Lab completed successfully
✓ All steps completed → No errors observed
Cleanup / Restore
• Revert to snapshot if needed
Design Reflection (VCDX)
Container workload troubleshooting on VCF demonstrates full-stack operational knowledge. VCDX panelists may present a Kubernetes performance scenario to test whether you can diagnose across VM, storage, and network layers — not just within Kubernetes.
Requirements
- Troubleshoot pod performance across the full stack
- Correlate Kubernetes events with infrastructure events
- Understand persistent volume performance on vSAN
Constraints
- Pod resource requests/limits must be configured for predictable performance
- Worker node VM sizing constrains pod density
- vSAN performance affects all PVC-backed pods
Assumptions
- kubectl access is available for Kubernetes diagnostics
- VCF Operations and vCenter provide correlated infrastructure metrics
Risks
- Misdiagnosing pod-level symptom when root cause is infrastructure
- Pod density exceeding worker node VM capacity causing noisy neighbor issues