Academy/vSphere Foundation 9.0 Support (2V0-18.25)/Lab: Diagnose Performance Issue with esxtop & Logs
This lab targets VCF 9.0

Lab: Diagnose Performance Issue with esxtop & Logs

VCF 9.0Intermediatevcp-foundation⏱ 135 min

Objectives

  • Simulate performance issue, use esxtop to identify bottleneck, collect logs, analyze and recommend fix.

Prerequisites

VCF lab environment deployed and operational

Lab Environment

Standard VCF lab environment for vSphere Foundation 9.0 Support (Support Specialist)

Tasks

Task 1 Lab: Diagnose Performance Issue with esxtop & Logs

esxtop is the essential real-time performance diagnostic tool for ESXi. Understanding the key metrics in each view (CPU, memory, network, disk) and what constitutes abnormal values is the foundation of performance troubleshooting.

Simulate performance issue, use esxtop to identify bottleneck, collect logs, analyze and recommend fix.

Step 1

SSH to ESXi host, run esxtop in CPU view

Step 2

Note current utilization (baseline)

Step 3

Initiate heavy I/O workload (stress-ng or fio on VM)

Step 4

Return to esxtop, observe latency and queue depth increasing

Step 5

Switch to disk view, identify bottleneck disk

Step 6

Switch to memory view, check for ballooning/swapping

Step 7

Generate vm-support bundle: vm-support -w /tmp/host.tgz

Step 8

Extract and review vmkernel.log for errors during workload

Step 9

Document findings: bottleneck cause, recommendation (add disk, upgrade NIC, etc.)

Validation Gate

Check: Using esxtop, identify: which VM has highest CPU ready time, whether any host memory pressure exists (ballooning/swapping), and current disk latency per datastore

Expected: CPU: %RDY <5% for all VMs (or identified high-RDY VMs for investigation). Memory: SWCUR=0 and MCTLSZ=0 (no pressure). Disk: GAVG <20ms per datastore (or identified high-latency datastores).

Common Errors

Reading esxtop %USED as VM CPU utilization
Fix: In esxtop CPU view: %USED = physical CPU time used by the VM (can exceed 100% on multi-vCPU VMs). %RDY (ready time) is more important — it shows time the VM wanted CPU but couldn't get it. %RDY > 5% indicates CPU contention. %CSTP (co-stop) > 3% indicates vSMP scheduling issues (reduce vCPU count).
Ignoring memory metrics beyond active/consumed
Fix: esxtop memory view shows: MCTLSZ (balloon driver reclaimed — host is under memory pressure), SWCUR (swapped to disk — performance is severely degraded), ZIP/UNZIP (compressed pages — moderate pressure). If SWCUR > 0 for any VM, the host is critically low on memory. MCTLSZ > 0 indicates ballooning — review memory allocation.
Not capturing esxtop batch mode output for historical analysis
Fix: Real-time esxtop shows current state only. For performance issues that are intermittent, use batch mode: 'esxtop -b -d 5 -n 720 > /tmp/esxtop.csv' (captures 5-second intervals for 1 hour). Analyze CSV in Excel or vscsiStats. Batch mode captures the data you need for Broadcom support cases.
Looking at only one resource dimension
Fix: Performance issues are often multi-dimensional: a VM with high CPU ready may actually be caused by storage latency (VM waiting on I/O completion holds CPU scheduling). Check all four dimensions: CPU (%RDY, %CSTP), Memory (SWCUR, MCTLSZ), Disk (GAVG/cmd latency), Network (dropped packets). The root cause is often in a different dimension than the symptom.

Final Validation

Lab completed successfully

✓ All steps completed → No errors observed

Cleanup / Restore

• Revert to snapshot if needed

Design Reflection (VCDX)

Performance troubleshooting methodology is tested in VCDX defense. Panelists may present a performance scenario and expect you to describe your diagnostic approach — esxtop is the first tool you should mention.

Requirements

  • Use esxtop to diagnose CPU, memory, disk, and network performance
  • Capture batch mode output for historical analysis
  • Correlate multi-dimensional metrics for root cause analysis

Constraints

  • esxtop shows real-time data only — use batch mode for intermittent issues
  • Performance diagnosis requires ESXi SSH access
  • esxtop output is complex — know key metrics per dimension

Assumptions

  • Baseline performance metrics are known for comparison
  • ESXi SSH access is available during troubleshooting

Risks

  • Misdiagnosing symptom dimension as root cause (CPU symptom with storage root cause)
  • Missing intermittent issues without batch mode capture

⚠ Known Pitfalls (from Community KB)

Treating %USED as the primary CPU health indicator — %RDY (ready time) is far more important for identifying CPU contention.
Checking only one resource dimension — performance issues are often cross-dimensional (disk latency causing CPU ready time increase).
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.