Performance Analysis with esxtop
Objectives
- Launch and navigate esxtop views (CPU, memory, disk I/O, host-level analysis)
- Interpret CPU metrics (%USED, %RDY, %CSTP, %SYS, %WAIT) and identify contention patterns
- Analyze memory reclamation (balloon, swapping, compression) and detect memory pressure
- Diagnose storage I/O latency and correlate with backend infrastructure
- Perform trend analysis using esxtop batch mode for baseline documentation
- Recommend evidence-based remediation actions based on metric analysis
- Document performance findings using RCAR methodology
Prerequisites
VCF lab environment deployed and operational with ESXi hosts
Required skills:
- SSH access to ESXi host
- Basic understanding of CPU scheduling and memory management
- Familiarity with performance baseline concepts
Lab Environment
Standard VCF lab environment for VCF 9.0 Support with multi-tier VM workloads
Tasks
Task 1 CPU Performance Analysis and Contention Detection
Master esxtop CPU view and interpret multi-dimensional contention signals to identify scheduling issues and recommend vCPU optimization.
SSH to an ESXi host in your lab. Launch esxtop by typing 'esxtop' and wait for the interactive interface to load.
Press 'c' to enter the CPU view (should already be the default). Examine the column headers. Identify which rows represent VMs vs. host kernel.
Locate a VM or workload with measurable CPU usage (ideally 30%+). Note its name, current %USED, %RDY, and %CSTP values. Document the current state in your lab notes.
Analyze the scenario: A VM shows %USED=65%, %RDY=20%, %CSTP=2%. Interpret each metric. High %RDY indicates CPU contention at the host level. %CSTP=2% on a multi-vCPU VM suggests scheduling delays for SMP alignment.
Press 'U' or 'v' to switch to a breakdown view showing per-vCPU details (if available in your esxtop version). Examine whether all vCPUs are equally loaded or if some idle while others busy.
Recommend remediation actions based on your analysis: (a) Reduce the VM's vCPU count by 1-2 cores to lower co-stop, (b) Enable CPU affinity if the workload is memory-bound, (c) Check and disable CPU affinity rules that may be forcing scheduling to specific cores, (d) Review DRS settings—ensure resource pools are not over-committed.
Document your CPU analysis in a structured format: Observation (metric values), Interpretation (what the metrics indicate), Action (specific change or next step), Rationale (why this action addresses the root cause).
Validation Gate
Check: CPU analysis complete with documented metrics and remediation plan
Expected: Lab notes show interpreted esxtop metrics, identified contention pattern, and evidence-based remediation recommendation
Common Errors
Task 2 Memory Performance Analysis and Pressure Detection
Interpret memory reclamation metrics (balloon, swapping, compression) to identify memory pressure and recommend resource allocation adjustments.
While in esxtop, press 'm' to switch to the memory view. Allow the view to render; it may take a few seconds.
Examine the memory view for a VM allocated 64GB. Note MCTLSZ, ACTVMB, SWCUR, and SWTGT values. Document the current state.
Introduce memory pressure scenario: Assume MCTLSZ=32GB (VM allocated 64GB, balloon reclaimed 32GB) and SWCUR=1.2GB. Interpret: Balloon is reclaiming 32GB (50% of allocation) and host is swapping 1.2GB. This indicates severe memory pressure.
Check ZIP and UNZIP metrics. High ZIP (compression) with high UNZIP (decompression) indicates the host is relying on memory compression as a stop-gap, which has CPU overhead. Estimate CPU cost: ~5% of host CPU per 100 ZIP/sec.
Recommend memory remediation: (a) Add physical RAM to the host(s), (b) Reduce the number of VMs, (c) Lower memory reservations to allow overcommit detection, (d) Migrate workloads to less-pressured hosts via DRS, (e) Audit guest OS memory settings for misconfiguration (e.g., huge pages, buffer pools).
Document memory pressure indicators and decision thresholds in your notes: (1) MCTLSZ > 10% of allocation = monitor closely. (2) MCTLSZ > 20% = action required. (3) SWCUR > 0 = emergency. (4) ZIP > 200/sec = investigate.
Validation Gate
Check: Memory analysis complete with interpreted pressure indicators and ranked remediation plan
Expected: Lab notes document memory metrics, pressure assessment, and actionable remediation with cost/benefit trade-offs
Common Errors
Task 3 Storage and Disk I/O Analysis
Diagnose storage latency using esxtop disk views and correlate metrics with vSAN or backend storage health.
From esxtop, press 'u' to switch to disk device view (physical/logical disk). Wait for the view to render.
Examine your storage device list. Identify the primary data store used by your lab VMs. Note the baseline DAVG. A healthy SSD-backed datastore should show DAVG < 5ms. HDD-backed should show DAVG < 10ms.
Introduce disk I/O stress: Run a heavy I/O workload on a guest VM (e.g., 'fio --name=randrw --ioengine=libaio --iodepth=32 --rw=randrw --bs=8k --runtime=300'). Monitor esxtop disk view in real-time.
Interpret the scenario: DAVG=45ms with baseline 5ms = 9x storage latency spike. Analyze root cause. In vSAN, check: (a) Is this a vSAN rebuild scenario (rebuilding a failed disk/host)? (b) Are vSAN resync operations active? (c) Is the storage network saturated (vSAN uses dedicated network)?
If your lab uses vSAN, SSH to a vSAN-enabled host and run vSAN-specific metrics: 'vsanstats -a' or 'vdq -l' (from vdq utility in vSAN Build 2.0+) to check disk group health, rebalance operations, and physical disk latency.
Press 'v' (VM disk view) to see disk latency from the VM perspective (GAVG). Identify which VM is experiencing the worst latency. GAVG should match or exceed DAVG (since GAVG includes VM OS scheduler overhead).
Recommend storage remediation based on your analysis: (a) For vSAN rebuild scenarios, wait for rebuild to complete and re-measure latency. (b) For persistent high DAVG, increase storage capacity (add disks to vSAN disk groups, upgrade SAN array). (c) Check network for vSAN: ensure vSAN network (usually VLAN 1647/1648) is not congested. (d) Consider QoS policies to throttle non-critical I/O during peak periods.
Validation Gate
Check: Storage I/O analysis complete with diagnosed latency spike and remediation recommendation
Expected: Lab notes document esxtop disk metrics, baseline vs. stressed state comparison, vSAN-specific diagnostics (if applicable), and tiered remediation plan
Common Errors
Task 4 Host-Level Analysis and Batch Mode Trending
Analyze host-level resource saturation and establish performance baselines using esxtop batch mode for trend analysis and documentation.
Return to esxtop interactive CPU view ('c'). Observe the host kernel metrics (typically listed first or in a specific row). Note host %USED and %SYS. Host %USED near 85%+ indicates resource saturation; %SYS > 8% indicates high system overhead (drivers, interrupts, page table management).
Interpret the scenario: Host %USED=85%, %SYS=8%. The host is near saturation, and system overhead is climbing. Possible causes: (a) Too many VMs for the hardware. (b) Driver inefficiency (NIC driver, storage controller driver). (c) Memory pressure causing high page-table walks. (d) Interrupt storms from overactive devices.
Exit esxtop interactive mode ('q'). Prepare to run esxtop in batch mode for trending. Batch mode captures metrics at regular intervals and outputs CSV data suitable for spreadsheet analysis.
Run esxtop in batch mode with the command: 'esxtop -b -d 5 -n 720 > /tmp/perf.csv'. This captures 720 samples at 5-second intervals (60 minutes of data). Let it run to completion.
Retrieve the CSV file to your workstation via SCP or FTP. Example: 'scp root@your-esxi-ip:/tmp/perf.csv ./perf-baseline.csv'. Open the file in a spreadsheet (Excel, Google Sheets, LibreOffice Calc).
In the spreadsheet, create a chart of key metrics over time: (1) Plot host %USED, %RDY, %SYS as a multi-series line chart. (2) Identify peak utilization and note the timestamp. (3) Look for trends: Is utilization increasing over the hour, or stable? (4) Calculate moving average (15-minute window) to smooth spikes.
Create a performance baseline document in your lab notes (or as a formal write-up). Include: (1) Host specifications (CPU sockets, cores, RAM). (2) Baseline metrics during steady-state (low load): %USED, %RDY, %SYS, DAVG, MCTLSZ. (3) Baseline metrics during peak load: same metrics at 75-85% utilization. (4) Thresholds for alerting (when to escalate).
Document batch mode methodology in your notes: When to run, how long to run, what to measure, and how to interpret results. Include a sample command and sample output interpretation.
Validation Gate
Check: Batch mode trending complete with baseline document and trend analysis interpretation
Expected: Lab notes contain esxtop batch command, CSV trend analysis, performance baseline document with thresholds, and methodology for future trending
Common Errors
Final Validation
Lab completed when all 4 tasks are executed, metrics are documented, and remediation recommendations are justified.
✓ CPU analysis with high-contention scenario analyzed and remediation recommended → Task 1 complete: %RDY, %CSTP metrics interpreted; vCPU adjustment documented with rationale
✓ Memory pressure scenario analyzed with reclamation metrics documented → Task 2 complete: MCTLSZ, SWCUR, ZIP metrics documented; memory pressure assessed; remediation ranked by priority
✓ Storage I/O latency diagnosed with baseline and stress comparisons → Task 3 complete: DAVG, KAVG, GAVG metrics documented; vSAN health checked (if applicable); storage remediation planned
✓ Host-level baseline established via batch mode trending → Task 4 complete: esxtop batch CSV captured, trending chart created, baseline document produced with alert thresholds
Cleanup / Restore
• Remove /tmp/perf.csv from ESXi host (or retain for 30 days as audit trail)
• Revert any VM I/O stress workloads (stop fio, stress-ng)
• Return host CPU/memory to baseline utilization (if VMs were intentionally overloaded)
• Restore baseline snapshots if any changes were made to VM configurations
Design Reflection (VCDX)
This lab emphasizes diagnostic methodology and evidence-based decision-making—core VCDX competencies. Panelists will probe: (1) How do you validate that your esxtop diagnosis matches actual VM performance issues? (2) When do you recommend quick fixes vs. architectural changes? (3) How do you balance performance tuning against operational risk? (4) Can you articulate the trade-offs between CPU overcommit and power efficiency? The RCAR framework below provides structure for panel discussion.
Requirements
- ESXi host(s) with operational VMs and I/O workloads
- SSH/CLI access to ESXi host
- Spreadsheet software for CSV analysis and charting
- Lab VMs running representative workloads (CPU, memory, disk I/O)
- vSAN deployment (optional, for advanced storage metrics) or iSCSI/NFS datastore
- Baseline performance data for the lab environment (to establish 'healthy' thresholds)
Constraints
- esxtop must be run directly on ESXi host; remote monitoring tools (vRealize Operations) provide different metrics and context
- Batch mode CSV captures are host-specific; trending across multiple hosts requires aggregation logic outside esxtop
- CPU/memory contention metrics are OS-kernel-dependent; different ESXi versions may report slightly different values
- vSAN-specific metrics (rebuild status, rebalance operations) require vSAN Skyline Health or vSAN Management REST API; not all metrics available in esxtop
- Interactive esxtop view requires SSH terminal session; cannot be automated or scheduled without batch mode
Assumptions
- Lab VMs are configured with realistic resource allocations (not over-provisioned with 100 vCPUs on 16-core host)
- Storage backend is functioning normally during the lab (no storage array failures masking remediation)
- Network connectivity between ESXi hosts and vSAN network is stable (for vSAN trending)
- Lab operator has basic Linux/Windows CLI knowledge to trigger I/O workloads and interpret OS-level metrics
- Remediation recommendations assume you have approval to modify VM configurations (vCPU count, memory allocation, DRS settings)
- Baseline 'healthy' values are organization-specific; lab uses industry standard thresholds (e.g., %RDY < 10%) as reference
Risks
- Risk: Over-aggressive VM scaling during CPU contention analysis may cause application downtime if vCPU is reduced too much — Mitigation: In a VCDX interview, outline the plan and revert immediately if issues emerge. Do not apply changes without supervisor approval in production.
- Risk: Batch mode trending captures can consume significant disk space if run for extended periods (e.g., 24-hour capture = 5-10 MB CSV) — Mitigation: Plan storage; use '-d 30' (30-second intervals) instead of '-d 5' for longer captures. Compress/delete old CSV files after analysis.
- Risk: Memory pressure scenario (MCTLSZ > 20%, SWCUR > 0) may trigger emergency performance degradation or guest OS OOM (Out of Memory) kill — Mitigation: In a lab environment, monitor guest OS closely during memory stress. Have snapshot/revert strategy ready if guest becomes unresponsive.
- Risk: Storage I/O stress (fio) may impact concurrent workloads if run on shared datastore; can cause noisy neighbor contention — Mitigation: Run I/O stress on a dedicated VM on dedicated datastore, or run during off-hours. Document timing in lab notes.
Self-Assessment Discussion Prompts
- Walk me through how you decided that 20% %RDY was the threshold for 'concerning contention.' Where does that number come from? Would you apply the same threshold to a DSS workload vs. an OLTP workload?
- You recommended reducing vCPU from 8 to 6. How would you validate that this change improves performance without harming application throughput? What metrics would you monitor post-change?
- In the memory pressure scenario (MCTLSZ=32GB, SWCUR=1.2GB), you recommended adding RAM as Priority 1. But hardware upgrades take weeks. What is your Priority 2 action to buy time while procurement happens?
- You observed DAVG spiking from 5ms to 45ms during I/O stress. Before blaming the storage array, what hypervisor/network factors would you investigate? How would you distinguish storage latency from network latency in a vSAN environment?
- Batch mode gave you 60 minutes of baseline data. Is that enough to define 'healthy' thresholds for the next quarter? What time window would you recommend for production baselining, and why?
- A VM shows %RDY=15%, DAVG=8ms, but the application team says performance is fine. How do you reconcile esxtop metrics with user-perceived performance? What layer are you missing?
- You're being asked to consolidate 20 more VMs onto this cluster. Based on your baseline, what's the maximum consolidated VMs this cluster can handle before %RDY exceeds 10%?