Advanced esxtop Performance Analysis
Objectives
- Diagnose multi-VM CPU contention using %RDY, %CSTP, and per-world scheduling analysis to isolate noisy-neighbor scenarios
- Evaluate NUMA topology awareness in esxtop and identify cross-NUMA memory access penalties affecting VM performance
- Analyze memory contention across competing VMs using balloon, swap, and compression metrics to determine cluster-wide overcommit severity
- Correlate esxtop storage latency (DAVG, KAVG, GAVG) with vSAN observer and vscsiStats for per-VM I/O profiling
- Build structured performance baselines from esxtop batch data and map them to Aria Operations dashboards for long-term trending
- Apply RCAR methodology to document performance findings and defend remediation strategies in a VCDX panel context
- Design a proactive monitoring framework that bridges real-time esxtop diagnostics with VCF 9.0 Aria Operations alerting
Prerequisites
VCF 9.0 lab environment deployed with vSAN-backed storage, multiple ESXi hosts, and Aria Operations configured
Prior labs: vcp-support-05
Required skills:
- SSH access to ESXi hosts and familiarity with esxtop interactive and batch modes
- Understanding of CPU scheduling, NUMA topology, and memory reclamation hierarchy
- Basic vSAN architecture knowledge (disk groups, fault domains, resync operations)
- Familiarity with Aria Operations (formerly vRealize Operations) dashboards and alerting
Lab Environment
Multi-host VCF 9.0 cluster with vSAN storage, minimum 2 ESXi hosts (dual-socket NUMA), 4+ VMs running mixed workloads (OLTP, batch, web tier), Aria Operations deployed for metric correlation
Tasks
Task 1 Multi-VM CPU Contention and NUMA-Aware Scheduling Analysis
Diagnose complex CPU contention across multiple competing VMs, identify NUMA-remote memory access penalties, and determine whether scheduling inefficiencies or genuine capacity exhaustion drive the contention.
SSH to an ESXi host running 4+ VMs. Launch esxtop and switch to CPU view ('c'). Press 'f' to add the NHN (NUMA home node), NRMEM (NUMA remote memory), and NLMEM (NUMA local memory) fields. Sort by %RDY descending ('R' then select %RDY). Document the top 5 world entries showing VM name, %USED, %RDY, %CSTP, %SYS, NHN, and NRMEM/NLMEM ratio.
Analyze the multi-VM contention pattern. With 4 VMs showing combined %USED near host capacity, identify which VM is the 'noisy neighbor.' Compare %RDY across VMs: the VM with highest %RDY relative to its %USED is being starved. Check %CSTP: if a VM has %CSTP >3%, its vCPU count is too high for the available physical cores, causing SMP co-scheduling delays.
Investigate the NUMA penalty on oltp-db-01. With NRMEM=38%, calculate the estimated latency impact. Assume local memory access = 80ns, remote = 140ns. Weighted average: (0.62 80) + (0.38 140) = 102.8ns per access vs. 80ns baseline. This is a 28.5% memory access latency increase. For a latency-sensitive OLTP workload, this is significant.
Design a remediation plan: (1) Reduce batch-report-02 vCPUs from 8 to 4 to lower its CPU footprint and reduce %CSTP on co-scheduled VMs. (2) Set CPU affinity or use vNUMA settings to pin oltp-db-01 to NUMA node 0. (3) Configure DRS rules to separate latency-sensitive VMs from batch workloads. (4) Document expected outcomes: %RDY target <5% for oltp-db-01, NRMEM target <10%.
Validate your diagnosis by checking the per-world CPU scheduling detail. Press 'e' to expand a VM and see individual vCPU worlds. Confirm that the oltp-db-01 vCPUs are spread across NUMA nodes (some worlds on node 0, some on node 1). Document the per-vCPU %RDY distribution to confirm uneven scheduling.
Validation Gate
Check: Multi-VM contention diagnosed with NUMA topology analysis and per-vCPU scheduling breakdown
Expected: Lab notes document noisy-neighbor identification, NUMA remote memory penalty calculation, per-vCPU contention asymmetry, and remediation plan with measurable success criteria
Common Errors
Task 2 Advanced Memory Contention and Overcommit Severity Assessment
Analyze memory contention across multiple competing VMs to determine cluster-wide overcommit severity, differentiate between healthy reclamation and critical pressure, and build a memory capacity model.
Switch to esxtop memory view ('m'). Press 'f' to ensure these fields are visible: MCTLSZ (balloon target), MCTLGT (balloon granted), SWCUR (current swap), SWTGT (swap target), ZIP/UNZIP (compression rates), CACHESZ (compression cache size), ACTVMB (active memory), GRANT (granted memory), OVHD (overhead memory). Document all 4+ VMs' memory states in a comparison table.
Analyze the reclamation hierarchy in action. The ESXi memory manager reclaims in this order: (1) Transparent Page Sharing (TPS) for identical pages, (2) Balloon driver requests guest OS to free pages, (3) Memory compression for infrequently accessed pages, (4) Host-level swap for last resort. Identify which stage each VM is at based on its metrics. web-app-03 with SWCUR=0.5GB has progressed to stage 4, indicating severe pressure.
Calculate the cluster-wide memory overcommit impact. Sum all MCTLSZ values to determine total balloon reclamation. Sum all SWCUR values for total swap. Calculate effective memory availability: Host Physical RAM - Host Overhead - Total SWCUR. If effective availability is negative, the cluster is in memory deficit. Document the overcommit ratio and deficit.
Examine memory overhead (OVHD) for each VM. Overhead memory is consumed by the hypervisor for VM state, page tables, and virtual hardware emulation. Large VMs with many vCPUs have higher overhead. Document OVHD for each VM and calculate total overhead as a percentage of host RAM. If total OVHD exceeds 5% of host RAM, it contributes to memory pressure.
Design a memory capacity model. Based on your analysis, determine: (1) Maximum VMs this host can sustain without swap (target: SWCUR=0 for all VMs). (2) Recommended memory reservation strategy (which VMs get reservations, at what level). (3) DRS memory threshold for triggering migration. (4) Aria Operations alert thresholds for proactive monitoring. Document as a capacity planning artifact.
Validation Gate
Check: Memory contention analysis complete with overcommit assessment, reclamation hierarchy mapping, and capacity planning model
Expected: Lab notes document per-VM memory states, reclamation stage classification, overcommit ratio calculation, overhead analysis, and a structured capacity model with reservation strategy
Common Errors
Task 3 Storage I/O Correlation with vSAN Metrics and Per-VM I/O Profiling
Correlate esxtop storage latency metrics with vSAN-specific diagnostics (vSAN observer, vscsiStats) to isolate whether I/O bottlenecks originate from the VM, the hypervisor, the vSAN network, or the physical disk tier.
Switch to esxtop disk device view ('u'). Identify the vSAN datastore devices. In vSAN, each disk group appears as a separate device. Press 'f' to add DAVG/cmd, KAVG/cmd, GAVG/cmd, and QAVG/cmd if not visible. Document baseline latency for each vSAN disk group: DAVG (total device latency including vSAN network), KAVG (kernel/VMkernel processing latency), and the calculated network component (DAVG - KAVG = approximate vSAN network + remote disk latency).
Switch to VM disk view ('v') to see per-VM I/O latency. Identify the VM with highest GAVG. Compare its GAVG to the device-level DAVG. If GAVG significantly exceeds DAVG, the VM guest OS is adding latency (guest I/O scheduler, filesystem journal, application-level queuing). Document per-VM GAVG, DAVG, READS/s, WRITES/s, and calculate the read/write ratio for each VM.
Enable vscsiStats for per-VM I/O profiling. SSH to the ESXi host and run: 'vscsiStats -s -w <world-id>' (get world ID from esxtop). Let it collect for 60 seconds, then run 'vscsiStats -p all -w <world-id>' to print the I/O histogram. This shows I/O size distribution, latency histogram, and outstanding I/O depth for the specific VM.
Correlate esxtop findings with vSAN observer data. Access vSAN performance service via vCenter (Monitor > vSAN > Performance). Compare the cluster-wide IOPS, throughput, and latency graphs with your esxtop per-host data. Identify discrepancies: if esxtop shows high DAVG on one host but vSAN observer shows low cluster latency, the issue is local to that host (disk group degradation, cache miss). If both show high latency, the issue is cluster-wide (network saturation, resync operations).
Design a storage remediation plan based on the multi-layer diagnosis. Address: (1) Immediate: Throttle batch-report-02 I/O using vSAN I/O limits (IOPS cap) to relieve cache pressure. (2) Short-term: Add SSD cache capacity to host-1 disk group 2 or rebalance objects across disk groups. (3) Long-term: Migrate write-heavy batch workloads to a dedicated vSAN storage policy with a stripe width of 2+ to distribute writes. (4) Monitoring: Configure Aria Operations storage alerts for vSAN cache hit ratio <80% and DAVG >15ms.
Validation Gate
Check: Storage I/O analysis complete with esxtop-to-vSAN correlation, per-VM I/O profiling via vscsiStats, and tiered remediation plan
Expected: Lab notes document device-level and VM-level latency comparison, vscsiStats I/O histograms, vSAN observer correlation, root cause isolation (local vs. cluster-wide), and policy-based remediation strategy
Common Errors
Task 4 Performance Baseline Construction and VCDX Defense Preparation
Build structured performance baselines from esxtop batch captures, map findings to Aria Operations dashboards for long-term trending, and prepare a VCDX-grade performance analysis document with RCAR methodology.
Run esxtop in batch mode capturing all resource types: 'esxtop -b -a -d 10 -n 360 > /tmp/advanced-baseline.csv'. The '-a' flag captures all counters (CPU, memory, disk, network, power). This produces a 1-hour capture at 10-second intervals. While batch mode runs, document the concurrent workload profile: which VMs are running, what applications they host, and what load they are under.
After batch capture completes, transfer the CSV to your workstation. Parse the CSV to extract key metrics per VM: average and p95 for %USED, %RDY, %CSTP (CPU); average and max for MCTLSZ, SWCUR (memory); average and p95 for DAVG, GAVG (storage). Calculate these statistics for the full hour and for 15-minute windows to identify temporal patterns.
Map your esxtop baseline metrics to Aria Operations (formerly vRealize Operations) dashboard widgets. In Aria Operations, navigate to the cluster dashboard and locate the corresponding metrics. Create or customize a dashboard with: (1) CPU contention heatmap (host %RDY over time), (2) Memory pressure gauge (cluster-wide balloon + swap), (3) Storage latency trend (vSAN DAVG 7-day rolling). Compare the last hour of Aria Operations data with your esxtop CSV baseline. Document any discrepancies in metric granularity or calculation methodology.
Configure Aria Operations alerts based on your baseline findings. Create custom alert definitions: (1) Warning: Any host %RDY >10% sustained for 15 minutes. (2) Critical: Any VM SWCUR >0 for 5 minutes. (3) Warning: vSAN DAVG >15ms sustained for 10 minutes. (4) Info: Batch ETL window detected (CPU spike 14:00-15:00 daily, suppress alerting during this window to reduce noise). Document alert thresholds, notification targets, and escalation procedures.
Compile a VCDX-grade performance analysis document using RCAR methodology. Structure: (1) Requirements: SLA targets for each workload tier. (2) Constraints: Hardware limits, licensing, budget. (3) Assumptions: Workload growth projections, planned changes. (4) Risks: Identified contention patterns with probability and impact. Include an executive summary (3 sentences), detailed findings (from Tasks 1-3), baseline metrics table, alert configuration, and capacity forecast (when will current hardware reach saturation at current growth rate).
Validation Gate
Check: Performance baseline constructed with statistical analysis, Aria Operations integration, alert framework, and RCAR document
Expected: Lab notes contain esxtop batch CSV analysis with windowed statistics, Aria Operations dashboard comparison, alert definitions with suppression logic, and VCDX-grade RCAR performance document with capacity forecast
Common Errors
Final Validation
Lab is complete when all 4 tasks demonstrate advanced esxtop analysis capabilities: multi-VM contention with NUMA awareness, memory overcommit modeling, vSAN I/O correlation, and baseline-driven capacity planning with Aria Operations integration.
✓ Multi-VM CPU contention diagnosed with NUMA topology analysis → Task 1 complete: Noisy-neighbor identified, NUMA remote memory penalty calculated, per-vCPU scheduling asymmetry documented, remediation plan with DRS rules and vNUMA recommendations
✓ Memory overcommit severity assessed across all VMs with reclamation hierarchy mapping → Task 2 complete: Per-VM reclamation stage classified, overcommit ratio calculated, memory reservation strategy defined, capacity model with break-even analysis
✓ Storage I/O correlated between esxtop, vscsiStats, and vSAN observer with root cause isolation → Task 3 complete: Device-level vs. VM-level latency decomposed, vscsiStats histograms captured, vSAN cache miss identified, storage policy remediation with IOPS limits designed
✓ Performance baseline constructed with Aria Operations integration and RCAR document → Task 4 complete: Batch CSV analyzed with windowed statistics, Aria Operations alerts configured with suppression windows, RCAR document produced with capacity forecast
✓ VCDX defense readiness demonstrated through structured analysis methodology → All tasks: Findings documented in RCAR format, remediation plans include success criteria, capacity forecasts include growth assumptions, and analysis distinguishes point-in-time diagnosis from long-term trending
Cleanup / Restore
• Remove /tmp/advanced-baseline.csv and any vscsiStats output files from ESXi hosts
• Stop vscsiStats collection if still running: 'vscsiStats -x -w <world-id>' for each monitored VM
• Revert any vSAN IOPS limits or storage policy changes applied during the lab
• Restore VM vCPU and memory configurations to pre-lab settings if modified during remediation testing
Design Reflection (VCDX)
This lab targets VCDX-level competency in performance diagnostics by requiring multi-tool correlation (esxtop + vscsiStats + vSAN observer + Aria Operations), NUMA-aware analysis that most administrators overlook, and the ability to translate raw metrics into business-relevant capacity forecasts. Panelists will specifically test: (1) Can you decompose a performance problem across the full stack (VM guest, hypervisor kernel, storage network, physical disk)? (2) Do you understand the statistical limitations of point-in-time metrics vs. trending data? (3) Can you design a monitoring architecture that bridges reactive troubleshooting with proactive capacity management? (4) How do you balance the cost of hardware upgrades against operational tuning? The RCAR framework provides the structured reasoning panelists expect.
Requirements
- Multi-host VCF 9.0 cluster with vSAN storage and mixed VM workloads
- SSH/CLI access to all ESXi hosts in the cluster
- Aria Operations instance connected to vCenter with historical data (minimum 7 days for trending)
- vscsiStats utility available on ESXi hosts (included in ESXi 6.5+)
- Spreadsheet or data analysis tool for CSV parsing and statistical calculations
- Lab VMs running representative workloads spanning CPU-intensive, memory-intensive, and I/O-intensive profiles
Constraints
- vscsiStats adds minor overhead (~1-2% CPU) while active; do not leave running on production hosts
- esxtop batch mode with '-a' flag generates large CSV files (15-25MB/hour); ensure sufficient local storage on ESXi host
- Aria Operations metric collection intervals (5-minute default) are coarser than esxtop (configurable to 2-second); direct comparison requires understanding of aggregation differences
- NUMA topology varies by hardware platform; analysis findings are host-hardware-specific and may not transfer across heterogeneous clusters
- vSAN observer and performance service require vSAN license; not available on vSAN Standard edition for all metrics
Assumptions
- Lab cluster represents a realistic production-like workload mix (not synthetic benchmarks only)
- ESXi hosts have dual-socket CPUs with distinct NUMA nodes (single-socket hosts will not demonstrate NUMA effects)
- vSAN is configured with default FTT=1 policy; analysis of write amplification uses this assumption
- Workload growth rate is estimable from historical Aria Operations data or business projections
- Lab operator has prior experience with esxtop fundamentals (covered in vcp-support-05)
- Aria Operations dashboards and alert definitions can be created/modified by the lab operator (requires appropriate RBAC permissions)
Risks
- Risk: vscsiStats left running on production hosts can consume CPU and memory over extended periods, impacting the workloads being measured — Mitigation: Always stop vscsiStats after data collection completes. Set a reminder or use 'timeout 300 vscsiStats -s -w <id>' to auto-terminate after 5 minutes.
- Risk: Applying vSAN IOPS limits during Task 3 may throttle production VMs if the lab shares a production vSAN cluster — Mitigation: Apply IOPS limits only to test VMs in the lab. Use a dedicated vSAN storage policy for lab VMs. Remove limits immediately after testing.
- Risk: Capacity forecast inaccuracies if growth assumptions are wrong. Over-estimating growth leads to premature hardware purchases; under-estimating leads to performance crises. — Mitigation: Present forecasts as ranges (optimistic, expected, pessimistic) with clearly stated assumptions. Revisit assumptions quarterly against actual Aria Operations trending data.
- Risk: Alert configurations created in Task 4 may fire in production and trigger unnecessary incident response if not properly scoped to the lab environment — Mitigation: Scope all alerts to lab cluster objects only. Use a dedicated Aria Operations alert notification group for lab alerts. Disable lab alerts after the exercise.
Self-Assessment Discussion Prompts
- You identified batch-report-02 as the noisy neighbor based on CPU consumption. But what if the batch job is business-critical with a hard SLA deadline? How do you resolve the conflict between two VMs that both need CPU resources simultaneously?
- Your NUMA analysis showed 38% remote memory access on oltp-db-01. The VM has 8 vCPUs on a host with 2x 8-core sockets. Should you reduce the VM to 8 vCPUs to fit one NUMA node, or keep 8 vCPUs and accept the cross-NUMA penalty? What factors determine this decision?
- You recommended an overcommit ratio of 1.5:1 for memory. A colleague argues that with TPS and balloon, 2:1 is safe. Walk me through the conditions under which 2:1 overcommit becomes dangerous, and how you would monitor for the transition from safe to dangerous.
- vscsiStats showed 15% of I/Os taking 20-100ms, but esxtop DAVG was only 6ms. Explain why the average hides the tail latency problem. How would you present this finding to an application team that only cares about average response time?
- Your Aria Operations alerts suppress notifications during the daily ETL window. A genuine hardware failure occurs at 14:15, during the suppression window. How would you redesign the alerting to catch hardware failures while still suppressing expected ETL-driven contention alerts?
- You forecasted CPU saturation in 14 months at 15% annual growth. The CTO asks: can we defer hardware investment by optimizing workloads instead? Walk me through the analysis you would perform to determine the maximum optimization potential before hardware becomes unavoidable.
- Compare your esxtop-based approach to Aria Operations predictive analytics (formerly vRops Capacity Analytics). When would you trust the automated forecast over your manual baseline analysis, and vice versa?
Extensions
Cross-Cluster NUMA Migration Impact Analysis
Extend the lab to simulate a vMotion of oltp-db-01 from a host with 2x 16-core sockets to a host with 2x 8-core sockets. Capture esxtop NUMA metrics before and after migration. Analyze how the VM's vNUMA topology changes when the physical NUMA geometry differs. Document the performance impact of NUMA topology mismatch after cross-cluster vMotion, and design DRS rules that prevent latency-sensitive VMs from migrating to hosts with incompatible NUMA geometry.
Automated Performance Anomaly Detection Pipeline
Build an automated pipeline using esxtop batch mode output, PowerCLI, and Aria Operations REST API. Script captures esxtop batch data hourly, parses CSV to extract key metrics, calculates rolling 7-day baselines with standard deviation bands, and pushes custom metrics to Aria Operations via REST API. Configure Aria Operations to alert when any metric exceeds 2 standard deviations from its rolling baseline. This extension demonstrates the bridge from manual diagnosis to automated monitoring infrastructure.
vSAN I/O Amplification Modeling and Policy Optimization
Extend the vSAN storage analysis by modeling write amplification under different storage policies (FTT=1 vs. FTT=2 vs. FTT=1 with RAID-5/6 erasure coding). Use vscsiStats to capture per-VM I/O profiles, then calculate the actual backend I/O generated under each policy. Build a cost-benefit matrix comparing data protection level, storage capacity consumption, and I/O throughput for each policy. Present recommendations as a VCDX-grade design decision with quantified trade-offs.
⚠ Known Pitfalls (from Community KB)
References
- vSphere Monitoring and Performance (vSphere 9.0)Tier 1 — Official
- NUMA and vNUMA Architecture in vSphereTier 1 — Official
- vSAN Performance Diagnostics and vscsiStats UsageTier 1 — Official
- Aria Operations for VMware Cloud Foundation MonitoringTier 1 — Official
- CPU Ready and Co-Stop Interpretation for SMP Virtual MachinesTier 1 — Official