Academy/VCF 9.0 Support (2V0-15.25)/Advanced esxtop Performance Analysis
This lab targets VCF 9.0

Advanced esxtop Performance Analysis

VCF 9.0Advancedvcp-foundationvcp-supportvcap-design⏱ 90 min

Advanced scenario-based esxtop analysis: multi-VM contention diagnosis, NUMA-aware scheduling, vSAN I/O correlation, and capacity-planning baselines. Builds on vcp-support-05 fundamentals.

Objectives

  • Diagnose multi-VM CPU contention using %RDY, %CSTP, and per-world scheduling analysis to isolate noisy-neighbor scenarios
  • Evaluate NUMA topology awareness in esxtop and identify cross-NUMA memory access penalties affecting VM performance
  • Analyze memory contention across competing VMs using balloon, swap, and compression metrics to determine cluster-wide overcommit severity
  • Correlate esxtop storage latency (DAVG, KAVG, GAVG) with vSAN observer and vscsiStats for per-VM I/O profiling
  • Build structured performance baselines from esxtop batch data and map them to Aria Operations dashboards for long-term trending
  • Apply RCAR methodology to document performance findings and defend remediation strategies in a VCDX panel context
  • Design a proactive monitoring framework that bridges real-time esxtop diagnostics with VCF 9.0 Aria Operations alerting

Prerequisites

VCF 9.0 lab environment deployed with vSAN-backed storage, multiple ESXi hosts, and Aria Operations configured

Prior labs: vcp-support-05

Required skills:

  • SSH access to ESXi hosts and familiarity with esxtop interactive and batch modes
  • Understanding of CPU scheduling, NUMA topology, and memory reclamation hierarchy
  • Basic vSAN architecture knowledge (disk groups, fault domains, resync operations)
  • Familiarity with Aria Operations (formerly vRealize Operations) dashboards and alerting

Lab Environment

Multi-host VCF 9.0 cluster with vSAN storage, minimum 2 ESXi hosts (dual-socket NUMA), 4+ VMs running mixed workloads (OLTP, batch, web tier), Aria Operations deployed for metric correlation

Tasks

Task 1 Multi-VM CPU Contention and NUMA-Aware Scheduling Analysis

This task goes beyond single-VM %RDY analysis to expose inter-VM contention dynamics and NUMA topology effects. Panelists will probe: How do you distinguish a noisy-neighbor VM from cluster-wide CPU saturation? What is the performance cost of cross-NUMA scheduling? When does NUMA affinity help vs. hurt? How would you present this analysis to a non-technical stakeholder?

Diagnose complex CPU contention across multiple competing VMs, identify NUMA-remote memory access penalties, and determine whether scheduling inefficiencies or genuine capacity exhaustion drive the contention.

Step 1

SSH to an ESXi host running 4+ VMs. Launch esxtop and switch to CPU view ('c'). Press 'f' to add the NHN (NUMA home node), NRMEM (NUMA remote memory), and NLMEM (NUMA local memory) fields. Sort by %RDY descending ('R' then select %RDY). Document the top 5 world entries showing VM name, %USED, %RDY, %CSTP, %SYS, NHN, and NRMEM/NLMEM ratio.

Table captured: e.g., 'oltp-db-01: %USED=72%, %RDY=18%, %CSTP=4.2%, NHN=0, NRMEM=38%, NLMEM=62%. batch-report-02: %USED=55%, %RDY=22%, %CSTP=1.1%, NHN=1, NRMEM=5%, NLMEM=95%.' The OLTP VM has 38% remote memory access indicating cross-NUMA scheduling, while the batch VM is NUMA-local.
NHN shows which NUMA node the VM is homed to. If NRMEM exceeds 20% of total memory, the VM is suffering cross-NUMA penalties. Each remote memory access adds 40-100ns latency compared to local access.
Step 2

Analyze the multi-VM contention pattern. With 4 VMs showing combined %USED near host capacity, identify which VM is the 'noisy neighbor.' Compare %RDY across VMs: the VM with highest %RDY relative to its %USED is being starved. Check %CSTP: if a VM has %CSTP >3%, its vCPU count is too high for the available physical cores, causing SMP co-scheduling delays.

Analysis: 'oltp-db-01 is the contention victim (%RDY=18% despite needing CPU). batch-report-02 is the noisy neighbor: it consumes 55% CPU with 8 vCPUs, driving cluster-wide %RDY up. oltp-db-01 has %CSTP=4.2% because its 8 vCPUs cannot co-schedule efficiently when batch-report-02 monopolizes cores on NUMA node 1.'
In a VCDX defense, articulate the difference: %RDY is a victim metric (VM waiting for CPU), while high %USED on another VM is the aggressor metric. Solving contention requires addressing the aggressor, not the victim.
Step 3

Investigate the NUMA penalty on oltp-db-01. With NRMEM=38%, calculate the estimated latency impact. Assume local memory access = 80ns, remote = 140ns. Weighted average: (0.62 80) + (0.38 140) = 102.8ns per access vs. 80ns baseline. This is a 28.5% memory access latency increase. For a latency-sensitive OLTP workload, this is significant.

NUMA impact documented: 'oltp-db-01 experiences 28.5% memory access latency overhead due to cross-NUMA scheduling. For an OLTP workload performing millions of memory operations per second, this translates to measurable query latency increase. Root cause: VM was vMotioned or DRS-placed without NUMA affinity consideration.'
ESXi normally respects NUMA boundaries automatically. Cross-NUMA scheduling often occurs when a VM has more vCPUs than a single NUMA node has physical cores, or after a vMotion that lands the VM on a host with different NUMA topology.
Step 4

Design a remediation plan: (1) Reduce batch-report-02 vCPUs from 8 to 4 to lower its CPU footprint and reduce %CSTP on co-scheduled VMs. (2) Set CPU affinity or use vNUMA settings to pin oltp-db-01 to NUMA node 0. (3) Configure DRS rules to separate latency-sensitive VMs from batch workloads. (4) Document expected outcomes: %RDY target <5% for oltp-db-01, NRMEM target <10%.

Remediation plan with success criteria: 'Action 1: Right-size batch-report-02 to 4 vCPU (matches single NUMA node width). Expected: cluster %RDY drops by 10-15 percentage points. Action 2: Verify vNUMA is exposing correct topology to oltp-db-01 (VM must see 2 NUMA nodes if allocated >1 socket worth of vCPUs). Action 3: Create DRS anti-affinity rule separating OLTP and batch tiers. Success metric: oltp-db-01 %RDY <5%, NRMEM <10%.'
Never hard-pin CPU affinity in production without understanding the trade-off: affinity prevents DRS from migrating the VM, which may cause worse contention during host maintenance. Use soft affinity (DRS rules) instead.
Step 5

Validate your diagnosis by checking the per-world CPU scheduling detail. Press 'e' to expand a VM and see individual vCPU worlds. Confirm that the oltp-db-01 vCPUs are spread across NUMA nodes (some worlds on node 0, some on node 1). Document the per-vCPU %RDY distribution to confirm uneven scheduling.

Per-vCPU analysis: 'oltp-db-01 vCPU-0 through vCPU-3: NHN=0, %RDY=8%. vCPU-4 through vCPU-7: NHN=1, %RDY=28%. The vCPUs on NUMA node 1 are heavily contended because batch-report-02 dominates that node.' This confirms the cross-NUMA scheduling is asymmetric and the remote-node vCPUs are the primary bottleneck.
Per-world expansion is essential for VCDX-level analysis. Surface-level VM metrics can mask that only half the vCPUs are contended. This granularity distinguishes a competent administrator from a VCDX candidate.

Validation Gate

Check: Multi-VM contention diagnosed with NUMA topology analysis and per-vCPU scheduling breakdown

Expected: Lab notes document noisy-neighbor identification, NUMA remote memory penalty calculation, per-vCPU contention asymmetry, and remediation plan with measurable success criteria

Common Errors

Treating all VMs with high %RDY as equally problematic. In a multi-VM scenario, one VM may be the aggressor (consuming disproportionate CPU) while others are victims (starved of CPU). Reducing the victim's vCPUs worsens the problem.
Fix: Always identify the aggressor first by correlating high %USED with high vCPU count. Remediate the aggressor (right-size or migrate) before touching victims.
Ignoring NUMA topology entirely. Many administrators diagnose CPU contention without checking NHN/NRMEM, missing a 20-30% latency penalty that explains application slowness even when %RDY is moderate.
Fix: Always add NUMA fields to esxtop CPU view. For any VM with >4 vCPUs, check whether vCPU count exceeds a single NUMA node's core count.
Setting hard CPU affinity to fix NUMA issues. Hard affinity prevents DRS migration, creating a maintenance nightmare. If the host needs patching, affinity-pinned VMs cannot evacuate.
Fix: Use vNUMA (automatically enabled for VMs with >8 vCPUs in vSphere 7+) and DRS affinity rules instead of hard CPU pinning. Let the scheduler optimize within guardrails.

Task 2 Advanced Memory Contention and Overcommit Severity Assessment

This task moves beyond single-VM balloon/swap analysis to evaluate the ESXi memory reclamation hierarchy under genuine multi-VM pressure. Panelists will probe: How do you prioritize which VM gets memory when the host is overcommitted? What is the performance difference between TPS, balloon, compression, and swap? How does transparent page sharing interact with large pages? When is overcommit acceptable vs. dangerous?

Analyze memory contention across multiple competing VMs to determine cluster-wide overcommit severity, differentiate between healthy reclamation and critical pressure, and build a memory capacity model.

Step 1

Switch to esxtop memory view ('m'). Press 'f' to ensure these fields are visible: MCTLSZ (balloon target), MCTLGT (balloon granted), SWCUR (current swap), SWTGT (swap target), ZIP/UNZIP (compression rates), CACHESZ (compression cache size), ACTVMB (active memory), GRANT (granted memory), OVHD (overhead memory). Document all 4+ VMs' memory states in a comparison table.

Memory comparison table: 'oltp-db-01: GRANT=64GB, ACTVMB=58GB, MCTLSZ=0, SWCUR=0. batch-report-02: GRANT=32GB, ACTVMB=12GB, MCTLSZ=8GB, SWCUR=0. web-app-03: GRANT=16GB, ACTVMB=14GB, MCTLSZ=2GB, SWCUR=0.5GB. dev-test-04: GRANT=8GB, ACTVMB=3GB, MCTLSZ=4GB, SWCUR=0.' Total host memory: 128GB. Total granted: 120GB. Overcommit ratio: 1.5:1 (VMs configured for 192GB total).
ACTVMB vs GRANT reveals actual memory utilization. batch-report-02 is allocated 32GB but actively uses only 12GB, making it a good balloon candidate. The hypervisor correctly targeted it for reclamation.
Step 2

Analyze the reclamation hierarchy in action. The ESXi memory manager reclaims in this order: (1) Transparent Page Sharing (TPS) for identical pages, (2) Balloon driver requests guest OS to free pages, (3) Memory compression for infrequently accessed pages, (4) Host-level swap for last resort. Identify which stage each VM is at based on its metrics. web-app-03 with SWCUR=0.5GB has progressed to stage 4, indicating severe pressure.

Reclamation hierarchy mapping: 'batch-report-02: Stage 2 (balloon active, 8GB reclaimed, no swap). web-app-03: Stage 4 (balloon active 2GB, swap active 0.5GB, critical). dev-test-04: Stage 2 (balloon 4GB, healthy since ACTVMB is only 3GB of 8GB allocated). oltp-db-01: Stage 0 (no reclamation, priority VM with reservation).' Assessment: web-app-03 is in critical memory pressure. The host has exhausted balloon and compression for this VM and resorted to swap.
Memory reservations change the reclamation order. A VM with a reservation will not be ballooned below its reservation. oltp-db-01 likely has a 64GB reservation, which is why it shows zero reclamation despite host pressure.
Step 3

Calculate the cluster-wide memory overcommit impact. Sum all MCTLSZ values to determine total balloon reclamation. Sum all SWCUR values for total swap. Calculate effective memory availability: Host Physical RAM - Host Overhead - Total SWCUR. If effective availability is negative, the cluster is in memory deficit. Document the overcommit ratio and deficit.

Overcommit analysis: 'Host physical: 128GB. Host overhead (vmkernel): 4GB. Available for VMs: 124GB. Total VM configured: 192GB. Overcommit ratio: 1.55:1. Total balloon: 14GB. Total swap: 0.5GB. Effective deficit: VMs need 192GB but only 124GB available, so 68GB must be reclaimed. Currently reclaiming 14.5GB (balloon + swap), meaning VMs are operating with 109.5GB effective memory.'
An overcommit ratio of 1.5:1 is common in production but risky for latency-sensitive workloads. The VCDX question is: What overcommit ratio is acceptable for this workload mix? Defend your answer with SLA requirements.
Step 4

Examine memory overhead (OVHD) for each VM. Overhead memory is consumed by the hypervisor for VM state, page tables, and virtual hardware emulation. Large VMs with many vCPUs have higher overhead. Document OVHD for each VM and calculate total overhead as a percentage of host RAM. If total OVHD exceeds 5% of host RAM, it contributes to memory pressure.

Overhead analysis: 'oltp-db-01: OVHD=380MB (64GB VM, 8 vCPU). batch-report-02: OVHD=220MB (32GB VM, 8 vCPU). web-app-03: OVHD=150MB (16GB VM, 4 vCPU). dev-test-04: OVHD=100MB (8GB VM, 2 vCPU). Total OVHD: 850MB (0.66% of 128GB host). Overhead is not a significant contributor to memory pressure in this scenario.'
OVHD becomes significant with very large VMs (512GB+ RAM) or Monster VMs with 64+ vCPUs. In those cases, overhead can consume 2-4GB per VM. Factor this into capacity planning for large-scale consolidation.
Step 5

Design a memory capacity model. Based on your analysis, determine: (1) Maximum VMs this host can sustain without swap (target: SWCUR=0 for all VMs). (2) Recommended memory reservation strategy (which VMs get reservations, at what level). (3) DRS memory threshold for triggering migration. (4) Aria Operations alert thresholds for proactive monitoring. Document as a capacity planning artifact.

Capacity model: 'Maximum VMs without swap: Reduce to 3 VMs (remove dev-test-04 or migrate to dev cluster). Memory reservations: oltp-db-01 = 48GB reservation (75% of allocation, protects against balloon). web-app-03 = 12GB reservation (prevents swap). DRS threshold: Trigger migration when host balloon exceeds 15% of physical RAM. Aria Operations alert: Warning at host memory usage >85%, Critical at >92% (swap imminent).'
A capacity model is a VCDX-grade deliverable. It shows you can translate esxtop findings into actionable infrastructure planning, not just reactive troubleshooting.

Validation Gate

Check: Memory contention analysis complete with overcommit assessment, reclamation hierarchy mapping, and capacity planning model

Expected: Lab notes document per-VM memory states, reclamation stage classification, overcommit ratio calculation, overhead analysis, and a structured capacity model with reservation strategy

Common Errors

Treating all balloon activity as a problem. Balloon reclamation on a VM with low ACTVMB (active memory much lower than allocated) is the hypervisor efficiently reclaiming unused memory. This is healthy.
Fix: Compare MCTLSZ to ACTVMB. If MCTLSZ < (GRANT - ACTVMB), the balloon is only reclaiming genuinely unused memory. Problem occurs when MCTLSZ > (GRANT - ACTVMB), meaning the balloon is reclaiming memory the VM is actively using.
Ignoring the interaction between memory reservations and DRS. VMs with large reservations reduce the host's available memory for other VMs, potentially forcing DRS to migrate non-reserved VMs more aggressively.
Fix: Set reservations strategically: only on VMs that genuinely need guaranteed memory (databases, latency-sensitive apps). Over-reserving defeats the purpose of consolidation.
Not accounting for transparent page sharing (TPS) limitations. Since vSphere 6.0, inter-VM TPS is disabled by default for security (Rowhammer vulnerability). Only intra-VM TPS (within a single VM) is active.
Fix: Do not assume TPS will save memory across VMs. Factor this into overcommit calculations. If you enable inter-VM TPS, document the security trade-off for VCDX panel discussion.
Recommending adding RAM without calculating the break-even point. A 128GB DIMM costs money; you need to justify that the cost is lower than migrating VMs or right-sizing allocations.
Fix: Calculate cost per GB of reclaimed swap vs. cost of RAM upgrade. Present both options to the panel with ROI analysis.

Task 3 Storage I/O Correlation with vSAN Metrics and Per-VM I/O Profiling

This task bridges esxtop disk analysis with vSAN-native tooling for deep I/O correlation. Panelists will probe: How do you distinguish vSAN network latency from disk latency? What does KAVG tell you that DAVG does not? How do you use vscsiStats to profile individual VM I/O patterns? When is the right time to add capacity vs. tune I/O policies?

Correlate esxtop storage latency metrics with vSAN-specific diagnostics (vSAN observer, vscsiStats) to isolate whether I/O bottlenecks originate from the VM, the hypervisor, the vSAN network, or the physical disk tier.

Step 1

Switch to esxtop disk device view ('u'). Identify the vSAN datastore devices. In vSAN, each disk group appears as a separate device. Press 'f' to add DAVG/cmd, KAVG/cmd, GAVG/cmd, and QAVG/cmd if not visible. Document baseline latency for each vSAN disk group: DAVG (total device latency including vSAN network), KAVG (kernel/VMkernel processing latency), and the calculated network component (DAVG - KAVG = approximate vSAN network + remote disk latency).

Baseline captured: 'vSAN disk group 1: DAVG=4.5ms, KAVG=0.8ms, network component=3.7ms. vSAN disk group 2: DAVG=5.1ms, KAVG=0.9ms, network component=4.2ms.' The high proportion of DAVG attributed to network component is normal for vSAN (writes must acknowledge across the network to remote hosts for data redundancy).
In vSAN, DAVG includes network round-trip time for replica acknowledgment. A 4-5ms DAVG with sub-1ms KAVG is healthy for a hybrid vSAN (SSD cache + HDD capacity). All-flash vSAN should show DAVG <2ms.
Step 2

Switch to VM disk view ('v') to see per-VM I/O latency. Identify the VM with highest GAVG. Compare its GAVG to the device-level DAVG. If GAVG significantly exceeds DAVG, the VM guest OS is adding latency (guest I/O scheduler, filesystem journal, application-level queuing). Document per-VM GAVG, DAVG, READS/s, WRITES/s, and calculate the read/write ratio for each VM.

Per-VM I/O profile: 'oltp-db-01: GAVG=8.2ms, DAVG=4.5ms, READS/s=2400, WRITES/s=800, R:W ratio=3:1. batch-report-02: GAVG=12.5ms, DAVG=5.1ms, READS/s=100, WRITES/s=3500, R:W ratio=1:35. web-app-03: GAVG=3.1ms, DAVG=2.8ms, READS/s=500, WRITES/s=50, R:W ratio=10:1.' batch-report-02 is write-heavy with GAVG much higher than DAVG, suggesting guest-level write queuing.
R:W ratio is critical for vSAN sizing. Write-heavy workloads amplify I/O because vSAN must write to cache + replicate across hosts. A 1:35 R:W ratio means batch-report-02 generates disproportionate vSAN network traffic.
Step 3

Enable vscsiStats for per-VM I/O profiling. SSH to the ESXi host and run: 'vscsiStats -s -w <world-id>' (get world ID from esxtop). Let it collect for 60 seconds, then run 'vscsiStats -p all -w <world-id>' to print the I/O histogram. This shows I/O size distribution, latency histogram, and outstanding I/O depth for the specific VM.

vscsiStats output for batch-report-02: 'I/O size distribution: 90% at 8KB, 5% at 64KB, 5% at 256KB. Latency histogram: 60% of I/Os complete in 1-5ms, 25% in 5-20ms, 15% in 20-100ms. Outstanding I/O: avg=24, max=64.' The 15% of I/Os taking 20-100ms are the long-tail latencies causing application slowness. The 8KB dominant I/O size confirms random write pattern.
vscsiStats provides granularity that esxtop cannot: I/O size distribution and latency percentiles. The long-tail (p95, p99) latencies are often invisible in esxtop averages but critical for SLA compliance.
Step 4

Correlate esxtop findings with vSAN observer data. Access vSAN performance service via vCenter (Monitor > vSAN > Performance). Compare the cluster-wide IOPS, throughput, and latency graphs with your esxtop per-host data. Identify discrepancies: if esxtop shows high DAVG on one host but vSAN observer shows low cluster latency, the issue is local to that host (disk group degradation, cache miss). If both show high latency, the issue is cluster-wide (network saturation, resync operations).

Correlation analysis: 'esxtop host-1: DAVG=45ms (spike). vSAN observer cluster latency: 6ms average. Discrepancy confirms the issue is local to host-1. vSAN observer shows host-1 disk group 2 with 95% cache miss rate and 40ms backend latency. Root cause: vSAN cache (SSD) on host-1 disk group 2 is saturated; write buffer is full, forcing destaging to capacity tier during I/O storm.'
vSAN observer (now integrated into vCenter 8.0+) provides the cluster-wide context that esxtop lacks. Always use both tools together: esxtop for per-host/per-VM granularity, vSAN observer for cluster-wide health.
Step 5

Design a storage remediation plan based on the multi-layer diagnosis. Address: (1) Immediate: Throttle batch-report-02 I/O using vSAN I/O limits (IOPS cap) to relieve cache pressure. (2) Short-term: Add SSD cache capacity to host-1 disk group 2 or rebalance objects across disk groups. (3) Long-term: Migrate write-heavy batch workloads to a dedicated vSAN storage policy with a stripe width of 2+ to distribute writes. (4) Monitoring: Configure Aria Operations storage alerts for vSAN cache hit ratio <80% and DAVG >15ms.

Tiered remediation plan: 'Immediate (today): Apply vSAN IOPS limit of 2000 to batch-report-02 VM storage policy. Short-term (this week): Add 1x 800GB SSD to host-1 disk group 2 to increase cache capacity by 100%. Long-term (next sprint): Create dedicated vSAN storage policy for batch workloads with FTT=1, stripe width=2, and IOPS reservation of 3000. Monitoring: Aria Operations dashboard with vSAN cache hit ratio widget, alert at <80%.'
vSAN storage policies are the architectural control plane. In a VCDX defense, demonstrate that you use policies (not just esxtop observations) to enforce performance guarantees. This bridges troubleshooting and design.

Validation Gate

Check: Storage I/O analysis complete with esxtop-to-vSAN correlation, per-VM I/O profiling via vscsiStats, and tiered remediation plan

Expected: Lab notes document device-level and VM-level latency comparison, vscsiStats I/O histograms, vSAN observer correlation, root cause isolation (local vs. cluster-wide), and policy-based remediation strategy

Common Errors

Blaming vSAN for high DAVG without checking KAVG. If KAVG is high (>5ms), the VMkernel is the bottleneck (driver issue, CPU contention affecting I/O processing), not the vSAN backend.
Fix: Always decompose: DAVG = KAVG + network/disk latency. If KAVG is the dominant component, investigate CPU contention or driver versions on the ESXi host.
Using esxtop average latency (DAVG) as the sole performance indicator. Averages hide long-tail latencies. A DAVG of 5ms may mask that 5% of I/Os take 50ms+.
Fix: Use vscsiStats for latency histogram analysis. The p95 and p99 latencies are what users actually feel. Report both average and percentile latencies in your analysis.
Ignoring write amplification in vSAN. A VM writing 1MB to vSAN with FTT=1 and stripe width=1 actually generates 2MB of I/O (1MB data + 1MB replica). With FTT=2, it generates 3MB. This is invisible in esxtop VM-level metrics but visible in vSAN observer.
Fix: Factor write amplification into capacity planning. If esxtop shows a VM writing 100MB/s, vSAN is actually processing 200-300MB/s depending on FTT policy. Check vSAN observer for actual backend throughput.

Task 4 Performance Baseline Construction and VCDX Defense Preparation

This task synthesizes all prior analysis into a defensible performance posture document. Panelists will probe: How do you define 'normal' for this environment? What statistical method do you use to set thresholds? How do you translate esxtop point-in-time data into operational monitoring? Can you present this to a CTO in 5 minutes?

Build structured performance baselines from esxtop batch captures, map findings to Aria Operations dashboards for long-term trending, and prepare a VCDX-grade performance analysis document with RCAR methodology.

Step 1

Run esxtop in batch mode capturing all resource types: 'esxtop -b -a -d 10 -n 360 > /tmp/advanced-baseline.csv'. The '-a' flag captures all counters (CPU, memory, disk, network, power). This produces a 1-hour capture at 10-second intervals. While batch mode runs, document the concurrent workload profile: which VMs are running, what applications they host, and what load they are under.

Batch capture initiated. Workload profile documented: 'Capture window: 14:00-15:00. oltp-db-01: Running TPC-C benchmark at 80% target load. batch-report-02: Executing weekly ETL pipeline (write-heavy). web-app-03: Serving synthetic HTTP load (500 req/s). dev-test-04: Idle (development hours ended).' CSV file size approximately 15-25MB for 1-hour all-counter capture.
The '-a' flag produces very wide CSV rows (500+ columns). This is intentional for comprehensive baselining. You will filter to relevant columns during analysis.
Step 2

After batch capture completes, transfer the CSV to your workstation. Parse the CSV to extract key metrics per VM: average and p95 for %USED, %RDY, %CSTP (CPU); average and max for MCTLSZ, SWCUR (memory); average and p95 for DAVG, GAVG (storage). Calculate these statistics for the full hour and for 15-minute windows to identify temporal patterns.

Statistical baseline: 'oltp-db-01 (1-hour): CPU %USED avg=68%, p95=82%. %RDY avg=12%, p95=22%. DAVG avg=6ms, p95=18ms. Memory MCTLSZ=0 (reserved). 15-min breakdown: 14:00-14:15 %RDY avg=8% (warm-up), 14:15-14:30 %RDY avg=15% (peak ETL overlap), 14:30-14:45 %RDY avg=10% (ETL completing), 14:45-15:00 %RDY avg=9% (steady state).'
The 15-minute windowed analysis reveals that the batch ETL job (14:15-14:30) drives the %RDY spike. This temporal correlation is invisible in hourly averages. Always decompose baselines into windows to identify workload interaction patterns.
Step 3

Map your esxtop baseline metrics to Aria Operations (formerly vRealize Operations) dashboard widgets. In Aria Operations, navigate to the cluster dashboard and locate the corresponding metrics. Create or customize a dashboard with: (1) CPU contention heatmap (host %RDY over time), (2) Memory pressure gauge (cluster-wide balloon + swap), (3) Storage latency trend (vSAN DAVG 7-day rolling). Compare the last hour of Aria Operations data with your esxtop CSV baseline. Document any discrepancies in metric granularity or calculation methodology.

Aria Operations comparison: 'Aria Operations reports cluster %RDY=10% (5-minute average) vs. esxtop %RDY=12% (10-second granularity). Discrepancy explained: Aria Operations smooths data to 5-minute intervals, masking sub-minute spikes visible in esxtop. Aria Operations memory balloon metric matches esxtop MCTLSZ within 2%. Storage DAVG in Aria Operations is 5.8ms vs. esxtop 6.0ms (within margin of collection interval difference).'
Aria Operations and esxtop measure the same underlying metrics but at different granularities. esxtop is the ground truth for point-in-time diagnosis; Aria Operations is the ground truth for trending. A VCDX candidate must understand when to use each tool.
Step 4

Configure Aria Operations alerts based on your baseline findings. Create custom alert definitions: (1) Warning: Any host %RDY >10% sustained for 15 minutes. (2) Critical: Any VM SWCUR >0 for 5 minutes. (3) Warning: vSAN DAVG >15ms sustained for 10 minutes. (4) Info: Batch ETL window detected (CPU spike 14:00-15:00 daily, suppress alerting during this window to reduce noise). Document alert thresholds, notification targets, and escalation procedures.

Alert configuration document: 'Alert 1: Host CPU Contention Warning. Condition: %RDY >10% for 15 min. Notification: Ops team Slack channel. Escalation: If not acknowledged in 30 min, page on-call engineer. Alert 2: VM Memory Swap Critical. Condition: SWCUR >0 for 5 min. Notification: Immediate page to on-call + VM owner. Alert 3: vSAN Latency Warning. Condition: DAVG >15ms for 10 min. Notification: Storage team email. Alert 4: Batch Window Suppression. Schedule: M-F 14:00-15:30, suppress Alert 1 for batch-report-02 host only.'
Alert suppression windows demonstrate operational maturity. Panelists will ask: How do you prevent alert fatigue while ensuring genuine issues are caught? The batch window suppression is a concrete example of signal-vs-noise management.
Step 5

Compile a VCDX-grade performance analysis document using RCAR methodology. Structure: (1) Requirements: SLA targets for each workload tier. (2) Constraints: Hardware limits, licensing, budget. (3) Assumptions: Workload growth projections, planned changes. (4) Risks: Identified contention patterns with probability and impact. Include an executive summary (3 sentences), detailed findings (from Tasks 1-3), baseline metrics table, alert configuration, and capacity forecast (when will current hardware reach saturation at current growth rate).

RCAR document: 'Executive Summary: The VCF cluster supports 4 VMs with a CPU overcommit of 2:1 and memory overcommit of 1.55:1. Peak CPU contention (%RDY=22%) occurs during the daily ETL window and impacts OLTP latency by 28%. Recommended actions: right-size batch VM vCPUs, add 64GB RAM per host, and implement workload scheduling separation. Capacity Forecast: At 15% annual workload growth, current hardware reaches CPU saturation (sustained %RDY >15%) in approximately 14 months. Memory saturation (sustained swap >0) in approximately 8 months without RAM upgrade.'
The capacity forecast is the most valuable output of this lab. It translates technical metrics into a business decision: when to invest in hardware. A VCDX panelist will challenge your growth assumptions. Be prepared to defend them with data.

Validation Gate

Check: Performance baseline constructed with statistical analysis, Aria Operations integration, alert framework, and RCAR document

Expected: Lab notes contain esxtop batch CSV analysis with windowed statistics, Aria Operations dashboard comparison, alert definitions with suppression logic, and VCDX-grade RCAR performance document with capacity forecast

Common Errors

Presenting raw esxtop numbers without statistical context. Saying '%RDY was 22%' without noting this was a p95 value during a specific 15-minute window misleads stakeholders into thinking contention is constant.
Fix: Always report averages AND percentiles (p50, p95, p99) with time windows. A p95 of 22% during ETL is very different from a sustained average of 22%.
Setting alert thresholds based on textbook values without validating against actual baseline. If your baseline %RDY average is 12%, alerting at 10% generates constant false positives.
Fix: Set thresholds at baseline p95 + margin. If p95 %RDY is 18%, set warning at 20% and critical at 30%. This captures genuine anomalies without alert fatigue.
Creating a capacity forecast without documenting assumptions. A forecast of '14 months to saturation' is meaningless without stating the assumed growth rate, workload mix, and planned changes.
Fix: Every forecast must include: growth rate assumption, confidence interval, and sensitivity analysis (what if growth is 2x expected?).
Treating Aria Operations and esxtop as interchangeable. Aria Operations aggregates and smooths data for trending, while esxtop provides raw real-time granularity. Using Aria Operations for acute diagnosis or esxtop for quarterly planning is using the wrong tool.
Fix: Define tool roles explicitly: esxtop = acute diagnosis and baseline capture. Aria Operations = long-term trending, alerting, and capacity planning. Cross-validate periodically to ensure consistency.

Final Validation

Lab is complete when all 4 tasks demonstrate advanced esxtop analysis capabilities: multi-VM contention with NUMA awareness, memory overcommit modeling, vSAN I/O correlation, and baseline-driven capacity planning with Aria Operations integration.

✓ Multi-VM CPU contention diagnosed with NUMA topology analysis → Task 1 complete: Noisy-neighbor identified, NUMA remote memory penalty calculated, per-vCPU scheduling asymmetry documented, remediation plan with DRS rules and vNUMA recommendations

✓ Memory overcommit severity assessed across all VMs with reclamation hierarchy mapping → Task 2 complete: Per-VM reclamation stage classified, overcommit ratio calculated, memory reservation strategy defined, capacity model with break-even analysis

✓ Storage I/O correlated between esxtop, vscsiStats, and vSAN observer with root cause isolation → Task 3 complete: Device-level vs. VM-level latency decomposed, vscsiStats histograms captured, vSAN cache miss identified, storage policy remediation with IOPS limits designed

✓ Performance baseline constructed with Aria Operations integration and RCAR document → Task 4 complete: Batch CSV analyzed with windowed statistics, Aria Operations alerts configured with suppression windows, RCAR document produced with capacity forecast

✓ VCDX defense readiness demonstrated through structured analysis methodology → All tasks: Findings documented in RCAR format, remediation plans include success criteria, capacity forecasts include growth assumptions, and analysis distinguishes point-in-time diagnosis from long-term trending

Cleanup / Restore

• Remove /tmp/advanced-baseline.csv and any vscsiStats output files from ESXi hosts

• Stop vscsiStats collection if still running: 'vscsiStats -x -w <world-id>' for each monitored VM

• Revert any vSAN IOPS limits or storage policy changes applied during the lab

• Restore VM vCPU and memory configurations to pre-lab settings if modified during remediation testing

Design Reflection (VCDX)

This lab targets VCDX-level competency in performance diagnostics by requiring multi-tool correlation (esxtop + vscsiStats + vSAN observer + Aria Operations), NUMA-aware analysis that most administrators overlook, and the ability to translate raw metrics into business-relevant capacity forecasts. Panelists will specifically test: (1) Can you decompose a performance problem across the full stack (VM guest, hypervisor kernel, storage network, physical disk)? (2) Do you understand the statistical limitations of point-in-time metrics vs. trending data? (3) Can you design a monitoring architecture that bridges reactive troubleshooting with proactive capacity management? (4) How do you balance the cost of hardware upgrades against operational tuning? The RCAR framework provides the structured reasoning panelists expect.

Requirements

  • Multi-host VCF 9.0 cluster with vSAN storage and mixed VM workloads
  • SSH/CLI access to all ESXi hosts in the cluster
  • Aria Operations instance connected to vCenter with historical data (minimum 7 days for trending)
  • vscsiStats utility available on ESXi hosts (included in ESXi 6.5+)
  • Spreadsheet or data analysis tool for CSV parsing and statistical calculations
  • Lab VMs running representative workloads spanning CPU-intensive, memory-intensive, and I/O-intensive profiles

Constraints

  • vscsiStats adds minor overhead (~1-2% CPU) while active; do not leave running on production hosts
  • esxtop batch mode with '-a' flag generates large CSV files (15-25MB/hour); ensure sufficient local storage on ESXi host
  • Aria Operations metric collection intervals (5-minute default) are coarser than esxtop (configurable to 2-second); direct comparison requires understanding of aggregation differences
  • NUMA topology varies by hardware platform; analysis findings are host-hardware-specific and may not transfer across heterogeneous clusters
  • vSAN observer and performance service require vSAN license; not available on vSAN Standard edition for all metrics

Assumptions

  • Lab cluster represents a realistic production-like workload mix (not synthetic benchmarks only)
  • ESXi hosts have dual-socket CPUs with distinct NUMA nodes (single-socket hosts will not demonstrate NUMA effects)
  • vSAN is configured with default FTT=1 policy; analysis of write amplification uses this assumption
  • Workload growth rate is estimable from historical Aria Operations data or business projections
  • Lab operator has prior experience with esxtop fundamentals (covered in vcp-support-05)
  • Aria Operations dashboards and alert definitions can be created/modified by the lab operator (requires appropriate RBAC permissions)

Risks

  • Risk: vscsiStats left running on production hosts can consume CPU and memory over extended periods, impacting the workloads being measured — Mitigation: Always stop vscsiStats after data collection completes. Set a reminder or use 'timeout 300 vscsiStats -s -w <id>' to auto-terminate after 5 minutes.
  • Risk: Applying vSAN IOPS limits during Task 3 may throttle production VMs if the lab shares a production vSAN cluster — Mitigation: Apply IOPS limits only to test VMs in the lab. Use a dedicated vSAN storage policy for lab VMs. Remove limits immediately after testing.
  • Risk: Capacity forecast inaccuracies if growth assumptions are wrong. Over-estimating growth leads to premature hardware purchases; under-estimating leads to performance crises. — Mitigation: Present forecasts as ranges (optimistic, expected, pessimistic) with clearly stated assumptions. Revisit assumptions quarterly against actual Aria Operations trending data.
  • Risk: Alert configurations created in Task 4 may fire in production and trigger unnecessary incident response if not properly scoped to the lab environment — Mitigation: Scope all alerts to lab cluster objects only. Use a dedicated Aria Operations alert notification group for lab alerts. Disable lab alerts after the exercise.

Self-Assessment Discussion Prompts

  1. You identified batch-report-02 as the noisy neighbor based on CPU consumption. But what if the batch job is business-critical with a hard SLA deadline? How do you resolve the conflict between two VMs that both need CPU resources simultaneously?
  2. Your NUMA analysis showed 38% remote memory access on oltp-db-01. The VM has 8 vCPUs on a host with 2x 8-core sockets. Should you reduce the VM to 8 vCPUs to fit one NUMA node, or keep 8 vCPUs and accept the cross-NUMA penalty? What factors determine this decision?
  3. You recommended an overcommit ratio of 1.5:1 for memory. A colleague argues that with TPS and balloon, 2:1 is safe. Walk me through the conditions under which 2:1 overcommit becomes dangerous, and how you would monitor for the transition from safe to dangerous.
  4. vscsiStats showed 15% of I/Os taking 20-100ms, but esxtop DAVG was only 6ms. Explain why the average hides the tail latency problem. How would you present this finding to an application team that only cares about average response time?
  5. Your Aria Operations alerts suppress notifications during the daily ETL window. A genuine hardware failure occurs at 14:15, during the suppression window. How would you redesign the alerting to catch hardware failures while still suppressing expected ETL-driven contention alerts?
  6. You forecasted CPU saturation in 14 months at 15% annual growth. The CTO asks: can we defer hardware investment by optimizing workloads instead? Walk me through the analysis you would perform to determine the maximum optimization potential before hardware becomes unavoidable.
  7. Compare your esxtop-based approach to Aria Operations predictive analytics (formerly vRops Capacity Analytics). When would you trust the automated forecast over your manual baseline analysis, and vice versa?

Extensions

Cross-Cluster NUMA Migration Impact Analysis

Extend the lab to simulate a vMotion of oltp-db-01 from a host with 2x 16-core sockets to a host with 2x 8-core sockets. Capture esxtop NUMA metrics before and after migration. Analyze how the VM's vNUMA topology changes when the physical NUMA geometry differs. Document the performance impact of NUMA topology mismatch after cross-cluster vMotion, and design DRS rules that prevent latency-sensitive VMs from migrating to hosts with incompatible NUMA geometry.

Automated Performance Anomaly Detection Pipeline

Build an automated pipeline using esxtop batch mode output, PowerCLI, and Aria Operations REST API. Script captures esxtop batch data hourly, parses CSV to extract key metrics, calculates rolling 7-day baselines with standard deviation bands, and pushes custom metrics to Aria Operations via REST API. Configure Aria Operations to alert when any metric exceeds 2 standard deviations from its rolling baseline. This extension demonstrates the bridge from manual diagnosis to automated monitoring infrastructure.

vSAN I/O Amplification Modeling and Policy Optimization

Extend the vSAN storage analysis by modeling write amplification under different storage policies (FTT=1 vs. FTT=2 vs. FTT=1 with RAID-5/6 erasure coding). Use vscsiStats to capture per-VM I/O profiles, then calculate the actual backend I/O generated under each policy. Build a cost-benefit matrix comparing data protection level, storage capacity consumption, and I/O throughput for each policy. Present recommendations as a VCDX-grade design decision with quantified trade-offs.

⚠ Known Pitfalls (from Community KB)

Analyzing CPU contention without NUMA awareness. An administrator may correctly identify %RDY=18% as problematic but miss that the root cause is cross-NUMA scheduling (adding 28% memory latency) rather than raw CPU shortage. The remediation for NUMA misalignment (right-size vCPUs to fit a NUMA node) is fundamentally different from the remediation for CPU shortage (add hosts or migrate VMs).
Using esxtop averages for SLA compliance reporting. esxtop DAVG and %RDY are time-averaged metrics that mask spike behavior. Reporting DAVG=6ms to an application team when p99 latency is 50ms creates a false sense of health. SLAs are typically violated by tail latencies, not averages.
Treating vSAN storage metrics in esxtop identically to traditional SAN metrics. In vSAN, DAVG includes network round-trip for replica acknowledgment, which does not exist in traditional SAN. A DAVG of 5ms in vSAN may be perfectly healthy, while the same DAVG in FC-SAN would indicate backend latency issues.
Building capacity forecasts from a single baseline capture. One hour of esxtop data during a specific workload pattern is not statistically representative. Forecasting from a single sample ignores daily, weekly, and seasonal workload variations that significantly affect capacity planning accuracy.

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.