Academy/VCF 9.0 Support (2V0-15.25)/Performance Analysis with esxtop
This lab targets VCF 9.0

Performance Analysis with esxtop

VCF 9.0Intermediatevcp-foundationvcp-support⏱ 90 min

Focus on deep metric interpretation and remediation methodology.

Objectives

  • Launch and navigate esxtop views (CPU, memory, disk I/O, host-level analysis)
  • Interpret CPU metrics (%USED, %RDY, %CSTP, %SYS, %WAIT) and identify contention patterns
  • Analyze memory reclamation (balloon, swapping, compression) and detect memory pressure
  • Diagnose storage I/O latency and correlate with backend infrastructure
  • Perform trend analysis using esxtop batch mode for baseline documentation
  • Recommend evidence-based remediation actions based on metric analysis
  • Document performance findings using RCAR methodology

Prerequisites

VCF lab environment deployed and operational with ESXi hosts

Required skills:

  • SSH access to ESXi host
  • Basic understanding of CPU scheduling and memory management
  • Familiarity with performance baseline concepts

Lab Environment

Standard VCF lab environment for VCF 9.0 Support with multi-tier VM workloads

Tasks

Task 1 CPU Performance Analysis and Contention Detection

Task progresses from basic metric recognition to scenario-based interpretation. Panelists will probe: How do you distinguish CPU contention from CPU-dependent application performance? What is the significance of %CSTP in a multi-socket VM? How would you validate your diagnosis before making changes?

Master esxtop CPU view and interpret multi-dimensional contention signals to identify scheduling issues and recommend vCPU optimization.

Step 1

SSH to an ESXi host in your lab. Launch esxtop by typing 'esxtop' and wait for the interactive interface to load.

esxtop displays the default CPU view with columns showing %USED, %RDY, %CSTP, %SYS, %WAIT, %IDLE. Screen header shows processor count and uptime. Legend available via 'h'.
If the screen appears garbled, press 'c' to explicitly switch to CPU view. Use 'q' to quit esxtop cleanly.
Step 2

Press 'c' to enter the CPU view (should already be the default). Examine the column headers. Identify which rows represent VMs vs. host kernel.

CPU view displays one row per VM plus rows for kernel domains. %USED should show current CPU usage, %RDY shows waiting-for-scheduler time, %CSTP shows co-stop time (SMP VMs waiting for all vCPUs to align).
Use 'f' to add/remove fields; for this lab ensure %USED, %RDY, %CSTP, %SYS, and %WAIT are visible. Familiarize yourself with the sort order (default is %USED descending).
Step 3

Locate a VM or workload with measurable CPU usage (ideally 30%+). Note its name, current %USED, %RDY, and %CSTP values. Document the current state in your lab notes.

Baseline snapshot captured: e.g., 'web-server-01: %USED=65%, %RDY=18%, %CSTP=0.5%'. Host %SYS should be between 3-6% under normal load.
If no VMs are running CPU workloads, manually trigger a CPU load: e.g., on the VM, run 'stress-ng --cpu 4 --timeout 300s' or equivalent for your OS.
Step 4

Analyze the scenario: A VM shows %USED=65%, %RDY=20%, %CSTP=2%. Interpret each metric. High %RDY indicates CPU contention at the host level. %CSTP=2% on a multi-vCPU VM suggests scheduling delays for SMP alignment.

Written analysis: 'Host is over-subscribed with CPU. The 20% ready time means this VM is waiting 20% of the time to be scheduled, reducing effective throughput. The 2% co-stop suggests the hypervisor is frequently delaying the VM to align all its vCPUs, typical when vCPU count is high relative to physical cores.'
Remember: %RDY > 10% is generally considered problematic. %CSTP > 1% on a single-socket VM is unusual; check vCPU count.
Step 5

Press 'U' or 'v' to switch to a breakdown view showing per-vCPU details (if available in your esxtop version). Examine whether all vCPUs are equally loaded or if some idle while others busy.

Per-vCPU breakdown may show asymmetric load distribution, indicating imbalanced threading or uneven workload distribution within the VM.
Not all esxtop versions expose per-vCPU granularity interactively. If unavailable, note this limitation for your VCDX panel discussion.
Step 6

Recommend remediation actions based on your analysis: (a) Reduce the VM's vCPU count by 1-2 cores to lower co-stop, (b) Enable CPU affinity if the workload is memory-bound, (c) Check and disable CPU affinity rules that may be forcing scheduling to specific cores, (d) Review DRS settings—ensure resource pools are not over-committed.

Remediation plan documented: 'Step 1: Reduce web-server-01 vCPU from 8 to 6. Step 2: Verify DRS cluster-level CPU %RDY drops below 10%. Step 3: Re-measure VM application latency to ensure no performance regression.'
Do NOT immediately apply these changes in a VCDX interview context—outline the plan and success criteria. Panelists are assessing your diagnostic reasoning, not your willingness to make changes.
Step 7

Document your CPU analysis in a structured format: Observation (metric values), Interpretation (what the metrics indicate), Action (specific change or next step), Rationale (why this action addresses the root cause).

Documented OIAR analysis for CPU task in your lab notes. Example: Observation: %RDY=20%, %CSTP=2%, vCPU count=8. Interpretation: Host CPU oversubscribed; SMP VM vCPU count misaligned. Action: Reduce vCPU to 6. Rationale: Reduce scheduling complexity.
This OIAR format mirrors VCDX scoring—panelists want to see evidence of systemic thinking.

Validation Gate

Check: CPU analysis complete with documented metrics and remediation plan

Expected: Lab notes show interpreted esxtop metrics, identified contention pattern, and evidence-based remediation recommendation

Common Errors

Confusing %RDY with %WAIT. %RDY = time waiting for CPU scheduler; %WAIT = time waiting for I/O completion. High %RDY + low %WAIT = CPU contention. High %WAIT + low %RDY = storage/network issue.
Fix: Before recommending vCPU changes, always check %WAIT. If %WAIT is high, the issue is I/O-bound, not CPU-bound. Reducing vCPU won't help.
Recommending vCPU reduction without considering application threading model. Some apps are multi-threaded; others are single-threaded but spawning many processes.
Fix: Ask: Is the VM's workload truly CPU-bound, or is it throttled by contention? Adjust vCPU proportionally to physical cores, not arbitrarily.
Ignoring %SYS (system overhead). High %SYS can indicate driver inefficiency, memory pressure causing page-table overhead, or excessive interrupt handling.
Fix: If %SYS > 8%, investigate memory pressure and driver versions before adjusting vCPU.

Task 2 Memory Performance Analysis and Pressure Detection

Task requires understanding of ESXi memory management layering: guest paging < balloon < compression < host swap. Panelists will ask: At what point does balloon activity become concerning? How do you distinguish memory pressure from memory efficiency? What is the performance cost of compression vs. ballooning vs. swapping?

Interpret memory reclamation metrics (balloon, swapping, compression) to identify memory pressure and recommend resource allocation adjustments.

Step 1

While in esxtop, press 'm' to switch to the memory view. Allow the view to render; it may take a few seconds.

Memory view displays columns including MCTLSZ (balloon-driven memory reclamation), SWCUR (current host swap usage), SWTGT (host swap target), ZIP/UNZIP (page compression/decompression rates), ACTVMB (active guest memory), GRANT (memory granted to VM), NOVER (memory overcommit ratio).
If memory columns don't fit the screen, press 'f' to customize fields. Ensure MCTLSZ, SWCUR, ZIP, and ACTVMB are visible.
Step 2

Examine the memory view for a VM allocated 64GB. Note MCTLSZ, ACTVMB, SWCUR, and SWTGT values. Document the current state.

Baseline: e.g., 'db-vm-01: Allocated=64GB, MCTLSZ=0GB, ACTVMB=48GB, SWCUR=0GB, SWTGT=0GB, ZIP=100/sec, UNZIP=5/sec'. This indicates healthy memory utilization: no balloon reclamation, no swap, minimal compression.
If you don't see compression/swap activity, intentionally increase VM workload (e.g., run 'memtester 50G' on the guest) to observe memory pressure responses.
Step 3

Introduce memory pressure scenario: Assume MCTLSZ=32GB (VM allocated 64GB, balloon reclaimed 32GB) and SWCUR=1.2GB. Interpret: Balloon is reclaiming 32GB (50% of allocation) and host is swapping 1.2GB. This indicates severe memory pressure.

Analysis: 'Host memory is exhausted. The ESXi hypervisor activated balloon driver to reclaim 32GB from this VM (reducing its effective memory to 32GB). Despite ballooning, the host still needs 1.2GB of swap space, indicating cluster-wide memory overallocation. This is a critical issue: VM performance will degrade significantly due to swap I/O latency.'
Balloon reclamation > 20% of VM allocation is a red flag. SWCUR > 0 is a critical indicator of overallocation.
Step 4

Check ZIP and UNZIP metrics. High ZIP (compression) with high UNZIP (decompression) indicates the host is relying on memory compression as a stop-gap, which has CPU overhead. Estimate CPU cost: ~5% of host CPU per 100 ZIP/sec.

If ZIP=500/sec, estimate ~25% of host CPU is consumed by compression overhead. This is unsustainable and indicates the need for hardware upgrade or workload rebalancing.
Compression is a last resort before swapping. If compression is high, treat it as a symptom of memory overallocation, not a solution.
Step 5

Recommend memory remediation: (a) Add physical RAM to the host(s), (b) Reduce the number of VMs, (c) Lower memory reservations to allow overcommit detection, (d) Migrate workloads to less-pressured hosts via DRS, (e) Audit guest OS memory settings for misconfiguration (e.g., huge pages, buffer pools).

Remediation plan: 'Priority 1: Add 32GB RAM to host. Priority 2: Migrate 2 non-critical VMs to another cluster. Priority 3: Set memory reservation on db-vm-01 to 48GB (actual peak usage) to trigger DRS migration if cluster memory < 50GB free.'
In a VCDX interview, provide a ranked plan with cost/benefit analysis. 'Add RAM' is expensive; demonstrate you've considered alternatives.
Step 6

Document memory pressure indicators and decision thresholds in your notes: (1) MCTLSZ > 10% of allocation = monitor closely. (2) MCTLSZ > 20% = action required. (3) SWCUR > 0 = emergency. (4) ZIP > 200/sec = investigate.

Reference table in lab notes with thresholds and recommended actions at each level. Example: 'If MCTLSZ=5GB on a 64GB VM and ZIP=150/sec, initiate DRS balancing but no hardware upgrade yet. If MCTLSZ=20GB and SWCUR=2GB, immediate escalation to management.'
These thresholds are guidelines, not absolutes. Adjust based on your organization's SLA and workload criticality.

Validation Gate

Check: Memory analysis complete with interpreted pressure indicators and ranked remediation plan

Expected: Lab notes document memory metrics, pressure assessment, and actionable remediation with cost/benefit trade-offs

Common Errors

Assuming MCTLSZ > 0 always means a problem. In highly consolidated environments, some ballooning is normal and expected.
Fix: Context matters. A 5GB balloon on a 64GB VM in a 256GB host is different from 32GB on a 64GB VM in a 256GB host. Compare MCTLSZ % to host free memory %.
Ignoring guest OS paging. If ACTVMB is low but VM is slow, the guest OS may be paging internally. esxtop can't see inside the guest.
Fix: Always cross-check: Open a guest console or SSH session and run 'free' (Linux) or 'wmic OS get TotalVisibleMemorySize,FreePhysicalMemory' (Windows) to see guest-level memory pressure.
Recommending memory upgrades without understanding the workload trend. Is memory pressure constant or spike-driven?
Fix: Use batch mode (see Task 4) to capture hourly trends. Spiky workloads may benefit from rightsizing instead of hardware upgrades.

Task 3 Storage and Disk I/O Analysis

Task covers data plane (disk latency) and control plane (vSAN metadata) diagnostics. Panelists will ask: How do you distinguish storage array latency from network latency in vSAN? What is the significance of QAVG vs. DAVG? When is storage latency a VM issue vs. a platform issue?

Diagnose storage latency using esxtop disk views and correlate metrics with vSAN or backend storage health.

Step 1

From esxtop, press 'u' to switch to disk device view (physical/logical disk). Wait for the view to render.

Disk view shows columns: DAVG (device average latency), KAVG (kernel average latency), GAVG (guest average latency), QAVG (queue average latency), CMDS/s (commands per second), READS/s, WRITES/s, BUSYQ, ACTV (active commands).
Latency in esxtop is displayed in milliseconds. If DAVG is not visible, press 'f' to add it.
Step 2

Examine your storage device list. Identify the primary data store used by your lab VMs. Note the baseline DAVG. A healthy SSD-backed datastore should show DAVG < 5ms. HDD-backed should show DAVG < 10ms.

Baseline recorded: e.g., 'vSAN-datastore: DAVG=3.2ms, READS/s=450, WRITES/s=120, QAVG=1.1, ACTV=2'. This indicates healthy storage with low latency and moderate queue depth.
Run 'iostat -x 1' on the ESXi host to cross-validate esxtop disk metrics with OS-level disk statistics.
Step 3

Introduce disk I/O stress: Run a heavy I/O workload on a guest VM (e.g., 'fio --name=randrw --ioengine=libaio --iodepth=32 --rw=randrw --bs=8k --runtime=300'). Monitor esxtop disk view in real-time.

Stressed state captured: e.g., 'vSAN-datastore: DAVG=45ms, QAVG=12.3, ACTV=28, READS/s=1200, WRITES/s=400'. DAVG jumped from 3.2ms to 45ms (14x increase), QAVG is high, ACTV is near queue capacity.
45ms is representative of a backend storage bottleneck. Compare to baseline to identify the magnitude of degradation.
Step 4

Interpret the scenario: DAVG=45ms with baseline 5ms = 9x storage latency spike. Analyze root cause. In vSAN, check: (a) Is this a vSAN rebuild scenario (rebuilding a failed disk/host)? (b) Are vSAN resync operations active? (c) Is the storage network saturated (vSAN uses dedicated network)?

Diagnostic analysis: 'DAVG spike is within hypervisor's disk scheduler but beyond storage array acceptable latency. Next step: Check vSAN metrics (vSAN cluster health, rebuild status, network capacity). If vSAN is healthy, suspect: backend SAN array degraded, network congestion, or insufficient queue depth on storage targets.'
esxtop DAVG is cumulative latency from VM I/O request to storage device response. It includes network latency for vSAN. KAVG measures hypervisor overhead; if KAVG is high relative to DAVG, the hypervisor is the bottleneck.
Step 5

If your lab uses vSAN, SSH to a vSAN-enabled host and run vSAN-specific metrics: 'vsanstats -a' or 'vdq -l' (from vdq utility in vSAN Build 2.0+) to check disk group health, rebalance operations, and physical disk latency.

vSAN metrics reveal: e.g., 'Disk group 1: Component resync in progress (15% complete). Physical disk latency: 4ms avg. Network latency to peer hosts: 1ms avg.' This explains the DAVG spike: vSAN is rebalancing components, consuming I/O capacity.
If vSAN is not available in your lab, document this as a prerequisite limitation and focus on generic disk latency interpretation.
Step 6

Press 'v' (VM disk view) to see disk latency from the VM perspective (GAVG). Identify which VM is experiencing the worst latency. GAVG should match or exceed DAVG (since GAVG includes VM OS scheduler overhead).

VM disk view: e.g., 'web-server-01: GAVG=48ms, DAVG=45ms, READS/s=800.' The 3ms difference (48-45) represents VM guest OS overhead. This is normal.
If GAVG is significantly higher than DAVG (e.g., GAVG=100ms, DAVG=45ms), the guest OS has queue depth limits or high interrupt handling overhead.
Step 7

Recommend storage remediation based on your analysis: (a) For vSAN rebuild scenarios, wait for rebuild to complete and re-measure latency. (b) For persistent high DAVG, increase storage capacity (add disks to vSAN disk groups, upgrade SAN array). (c) Check network for vSAN: ensure vSAN network (usually VLAN 1647/1648) is not congested. (d) Consider QoS policies to throttle non-critical I/O during peak periods.

Remediation plan: 'Step 1: Wait for vSAN rebuild to complete (est. 2 hours). Step 2: Re-measure DAVG to confirm recovery. Step 3: If DAVG remains > 10ms, request 2 additional storage disks per host to increase vSAN IOPS capacity. Step 4: Implement VM-level I/O QoS (1000 IOPS cap on reporting VM, 5000 IOPS cap on production VM).'
Remediation for storage is often slow (hardware procurement, vSAN rebuilds). In a VCDX interview, emphasize patience and data-driven decision-making.

Validation Gate

Check: Storage I/O analysis complete with diagnosed latency spike and remediation recommendation

Expected: Lab notes document esxtop disk metrics, baseline vs. stressed state comparison, vSAN-specific diagnostics (if applicable), and tiered remediation plan

Common Errors

Confusing DAVG with GAVG. DAVG is device-level latency (what the storage returns); GAVG is guest-level (what the VM observes). High GAVG with low DAVG indicates VM guest OS throttling, not storage issue.
Fix: Always check both metrics. If GAVG > 3*DAVG, suspect guest OS paging or I/O scheduler misconfiguration.
Ignoring QAVG. A high DAVG with low QAVG suggests the storage is busy but not backlogged—likely a transient spike. High QAVG with high DAVG indicates sustained contention and backlog.
Fix: QAVG trend is more important than DAVG snapshot. Monitor QAVG over time to distinguish spikes from sustained load.
Not considering vSAN-specific latency sources. vSAN adds network latency for replica acknowledgments. A 45ms DAVG in vSAN includes network round-trip time; don't assume it's array latency.
Fix: In vSAN environments, always check network latency first (vmkping between transport nodes). If network is healthy, then suspect disk or rebalance operations.

Task 4 Host-Level Analysis and Batch Mode Trending

Task emphasizes systematic trending and documentation. Panelists will ask: How do you define a healthy baseline? What time window is needed for statistically valid trending? How would you design a monitoring strategy based on batch data? Can you spot anomalies in trend data?

Analyze host-level resource saturation and establish performance baselines using esxtop batch mode for trend analysis and documentation.

Step 1

Return to esxtop interactive CPU view ('c'). Observe the host kernel metrics (typically listed first or in a specific row). Note host %USED and %SYS. Host %USED near 85%+ indicates resource saturation; %SYS > 8% indicates high system overhead (drivers, interrupts, page table management).

Host metrics captured: e.g., 'Host CPU: %USED=82%, %SYS=7.5%, %IDLE=10.5%'. The 82% utilization is near saturation; further VM consolidation risks contention. The 7.5% system overhead is within acceptable range.
Host %USED includes all VMs plus hypervisor overhead. Unlike VM %USED, it should not exceed 85% under normal operations; leave headroom for spikes.
Step 2

Interpret the scenario: Host %USED=85%, %SYS=8%. The host is near saturation, and system overhead is climbing. Possible causes: (a) Too many VMs for the hardware. (b) Driver inefficiency (NIC driver, storage controller driver). (c) Memory pressure causing high page-table walks. (d) Interrupt storms from overactive devices.

Analysis: 'Host approaching saturation. The 8% system overhead is elevated. Recommend: (1) Check driver versions (NIC, storage). (2) Review memory pressure (if memory is ballooned/swapping, page-table overhead increases). (3) Reduce VM count by 2-3 non-critical VMs or enable power management (C-states) if available.'
Don't confuse saturation with utilization. 85% host %USED is not dangerous if load is predictable. 95% is dangerous if spikes occur.
Step 3

Exit esxtop interactive mode ('q'). Prepare to run esxtop in batch mode for trending. Batch mode captures metrics at regular intervals and outputs CSV data suitable for spreadsheet analysis.

esxtop exited cleanly. You're back at the ESXi shell prompt.
Batch mode output is CSV; you'll import it into a spreadsheet in the next step.
Step 4

Run esxtop in batch mode with the command: 'esxtop -b -d 5 -n 720 > /tmp/perf.csv'. This captures 720 samples at 5-second intervals (60 minutes of data). Let it run to completion.

esxtop batch mode runs silently. After ~60 minutes, /tmp/perf.csv will contain 720 rows of CSV data with timestamps and all esxtop metrics (CPU, memory, disk, network, virtual machine uptime).
Use '-n 720' for 1 hour. For daily trending, use '-n 17280' (24 hours). For larger captures, redirect to a host with ample disk space (e.g., /vmfs/volumes/datastore).
Step 5

Retrieve the CSV file to your workstation via SCP or FTP. Example: 'scp root@your-esxi-ip:/tmp/perf.csv ./perf-baseline.csv'. Open the file in a spreadsheet (Excel, Google Sheets, LibreOffice Calc).

CSV file downloaded to local workstation. When opened in a spreadsheet, rows represent time samples, columns represent metrics (timestamps, %USED, %RDY, %SYS, MCTLSZ, DAVG, etc.). Data is ready for analysis.
The first row may contain header metadata from esxtop; import carefully to avoid treating metadata as data.
Step 6

In the spreadsheet, create a chart of key metrics over time: (1) Plot host %USED, %RDY, %SYS as a multi-series line chart. (2) Identify peak utilization and note the timestamp. (3) Look for trends: Is utilization increasing over the hour, or stable? (4) Calculate moving average (15-minute window) to smooth spikes.

Trend chart generated showing: Host %USED varies between 50-85% over the hour, with peak at 10:35 AM. %RDY spikes to 18% around peak utilization, then returns to baseline. %SYS remains stable at 6-8% throughout. Moving average smooths out sub-minute spikes.
Trends reveal patterns. If %USED increases linearly over the hour, the workload is growing; expect saturation. If %USED spikes at regular intervals, there's a periodic job (backups, reports) that contends for resources.
Step 7

Create a performance baseline document in your lab notes (or as a formal write-up). Include: (1) Host specifications (CPU sockets, cores, RAM). (2) Baseline metrics during steady-state (low load): %USED, %RDY, %SYS, DAVG, MCTLSZ. (3) Baseline metrics during peak load: same metrics at 75-85% utilization. (4) Thresholds for alerting (when to escalate).

Baseline document: 'Host: 2x Intel Xeon 16-core CPUs, 256GB RAM. Steady-state: %USED=45%, %RDY=3%, DAVG=4ms, MCTLSZ=0GB. Peak load (85% utilization): %RDY should stay < 10%, DAVG < 15ms, MCTLSZ < 5GB for health. Alert threshold: %RDY > 15% or DAVG > 20ms triggers investigation.'
This baseline document is your reference for troubleshooting. Store it in your lab documentation for future comparisons.
Step 8

Document batch mode methodology in your notes: When to run, how long to run, what to measure, and how to interpret results. Include a sample command and sample output interpretation.

Methodology documented: 'Run esxtop -b -d 5 -n 17280 to capture 24-hour trends. Import CSV to spreadsheet. Calculate hourly averages and identify peak utilization window. If peak utilization > 85%, correlate with workload scheduling (backups, reports) and plan load balancing or capacity upgrade.'
This methodology is a VCDX-style artifact: it demonstrates that you approach performance analysis systematically, not reactively.

Validation Gate

Check: Batch mode trending complete with baseline document and trend analysis interpretation

Expected: Lab notes contain esxtop batch command, CSV trend analysis, performance baseline document with thresholds, and methodology for future trending

Common Errors

Running batch mode for insufficient time. 5 minutes of data is not enough to identify patterns. 60+ minutes is necessary for meaningful trending.
Fix: Always run at least 1 hour of batch mode; for weekly trends, capture 24 hours. Use '-d 10' or '-d 30' for longer captures (reduces file size).
Forgetting to account for workload timing. If batch mode captures a time window when no workloads are running, the baseline is artificially low.
Fix: Run batch mode during normal business hours when your lab VMs are operational. Document the time window and workload characteristics in your baseline report.
Misinterpreting CSV column order. esxtop batch mode output column order is not intuitive; without headers, metrics are hard to identify.
Fix: Open the CSV in a spreadsheet and manually add headers based on esxtop interactive view. Use 'esxtop -b -h' (if available) to get column descriptions.

Final Validation

Lab completed when all 4 tasks are executed, metrics are documented, and remediation recommendations are justified.

✓ CPU analysis with high-contention scenario analyzed and remediation recommended → Task 1 complete: %RDY, %CSTP metrics interpreted; vCPU adjustment documented with rationale

✓ Memory pressure scenario analyzed with reclamation metrics documented → Task 2 complete: MCTLSZ, SWCUR, ZIP metrics documented; memory pressure assessed; remediation ranked by priority

✓ Storage I/O latency diagnosed with baseline and stress comparisons → Task 3 complete: DAVG, KAVG, GAVG metrics documented; vSAN health checked (if applicable); storage remediation planned

✓ Host-level baseline established via batch mode trending → Task 4 complete: esxtop batch CSV captured, trending chart created, baseline document produced with alert thresholds

Cleanup / Restore

• Remove /tmp/perf.csv from ESXi host (or retain for 30 days as audit trail)

• Revert any VM I/O stress workloads (stop fio, stress-ng)

• Return host CPU/memory to baseline utilization (if VMs were intentionally overloaded)

• Restore baseline snapshots if any changes were made to VM configurations

Design Reflection (VCDX)

This lab emphasizes diagnostic methodology and evidence-based decision-making—core VCDX competencies. Panelists will probe: (1) How do you validate that your esxtop diagnosis matches actual VM performance issues? (2) When do you recommend quick fixes vs. architectural changes? (3) How do you balance performance tuning against operational risk? (4) Can you articulate the trade-offs between CPU overcommit and power efficiency? The RCAR framework below provides structure for panel discussion.

Requirements

  • ESXi host(s) with operational VMs and I/O workloads
  • SSH/CLI access to ESXi host
  • Spreadsheet software for CSV analysis and charting
  • Lab VMs running representative workloads (CPU, memory, disk I/O)
  • vSAN deployment (optional, for advanced storage metrics) or iSCSI/NFS datastore
  • Baseline performance data for the lab environment (to establish 'healthy' thresholds)

Constraints

  • esxtop must be run directly on ESXi host; remote monitoring tools (vRealize Operations) provide different metrics and context
  • Batch mode CSV captures are host-specific; trending across multiple hosts requires aggregation logic outside esxtop
  • CPU/memory contention metrics are OS-kernel-dependent; different ESXi versions may report slightly different values
  • vSAN-specific metrics (rebuild status, rebalance operations) require vSAN Skyline Health or vSAN Management REST API; not all metrics available in esxtop
  • Interactive esxtop view requires SSH terminal session; cannot be automated or scheduled without batch mode

Assumptions

  • Lab VMs are configured with realistic resource allocations (not over-provisioned with 100 vCPUs on 16-core host)
  • Storage backend is functioning normally during the lab (no storage array failures masking remediation)
  • Network connectivity between ESXi hosts and vSAN network is stable (for vSAN trending)
  • Lab operator has basic Linux/Windows CLI knowledge to trigger I/O workloads and interpret OS-level metrics
  • Remediation recommendations assume you have approval to modify VM configurations (vCPU count, memory allocation, DRS settings)
  • Baseline 'healthy' values are organization-specific; lab uses industry standard thresholds (e.g., %RDY < 10%) as reference

Risks

  • Risk: Over-aggressive VM scaling during CPU contention analysis may cause application downtime if vCPU is reduced too much — Mitigation: In a VCDX interview, outline the plan and revert immediately if issues emerge. Do not apply changes without supervisor approval in production.
  • Risk: Batch mode trending captures can consume significant disk space if run for extended periods (e.g., 24-hour capture = 5-10 MB CSV) — Mitigation: Plan storage; use '-d 30' (30-second intervals) instead of '-d 5' for longer captures. Compress/delete old CSV files after analysis.
  • Risk: Memory pressure scenario (MCTLSZ > 20%, SWCUR > 0) may trigger emergency performance degradation or guest OS OOM (Out of Memory) kill — Mitigation: In a lab environment, monitor guest OS closely during memory stress. Have snapshot/revert strategy ready if guest becomes unresponsive.
  • Risk: Storage I/O stress (fio) may impact concurrent workloads if run on shared datastore; can cause noisy neighbor contention — Mitigation: Run I/O stress on a dedicated VM on dedicated datastore, or run during off-hours. Document timing in lab notes.

Self-Assessment Discussion Prompts

  1. Walk me through how you decided that 20% %RDY was the threshold for 'concerning contention.' Where does that number come from? Would you apply the same threshold to a DSS workload vs. an OLTP workload?
  2. You recommended reducing vCPU from 8 to 6. How would you validate that this change improves performance without harming application throughput? What metrics would you monitor post-change?
  3. In the memory pressure scenario (MCTLSZ=32GB, SWCUR=1.2GB), you recommended adding RAM as Priority 1. But hardware upgrades take weeks. What is your Priority 2 action to buy time while procurement happens?
  4. You observed DAVG spiking from 5ms to 45ms during I/O stress. Before blaming the storage array, what hypervisor/network factors would you investigate? How would you distinguish storage latency from network latency in a vSAN environment?
  5. Batch mode gave you 60 minutes of baseline data. Is that enough to define 'healthy' thresholds for the next quarter? What time window would you recommend for production baselining, and why?
  6. A VM shows %RDY=15%, DAVG=8ms, but the application team says performance is fine. How do you reconcile esxtop metrics with user-perceived performance? What layer are you missing?
  7. You're being asked to consolidate 20 more VMs onto this cluster. Based on your baseline, what's the maximum consolidated VMs this cluster can handle before %RDY exceeds 10%?
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.