VCF 9.0 Support (2V0-15.25)
Comprehensive study guide for the VMware Cloud Foundation 9.0 Support exam (2V0-16.25). Covers VCF and vSphere Foundation architecture, installation, operations, upgrades, licensing, workload domains, VCF Automation, Broadcom support processes, and deep troubleshooting scenarios. Enriched with content from the VCFFTS9 official lecture manual (933-page course material).
Exam Blueprint Weights
Version Evolution
VCFFTS9 is VCF 9.0 focused. Covers both VCF Private Cloud and vSphere Foundation. Key VCF 9.0 changes: VCF Installer replaces Cloud Builder, core-based licensing replaces socket-based, VCF Operations console is integrated, ESXi host prep via VCF Installer wizard. Course includes upgrade paths from VCF 5.x and vSphere-to-VCF conversion/import.
Learning Outcomes
- Articulate VCF and vSphere Foundation use cases, differentiating between them
- Deploy a VCF fleet using both the Installer wizard and JSON deployment specification
- Navigate VCF Operations console to monitor storage, compute, and security health
- Use Log Assist for log analysis and troubleshooting
- Prepare and configure ESXi hosts for VCF deployment
- Explain vSAN architecture and its role in VCF storage
- Describe NSX core components and their role in VCF networking
- Plan VCF 5.x to 9.0 upgrades and vSphere-to-VCF conversions
- Understand the VCF 9.0 licensing model (core-based subscription)
- Create and manage workload domains and network pools
- Describe VCF Automation capabilities
- Navigate Broadcom Wolken case management for support cases
- Support Readiness:Before calling support, have ready: (1) exact error, (2) SoS bundles from all affected hosts, (3) vCenter logs (last 24 hours), (4) list of recent changes (dates, names, approvals),
- Support Exam Insight:Panelists (or cert exams) may ask "You have case SID-12345 for customer Y; their environment is vSAN 9.0.1, vCenter 9.0 GA. Which KB articles are applicable?" You must recognize v
- Obj 5.2: Plan and execute VCF 5.x to 9.0 upgrades — pre-checks, bundle management, rolling upgrade sequence
- Obj 5.3: Create and scale workload domains — commission hosts, add/remove clusters, import vSphere environments
- Obj 5.11: Configure HCX and execute workload migrations — RAV, bulk migration, cold migration, L2 extension
VCFFTS9 Exam Strategy & Lecture Manual Overview#
The VMware Cloud Foundation Fundamentals for Technical Support (VCFFTS9) exam covers 15 modules across 933 pages of lecture material. The exam targets Technical Support Engineers and focuses on understanding VCF architecture, installation, upgrades, operations, and support processes.
Key exam areas by weight (based on module page count): Compute Core Concepts (156 pages — largest), VCF Upgrades (168 pages — second largest), VCF Automation (80 pages), Storage Core Concepts (74 pages), VCF Operations (68 pages), Networking Core Concepts (62 pages), Technical Support Fundamentals (61 pages).
Study strategy: prioritize Compute (Module 4) and Upgrades (Module 9) as they constitute 35% of the total content. The support-specific modules (Wolken case management, KB articles, Log Assist) are unique to this exam and not covered by other VMware certifications.
Key Takeaways
- 933-page lecture manual across 15 modules — prioritize by page count
- Compute (156pp) and Upgrades (168pp) = 35% of content
- Support-specific content (Wolken, KB, Log Assist) is unique to VCFFTS9
- VCF 9.0 licensing model is heavily tested — understand file vs keys transition
- Cross-reference with existing Academy sections for deeper coverage
VCF Troubleshooting Methodology#
Structured Troubleshooting Approach
Effective troubleshooting follows a systematic process: identify the problem precisely, isolate the failing component, collect evidence, form hypotheses, test, and document the solution for future reference.
Troubleshooting Framework (The Four Ps):
Problem Statement:
"VMs cannot be deployed on cluster-1" is clear. "The system is slow" is vague. Quantify: "Deploy times increased from 2min to 10min starting Tuesday 3pm".
Planes Affected:
Which plane is failing? Management (vCenter down → cannot manage), Control (NSX controller cluster degraded → networking flakes), Data (ESXi host PSOD → VMs crash). Identify plane first to narrow scope.
Peripheral Devices:
Does the issue trace to storage (vSAN), network (NSX, physical NIC), compute (CPU/memory/hypervisor), or external dependencies (NTP, DNS)?
Past Changes:
What changed recently? Patch Tuesday, host reboot, network config update, SRM test failover? Track all changes in a log.
Quick Triage Questions:
Is management plane up? (Can you log into vCenter?)
Is control plane healthy? (NSX Manager cluster responding?)
Which hosts are affected? (All in cluster? One rack? One host?)
When did it start? (Correlate with change logs, maintenance windows, upgrades)
Is it intermittent or consistent? (Intermittent = harder to diagnose; may be time/load-dependent)
Log File Locations & Log Collection
VCF logs are distributed across multiple components. Know where to find them quickly.
SDDC Manager & VCF Lifecycle Logs:
SDDC Manager (runs on vCenter VM, VCF orchestration)
/var/log/vmware/vcf/sddc-manager/sddc-manager.log
/var/log/vmware/vcf/lcm/ (Lifecycle Manager - patches, upgrades)
- lcm.log (main LCM process)
- lcm-upgrade.log (upgrade-specific events)
- pre-check.log (pre-upgrade validation failures)
VCF Operations Platform
/var/log/vmware/vcf/operations/
- vrops-suite-install.log
- vrli-collector.log (vRealize Log Insight)
On SDDC Manager, check bundle cache:
/nfs/vmware/sddc-manager/bundles/ (downloaded patch bundles)
vCenter Logs:
vCenter Server (on vCenter VM)
- /var/log/vmware/vcenter-server/catalina.log (main web container)
- /var/log/vmware/vcenter-server/vmware-vcenterd.log (core VC process)
- /var/log/vmware/vcenter-server/vcu-service.log (vSAN, storage)
vCenter Service Control
/var/log/vmware/vpxd/vpxd.log (vCenter daemon)
SSO & authentication
/var/log/vmware/sso/sso.log
/var/log/vmware/sso/sts.log (Secure Token Service)
Event database (contains all VC events for 30 days)
/var/lib/vmware/vpxd/events/events.db (query via vpxd logs)
ESXi Logs (on each host via SSH or DCUI):
ESXi kernel and vmkernel logging
/var/log/vmkernel (main kernel log, circular ~4MB)
/var/log/vmkernel.0.gz (compressed backup)
ESXi system messages
/var/log/syslog.log
vSAN-specific
/var/log/vsand.log (vSAN daemon)
/var/log/vmkwarning (kernel warnings, SCSI errors)
vMotion & hostd
/var/log/hostd.log (ESXi agent managing VMs)
Storage/SCSI
/var/log/vobd.log (Virtual Object Binding Daemon, vSAN coordination)
NSX Manager & Control Plane Logs:
NSX Manager (3-node cluster)
- /var/log/vmware/nsx-manager/manager.log
- /var/log/vmware/nsx-manager/nsx-manager.log (main API)
- /var/log/vmware/nsx-manager/audit.log (all API operations)
NSX central control plane (runs inside the NSX Manager appliance since NSX-T 2.4)
/var/log/vmware/nsx-controller/ (CCP logs on each NSX Manager node)
/var/log/vmware/nsx-controller/controller.log
NSX Edge (logical router)
/var/log/vmware/nsx-edge/ (on edge VM)
/var/log/vmware/nsx-edge/syslog
NSX Kernel Module (on ESXi hosts)
/var/log/vmkernel (filter for nsx-gen logs, DFW entries)
Support and Serviceability (SoS) Utility
SoS is a VMware tool that collects logs, diagnostics, and configuration from ESXi hosts and vCenter for support bundles.
Accessing SoS on ESXi:
SSH to ESXi host
ssh root@esxi-host
Run SoS to generate bundle
/sbin/sos
Output: sosreport-[hostname]-[date].tar.gz in /var/log/
Size: typically 50-200MB depending on verbosity
List SoS bundles
ls -lah /var/log/sosreport*
Securely transfer to support workstation
scp /var/log/sosreport*.tar.gz support-workstation:/tmp/
Log Bundles from vCenter:
vCenter has built-in export (Support -> Request Log Bundle)
Or via SSH to vCenter:
/usr/lib/vmware-logsummary/logsummary.pl > /tmp/logsummary.txt
Collect vCenter diagnostics:
/usr/lib/vmware-logsummary/vmware-support-dumps.py \
--output /tmp/vcenter-dump.tar
SDDC Manager logs (if available on Linux mgmt VM)
tar czf sddc-manager-logs.tar.gz /var/log/vmware/vcf/
VMware Support Engagement Process
When to escalate to VMware and how to prepare for support call.
Critical issues:
Production outage, data corruption, security breach → open P1 case immediately via VMware Support portal (support.vmware.com). High severity: Feature not working, performance degraded, recurring errors → open P2/P3 case, provide initial diagnostics.
Information needed for support case:
Exact error messages, timestamps (convert
Key Takeaways
- Support Readiness:Before calling support, have ready: (1) exact error, (2) SoS bundles from all affected hosts, (3) vCenter logs (last 24 hours), (4) list of recent changes (dates, names, approvals), (5) network diagram showing which subnets are affected.
vSphere Troubleshooting#
ESXi Host Issues
Purple Screen of Death (PSOD):
ESXi kernel panic. Host stops executing VMs and displays purple error code.
Common causes:
Hardware failure (CPU, memory), firmware bug (BIOS update available), driver incompatibility (NIC, HBA), VM code triggering hypervisor bug, memory corruption.
Diagnosis:
Check ESXi console; note the error code and memory dump address. Collect PSOD dump via
esxcli system dump file list
.
Recovery:
Reboot host (VMware HA restarts VMs on other hosts). Isolate host from cluster, run diagnostics (hardware test, firmware update). If PSODs recur, RMA hardware.
Prevention:
Keep ESXi, firmware, and drivers at latest patched levels. Use hardware support matrix to verify components.
Network Connectivity Loss:
Host cannot reach vCenter, storage, or other hosts.
Diagnose network on ESXi:
esxcli network ip interface list
Shows vMotion, management, vSAN port groups and IPs
esxcli network ip route ipv4 list
Verify routes (should have default gateway)
ping 8.8.8.8 (or internal gateway IP)
Test connectivity
esxcli hardware nic list
List physical NICs; note status (Up/Down)
ethtool vmnic0
Check link speed, autoneg status
esxcli network vport get
List virtual ports on physical NIC; high dropped count = QoS saturation
Check vSwitch/DVS configuration:
esxcli network vswitch standard list
esxcli network dvswitch dvport list -d 'dvswitch-id'
Storage Disconnects:
LUNs disappear or go offline intermittently.
Symptoms:
Datastore goes "Inaccessible" or "Unmounted" in vCenter. VMs freeze. Error: "No accessible replicas".
Causes:
SAN fabric loop, switch port flapping, HBA firmware mismatch, SCSI reservation conflict, storage array failover.
Diagnosis:
Check ESXi vmkernel log for SCSI errors:
cat /var/log/vmkernel | grep -i scsi
. Check SAN switch logs for port resets. Verify HBA firmware is up-to-date and matches SAN vendor compatibility list.
Recovery:
Rescan storage:
esxcli storage core adapter rescan --adapter=vmhba0
. If LUN still missing, verify zoning on FC switch. Reboot host if needed.
vCenter Server Issues
vCenter Services Not Starting:
Web UI unreachable, cannot connect to vCenter.
Check service status on vCenter VM:
- systemctl status vmware-vcenterd
- systemctl status vmware-sso
- systemctl status vmware-vpxd (vCenter daemon)
View recent errors:
journalctl -u vmware-vcenterd -n 50 --no-pager
Check connectivity to database:
nc -zv 'database-host' 5432 (Postgres, or 1433 for MSSQL)
Restart services (in order):
- systemctl stop vmware-vcenterd
- systemctl stop vmware-sso
- systemctl start vmware-sso
- systemctl start vmware-vcenterd
If still failing, check disk space:
df -h /storage/vcenterd
vCenter needs >20GB free for operation
Database Issues:
vCenter cannot connect to its database, or database is slow/corrupt.
Embedded Postgres (default):
vCenter 8.x uses embedded Postgres. Check size:
du -sh /storage/db
. If >50GB, consider external DB.
Database locks:
If DB is locked, other vCenter operations hang. Check active sessions:
SHOW max_connections;
in Postgres.
Event database vacuum:
Event DB (events.db) grows large over time. Trim old events: vCenter → Administration → System Configuration → Event Database → Compact.
Backup database regularly:
Use vCenter native backup (doesn't require DB backup if using embedded Postgres; backup tool handles it).
SSO (Single Sign-On) Problems:
Cannot log into vCenter; "Authentication Failed".
Causes:
SSO service down, local user database corrupted, LDAP/AD connection broken, certificate expired.
Check SSO:
systemctl status vmware-sso
. If down, restart (may take 5 min to initialize).
Certificate check:
openssl x509 -in /etc/vmware-sso/certs/sso.crt -text -noout | grep -A2 "Validity"
. If expired, request new cert from CA.
LDAP test:
If using AD, test connectivity to DC:
ldapsearch -H ldap://ad-server:389 -D 'cn=user,dc=corp,dc=local' -w password
. If fails, verify AD reachability and credentials.
Reset local admin:
If AD is down, still need to log in with local administrator@vsphere.local account. Recovery option available at vCenter login page.
VM Issues
VM Cannot Power On:
VM stays in "Powering On" or fails with error.
Check vCenter task history:
Recent Tasks panel → look for VM power-on task with error message
Common errors and causes:
"No compatible hosts found" → Admission control failed (HA insufficient capacity)
Fix: Check HA admission control settings; may need to power off non-critical VM
"Cannot open disk" → VMDK is locked or storage is inaccessible
Fix: Check storage connectivity; if lock persists, restart hostd:
/etc/init.d/hostd restart
"Insufficient memory" → Host RAM full
Fix: Power off non-critical VMs; check VM ballooning in esxtop
"VM config file missing" → VM registered but datastore moved/renamed Fix: Re-register VM from datastore (Datacenter → Storage → browse datastore → .vmx file)
Snapshot Consolidation Failure:
VM has delta disk snapshots that won't consolidate/delete.
Ca
Key Takeaways
- Purple Screen of Death (PSOD) decoding: BugCheck codes (0x11, 0x21, etc.) pinpoint root cause; always capture backtrace for support escalation.
- Storage path failover: vSAN resyncs automatically; monitor resync rate and adjust throttle if impacting production I/O.
- VM power-on failures: admission control, memory balloon, and snapshot consolidation are top 3 causes; triage in order.
vSAN Troubleshooting#
vSAN Health Check & Diagnostics
vSAN health check performs deep diagnostics on cluster, storage, network, and disk health. Interpret failures and remediate systematically.
Running vSAN Health Check:
vCenter → Cluster → Monitor → vSAN → Health
Click "Run Health Check Now"
CLI method (on any ESXi host in vSAN cluster):
esxcli vsan health cluster list
Shows: Storage health, network health, limit health, host health
esxcli vsan health object list -u '
'
Detailed health for specific object
esxcli vsan cluster get
Cluster status, mode (OSA or ESA), current capacity
Health Check Categories & Failures:
Storage health:
Issues with disk groups (bad disks, RAID failures). Fix: Replace failed disk, rebuild vSAN data.
Network health:
Latency between hosts >5ms (stretched cluster), lost packets (MTU mismatch). Fix: Adjust network config, enable MTU 9000 on switches.
Limit health:
Cluster resource limits hit (max VMs, max objects, dedup memory). Fix: Reduce object count, upgrade hardware, enable compression.
Host health:
ESXi host unable to contribute storage (disk errors, not claimed). Fix: Re-claim disks, or replace host.
Disk Group & Capacity Issues
Disk Group Failures:
One or more disks in a disk group are failed or degraded.
List disk groups on host:
esxcli vsan storage list
Shows disks, RAID status, capacity
If disk is failed (RED):
esxcli vsan storage remove --disk='naa.id' --force
Removes disk from disk group; triggers rebuild on remaining disks
Replace the physical disk (DPM mode, or manually power-cycle), then re-add:
esxcli vsan storage claim --disk='naa.id'
Claims new disk back into disk group
Monitor rebuild progress:
esxcli vsan resync get
Shows objects resyncing, ETA
Capacity Warnings:
vSAN cluster is 80%+ full.
Symptom:
Health check warns "Capacity limit reached", new objects cannot be placed.
Root cause:
Insufficient capacity (add hosts/disks), dedup/compression not enabled (data is not compressing as expected), orphaned objects (snapshot leftovers, failed VMs).
Remediation:
(1) Identify large VMs/snapshots; delete unnecessary snapshots, (2) Add new disk group to host or new host to cluster, (3) Enable compression if not on ESA (if using OSA), (4) Analyze dedup ratio to confirm effectiveness.
vSAN Resync & Object Health
Resync Operations:
After disk/host failure, vSAN automatically resyncs data to maintain FTT.
Monitor active resyncs:
esxcli vsan resync get
Shows: component syncing, progress %, ETA, network traffic
During large resync (many objects):
Network bandwidth consumed can be 50-80% of available; may impact VMs
Solution: Reduce resync priority (if not emergency)
esxcli vsan cluster get
Shows resyncing objects, total objects, rebuild completion time
Object Health States:
vSAN reports objects in various states (healthy, degraded, inaccessible).
Healthy (Green):
Object has all replicas, FTT met. No action needed.
Degraded (Yellow):
Object has lost a replica (host down), but still accessible. Resync is in-progress. Wait for resync to complete, or manually trigger rebuild if stalled.
Inaccessible (Red):
Object has lost quorum (>FTT failures). VM cannot access disk. Immediate action needed: bring hosts online, or rebuild from backup.
Absent:
Object deleted intentionally (snapshot consolidation, VM deletion).
Network Partition Handling
Network Partition (Split-Brain):
Cluster splits into multiple non-communicating partitions. vSAN elects partition with majority hosts to continue; minority partitions stop.
Symptom:
VMs on minority partition freeze or are stopped. Health check shows high latency, packet loss to other hosts.
Diagnosis:
Check network connectivity between hosts:
- ping each-host
- . Check vSAN network latency:
- esxcli vsan cluster get
.
Recovery:
Fix network issue (reconnect switch, repair cable). vSAN automatically remerges partitions (quorum retaken by full cluster). Monitor resync during merge.
Data consistency:
If partition lasted
1 day, verify data integrity after merge.
Witness Node in Stretched Cluster
Witness Replacement:
Witness node is down or replaced; cluster loses tiebreaker.
Symptom:
Health check fails; stretched cluster cannot fail over sites.
Fix witness placement:
Witness must be at third location (not co-located with either primary site). Re-configure witness in vCenter: Cluster → Configure → vSAN → Stretched Cluster.
Replace witness VM:
Delete old witness VM, deploy new witness VM on different host. Add to vSAN as witness role.
vSAN Performance Troubleshooting
HCIBench (Hyper-Converged Infrastructure Benchmark):
Performance testing tool for vSAN.
Download HCIBench from VMware site
Deploy on vSAN cluster to measure:
- Sequential read/write IOPS
- Random read/write IOPS
- Latency (avg, p50, p99, p100)
Typical vSAN all-flash (ESA) performance:
- Throughput: 100K - 500K IOPS (depends on disk count, CPU)
- Latency: <1ms for cache hits, <5ms for capacity tier
If performance is below baseline, c
Key Takeaways
- vSAN health check must pass all categories (cluster, storage, network, limits) before production deployment changes.
- Component state transitions: healthy → degraded → inaccessible; resync queue depth and rebuild rate determine recovery time.
- Dedup/compression performance: monitor compression ratio; ratios <1.5:1 indicate dedup not working as expected, verify configuration.
NSX Troubleshooting#
Control Plane vs Data Plane Failures
Control Plane (NSX Manager, Controllers):
Manages policy, routing decisions. If down, cannot create new policies or routes, but existing traffic continues.
Failure symptom:
Cannot log into NSX Manager, cannot create new segments or rules.
Check NSX Manager:
kubectl get pods -A
(NSX Manager is containerized). Verify 3 manager nodes all healthy.
Fix:
Restart NSX Manager services (do not reboot all 3 simultaneously; stagger restarts). Control plane resyncs with data plane ~1 min after recovery.
Data Plane (Kernel Module on ESXi, Edge VMs):
Forwards actual packets. If down, traffic is dropped.
Failure symptom:
VMs cannot communicate (e.g., web cannot reach app tier). NSX Manager is up, but traffic is dropped.
Check data plane:
SSH to ESXi host, verify NSX kernel module is loaded:
- lsmod | grep nsx
- . If not loaded, reload:
- vmkload_mod nsx_im
.
Fix:
Reboot host (VMs migrate via vMotion), or reload kernel module. If reload fails, reboot ESXi host (HA restarts VMs).
Tunnel Endpoint (TEP) Connectivity
TEP is the IP address used for overlay encapsulation (GENEVE). TEP failure = overlay traffic drops.
Check TEP IP on host:
esxcli network ip interface list
Look for "vmk-nsx" or "TEP" port group
Ping TEP from remote host:
ping
If fails, TEP VLAN is down or unreachable
Verify TEP VLAN connectivity:
esxcli network vswitch standard portgroup list
Confirm TEP VLAN is tagged correctly
Rebuild TEP if corrupted:
vCenter → Cluster → Configure → NSX → Reapply NSX Config
This rebuilds kernel modules and TEP network
DFW Troubleshooting
Traceflow Tool:
Test if traffic is allowed by DFW rules.
NSX Manager UI → Troubleshooting → Traceflow
Or CLI:
- nsx-cli --user admin -p
- --hostname '
- '
- API request: POST /api/v1/logical-ports/
- /traceflow
Traceflow results:
DROPPED = DFW rule rejected traffic
FORWARDED = Allowed by rule
UNDELIVERABLE = Destination unreachable (routing issue)
DFW Rule Debugging:
Rule not matching:
Check rule priority (top to bottom, first match wins). Check source/destination scope (segment, VM tag, group name must be exact).
Traffic still passing when rule says DENY:
Another rule above may be allowing it. Review rule set order.
Performance impact:
Too many rules (>200) causes CPU overhead on hosts. Consolidate rules using tags/groups.
Packet Capture for DFW Analysis:
Capture packets at host level (ESXi kernel module):
esxcli system packet list
Shows active packet captures on host
Start capture on specific VM port:
esxcli system packet capture start --tap=
--filter='src 10.0.0.5'
Captures traffic to/from 10.0.0.5
Stop capture and analyze:
esxcli system packet capture stop
Output in pcap format; transfer to Wireshark for analysis
Edge Node Issues
Edge VM Down:
Tier-0 or Tier-1 edge VM crashed, routing fails.
Check edge status:
NSX Manager → Fabric → Nodes → Edge Nodes. Verify "Status" is green.
Restart edge:
Power off/on edge VM, NSX Manager auto-fails traffic to secondary edge (if HA configured).
HA failover:
If using active-active edges, traffic automatically shifts to healthy edge within seconds.
BGP & Routing Failures
BGP Neighbor Down:
NSX Tier-0 cannot establish BGP session with physical router.
Check BGP status on edge:
- nsx-cli --user admin --hostname '
- '
- show routing bgp
Output shows:
BGP neighbor state: Established (good), NotConnected (bad)
If NotConnected, check:
- IP reachability: ping neighbor IP from edge
- TCP/179: telnet neighbor-ip 179 (BGP uses TCP 179)
- BGP credentials: verify AS numbers, router IDs match config
Restart BGP service on edge:
nsx-cli> restart service bgp
NSX Manager Cluster Health
Multi-Node NSX Manager:
3-node cluster for HA. If nodes lose quorum, cluster stops responding.
Check cluster status:
kubectl get nodes
(all 3 should be Ready).
kubectl get pods -A
(all pods should be Running).
Node down:
If 1 of 3 nodes is down, cluster continues with 2 nodes. But if 2 nodes down, no quorum = cluster offline.
Recovery:
Bring down nodes one at a time (allow cluster to resync between). Stagger restarts by 5 minutes.
Key Takeaways
- Control plane vs data plane: Manager down = cannot create new rules; kernel module down = traffic drops. Diagnose separately.
- TEP connectivity: VLAN and MTU (>=1600) must be correct; vmkping TEP IP validates overlay encapsulation path.
- DFW traceflow: test traffic through exact rule set; rule priority matters (top-to-bottom match), verify with traceflow before troubleshooting.
VCF Lifecycle Troubleshooting#
SDDC Manager Troubleshoot — Lifecycle Failures & Rollback
When you need to troubleshoot SDDC Manager lifecycle operations, follow this structured approach.
VCF Upgrade Failure:
SDDC Manager upgrade halts with error, partial completion.
Check upgrade status:
SDDC Manager → Lifecycle Management → Upgrades. View task details and error message.
Common failures:
Pre-check failed (host configuration issue), bundle download failed (network/storage), timeout during upgrade (service hung), certificate expired mid-upgrade.
Rollback procedure:
If upgrade partially completed, rollback by re-running upgrade with previous bundle. SDDC Manager maintains previous version snapshots (for 7 days) to enable rollback.
Manual rollback:
If auto-rollback fails, manually restore snapshot of SDDC Manager/vCenter VM from previous version date.
Bundle Download Failures
VCF bundles are large (>1GB) and require stable connectivity to download.
Check bundle cache on SDDC Manager:
/nfs/vmware/sddc-manager/bundles/
Lists downloaded bundles and their status
If download is slow or fails:
- Check network connectivity: ping vmware.com (internet access)
- Check DNS: nslookup download.vmware.com
- Check NFS storage: df -h /nfs/vmware/sddc-manager/
(needs 100GB+ free for bundle cache)
- Retry download with longer timeout (GUI or CLI)
Manual bundle upload (if download not available):
Upload .tar.gz file to /nfs/vmware/sddc-manager/bundles/
scp bundle.tar.gz admin@sddc-manager:/nfs/vmware/sddc-manager/bundles/
Pre-Check Failures & Remediation
Pre-Upgrade Health Check:
SDDC Manager validates environment before starting upgrade.
Common Pre-Check Failures:
vCenter certificate expired:
Renew certificate before upgrade.
openssl x509 -in /path/to/cert -text -noout | grep Validity
.
DNS/NTP misconfigured:
All VMs must have correct NTP offset (<1 second). Check:
timedatectl
on ESXi hosts and vCenter.
Cluster admission control failure:
Not enough failover capacity for VMs to migrate during upgrade. Enable DPM (power down underutilized hosts) or temporarily suspend HA during upgrade window.
Storage connectivity issue:
Pre-check validates datastore accessibility. If LUN is intermittently unreachable, fix SAN connectivity first.
License incompatibility:
New VCF version requires different license keys. Verify license capacity (e.g., vSAN advanced for stretched cluster) before upgrade.
Certificate Expiry Issues
Certificates used by vCenter, NSX Manager, and SDDC Manager expire if not renewed.
Check vCenter certificate validity:
openssl x509 -in /etc/vmware-sso/certs/sso.crt -text -noout | grep -A2 Validity
Check SDDC Manager certificate:
ssh admin@sddc-manager
openssl x509 -in /etc/ssl/certs/sddc-manager.crt -text -noout
NSX Manager certificate (on NSX Manager VM):
- kubectl exec -it
- \
- openssl x509 -in /etc/ssl/certs/nsx-cert.crt -text -noout
Renewal process:
- Generate CSR (Certificate Signing Request)
- Submit to internal CA or Let's Encrypt
- Import renewed certificate
- Restart service (may cause brief downtime)
Automated renewal (production environments):
Configure certificate auto-renewal with monitoring alerts at 90 days before expiry
Drift Remediation
Configuration Drift:
Desired state (defined in SDDC Manager) diverges from actual state (on hosts/vCenter/NSX).
Causes of drift:
Manual configuration change bypassing SDDC Manager (e.g., admin directly edited NSX rule, or vCenter network)
Patch applied that overwrites configuration
Component rollback during failed upgrade
Detecting drift:
SDDC Manager → Lifecycle Management → Drift Detection
Runs automated check comparing desired vs actual state
Manual check (CLI):
sddc-manager-cli drift-detect --domain '
'
Returns list of divergences
Remediation:
Non-critical drift:
Accept and update desired state, or manually fix the component to match desired state.
Critical drift:
Re-apply configuration via SDDC Manager to force convergence. Involves brief service interruption.
Prevention:
Enforce change control; all infrastructure changes must go through SDDC Manager, not manual edits.
Key Takeaways
- Pre-check failures: certificate expiry, DNS/NTP misconfig, and license incompatibility are most common; fix all before retrying upgrade.
- Bundle download: verify network access to download.vmware.com, check NFS storage capacity (100GB+ needed), enable long timeouts for large bundles.
- Drift remediation: compare desired state (SDDC Manager) vs actual state (hosts/vCenter); re-apply config if critical drift detected.
Performance Troubleshooting#
esxtop: CPU Metrics Deep Dive
esxtop is the primary performance diagnostic tool on ESXi. It samples CPU, memory, network, and storage metrics in real-time.
Running esxtop:
ssh root@esxi-host
esxtop
Interactive mode; press 'c' for CPU, 'm' for memory, 'n' for network, 'd' for disk
Press '?' for help
To capture data for later analysis:
esxtop -b -c '/tmp/esxtop.csv'
Batch mode, writes to CSV (can be analyzed in Excel)
CPU Metrics Explained:
CPU Tab Columns:
%USED = CPU utilization as seen by VM (includes ready + co-stop)
%RDY = Ready time: VM waiting for physical CPU, not running.
High %RDY (>10%) = CPU overallocation (too many vCPUs for pCPUs)
Action: Reduce vCPU count, or reduce vCPU overcommit ratio
%CSTP = Co-Stop: VM waiting for all vCPUs to align for parallel execution.
High %CSTP (>5%) = vCPU topology mismatch (e.g., VM has 4 vCPU but only 2-core socket)
Action: Reduce vCPU count to match socket count, or pin to NUMA nodes
%MLMTD = Limit: VM throttled due to CPU limit set
High %MLMTD = configured CPU limit is too low
Action: Increase CPU limit (vCenter → VM settings → CPU limit)
- %SYS = System (hypervisor kernel) CPU utilization
- High %SYS (>10%) = kernel overhead (DRS, vSAN, NSX processing)
- Action: Verify DRS/vSAN/NSX not consuming excessive cycles; check logs
MHZALLOC = Allocated MHz (sum of all VM vCPU MHz)
MHZUSED = Actual MHz consumed
Example diagnosis:
CPU cores: 32, MHz: 3000 = 96,000 MHz capacity
- MHZALLOC: 120,000 MHz (125% overallocation)
- VMs report low performance, high %RDY
- Fix: Power off 1 VM or reduce vCPU allocation
Memory Metrics
Memory Troubleshooting with esxtop:
Memory Tab Columns:
- MCTLSZ = Memory controlled by balloon driver (VM thinks it has less RAM)
- High MCTLSZ = host memory pressure; balloon reclaiming memory from VMs
- Action: Reduce VM memory allocation, add more RAM to host, or migrate VM
SWCUR = Current swapped pages (VM pages written to disk)
Any non-zero SWCUR = SEVERE performance impact
Action: Immediately power off non-critical VM, add more host RAM
Note: Storage I/O for swapping is ~100x slower than physical RAM
ZIP = Compressed pages (memory compression)
High ZIP (>5% of allocated) = useful, helps fit more VMs in RAM
Some compression overhead; if very high, consider dedup/compression
MEMCTL = Memory allocation target
VM balloon driver targets this; if ballooning, VM swap increases
Action: Verify VM has sufficient reserved memory
Example diagnosis:
Host RAM: 512GB, MCTLSZ: 50GB (VM memory ballooning)
SWCUR: 2GB (2GB is swapped to disk - CRITICAL)
Solution: Migrate VMs to another host, or power off non-critical VM
Prevention: Reserve 10% of host RAM for hypervisor, only allocate 90% to VMs
Network & Storage Metrics
Network Metrics (Press 'n' in esxtop):
Network Tab Columns:
- droppedRx = Packets dropped on receive (physical NIC to kernel)
- High value = NIC RX ring buffer overflow (network saturation)
- Action: Check physical NIC utilization (ethtool), enable TSO/LRO offload
- droppedTx = Packets dropped on transmit (kernel to physical NIC)
- High value = TX queue overflow, or network congestion
- Action: Check uplink bandwidth, enable jumbo frames (MTU 9000)
- retrans = TCP retransmissions
- High retrans = network quality issues (latency, jitter, loss)
- Action: Check switch/cable for errors (SNMP counters on switch)
- For vSAN: retrans impacts replication; fix immediately if high
Storage Metrics (Press 'd' in esxtop):
Disk Tab Columns:
DAVG = Average I/O latency to device (milliseconds)
Baseline: vSAN <5ms, NFS <10ms, FC/iSCSI <5ms
If DAVG > baseline: storage is slow (array congestion, bad disk, network issue)
Action: Check storage array health, verify IOPS/throughput not maxed
- KAVG = Average I/O latency at kernel level (ESXi scheduler latency)
- Should be <1ms; if high, indicates kernel contention
- Action: Check CPU %RDY (if high, CPU contention causes I/O stalls)
- GAVG = Guest latency (DAVG - KAVG)
- This is what the VM perceives
- High GAVG = storage device is slow
Outstanding I/O = Number of I/Os in flight
High value (>20 for HDD, >100 for SSD) = storage queue is full
Action: Reduce workload, or add storage capacity
Example diagnosis:
- DAVG: 50ms (should be 5ms for vSAN)
- Outstanding I/Os: 200
- Storage array showing 95% utilization
- Solution: Migrate data to less-congested datastore, or upgrade array
vRealize Operations & VCF Operations Metrics
vRealize Operations (vROps):
Centralized monitoring across VCF.
Metrics available:
CPU, memory, network, storage, VM latency, application performance.
Anoma
Key Takeaways
- esxtop CPU metrics: %RDY > 10% indicates overallocation; %CSTP > 5% suggests vCPU topology mismatch. Drill into VM-level CPU contention.
- Memory ballooning: non-zero SWCUR (swapped pages) is emergency condition, immediately power off VMs to restore performance.
- Storage latency baseline: vSAN <5ms, NFS <10ms; DAVG spikes = array congestion, check outstanding I/O count and reduce workload.
Support Engineering Fundamentals#
VMware Support Case Workflow & SLAs
Severity Definitions (Critical to Low):
- Severity
- Definition
- Example
- Response SLA (Premier)
- Workaround Available
- Severity 1 (Critical)
- Total service loss; multiple systems down; customer business halted
- Entire vSAN cluster offline; vCenter unreachable; all VMs lost network
- 15 minutes (24/7)
- No (urgent escalation required)
- Severity 2 (High)
- Significant function impaired; partial service loss; performance degradation >50%
- vSAN single host down; vMotion broken; DRS not working
- 1 hour (24/7)
- Partial (e.g., manual DRS workaround)
- Severity 3 (Medium)
- Minor function impaired; isolated issue; workaround available
- NSX segment slow to create; vCenter event log filling (cleanup available)
- 4 hours (business hours)
- Yes (clear workaround)
- Severity 4 (Low)
- Cosmetic issue; feature request; general info request
- "How do I configure DRS affinity rules?"; "Dashboard page slow to load"
- 24 hours (business hours)
- N/A (informational)
Support Tiers & Offerings
Premier Support:
24/7, 15 min – 4 hour SLA depending on severity, phone access to L2/L3 engineers, quarterly health checks.
Production Support:
Business hours (8–5 M–F), 4 hour – 24 hour SLA, portal & email only (no phone).
Basic Support:
Business hours, no SLA guarantees, portal/email, limited to general questions (not production debugging).
Case Lifecycle & Engagement
Customer submits case → Broadcom Support Portal ↓ Tier 1 (Front-line triage): - Collect info (environment details, error logs, reproduce steps) - Run diagnostics (KB matching, basic troubleshooting) - If simple → resolve; if complex → escalate
↓ (if escalated)
Tier 2 (Senior engineer):
- Deep troubleshooting (log analysis, performance profiling)
- May request SoS bundle, GuestInfo dumps
- Engage with product engineering if needed
- Typical resolution: 2–5 days for Sev 2
↓ (if still unresolved)
Tier 3 / Engineering escalation:
- Product engineering team (VMware internal)
- Potential bug fix, workaround, or design review
- May require patch release (if critical)
- Timeline: 1–4 weeks for bug fix
Case resolution:
- Solution documented, case closed
- Customer feedback survey, NPS (Net Promoter Score)
Broadcom Support Portal vs. Legacy vmware.com
After 2023 acquisition, VMware moved to Broadcom support portal:
Broadcom portal:
https://support.broadcom.com (new standard).
Legacy vmware.com:
Still accessible for older contracts; being phased out.
Key differences:
Unified support for all Broadcom products (VMware + networking); single sign-on (SSO); AI-powered KB search; case collaboration tools.
TSE (Technical Services Engineer) Engagement Pattern
For high-value customers or extended deployments:
Assigned TSE:
Dedicated engineer assigned for the engagement duration.
Proactive monitoring:
TSE monitors environment health, anticipates issues, offers optimization recommendations.
Scheduled check-ins:
Weekly/monthly syncs to review logs, alert trends, capacity planning.
Escalation path:
Direct phone/email access (vs. ticket queue); faster response.
Escalation Triggers & L2/L3 Handoff
Reading KB Articles for Build/Version Applicability
KB article structure (Broadcom/VMware):
Title:
Concise, includes key terms (e.g., "vCenter password change fails on VCF 9.0").
Applicable Products/Versions:
CRITICAL section at top. Example: "Applies to: vSAN 7.0 U1–9.0.1; does NOT apply to vSAN 6.x or 9.0.2+".
Symptoms:
Observable behavior; reproduction steps.
Root Cause:
Technical explanation (bug, config issue, limitation).
Resolution:
Fix (upgrade/patch), workaround, or configuration change.
Affected Builds:
Specific version ranges (e.g., "vCenter 6.7 GA through 6.7 U3d"; "Fixed in 7.0 U2").
Tip:
Key Takeaways
- Support Exam Insight:Panelists (or cert exams) may ask "You have case SID-12345 for customer Y; their environment is vSAN 9.0.1, vCenter 9.0 GA. Which KB articles are applicable?" You must recognize version-specific fixes and cross-reference documentation. Know the common breakfix versions (e.g., vS
Log Collection Deep Dive (SoS Utility)#
SoS Utility: Complete Options & Usage
The SoS (Supportability and Serviceability) utility collects diagnostics from vSAN, vCenter, NSX and VCF infrastructure. It runs on the SDDC MANAGER appliance (also available on Cloud Builder) — NOT on an ESXi host. SSH in as the 'vcf' user; Fix-It-Up options require su to root and running ./sos from /opt/vmware/sddc-support.
Basic syntax:
sudo /opt/vmware/sddc-support/sos --help
Core options:
- --vc-logs Collect vCenter logs
- --esx-logs Collect ESXi vmkernel, hostd logs
- --psc-logs Collect Platform Services Controller (vCenter embedded)
- --sddc-manager-logs Collect SDDC Manager logs (VCF-specific)
- --nsx-logs Collect NSX Manager/Controller/Edge logs
- --vsan-logs Collect vSAN cluster diagnostics, cmmds database
- --health-check Run vSAN health checks (redundancy, disk state, capacity)
- --all Collect everything (useful for critical issues)
Example command:
sos.py -f /var/tmp/sos-$(hostname)-$(date +%Y%m%d).tar.gz \
--vc-logs --esx-logs --vsan-logs --health-check
Log Bundle Size Management
SoS bundles can grow large (500 MB – 2 GB depending on cluster size and debug detail level).
Size estimation:
- vSAN logs (10-node cluster): ~100–200 MB
- vCenter logs (7 days): ~50–100 MB
- NSX Manager/Edge logs: ~50–100 MB
- SDDC Manager logs: ~30–50 MB
- Total typical bundle: ~250–450 MB
To limit bundle size:
- --max-days 3 Collect only last 3 days (instead of 7)
- --compress gzip Use gzip compression (default)
- --exclude-vmdk-logs Skip verbose datastore logs (rarely needed)
Log Rotation & Retention
ESXi vmkernel.log rotation:
Default: rotated daily; 10 rotations kept (10 days of logs).
Location: /var/log/vmkernel.log, vmkernel.log.1, vmkernel.log.2, etc.
Config:
- esxcli system syslog config get
- Allowed: /var/log, /scratch/log (on ramdisk), /vmfs/volumes/datastore/logs
- Size limit: configurable per log file
- vpxd.log (vCenter):
- /var/log/vmware/vpxd/vpxd.log
- Rotation: daily + size-based; 30 days retained by default
- Monitor with: du -sh /var/log/vmware/
Syslog Forwarding Architecture
Configure ESXi, vCenter, NSX to forward logs to central SIEM for analysis:
Option 1: vSphere Cluster with Log Intelligence (VCF Ops)
ESXi/vCenter/NSX → syslog → VCF Ops Log Intelligence (in-cluster)
Queries: search logs by keyword, filter by severity, timeline analysis
Retention: 30 days (configurable)
Option 2: External Splunk
esxcli system syslog config set --loghost siem.corp.local:514
ESXi → syslog (UDP/TCP) → Splunk TCP input → indexed
Advantage: enterprise SIEM; longer retention; cross-system correlation
- Option 3: Elasticsearch/Logstash/Kibana (ELK)
- Similar to Splunk; OSS alternative
- Configure syslog forwarder on vCenter/NSX to send to logstash input
Option 4: QRadar (IBM)
Syslog + NSX event API → QRadar
Enhanced insights for security use case
Log Redaction for PII (Personally Identifiable Information)
Before sending logs to external support or SIEM:
Redact (before sending to support):
- Customer names: esxi-1.acme.com → esxi-1.REDACTED.com - IP addresses: 192.168.1.100 → 192.168.REDACTED.100 (optional) - Personal email: john.smith@acme.com → [email protected]
- Credit card numbers: 1234-5678-XXXX-XXXX (last 4 only)
Tools:
- Manual: grep -o 'EMAIL_PATTERN\|IP_PATTERN' logfile | sed 's/\..*@/@REDACTED/'
- Automated: Splunk data anonymization, ELK ingest filter
VMware SoS default:
- Does NOT redact by default (be careful when sharing)
- Use --redact-pii flag if available (newer versions)
Log File Location Reference by Component
- Component
- Log File
- Use Case
- Key Fields to Check
- ESXi vmkernel
- /var/log/vmkernel.log
- PSOD, kernel panic, hardware errors, memory correctable errors (CXE), NMI
- BugCheck code, backtrace, module name (vsan, nsx, etc.)
- ESXi hostd
- /var/log/hostd.log
- VM lifecycle (power-on/off), device hotplug, virtsched, DRS decisions
- "Error", "Died", "migration", "schedule"
- ESXi vpxa
- /var/log/vpxa.log
- vCenter agent communication, config drift, HA heartbeat
- "Lost connection", "Heartbeat", "Quarantine"
- vCenter vpxd
- /var/log/vmware/vpxd/vpxd.log
- Inventory operations, permissions, tasks, alarms, event storm
- "Exception", "Failed task", "Event flood"
- NSX Manager
- /var/log/vmware/nsx/manager.log
- DFW rule push, segment creation, federation sync
- "Replication lag", "Failed sync", "Policy push"
- vSAN cmmds
- /var/log/vmware/vsan/cmmds.log
- vSAN cluster membership, component state changes, rebuilds
- "Disk missing", "Resync", "Component degraded"
- dmesg (kernel buffer)
- /var/log/dmesg (or
dmesgcommand) - Early boot errors, driver issues, hardware detection
- device failures, CPU/memory errors during boot
Symptom-to-Log-File Quick Reference
Symptom: "VMs won't power on after host failure"
Check: hostd.log (virtsched pressure), vmkernel.log (memory pressure), vpxa.log (HA decision)
Command: grep -i "error\|failed\|schedule" /var/log/hostd.log | tail -20
Symptom: "vSAN object missing component, slow resync"
Check: cmmds.log, esxcli vsan command output
Command: esxcli vsan storage list | grep -i "degrade
Key Takeaways
- SoS bundle selection: --vsan-logs for vSAN issues, --nsx-logs for networking, --sddc-manager-logs for lifecycle; combine as needed.
- Log file locations: /var/log/vmkernel on ESXi, /var/log/vmware/vpxd on vCenter, /var/log/vmware/nsx-manager on NSX; grep strategically.
- Log bundle size: typical 500MB-2GB; compress with gzip for support upload, exclude verbose logs (--exclude-vmdk-logs) if size exceeds storage.
ESXi & Hardware Troubleshooting#
PSOD (Purple Screen of Death) Decoding
A PSOD is an ESXi kernel panic. The screen displays a stack trace (backtrace) and BugCheck code that pinpoints the root cause.
Common BugCheck codes:
- BugCheck Code
- Meaning
- Common Causes
- Mitigation
- 0x1 (Page Fault)
- Kernel accessed unmapped memory
- Driver bug, memory corruption, bad RAM
- Update driver/firmware, run memtest86
- 0x11 (Heartbeat Timeout)
- ESXi lost coredump partition responsiveness (deadlock)
- Hung I/O (SAN timeout), BIOS hang, infinite loop in kernel
- Check SAN/network connectivity, BIOS/firmware updates
- 0x1D (NMI – Non-Maskable Interrupt)
- Hardware triggered panic (IPMI watchdog, external signal)
- Hardware error (CPU, chipset), IPMI misconfiguration
- Check event logs (IPMI, BMC), update BIOS
- 0x21 (Memory Corruption)
- Detected memory corruption (ECC error, bit flip)
- Faulty DIMM, bit flip (soft error), unstable power
- Replace RAM DIMM, check power supply, monitor ECC errors
- 0x32 (vSAN Bug)
- vSAN driver issue
- vSAN driver bug (specific version), HCL violation, disk failure
- Upgrade vSAN, check HCL, check disk SMART status
Reading a PSOD backtrace:
Example PSOD screen:
BugCheck: 0x11
Backtrace:
0xffffffff810xxxxx: kernel_panic+0x100
- 0xffffffff8107xxxx: esxcli_vsan_command_handler+0x200
- 0xffffffff812xxxxx: vsan_io_handler+0x50
- 0xffffffff810xxxxx: generic_dispatch+0x80
Interpretation:
- BugCheck 0x11 = Heartbeat timeout
- vsan_io_handler is the top frame → hang is in vSAN I/O
- Likely cause: SAN array unresponsive, I/O queue overflow, network partition
Action:
- Check SAN connectivity (ping SAN IP)
- Monitor vmkernel.log for "Device offline" or "LUN timeout" before PSOD
- Review SAN array logs for I/O errors/latency spikes
- If vSAN cluster: check vSAN resync queue, rebuild rate limits
Hardware Compatibility & VCG Validation
Every ESXi host must be on the VMware Compatibility Guide (VCG) for that hardware model:
Check VCG at: https://compatibilityguide.broadcom.com/
Input: Server model (e.g., Dell PowerEdge R6715)
Output: Supported ESXi versions, required firmware/driver versions
Example valid entry:
Dell R6715 + ESXi 9.0 + Broadcom RAID adapter driver 12.16.10.0
= Certified; fully supported
Invalid entry:
Dell R6715 + ESXi 9.0 + Broadcom driver 11.xxx (old)
= Not certified; no support
Validation check (on ESXi):
esxcli system hardware info get | grep "System Model"
→ Record model → Cross-check against VCG
Firmware & Driver Stack Validation
Critical firmware components to monitor:
BIOS:
Current: dmidecode | grep -i "bios version"
Check: Match against VCG recommended version
Risk: Old BIOS → memory controller issues, CPU power management bugs
BMC / iLO / iDRAC firmware:
Current: ilorest get /rest/v1/Systems/1/ | grep FirmwareVersion
Recommendation: Keep within 2 major versions of BIOS
Risk: Out-of-date BMC → IPMI command failures, NMI watchdog misconfiguration
NIC firmware:
esxcli hardware pci get | grep -i "ethernet\|nic"
→ Get PCI IDs; check driver version esxcli software vib list | grep -i "network" Risk: Old driver → packet loss, vMotion failures
RAID adapter firmware:
storcli64 /c0 show all | grep -i version
Recommendation: Match VCG for RAID version
Risk: Outdated firmware → controller hangs, data corruption
QuickBoot & Coredump Configuration
QuickBoot reduces reboot time from 5–10 minutes to 30 seconds by skipping BIOS POST. However, it skips memory validation:
Enable QuickBoot:
- SSH into ESXi host
- esxcli system settings advanced set -o /UserVars/QuickBootEnabled -i 1
- Reboot
Disable QuickBoot (if memory errors suspected):
esxcli system settings advanced set -o /UserVars/QuickBootEnabled -i 0
Coredump configuration:
esxcli system coredump partition get
→ Shows current coredump partition (usually /dev/sdX:1)
If no coredump configured:
esxcli system coredump partition set -p /dev/sdX:1
(Must be on local storage, min 110 MB)
Coredump file location after PSOD:
/scratch/coredumps/ (if VMFS datastore mounted)
OR entire memory to disk (if partition full, memory may be lost)
Recommendation: Always configure coredump on local storage for PSOD debugging
ESXCLI Deep Troubleshooting Commands
Maintenance mode (enter/exit):
Enter maintenance mode (evacuates VMs via vMotion):
esxcli system maintenanceMode set --enable true
Exit maintenance mode:
esxcli system maintenanceMode set --enable false
Check status:
esxcli system maintenanceMode get
Output: Enabled/Disabled
Storage path analysis (SAN connectivity):
List storage paths:
esxcli storage nmp satp rule list
→ Shows multipathing policy per LUN
Check path state:
esxcli storage nmp path list
→ Displays each path (LUN path) and active/standby state
Rescan storage:
esxcli storage core adapter rescan --adapter vmhba2
→ Re-discovers LUNs attached to RAID adapter
Monitor path failures:
grep -i "path lost\|device offline" /var/log/vmkernel.log
Net
Key Takeaways
- VCG (VMware Compatibility Guide): always verify server model + firmware versions match approved combinations before deployment.
- Coredump configuration: network coredump (ISCSI/NFS) preferred; local coredump requires 1GB+ partition and is useful only if host reboots.
- Firmware updates: RAID, NIC, SSD firmware must follow VCG sequence; disable QuickBoot before BIOS updates to validate memory POST.
vCenter / VCSA Troubleshooting#
VCSA Filesystem Layout
VCSA (vCenter Server Appliance) is a Linux VM with specific mount points:
Mount point layout:
- /storage/db PostgreSQL database, vCenter Inventory
- /storage/log Log files (vpxd, vmdir, sso, etc.)
- /storage/seat License seat database
- /root System root (OS)
- /tmp Temporary files (16 GB max on 32 GB VCSA instance)
Disk size allocation (typical 32 GB VCSA):
- OS (/) : 12 GB
- /storage/db : 10 GB
- /storage/log : 5 GB
- /storage/seat : 2 GB
- /tmp : 3 GB
Monitoring disk:
SSH into VCSA → df -h
Alert if /storage/db > 80% (database growing rapidly)
Alert if /storage/log > 90% (log rotation not working)
Expand disk:
- Shut down VCSA
- Add disk space (hypervisor) to VCSA VM
- Boot VCSA; resize filesystem: resize2fs /dev/sda3
- Verify: df -h
VCSA Service Architecture
Core services running on VCSA:
- Service
- Process Name
- Function
- Log File
- vpxd
- vpxd (main process)
- Core vCenter server; inventory, tasks, events
- /var/log/vmware/vpxd/vpxd.log
- vmdir (IDM)
- vmdir
- Identity Management; stores SSO users, groups
- /var/log/vmware/vmdir/syslog
- sts-idm
- sts-idm
- Security Token Service; OAuth/SAML tokens
- /var/log/vmware/sts-idm/sts.log
- vsphere-ui
- vsphere-ui (node.js web UI)
- Web interface (https://vcenter-ip/ui)
- /var/log/vmware/vsphere-ui/
- eam
- eam
- Extensibility and Manageability Agent
- /var/log/vmware/eam/
- vapi
- vapi-endpoint
- REST/vSphere API endpoint
- /var/log/vmware/vapi/vapi.log
Service control commands:
Check status:
service-control --status --all
Output: [RUNNING], [STOPPED], [UNKNOWN]
Restart single service (e.g., vpxd):
service-control --restart vsphere-client
service-control --restart vpxd
Stop all services (for maintenance):
service-control --stop --all
Start all services:
service-control --start --all
vmafd & SSO Troubleshooting
vmafd: VMware Authentication Framework Daemon
Manages VCSA certificates (VMCA).
Responsible for vmdir (LDAP) connectivity.
If vmafd fails → vmdir replication breaks → SSO fails.
Check certificate status:
Get certificate status:
vmafd-cli get-cert-status --server-name localhost
→ Shows cert chain, expiry dates, CA cert fingerprint
Get machine certificate expiry:
/usr/lib/vmware-vmafd/bin/dir-cli instance list
→ Lists all vCenter instances in domain
Certificate renewal (before expiry):
/usr/lib/vmware-vmafd/bin/vecs-cli entry list --store TRUSTED_ROOTS
→ If cert expires in <30 days, renew:
/usr/lib/vmware-vmafd/bin/dir-cli instance new --server VCENTER_FQDN
Monitor certificate age:
openssl x509 -in /etc/vmware-vpx/ssl/rui.crt -noout -dates
→ Shows notBefore, notAfter
vmdir replication issues:
Check vmdir service:
systemctl status vmdir
→ Should be [active (running)]
Check vmdir replication (if HA Linked Mode):
ldapsearch -h localhost -p 389 -b "cn=config" | grep -i "serverID\|partner"
→ Shows replication partners
Monitor replication lag:
vmdir-cli database check
→ Reports any replication issues
Fix replication (if broken):
- Restart vmdir: service-control --restart vmdir
- Monitor logs: tail -f /var/log/vmware/vmdir/syslog | grep -i "repl\|error"
PostgreSQL Basics & VCDB Management
vCenter Database (VCDB) is PostgreSQL:
Connect to VCDB:
/opt/vmware/vpostgres/bin/psql -U postgres -h localhost vcdb
(Requires VCSA root access via SSH)
Check database size:
\l+ vcdb
→ Shows vcdb size (e.g., "2045 MB")
Query event table size:
SELECT tablename, pg_size_pretty(pg_total_relation_size(schemaname||'.'||tablename))
FROM pg_tables WHERE tablename = 'vpx_event' ORDER BY pg_total_relation_size DESC;
→ Shows vpx_event table size
If vpx_event too large (>500 MB):
Cleanup old events:
UPDATE vpx_event SET created_time = created_time - INTERVAL '7 days' WHERE created_time < NOW() - INTERVAL '90 days';
VACUUM ANALYZE vpx_event;
OR use vpxd-ctl cleanup:/usr/lib/vmware-vpxd/bin/vpxd-ctl cleanuptasks
vpxd-ctl commands:
Check vpxd database connections:
/usr/lib/vmware-vpxd/bin/vpxd-ctl cleanupevents
Query events by type:
/usr/lib/vmware-vpxd/bin/vpxd-ctl querytask --filter "task_type LIKE 'VmPowerOn'"
→ Lists PowerOn tasks
HA Quorum & Admission Control
vCenter HA (if deployed in HA cluster):
vCenter HA requires 3+ VCSA instances in a cluster (primary + 2 passive replicas).
Quorum algorithm: majority (2 of 3) must be healthy for cluster to function.
If 1 replica down: cluster still operational.
If 2 replicas down: cluster loses quorum → read-only mode, no new tasks.
Monitor HA status:
/usr/lib/vmware-common-javautil/bin/jcli ha status
→ Shows [Active], [Passive], quorum state
Key Takeaways
- VCSA filesystem: /storage/db growth indicates event database bloat; trim events >90 days old to reclaim space, monitor free space >= 20%.
- vmafd certificate expiry: critical for vmdir replication; monitor certificates at 60+ days before expiry, renew via certificate-manager tool.
- SSO domain joins: external AD requires domain controller reachability (ping, LDAP 389); import root CA certificate for validation.
vSAN Troubleshooting at Scale#
Object-Level Analysis: cmmds-tool & vsan-dump
vSAN objects:
Each VM disk is a vSAN object; objects are composed of components (replicas, parity).
List vSAN objects:
esxcli vsan cluster get
→ Shows cluster info (UUID, member count, health)
Get object stats:
esxcli vsan object list
→ Lists all objects with UUID, size, policy, health
Check specific object:
esxcli vsan object get --objid <uuid>
→ Shows object components, state (healthy, degraded, etc.)
Deep dive (cmmds-tool):
/usr/lib/vmware-cmmds/bin/cmmds-tool -i /etc/vmware/cmmds/cmmds-config
→ Interactive tool; list all CMMDS objects
cmmds-tool> find --type LOM_OBJECT --all
→ Lists all logical object mappings (vSAN objects)
vsan-dump (health analysis):
/usr/lib/vmware-vsan/bin/vsan-dump -m full
→ Comprehensive dump; useful for support escalation
DOM Layer Troubleshooting
DOM: Distributed Object Manager — vSAN's data coherency layer.
Manages object placement, replication, rebuilds.
If DOM slow → resync stalls, object state changes delayed.
Check DOM health:
esxcli vsan observer querydomstatus
→ Shows DOM load, queue depth, resync activity
Monitor DOM activity:
/usr/lib/vmware-vsanclibs/bin/observer-cli getClusterDomStatus
→ Real-time DOM queue status
If DOM slow (high queue depth):
- Check: esxcli vsan resync get
→ Shows rebuild rate, resync tasks
2. If rebuild too slow: esxcli vsan resync set --rebuild-rate 1000
→ Increase from default 100 MB/s to 1000 MB/s
3. Monitor: esxcli vsan observer getclustertasks
→ Watch task completion rateLSOM Disk State Transitions & Diagnosis
LSOM: Log-Structured Object Manager — vSAN's I/O layer.
Disk state machine:
- State
- Meaning
- Next State
- Action if Stuck
- Healthy
- Disk operational, I/O OK
- Absent (on failure)
- Monitor
- Absent
- Disk not visible (unplugged, offline)
- Healthy (re-seated), Degraded (data loss)
- Check cable; reboot; replace disk if bad
- Degraded
- Disk I/O errors; rebuild in progress
- Healthy (if rebuild succeeds), Dead (if fails)
- Monitor resync; speed up if tolerable
- Dead
- Disk permanently failed; rebuild failed or not attempted
- Replace disk
- Physical replacement; rescan; re-enter cluster
Disk state check:
List all disks and state:
esxcli vsan storage list
→ Shows disk UUID, type (cache/capacity), state, errors
Check disk errors (SMART):
esxcli storage core device list | grep -i "vmhba\|device"
→ Get device name (e.g., t10.ATA____Samsung_SSD_850___SERIAL)
esxcli storage core device smart get --device-name <name>
→ Shows SMART health, error count
Monitor rebuild status:
esxcli vsan observer resyncstatus
→ Shows resync rate, objects remaining, ETA
Resync Rate & Rebuild Performance Tuning
Resync math:
Rebuild time = (Total data to rebuild) / (Resync rate) / (Parallel rebuild threads)
Example:
Disk failure: 4 TB capacity disk lost
Policy: FTT2 (3 copies) → only 1/3 of data needs rebuild
- Data to rebuild: 4 TB / 3 = 1.3 TB
- Resync rate: default 100 MB/s
- Parallel threads: ~3 (typical)
Time = 1.3 TB / 100 MB/s / 3 = ~4.3 hours
To speed up: esxcli vsan resync set --rebuild-rate 500
New time = 1.3 TB / 500 MB/s / 3 = ~0.9 hours
Warning: Higher rebuild rate increases I/O load on cluster (can impact VMs)
Recommendation: During business hours, limit to 200–300 MB/s; overnight, 500+ MB/s
Stretched Cluster Partition Scenarios & Recovery
Split-brain scenario: Network partition between sites.
Stretched cluster (6 nodes + witness):
- Site A: 3 nodes + witness
- Site B: 3 nodes
- Network partition between sites (link down)
vSAN quorum logic:
Site A: 3 nodes + 1 witness = 4 quorum votes → has quorum; continues Site B: 3 nodes (no witness) → no quorum; fenced (read-only mode)
Site A continues serving VMs; Site B blocks writes.
When network heals: Site B re-syncs from Site A (data reconciliation).
Recovery procedure:
1. Detect partition: esxcli vsan cluster get → shows member state
- If Site B fenced (no quorum):
- a. Verify Site A is healthy and has all data
- b. Restore network link (check switch, fiber status)
- c. Site B will auto-recover (CMMDS full resync)
- Monitor resync: esxcli vsan observer resyncstatus
(May take hours for full data sync)
- Validate: All objects must achieve "OK" health status
vSAN Observer & Performance Service
vSAN Observer (legacy UI tool, still useful):
Access via vCenter web UI:
Monitor → vSAN → Performance Service
Real-time metrics:
- Cluster IOPS, throughput (MB/s), latency
- Per-disk IOPS, outstanding I/Os
- Resync queue depth, rebuild rate
- Component state summary
Interpretation:
IOPS spike → application load increase or rebuild in progress Latency spike >50ms → storage congestion (add disks) or network latency Queue depth >20 → potential I/O bottleneck (rebuild too slow)
Historical data:
Performance Service retains 7 days (default); can be configured to 30 days
Useful for capacity planning, trend analysis
ESXi vSAN Debug Commands
Advan
Key Takeaways
- DOM (Distributed Object Manager) load: monitor observer queue depth; if > 20, rebuild rate too slow; increase rebuild rate temporarily.
- CMMDS (Cluster Monitoring): gossip protocol propagates state within 1-2 seconds; partition scenarios require quorum to continue I/O.
- Stretched cluster witness: witness node must be accessible to all sites via network; witness failure triggers minority partition fencing.
NSX Troubleshooting Deep Dive#
Control Plane Diagnostics: Manager & Controller Status
NSX Manager (management plane):
SSH into NSX Manager:
Default: https://<nsx-mgr-ip>
CLI: ssh admin@<nsx-mgr-ip>
Check manager cluster status:
get managers
→ Shows primary, secondary, standby status; heartbeat age
Check version:
get version
→ NSX version (e.g., 4.2.1.0)
Check control-plane (CCP) connectivity from a transport node:
get controllers # on ESXi use: nsxcli -c get controllers
→ Shows the NSX Manager nodes hosting CCP and session status (there is no separate controller appliance since NSX-T 2.4)
Health check:
get cluster status
→ Reports manager replication, controller connectivity, edge status
Controller cluster status:
On NSX Manager, check controller health:
get controllers
→ Output: Controller IP, status [up/down], role [active/standby]
If controller down:
- SSH controller: ssh admin@controller-ip
- Check service: systemctl status nsxcontroller
- If not running: systemctl start nsxcontroller
- Monitor logs: tail -f /var/log/nsxcontroller/nsxcontroller.log
Join new controller to cluster:
On manager:
add controllers <controller-ip>
Typical controller cluster: 1–3 controllers (1 for small, 3 for large)
Data Plane Diagnostics: Logical Switches & Forwarding
Check logical segments (via nsxcli on edge):
SSH into NSX Edge (data plane):
ssh admin@<edge-ip>
List logical switches (segments):
get logical-switches
→ Shows segment UUID, VLAN ID, vxlan ID, status
List interfaces:
get interfaces
→ Shows management, TEP, uplink interfaces
Check uplink status:
get interfaces | grep -i uplink
→ Verify uplink linked status (up/down)
Monitor packet forwarding:
get bfd-config
→ BFD heartbeat detection (for failover convergence)
Check firewall rules (DFW):
get firewall rules
→ Lists all rules, hit counts, connection tracking
Traceflow for DFW debugging:
Use NSX Manager UI:
Troubleshooting → Traceflow
Specify:
- Source VM: web-server-01
- Dest VM: db-server-01
- Protocol: TCP, Port: 5432
Result:
Path: web-01 → DFW (allow/deny) → VXLAN overlay → ECMP → db-01
Shows exactly which DFW rule matched (or if denied)
If traffic blocked by DFW:
- Check DFW rules: which rule denied it?
- Verify source/dest tags match rule (may be tag mismatch)
- Test with temporary allow-all rule
- Monitor: get firewall rules (shows hit count; verify rule is being evaluated)
Packet Capture on Edge (pktcap-uw)
Deep packet inspection on NSX edge:
Start packet capture (on edge):
pktcap-uw --interface eth0 --file /tmp/edge-capture.pcap
(Ctrl+C to stop)
Or capture with filter:
pktcap-uw --interface eth0 --file /tmp/web.pcap \
'net 10.0.1.0/24 and tcp port 443'
→ Only capture HTTPS traffic from web subnet
Transmit pcap to external for Wireshark analysis:
scp /tmp/edge-capture.pcap analyst@workstation:/tmp/
Analyze on Wireshark:
File → Open → edge-capture.pcap → View packet header, payload, stream analysis → Debug: verify TCP handshake, application data, retransmits
Common captures:
- No SYN-ACK from server → server down or filtering - RST packets → connection forcefully closed (firewall rule?) - Retransmits > 2% → packet loss or congestion
Edge Service Diagnostics: BGP & Routing
Tier-0 BGP health:
Check BGP neighbor status:
On edge: get bgp neighbor
→ Shows neighbor AS, IP, state [Established/Down/Connect]
If neighbor down:
- Check uplink connectivity: ping <neighbor-ip>
- Verify BGP timers: get bgp timers (should be 3s keepalive, 9s holdtime)
- Check AS numbers: must match if iBGP, must differ if eBGP
- Review logs: tail -f /var/log/quagga/bgpd.log
Redistribute routes (OSPF ↔ BGP):
If routes not appearing in routing table:
- Check route-map: get route-map | grep redistribute
- Verify network statement: get network | grep <prefix>
- Test redistribution: clear bgp * soft (reload routes)
Monitor BGP session:
show bgp ipv4 unicast neighbors <neighbor-ip>
→ Detailed session stats, prefixes learned/advertised
Tier-1 routing table:
On T1 service router (edge or ESXi T1 instance):
get routing-table
→ Shows all routes (connected, static, dynamic)
If route missing:
- Verify connectivity to T0: ping t0-ip
- Check T1 advertise settings: (should advertise T1 subnets to T0)
3. Test re-advertise: Firewall → Network Services → Route advertisement (toggle off/on)
Monitor latency (if routing slow):
ping -c 5 <remote-tier>
→ If latency spiky, check controller load (get manager status)
TEP Tunnel State & vxlan Diagnostics
TEP: Tunnel Endpoint — each ESXi/Edge has a TEP IP for vxlan overlay.
Check TEP status (on ESXi):
nsxcli -c get vtep
→ Shows TEP IP, MAC, role [vtep/edge], status
Check tunnel status:
nsxcli -c get tunnel
→ Lists all tunnel connections, state [Up/Down], latency
If tunnel down:
- Verify TEP IP reachable: vmkping -I vmk-nsx <remote-tep-ip>
- Check MTU: must be ≥1600 (Geneve minimum; 1700 recommended and the NSX default Global TEP MTU)
vmkping -I vmk-nsx -s 1572 -d <remote-tep-ip> # -d = do not fragment
(1572 payload + 28 ICMP/IP headers = 1600 on the wire)
- Review logs:
Key Takeaways
- Manager cluster health: all 3 managers must be running and replicating; 1 manager down = degraded, 2+ down = quorum loss.
- BGP neighbor state: verify peer IP reachable, TCP/179 open, AS numbers correct; check route-map for redistribute policy.
- Packet capture (pktcap-uw) on edge: filter by source/dest IP, protocol; export pcap for Wireshark analysis to debug DFW or routing.
VCF Upgrade & vSphere-to-VCF Conversion (Objectives 5.2–5.3)#
VCF 5.x to 9.0 Upgrade Path (Obj 5.2)
VCF upgrades follow a strict sequence orchestrated by SDDC Manager's Lifecycle Manager (LCM):
Upgrade Sequence (per Broadcom KB 390634 for VCF 9.0):
- SDDC Manager (must be upgraded first)
- VCF Operations + fleet management appliance
- VCF Automation, VCF Operations for Networks
- VADP-based backup, vReplication, SRM, AVI
- HCX
- NSX Manager cluster (3-node rolling upgrade)
- vCenter Server (VCSA upgrade)
- VCF Identity Broker (new in 9.0) + Supervisor
- vSAN Witness Host
- ESXi hosts (rolling upgrade via vLCM) + NSX Finalize
11. VMware Tools → Virtual HW → vSAN on-disk format → vSAN File Service
Pre-Upgrade Checks:
├── Run SDDC Manager pre-check: SDDC Manager UI → Lifecycle Management → Pre-check ├── Verify bundle compatibility: ensure all component bundles match target VCF version ├── Check vSAN health: all health checks must pass before host upgrades ├── Verify NTP synchronization across all components ├── Backup SDDC Manager database, vCenter, NSX Manager configs ├── Ensure N+1 host capacity in each cluster for rolling upgrade └── Review VMware Interoperability Matrix for component version compatibility
Upgrade Bundles:
SDDC Manager downloads upgrade bundles from VMware depot (or offline depot for air-gapped):
├── Bundle types: SDDC Manager, vCenter, ESXi, NSX, vSAN witness ├── Bundle validation: SHA-256 checksum verified before staging ├── Staging: bundles staged to target hosts before maintenance window └── Rollback: SDDC Manager supports rollback for most components if upgrade fails
Common Upgrade Issues:
├── Pre-check failures: resolve ALL pre-check warnings before proceeding ├── Bundle download timeout: verify depot URL reachability and proxy settings ├── ESXi host stuck in maintenance mode: check vSAN data evacuation status ├── NSX upgrade compatibility: NSX version must be compatible with target vCenter version └── Post-upgrade vSAN resync: monitor resync completion before upgrading next host
vSphere-to-VCF 9.0 Conversion (Obj 5.2)
Converting standalone vSphere to VCF (brownfield adoption):
Conversion Process:
- Deploy SDDC Manager alongside existing vCenter
- Run VCF Import Tool to discover existing infrastructure
- Validate: ESXi version compatibility, networking (vDS required, no VSS), DNS/NTP
- Import existing vCenter + cluster as Management Domain
- Deploy NSX Manager (not present in standalone vSphere)
- Configure VCF networking overlays
- Validate all VCF services operational
Conversion Requirements:
├── ESXi hosts must be at minimum supported version for VCF 9.0 ├── vCenter must be compatible version (upgrade first if needed) ├── Networking must use vDS (migrate from VSS before conversion) ├── vSAN must be configured (VCF requires vSAN for management domain) ├── DNS forward/reverse resolution for all components ├── NTP synchronized across all hosts and management VMs └── Minimum 4 hosts for management domain (3 for vSAN + 1 for N+1)
VCF Workload Domains — Create, Scale, Import (Obj 5.3)
Creating Workload Domains:
SDDC Manager UI → Workload Domains → Create: ├── Type: VI (vSphere-based) or VCF (full stack) ├── Cluster configuration: host count, vSAN storage, networking ├── Network Pool: select or create network pool for vMotion, vSAN, TEP IP ranges ├── DNS/NTP: inherit from management domain or specify custom ├── NSX: auto-deployed per workload domain (NSX Edge optional based on topology) └── Validation: SDDC Manager runs full validation before deployment (~15-30 min)
Scaling Workload Domains:
Add Cluster:
├── SDDC Manager → Workload Domain → Add Cluster ├── Requires: commissioned hosts in SDDC Manager inventory ├── vSAN cluster: minimum 3 hosts per cluster ├── NSX transport nodes auto-configured for new cluster └── Network Pool assignment for new cluster's traffic
Remove Cluster:
├── Evacuate all VMs from cluster (DRS/manual migration) ├── SDDC Manager → Workload Domain → Cluster → Remove ├── Hosts returned to SDDC Manager free pool └── NSX transport node config removed from hosts
Add/Remove ESXi Hosts:
├── Commission new host: SDDC Manager → Hosts → Commission → provide credentials ├── Host validated: hardware compatibility, ESXi version, network config ├── Add to cluster: SDDC Manager → Workload Domain → Cluster → Add Host ├── vSAN disk auto-claim on new host ├── Remove host: enter maintenance mode → full vSAN data migration → decommission └── Decommissioned host returns to free pool or can be removed from inventory
Import Existing vSphere Environment:
├── VCF Import Tool scans existing vCenter and clusters ├── Validates: version compatibility, networking (vDS required), vSAN status ├── Converts existing cluster into VCF Workload Domain ├── Non-destructive: existing VMs continue running during import ├── Post-import: deploy NSX, configure VCF networking overlays └── Limitations: VSS-based environments must migrate to vDS first
VCF Operations — Observability Workbench & Log Assist (Obj 5.9)
The Observability Workbench is a VCF 9.0 feature in VCF Operations that provides correlated troubleshooting across metrics, logs, and events in a single view.
Observability Workbench:
├── Unified view: Select any VCF object (host, VM, cluster) → see metrics timeline + event timeline + log entries ├── Correlation: Click a metric anomaly → see concurrent logs and events → identify root cause ├── Cross-component: Correlates vCenter events, ESXi logs, vSAN health, NSX alerts in one timeline ├── Time range: Adjustable time window for investigation (last 1h, 6h, 24h, custom) └── Export: Export correlated data for support case attachment
Log Assist:
├── AI-powered log analysis: Select complex error messages → Log Assist explains in plain language ├── Root cause suggestion: Identifies likely cause based on log patterns ├── Remediation steps: Suggests specific fix actions based on error context ├── Log bundle generation: VCF Operations → Administration → Support → Create Support Bundle └── Upload: Attach generated bundle to Broadcom support case via Wolken portal
VCF Operations for Networks:
├── Network topology visualization: See NSX segments, gateways, DFW rules graphically ├── Flow analysis: Identify VM-to-VM traffic patterns, detect anomalies ├── Path analysis: Trace packet path from source to destination through NSX overlay ├── Intent verification: Compare intended network policy with actual forwarding state └── Troubleshooting: Identify dropped flows, misconfigured rules, routing loops
Key Takeaways
- VCF 9.0 upgrade sequence (per KB 390634): SDDC Manager → HCX → NSX → vCenter → VCF Identity Broker → vSAN Witness → ESXi → NSX Finalize → VMware Tools → Virtual HW → vSAN ODF → vSAN Fileservice. Always upgrade SDDC Manager first.
- Pre-upgrade pre-checks are mandatory — resolve ALL warnings before proceeding. Common blocker: vSAN health check failures.
- vSphere-to-VCF conversion requires vDS (no VSS), compatible ESXi version, and vSAN. Deploy SDDC Manager alongside existing vCenter.
- Workload domain creation requires commissioned hosts and configured network pools. Minimum 3 hosts per vSAN cluster.
- Host commissioning via SDDC Manager validates hardware, ESXi version, and network before allowing cluster addition.
HCX Workload Migration for VCF (Objective 5.11)#
HCX Overview & Architecture
HCX (Hybrid Cloud Extension) enables large-scale workload migration between sites — on-premises to cloud, VCF to VCF, or vSphere to VCF. Key for Obj 5.11.
HCX Components:
├── HCX Manager: management plane, deployed as OVA on source and destination ├── HCX Connector: source-side component (on-premises or legacy site) ├── HCX Cloud: destination-side component (target VCF or VMC environment) ├── Service Mesh: network tunnel between sites for migration traffic ├── IX (Interconnect): WAN optimization and encryption appliance pair ├── NE (Network Extension): L2 stretch appliance pair (extends VLANs across sites) └── WO (WAN Optimization): optional, compresses/deduplicates migration traffic
HCX Configuration (Obj 5.11)
Deployment Steps:
- Deploy HCX Manager OVA on source vCenter
- Activate HCX with license key (from Broadcom portal)
- Configure HCX Connector: pair with destination HCX Cloud
4. Site Pairing: source HCX ↔ destination HCX (requires network connectivity on port 443) 5. Create Compute Profile: define source/destination resources (clusters, datastores, networks) 6. Create Service Mesh: select appliance types (IX, NE, WO) and uplink networks 7. Deploy Service Mesh appliances (auto-deployed as VMs on both sites) 8. Verify tunnel status: Service Mesh → Tunnel Status → all tunnels green
Network Requirements:
├── Port 443 (HTTPS): management traffic between HCX managers ├── Port 4500 (UDP): IPSec tunnel for IX appliances ├── Port 8000 (TCP): vMotion traffic through IX tunnel ├── Port 31031 (TCP): bulk migration traffic ├── MTU 1350 minimum for tunnel overhead (1500 standard, 150 byte HCX header) └── Bandwidth: minimum 100 Mbps WAN, recommended 1 Gbps for bulk migrations
HCX Workload Migration Types (Obj 5.11)
Bulk Migration (HCX vMotion):
├── Moves VMs with zero downtime (warm migration) ├── Uses vMotion over WAN tunnel — no re-IP required if L2 extension in place ├── Default: up to 600 concurrent Bulk/RAV migrations per HCX Manager; scales to 1,000 in VCF environments with medium/large HCX Manager sizing (HCX 4.10+) ├── Pre-replication: initial sync while VM is running, final sync during cutover ├── Cutover window: typically <1 second for final memory sync └── Best for: large-scale datacenter migrations (100+ VMs)
RAV (Replication Assisted vMotion):
├── Combines bulk replication with vMotion final switchover ├── Phase 1: replicates VM disks over WAN (delta sync, dedup, compression) ├── Phase 2: final vMotion switchover with <1s downtime ├── Advantages: reduces WAN bandwidth vs pure vMotion, handles large disks efficiently ├── Supports migration scheduling (maintenance windows) └── Best for: VMs with large disks (>500 GB), WAN bandwidth-constrained environments
Cold Migration:
├── Powers off VM at source, copies VMDK to destination, powers on at destination ├── Downtime: full duration of copy (depends on disk size and WAN speed) ├── Simplest migration type — no running state to transfer ├── Required for: VMs with raw device mappings (RDMs), physical compatibility mode └── Best for: non-production VMs, VMs that can tolerate extended downtime
OS Assisted Migration (OSAM):
├── Agent-based migration for physical-to-virtual (P2V) or cross-platform ├── Installs HCX Sentinel agent on source machine ├── Replicates at block level, converts to VM on destination ├── Supports: Windows, Linux source machines (physical or virtual) └── Best for: legacy workloads on non-VMware platforms
HCX Troubleshooting for Support
Common Issues:
├── Service Mesh deployment fails: check network connectivity, port 443 open, DNS resolution ├── Tunnel status shows 'Degraded': verify MTU settings, WAN latency (<150ms recommended) ├── Migration fails at 80%: usually disk I/O bottleneck — check storage latency on both sides ├── vMotion switchover timeout: insufficient bandwidth or high memory change rate on VM ├── L2 Extension MAC conflict: duplicate MAC addresses across sites — use HCX mobility-optimized networking (MON) └── License expiry: HCX requires valid subscription — check Broadcom portal for renewal
Migration Planning Best Practices:
├── Wave planning: group VMs by application affinity (migrate app tiers together) ├── Pre-migration validation: run HCX Migration Validation to check VM compatibility ├── Network dependency mapping: identify VM-to-VM network flows to avoid breaking application chains ├── Rollback plan: HCX supports reverse migration if issues detected post-cutover ├── Performance monitoring: use HCX Dashboard to track migration progress, errors, bandwidth └── Post-migration validation: verify VM health, network connectivity, application functionality
Key Takeaways
- HCX Service Mesh creates IX/NE/WO appliance pairs between sites — requires port 443, 4500, 8000, 31031 open.
- RAV (Replication Assisted vMotion) is the preferred migration type for large VMs — combines disk replication with <1s vMotion cutover.
- Bulk migration supports up to 600 concurrent VMs per HCX Manager at default config (1,000 in VCF with large sizing) — use wave planning to group VMs by application affinity.
- L2 Network Extension stretches VLANs across sites — enables migration without re-IP. Mobility Optimized Networking (MON) prevents tromboning.
- HCX troubleshooting: check tunnel status, MTU (min 1350), WAN latency (<150ms), and port connectivity before migration.
Exam Mapping: 2V0-15.25 — VMware Cloud Foundation 9.0 Support (2V0-15.25)
- See VCF 9.0 Support exam blueprint for detailed objectives
Labs in This Section
VMware Cloud Foundation Overview
VCF 9.0BeginnervSphere Foundation Overview
VCF 9.0BeginnerCompute Core Concepts — ESXi Host Architecture
VCF 9.0BeginnerStorage Core Concepts — vSAN for VCF
VCF 9.0BeginnerNetworking Core Concepts — NSX Architecture
VCF 9.0BeginnerVMware Cloud Foundation Installation — Deployment Wizard & JSON Spec
VCF 9.0IntermediatevSphere Foundation Installation
VCF 9.0IntermediateGate: vcffts9-03
VMware Cloud Foundation Upgrades — 5.x to 9.0 & Conversion
VCF 9.0IntermediateGate: vcffts9-03
License Management — VCF 9.0 Licensing Model
VCF 9.0BeginnerCreating Workload Domains & Network Pools
VCF 9.0IntermediateVCF Operations Console — Monitoring Storage, Compute & Security
VCF 9.0IntermediateGate: vcffts9-07
VCF Operations for vSphere Foundation
VCF 9.0BeginnerVCF Automation — Purpose & Functionality
VCF 9.0BeginnerTechnical Support Fundamentals — Broadcom Wolken Case Management
VCF 9.0BeginnerVCF Operations HCX — Workload Mobility & Migration Troubleshooting
VCF 9.0IntermediatevCenter Service Recovery
VCF 9.0IntermediatevSAN Disk Failure Recovery
VCF 9.0IntermediateDFW Rule Debugging & Traceflow
VCF 9.0IntermediateUpgrade Failure & Remediation
VCF 9.0AdvancedPerformance Analysis with esxtop
VCF 9.0IntermediatevSAN Failure & Recovery Scenario
VCF 9.0AdvancedNSX Connectivity Troubleshooting
VCF 9.0AdvancedUpgrade Troubleshooting Scenario
VCF 9.0IntermediateAdvanced esxtop Performance Analysis
VCF 9.0AdvancedAvi Load Balancer Troubleshooting
VCF 9.0Intermediate📝 Quiz — VCF 9.0 Support
Section 1 — Troubleshooting Methodology and Tools
- esxcli sos bundle create
- python /opt/vmware/sddc-support/sos --domain-name <name>
- sddc-manager collect --logs
- vcf-support-bundle --domain <name>
- /var/log/messages
- /var/log/vmware/vcf/
- /opt/vmware/logs/
- /var/lib/vmware/sddc/
- SDDC Manager UI Dashboard
- VCF Operations for Logs with unified indexing
- vSphere Client Tasks & Events tab
- esxtop
- A screenshot of the alert
- The SOS diagnostic bundle
- A list of installed VMware Tools versions
- The vCenter configuration backup
- ping -s 8972 <ip>
- vmkping -I vmk2 -s 8972 -d <ip>
- esxcli network diag ping --host <ip>
- tcpdump -n host <ip>