Academy/VCF 9.0 Support (2V0-15.25)/Compute Core Concepts — ESXi Host Architecture
This lab targets VCF 9.0

Compute Core Concepts — ESXi Host Architecture

VCF 9.0Beginnersupportadmin⏱ 60 min

VCFFTS9 course content — VCF 9.0 focused

Objectives

  • Describe ESXi host architecture, security features, and boot bank layout
  • List hardware requirements for ESX 9.0 and vSAN
  • Perform ESXi host preparation for VCF (network, certificates, NTP)
  • Explain vCenter Appliance architecture and the vpxd→vpxa→hostd communication chain
  • Describe vCenter SSO and Enhanced Linked Mode (up to 15 vCenter instances)
  • Differentiate VSS from VDS and explain NIOC
  • Identify snapshot types by datastore (VMFSsparse, SEsparse, vsanSparse)
  • Explain vMotion including vGPU migration (up to 60 Gbps)
  • Describe vSphere Lifecycle Manager image-based management in 9.0

Prerequisites

Active Holodeck VCF 9.0 lab or access to VCF documentation

Required skills:

  • Basic VMware terminology

Tasks

Task 1 ESXi Host Architecture & Configuration

ESXi host preparation is where most VCF deployments fail. Certificate regeneration, NTP sync, and DNS resolution must be perfect before Cloud Builder can commission hosts. A single misconfigured host blocks the entire bring-up.

Understand ESXi architecture, HW requirements, and VCF-specific host preparation

Step 1
ESXi Architecture

ESX is a bare-metal hypervisor. Security features: host-based firewall, memory hardening, kernel module integrity, TPM 2.0, UEFI secure boot, encrypted core dumps. Boot bank layout: /bootbank, /altbootbank (dual-bank for patching), /productLocker (VMware Tools ISOs), /var/core (core dumps). Small disk footprint with quick boot for faster patching.

Step 2
Hardware Requirements

ESX 9.0 minimum requirements: supported server platform (check VMware Compatibility Guide), 2+ CPU cores, 8GB RAM (12GB production recommended), 1+ GbE NIC, 32GB persistent boot disk. For vSAN: disks and controllers must be compatible per VMware Compatibility Guide. Memory Tiering: PCIe NVMe devices can serve as second memory layer for more in-memory computing.

[HOLODECK NOTE] In Holodeck, ESXi runs as nested VMs on a physical host. Hardware compatibility checks (VMware Compatibility Guide) do not apply — nested ESXi uses virtual hardware. CPU features are passed through from the physical host. Memory is allocated from the physical host's RAM pool — plan 64GB+ physical RAM for a minimal VCF deployment in Holodeck.

Step 3
ESXi Host Preparation for VCF

For management domain: interactive ESX installation on all hosts. Post-install config via DCUI and Host Client: (1) Management network — adapter selection, VLAN ID, IPv4/IPv6, hostname, DNS; (2) VM Network port group — set VLAN ID on vSS for Cloud Builder connectivity; (3) Certificate regeneration — run /sbin/generate-certificates then restart hostd/vpxa/rhttpproxy; (4) NTP configuration (UDP 123) or PTP (UDP 319/320 for microsecond accuracy). Ensure no Custom DNS Suffixes defined.

Step 4
Remote Access & Security

ESX firewall activated by default — blocks all except enabled services. SSH and Shell managed by admin users. Lockdown mode: host accessible only via DCUI or through vCenter (no direct remote login). Services management: START/STOP/RESTART, startup policy (manual, start with host, start/stop with port usage).

Validation Gate

Check: Verify ESXi host is VCF-ready: FQDN matches DNS (forward + reverse), certificates regenerated, NTP synchronized, no Custom DNS Suffix

Expected: All four checks pass. Host is ready for Cloud Builder commissioning.

Common Errors

Forgetting to regenerate ESXi certificates after hostname change
Fix: After setting hostname and DNS, run: /sbin/generate-certificates && /etc/init.d/hostd restart && /etc/init.d/vpxa restart. Without this, Cloud Builder's certificate validation fails during host commissioning with 'certificate hostname mismatch'.
Leaving Custom DNS Suffix configured on ESXi host
Fix: A Custom DNS Suffix in ESXi DCUI causes FQDN resolution mismatches. Cloud Builder expects the host FQDN to match forward and reverse DNS exactly. Remove any Custom DNS Suffix before commissioning.
NTP not synchronized before bring-up
Fix: VCF bring-up requires all hosts within 30 seconds of NTP source. Use 'esxcli system time get' and verify NTP service is running. Time drift causes certificate validation failures and vSAN cluster formation issues.
Using DHCP instead of static IP for ESXi management
Fix: VCF requires static IP assignment for all ESXi management interfaces. DHCP-assigned addresses cause host commissioning failures because SDDC Manager cannot reliably reach hosts if IPs change.

Task 2 vCenter Architecture & Management

Understanding the vpxd → vpxa → hostd communication chain is essential for troubleshooting connectivity issues between vCenter and ESXi hosts. When a host shows 'disconnected' in vCenter, the root cause is always in this chain.

Understand vCenter Appliance architecture, services, and SSO

Step 1
vCenter Appliance Architecture

vCenter Appliance = prepackaged Linux VM: Photon OS + PostgreSQL database + vCenter services. All services on single VM. Services include: vCenter, vSphere Client, Authentication Services, License service, Content Library, vSphere Lifecycle Manager. Deploy on existing ESX host, select appliance size for environment + storage size for DB.

Step 2
ESXi-vCenter Communication
Communication chain: vSphere Client → vpxd (vCenter) → vpxa (host agent, auto-started when host added) → hostd (host daemon). hostd manages all local operations: VM creation, power state, storage visibility. Direct host access (bypassing vCenter) communicates with hostd directly via VMware Host Client.
Step 3
vCenter SSO & Enhanced Linked Mode

SSO provides centralized authentication: identity sources (AD, LDAP, local OS), security tokens, site/domain concepts. Enhanced Linked Mode: up to 15 vCenter instances in single SSO domain for unified inventory, role management, and search across data centers. IMPORTANT: Enhanced Linked Mode (ELM) is DEPRECATED in vCenter 9.0 and will be removed in a future release. ELM is still supported in VCF 9.0 only to allow smooth upgrades of existing deployments. New deployments should use VCF Operations Fleet Management for multi-instance management instead of ELM.

Validation Gate

Check: Trace a VM power-on request from vSphere Client through vpxd, vpxa, and hostd

Expected: vSphere Client → vpxd (vCenter service, validates permissions, updates DB) → vpxa (host agent on target ESXi, receives task) → hostd (executes power-on, updates VM state) → status propagates back up the chain

Common Errors

Confusing Enhanced Linked Mode deprecation with immediate removal
Fix: ELM is deprecated in vCenter 9.0, not removed. Existing ELM deployments continue to function. However, new VCF 9.0 deployments should NOT configure ELM — use VCF Operations Fleet Management instead for multi-instance visibility.
Not understanding SSO domain scope for RBAC
Fix: SSO domain defines the RBAC boundary. Permissions assigned in one vCenter within an SSO domain apply across all linked vCenters. In VCF, each workload domain has its own vCenter, but they share the SSO domain — understand the RBAC implications.

Task 3 vSphere Networking, VMs & Clusters

Snapshot management is a common operational pain point. Understanding snapshot types per datastore is critical for troubleshooting performance issues caused by snapshot chains and for backup architecture design.

Cover networking constructs, VM operations, and cluster capabilities

Step 1
vSphere Networking

VSS (Standard Switch): per-host, manual config. VDS (Distributed Switch): centrally managed across hosts, supports port groups spanning hosts, NetFlow, port mirroring, LACP. VMkernel adapters for: management, vMotion, vSAN, fault tolerance logging, provisioning traffic. NIC teaming: load-based, source port ID, source MAC hash, explicit failover. Network I/O Control (NIOC): bandwidth shares and reservations per traffic type.

[HOLODECK NOTE] Holodeck uses virtual switches on the physical host to provide connectivity to nested ESXi. VDS in nested ESXi connects to virtual port groups on the physical host's vSwitch. NIC teaming policies in nested ESXi operate on virtual NICs — failover and load balancing behavior differs from physical multi-NIC configurations. NIOC bandwidth reservations have no real effect in nested environments.

Step 2
Virtual Machine Operations

VM files: .vmx (config), .vmdk (disk descriptor), -flat.vmdk (data), .nvram (BIOS/EFI), .vmsd (snapshot list), .vmsn (snapshot state), .vmem (memory state). Snapshot types by datastore: VMFSsparse (VMFS5 <2TB, 512-byte blocks), SEsparse (VMFS6, 4KB blocks, space-efficient with unmap), vsanSparse (vSAN ESA, delta objects, 4MB blocks). CBT (Changed Block Tracking): VMkernel feature for incremental backups — tracks changed blocks, reduces restore time.

Step 3
Cluster Operations & Live Migration

Cluster: up to 96 ESX hosts (64 with vSAN). DRS: automated load balancing. HA: host monitoring, admission control, VM restart on failure. vMotion: live migration with zero downtime. vGPU vMotion: multi-parallel TCP connections up to 60 Gbps (6x faster), zero-copy technique, precopy for minimal downtime. EVC (Enhanced vMotion Compatibility): CPU feature masking for heterogeneous clusters. Migration types: compute only, storage only, both, cross-vCenter export.

[HOLODECK NOTE] In Holodeck's nested environment: vMotion works but bandwidth is limited by virtual NIC throughput, not physical 25GbE. The 60 Gbps vGPU vMotion figure applies to bare-metal with dedicated NICs — Holodeck nested NICs will not approach this. vSAN cluster limits (64 hosts) are theoretical — Holodeck typically runs 3-4 nested hosts due to RAM constraints.

Step 4
Content Libraries & vSphere Lifecycle Manager

Content libraries: local (admin-controlled), published (available for subscription), subscribed (sync from publisher). Templates stored as OVF or VM templates. vSphere 9.0: library migration to new datastore supported. vSphere Lifecycle Manager: image-based management (no more baselines in 9.0), desired state model — cluster image includes base ESXi + vendor add-on + components + firmware.

Validation Gate

Check: Name the snapshot type for each datastore: VMFS5, VMFS6, vSAN ESA

Expected: VMFS5 (<2TB): VMFSsparse (512-byte blocks). VMFS6: SEsparse (4KB blocks, space-efficient with UNMAP). vSAN ESA: vsanSparse (4MB blocks, delta objects within vSAN object model).

Common Errors

Running VMs on snapshots for extended periods in production
Fix: Snapshots are for short-term use (backups, patching windows). A snapshot chain growing beyond 3 levels causes exponential I/O overhead — each read must traverse the chain. Monitor snapshot age via VCF Operations alerts. vSAN ESA uses vsanSparse with 4MB blocks for better performance than legacy formats.
Confusing vMotion bandwidth in nested vs bare-metal
Fix: The 60 Gbps vGPU vMotion figure applies to bare-metal with dedicated 25GbE NICs and RDMA. In Holodeck nested environments, vMotion bandwidth is limited by virtual NIC throughput (typically 1-10 Gbps). Do not cite nested lab migration times as production performance.
Not understanding image-based lifecycle management in vSphere 9.0
Fix: vSphere 9.0 removes baseline-based patching entirely. All host updates are image-based via vSphere Lifecycle Manager (vLCM). The cluster image includes: base ESXi + vendor add-on (firmware/drivers) + components. This is a significant operational change from legacy Update Manager baselines.

Design Reflection (VCDX)

Compute architecture forms the foundation of every VCF design. VCDX panelists test whether you understand ESXi host sizing, cluster limits, and lifecycle management. The shift to image-based management in vSphere 9.0 is a frequent discussion point.

Requirements

  • Understand ESXi architecture, security features, and VCF preparation
  • Master vCenter communication chain for troubleshooting
  • Know cluster limits, migration types, and snapshot behavior

Constraints

  • ESXi 9.0 requires image-based lifecycle management (no baselines)
  • ELM deprecated — use Fleet Management for multi-instance
  • Cluster limit: 96 hosts (64 with vSAN)

Assumptions

  • Hardware meets VMware Compatibility Guide requirements
  • NTP, DNS, and certificates are configured correctly before bring-up

Risks

  • Snapshot chain growth causing I/O performance degradation
  • Certificate mismatch blocking host commissioning
  • Confusing nested lab performance with production benchmarks

⚠ Known Pitfalls (from Community KB)

Quoting nested lab performance numbers (vMotion speed, I/O throughput) as production expectations — always distinguish Holodeck measurements from bare-metal baselines.
Forgetting that vSphere 9.0 removes baseline-based patching — all host lifecycle is image-based via vLCM.
Overlooking ESXi certificate regeneration as a pre-commissioning step — this is the #1 cause of Cloud Builder failures in the field.

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.