VCF Operations HCX — Workload Mobility & Migration Troubleshooting
Objectives
- Understand HCX architecture and components in VCF 9.0
- Configure HCX site pairing and network extension
- Identify common HCX migration types and troubleshooting scenarios
- Troubleshoot HCX configuration and migration failures
Prerequisites
Active Holodeck VCF 9.0 lab with HCX deployed, or access to VCF/HCX documentation
Prior labs: vcffts9-01, vcffts9-05
Required skills:
- Basic VMware networking
- Understanding of VCF architecture
Tasks
Task 1 HCX Architecture & Configuration
Understand HCX components and configuration for VCF workload mobility
VCF Operations HCX (formerly Hybrid Cloud Extension) enables workload mobility between VCF environments and across sites. Core components: HCX Manager (management appliance, deployed per site), HCX Connector (source site, connects to vCenter), HCX Cloud (destination site). Service Mesh: automated deployment of HCX appliances (Interconnect, WAN Optimization, Network Extension) that create the data plane for migrations. Site Pairing: establishes trust between source and destination HCX Managers — requires network connectivity on port 443, valid certificates, and admin credentials on both sides.
[HOLODECK NOTE] HCX can be deployed in Holodeck for functional testing of site pairing and migration workflows. However, WAN Optimization appliance performance is not representative of production — Holodeck's virtual networking has no real WAN latency or bandwidth constraints to optimize. For RAV (Replication Assisted vMotion) testing, both source and destination Holodeck environments need sufficient RAM to run the migration appliances.
[SOURCE NOTE] The VCFFTS9 lecture manual covers VCF Operations HCX at overview level only (Module 2: Use Cases). Detailed HCX architecture, Service Mesh, and migration type content in this lab is sourced from HCX documentation and general VMware training materials, not the VCFFTS9 lecture manual. This content addresses exam objective 5.11 which requires HCX troubleshooting knowledge.
Network Extension: stretches L2 networks between sites — VMs keep IP addresses during migration. Requires Network Extension appliance pair (source + destination). Proximity Routing prevents tromboning by advertising the migrated network to the destination site's gateway. Migration types: (1) Cold Migration — powered-off VMs, uses OVF export/import; (2) vMotion — live migration, zero downtime, single VM at a time, zero downtime, performance dependent on WAN bandwidth; (3) Bulk Migration — scheduled, brief downtime during switchover, parallel migrations (up to 300 per HCX Manager (scalable to 600 in HCX 4.7+)); (4) Replication Assisted vMotion (RAV) — combines replication for bulk data transfer with vMotion for final switchover, best for large VMs over WAN (handles large VMs efficiently via replication-based data transfer); (5) OS Assisted Migration — for non-vSphere sources. For exam: understand which migration type to recommend based on VM size, downtime tolerance, and WAN bandwidth.
Common HCX troubleshooting scenarios for exam: (1) Site pairing failures — check DNS resolution between HCX Managers, port 443 connectivity, certificate trust chain, and admin credentials; (2) Service Mesh deployment failures — verify uplink network profiles (management, vMotion, replication networks), check that required IP pools have available addresses, verify MTU settings (jumbo frames recommended for replication traffic); (3) Migration failures — RAV: check replication network bandwidth and latency, verify source VM has no unsupported devices (raw disk mappings, USB passthrough); vMotion: verify vMotion VMkernel connectivity between sites, check EVC compatibility mode; Bulk: check that replication network IP pool is not exhausted; (4) Network Extension issues — verify L2 extension appliance pair health, check uplink VLAN trunking, verify proximity routing configuration if VMs cannot reach default gateway after migration; (5) HCX activation failures — verify license key is valid and has not exceeded the allowed number of site pairings.
Validation Gate
Check: When should you use HCX vMotion vs HCX RAV vs HCX Cold Migration?
Expected: HCX vMotion: single critical VM, zero downtime required. HCX RAV: bulk migration with minimal downtime (replication + final vMotion switchover). HCX Cold: VM can be powered off during migration, simplest approach. HCX Bulk: large-scale parallel migrations with scheduled switchover windows.
Common Errors
Task 2 HCX in VCF 9.0 Context
Understand HCX integration with VCF fleet architecture and upgrade considerations
In VCF 9.0, HCX integrates with the fleet architecture: HCX Manager is deployed per VCF Instance, managed via SDDC Manager lifecycle management. HCX supports migrations between VCF Instances within a fleet, between VCF and standalone vSphere, and between on-premises VCF and VMware Cloud on hyperscalers. HCX licensing in VCF 9.0 still uses per-instance license keys (not yet on the VCF 9.0 license file model — see Lab 9, 'Products Not Yet on VCF 9.0 License Model'). For VCF 5.x to 9.0 upgrades: HCX should be upgraded after the core VCF stack (Phase 2 completion) — consult the VCF Transition Sequence Library for exact version compatibility.
Migration planning considerations specific to VCF: (1) NSX compatibility — ensure NSX security policies and micro-segmentation rules are replicated or mapped at the destination before migration; (2) vSAN storage — migrating VMs between vSAN clusters requires sufficient destination capacity and compatible storage policies; (3) Network pools — destination workload domain must have appropriate network pools configured before receiving migrated VMs; (4) Certificate management — HCX Manager certificates must be valid and trusted by both source and destination vCenter/NSX; (5) Bandwidth planning — RAV migrations consume replication network bandwidth; plan for 70% utilization to avoid impacting production traffic.
[HOLODECK NOTE] HCX migration testing between two Holodeck environments requires two separate Holodeck deployments with routable connectivity between them. This demands substantial physical resources (2x the RAM/CPU/storage of a single deployment). For VCDX study purposes, focus on understanding migration type selection criteria and troubleshooting workflows rather than attempting full cross-site HCX lab scenarios in Holodeck.
Validation Gate
Check: What are the three most common causes of HCX migration failure?
Expected: 1. Replication sync not converging (high change rate + insufficient bandwidth). 2. Network connectivity issues between IX appliances (firewall rules, MTU). 3. Source VM compatibility issues (unsupported virtual hardware version, incompatible guest OS).
Common Errors
Design Reflection (VCDX)
Migration strategy is a key VCDX topic — panelists test whether you understand migration types, can estimate migration timelines based on VM count and bandwidth, and have a rollback plan for failed migrations. HCX knowledge demonstrates practical migration architecture skills.
Requirements
- Plan migration strategy using appropriate HCX migration types
- Size HCX appliances for target migration throughput and concurrency
- Design network extension and cutover strategy
Constraints
- HCX vMotion is one-VM-at-a-time — not suitable for bulk
- Network extension adds WAN latency — not a permanent architecture
- Source vSphere version must be compatible per HCX matrix
Assumptions
- Sufficient WAN bandwidth for replication sync to converge
- HCX licenses are available for source and destination sites
Risks
- Replication sync never converging due to high change rate
- Network extension becoming permanent instead of temporary migration aid
- Application failures at destination due to untested dependencies
⚠ Known Pitfalls (from Community KB)
References
- HCX DocumentationTier 1 — Official
- HCX Migration Planning GuideTier 1 — Official