Academy/Holodeck Lab Setup & Operations/Deploy a VI Workload Domain on Holodeck
This lab targets VCF 9.0.2

Deploy a VI Workload Domain on Holodeck

VCF 9.0.2Advancedadminarchitectvcdx⏱ 150 min

Applies to VCF 9.0.0 through 9.0.2. Workload domain architecture is consistent across the 9.0.x family.

Objectives

  • Commission additional ESXi hosts into SDDC Manager and prepare them for workload domain use
  • Deploy a VI (Virtual Infrastructure) Workload Domain using SDDC Manager, including compute, storage (vSAN), and NSX edge clusters
  • Validate workload domain health across vCenter, NSX, and SDDC Manager
  • Deploy and validate a test VM workload in the workload domain with networking and storage placement
  • Perform basic workload domain lifecycle operations: scale-out, scale-in, and cleanup
  • Articulate the VCDX design implications of multi-domain architecture, failure isolation, and resource sharing in nested environments

Prerequisites

Holodeck instance with management domain deployed and healthy (holodeck-02 completed). SDDC Manager, vCenter, NSX Manager, and 4x ESXi hosts all Active/operational. Clean snapshot 'holodeck-02-complete' should exist. At least 2 additional uncommissioned ESXi hosts available in Holodeck (part of the standard Holodeck 2.1.x deployment that can be reserved for workload domain).

Prior labs: holodeck-02

Required skills:

  • SDDC Manager workload domain administration (from holodeck-02)
  • vCenter cluster and vSAN configuration
  • NSX transport nodes and edge cluster concepts
  • vSAN stretched/multi-cluster design considerations
  • Understanding of VCF licensing (evaluation licenses include WLD capability)
📸 Starting State: S3-52 — Management Domain Deployed (VCF 5.2)

Lab Environment

Single Holodeck pod with management domain already deployed. This lab adds a workload domain using: 2-3 additional nested ESXi hosts (to be commissioned), a dedicated vSAN storage cluster, NSX edge cluster (2-3 edge nodes), isolated networks (vMotion, vSAN, TEP). Workload domain is separate from management domain with distinct failure boundaries.

graph TB
  MDCluster[Management Cluster<br/>4x ESXi]:::mgmt
  WLDCluster[WL Cluster<br/>2-3x ESXi]:::wld
  SDDC[SDDC Manager 10.1.1.4]:::mgmt
  VC1[vCenter 10.1.1.6]:::mgmt
  NSX[NSX Manager 10.1.1.10]:::mgmt
  EdgeMgmt[Edge Cluster<br/>Management]:::mgmt
  EdgeWLD[Edge Cluster<br/>Workload]:::wld
  Network[Network Pool<br/>vMotion/vSAN/TEP]:::wld
  SDDC --> MDCluster
  SDDC --> WLDCluster
  SDDC --> Network
  VC1 --> MDCluster
  VC1 --> WLDCluster
  NSX --> EdgeMgmt
  NSX --> EdgeWLD
  WLDCluster --> Network
  classDef mgmt fill:#003d82,stroke:#0055cc,color:#fff
  classDef wld fill:#7cb342,stroke:#558b2f,color:#fff

IP Addressing

NetworkPurposeVLAN
10.1.1.0/20Management network (SDDC Manager, vCenter, NSX, Cloud Builder)VLAN 1644 (shared)
10.1.1.0/24Management cluster ESXi hostsVLAN 1644
10.1.10.0/24Workload cluster ESXi hostsVLAN 1649
10.1.11.0/24Workload vMotionVLAN 1650
10.1.12.0/24Workload vSANVLAN 1651
10.1.13.0/24Workload NSX Host TEPVLAN 1652
10.1.14.0/24Workload NSX Edge TEPVLAN 1653
192.168.1.0/24Workload VM segment (test workload)VLAN 1701

Credentials

SystemUsernamePassword
SDDC Manageradministrator@vsphere.localSet during management domain deployment (holodeck-02)
vCenter Serveradministrator@vsphere.localSame as SDDC Manager SSO
NSX ManageradminSet in bring-up spec during management domain deployment. API password is VMware123!VMware123! (doubled)
Nested ESXi Hosts (WLD)rootSame as management domain hosts (set in config.json)
API Authentication (NSX/SDDC Manager)administrator@vsphere.localAPI auth requires VMware123!VMware123! (doubled password, not single)

Tasks

Task 1 Commission additional ESXi hosts for the workload domain

manageability

Workload domain hosts must be registered with SDDC Manager before cluster creation. This mirrors the Day 2 operation where new hardware is added to a VCF environment. In a production VCDX design, the question 'How do you expand a VI workload domain?' goes directly to commissioning and cluster membership strategies.

Step 1

Open SDDC Manager: https://sddc-manager.site-a.vcf.lab. Navigate to Inventory > Hosts.

Host inventory shows 4 commissioned hosts from the management domain (ESXi-01 through ESXi-04 with IPs 10.1.1.101-104), all Active
Step 2

In SDDC Manager, navigate to Inventory > Hosts > Add Host. Select 'Manual' host addition (do not auto-discover). Configure the first workload host: Hostname/FQDN: esxi-wld-01.site-a.vcf.lab (or esxi-05.site-a.vcf.lab), IP: 10.1.10.101, Root password: (same as management hosts from config.json), Network timeout: 60 seconds.

Form accepts the hostname and IP. Connection test passes.
Do NOT use auto-discovery in Holodeck — it may pick up management domain hosts. Manual addition ensures you select the correct uncommissioned host.
Step 3

Click 'Add' to commission the first host. Monitor SDDC Manager > Home > Task Queue until the host commissioning task completes (usually 5-10 minutes).

First workload host transitions from Provisioning to Active status in Inventory > Hosts
While waiting, proceed to commission the second host in a separate browser tab to save time.
Step 4

Repeat Steps 2-3 for a second workload host: Hostname: esxi-wld-02.site-a.vcf.lab (or esxi-06.site-a.vcf.lab), IP: 10.1.10.102. For a larger workload domain (advanced), add a third: esxi-wld-03, IP: 10.1.10.103.

2-3 new hosts listed in Inventory > Hosts with Active status and all health indicators green
Step 5

Create a network pool for workload domain use. Navigate to Inventory > Network Pools. Click 'Add Network Pool'. Create a pool named 'wld-netpool' with the following address ranges (use the IP ranges from the environment section above): vMotion: 10.1.11.0/24 (VLAN 1650), vSAN: 10.1.12.0/24 (VLAN 1651), NSX Host TEP: 10.1.13.0/24 (VLAN 1652), NSX Edge TEP: 10.1.14.0/24 (VLAN 1653).

Network pool 'wld-netpool' created and shows in the Network Pool inventory with all ranges defined
Network pool IP ranges MUST NOT overlap with management domain ranges (10.1.1-8.x). Thread 52 of the community KB documents a scenario where an overlooked network conflict caused workload domain deployment to hang at 'Validating NSX Transport Node configuration'.

Validation Gate

Check: SDDC Manager Inventory > Hosts shows 6-7 total hosts (4 management + 2-3 workload), all Active. Inventory > Network Pools shows 'wld-netpool' with all four address ranges defined.

Expected: Hosts commissioned and network pool configured with no overlap warnings

Common Errors

Host commission fails with 'Connection refused' or 'Host not responding'
Cause: The ESXi host IP is incorrect, the host is not booted, or networking is misconfigured (VLAN mismatch or routed network unreachable)
Fix: Verify the ESXi host is powered on (check in vSphere Client or Holodeck console). Test connectivity: ping 10.1.10.101 from HoloConsole. Verify VLAN 1649 is routed/accessible from HoloRouter.
📋 KB: Thread 52: Unable to deploy additional Workload Domain (host commissioning prerequisite)
Host commission hangs in 'Provisioning' state for >15 minutes
Cause: SDDC Manager is backlogged or the host's management VMkernel is slow to respond. Resource contention on the physical host.
Fix: Check SDDC Manager Task Queue for other running tasks. Check physical host CPU/memory utilization. Retry the host commission. If persistent, check SDDC Manager service logs: SSH sddc-manager, sudo tail -f /var/log/domainmanager/sddc-manager-*.log
Network pool creation fails with 'VLAN is already in use' or 'IP range overlaps'
Cause: The VLAN or IP range already exists in the network pool configuration
Fix: Verify the IP ranges do not overlap with management domain (10.1.1.0/24 through 10.1.8.0/24 are reserved). Use different VLAN IDs (1649-1653 should be clear).

Task 2 Create the VI Workload Domain via SDDC Manager

availability

Workload domain creation is the pivotal moment where VCF separates concerns: management operations (SDDC Manager, vCenter, NSX) remain in the management domain; user workloads run in the workload domain with its own vCenter cluster, vSAN store, and NSX transport nodes. VCDX panelists will probe your understanding of the cluster membership, storage architecture, and NSX edge placement decisions.

Step 1

In SDDC Manager, navigate to Workload Domains > Add Domain. Name the domain 'workload-domain-1'. Select 'VI' type (Virtual Infrastructure — the standard workload domain type that hosts VMs).

Add Workload Domain wizard opens with the VI template selected
Step 2

Configure the cluster membership. Select the 2-3 commissioned workload hosts (esxi-wld-01, esxi-wld-02, and optionally esxi-wld-03) to form the workload cluster. These hosts will be moved into a new vSAN-backed cluster named 'wld-cluster-1'.

Wizard shows the selected hosts highlighted in green, ready for cluster creation
Step 3

Configure storage: Select 'vSAN' as the storage type. Set vSAN FTT (Fault Tolerance) to 1 (RAID-1) to match the management domain. This ensures consistent data protection across both domains. Configure vSAN network: Select the vSAN range from 'wld-netpool' (10.1.12.0/24, VLAN 1651).

vSAN configuration shows FTT=1, deduplication enabled (optional but recommended), vSAN network mapped to 10.1.12.0/24
FTT=0 (no redundancy) is not recommended for any production-like configuration — even in a nested lab, it teaches bad habits. If you have only 2 hosts, FTT=1 with 2 hosts is supported (mirror + witness on a third site is required for larger clusters, not applicable here). The community thread 'vSAN FTT configuration for workload domain' discusses this nuance.
Step 4

Configure NSX: Enable 'Deploy NSX Edge Cluster' (critical for workload domain isolation). Select the edge network range from 'wld-netpool' (10.1.14.0/24, VLAN 1653) for the edge cluster. Set edge node count to 2 (minimum for production-grade HA, though 1 is acceptable in labs to conserve resources). Configure vMotion network: use 10.1.11.0/24 (VLAN 1650) from the network pool.

NSX configuration shows edge cluster deployment enabled, 2 edge nodes planned, networks assigned correctly
Step 5

Select principal storage policy. Accept the default 'vSAN Default Storage Policy' which uses FTT=1 — this ensures workload domain VMs use the same protection as management domain.

Storage policy selected and shown in the review pane
Step 6

Review the entire configuration. The summary should show: cluster: 2-3 hosts, storage: vSAN FTT=1, NSX: edge cluster with 2 nodes, networks: vMotion, vSAN, NSX Host TEP, NSX Edge TEP all mapped. Click 'Review and Deploy'.

Configuration summary appears, all fields populated with green checkmarks
Step 7

Confirm deployment. Click 'Deploy'. Monitor the deployment in SDDC Manager > Home > Task Queue. The workflow includes: (1) Cluster creation in vCenter, (2) vSAN setup, (3) NSX transport node configuration, (4) Edge cluster deployment. Full deployment typically takes 30-60 minutes.

Task Queue shows 'Create Workload Domain' in progress. Major sub-steps update: 'Creating cluster', 'Configuring vSAN', 'Adding transport nodes', 'Deploying edge cluster'.
Do NOT cancel the deployment mid-way. If it fails, you will need to manually clean up VMs and remove the failed domain before retrying. Common failure points: vSAN formation timeout (resource contention), NSX transport node installation (VIB deployment timeout), edge cluster ready check (network unreachable).
Step 8

Wait for deployment completion. When the task finishes, verify the Workload Domains inventory shows 'workload-domain-1' with status 'Active'.

Workload domain listed in SDDC Manager with status: Active, showing 2-3 hosts and vSAN cluster health green

Validation Gate

Check: SDDC Manager: Workload Domains > workload-domain-1 = Active. Inventory > Hosts shows the workload hosts now assigned to a new cluster (no longer unassigned).

Expected: Workload domain created and active in SDDC Manager, cluster visible in vCenter, vSAN formed, NSX edge cluster operational

Common Errors

Domain creation hangs at 'Configuring NSX transport nodes' for >20 minutes
Cause: NSX VIB installation is slow (common in nested environments due to resource contention). Network timeouts between NSX Manager and workload hosts.
Fix: Check NSX Manager logs: SSH nsx-manager, check /var/log/proton/nsx-manager.log. Verify network connectivity between management and workload clusters (ping NSX from workload ESXi, ping workload ESXi from NSX). Increase NSX Manager timeout in SDDC Manager configuration if needed.
Edge cluster deployment fails with 'Insufficient resources' or 'Edge VM placement failure'
Cause: The workload cluster lacks sufficient memory/CPU for edge VMs (each edge VM requires 4 vCPU, 8-12 GB RAM). Physical host resource contention.
Fix: Reduce edge node count to 1 (acceptable for labs). Verify physical host has sufficient free RAM (check HoloConsole: free -h or vSphere Client resource utilization). Shut down unnecessary VMs on the physical host.
Task fails with 'Licensing check failed' or 'Workload domain license limit reached'
Cause: VCF evaluation license has reached the workload domain limit. Holodeck includes 1-2 WLD evaluation licenses.
Fix: Check SDDC Manager > Administration > Licensing. Request additional evaluation licenses from Broadcom. If only 1 WLD is licensed, delete this WLD and try again to ensure you have capacity (see Task 5 for cleanup).
vSAN formation fails with 'disk not available' or 'cannot find disks'
Cause: The workload domain hosts do not have unformatted disks available, or disks are already claimed by management domain vSAN
Fix: This is a Holodeck configuration issue — the toolkit should have reserved disks for each host. Verify in vSphere Client > workload host > Configure > Storage > Storage Devices. Look for unclaimed disks. If none exist, you may need to re-prepare the Holodeck instance with separate disk allocations.

Task 3 Validate Workload Domain health

manageability

Post-creation validation mirrors production Day 2 operations. You must know how to independently verify each component's health (vCenter cluster, vSAN, NSX fabric, SDDC Manager) — this is essential VCDX knowledge. Panelists will ask: 'If workload domain creation succeeded but a workload fails to start, how would you diagnose?'

Step 1

Open vCenter: https://vcenter.site-a.vcf.lab. Navigate to the workload domain cluster (should be named 'wld-cluster-1' or similar). Verify: (a) Cluster shows 2-3 connected hosts with green status, (b) vSAN health is green (no warnings about disk or network issues).

vCenter cluster view shows all workload hosts connected and vSAN health summary green
Take a screenshot of the vSAN dashboard — this is a key artifact for VCDX design documentation.
Step 2

In vCenter, verify DRS and HA are enabled on the workload cluster. Navigate to the cluster > Configure > vSphere DRS. Confirm DRS is 'Enabled' with 'Fully Automated' migration threshold. Confirm HA is enabled with admission control set (typically 'Reservation based').

DRS Enabled / Fully Automated, HA Enabled with admission control configured
Step 3

In SDDC Manager, navigate to Inventory > Workload Domains > workload-domain-1. Verify status is 'Active' and all health indicators are green. Check the 'Summary' tab: confirms 2-3 hosts, vSAN datastore shows as 'Configured', edge cluster status shows.

Workload domain summary shows Active status with all components healthy
Step 4

In NSX Manager: https://nsx-manager.site-a.vcf.lab. Navigate to System > Fabric > Nodes > Host Transport Nodes. Verify all workload domain ESXi hosts are listed with 'Configuration State: Success' and 'Transport Node Status: Up'.

2-3 workload domain transport nodes, all Success / Up
Step 5

In NSX Manager, navigate to System > Fabric > Nodes > Edge Transport Nodes. Verify the workload domain edge cluster is listed (should be named something like 'wld-edge-cluster-1') with all edge nodes showing 'Up' status.

Edge cluster listed with all edge nodes operational
Step 6

In NSX Manager, navigate to Networking > Tier-0 Gateways. Verify a Tier-0 gateway for the workload domain exists (may be named 'Tier0-WLD' or similar) and is in 'Active' status with service routers attached to workload edge nodes.

Workload Tier-0 gateway visible and active, routing service operational on workload edge cluster
Step 7

Test network connectivity between domains. From a management domain host, ping a workload domain host on the NSX Host TEP network (10.1.13.x). Expected: reply in <5ms (TEP tunnels are low-latency). From SDDC Manager or vCenter UI, there should be no warnings about cross-domain communication.

Ping replies received, latency <10ms

Validation Gate

Check: vCenter: Workload cluster shows 2-3 connected hosts, vSAN green, DRS/HA enabled. SDDC Manager: workload domain Active. NSX: all workload transport nodes Success/Up, edge cluster Up, Tier-0 active. Network connectivity confirmed.

Expected: All workload domain components healthy and interconnected with no degradation

Common Errors

vCenter workload cluster shows '1/2 hosts connected' or 'host disconnected'
Cause: A workload host lost connection to vCenter — usually a network interruption, ESXi management VMkernel timeout, or vCenter service restart
Fix: Check the host's management network connectivity: ping 10.1.1.6 (vCenter) from the ESXi host console. Restart the ESXi host's management network: in ESXi shell, esxcli network ip interface set -i vmk0 -I 10.1.10.101 (or applicable IP). Reconnect the host in vCenter.
NSX transport nodes show 'Install Failed' or 'Not Ready'
Cause: NSX VIB installation did not complete cleanly on one or more workload hosts
Fix: In NSX Manager, select the failed transport node and click 'Resolve'. NSX will retry the VIB installation. If repeated failures, SSH to the ESXi host and check: esxcli software vib list | grep nsx to see actual VIB version.
vSAN shows 'Object replication failure' or 'Reduced redundancy' warning
Cause: A workload host's vSAN network is misconfigured or a disk failed
Fix: In vSAN Health, go to Physical Disks and verify all disks are 'Healthy'. Check vSAN network: vSphere Client > host > Configure > VMkernel adapters, verify vSAN vmk is bound to correct NIC and VLAN 1651 is routed.

Task 4 Deploy a test workload into the workload domain

security

A workload domain is only proven by actually running workloads. This task teaches you the difference between a 'ready' domain and a 'proven' domain. VCDX candidates must demonstrate that they understand failure isolation: a workload domain VM failure should not impact the management domain, and vice versa.

Step 1

In NSX Manager, create a new segment (logical network) for workload VMs. Navigate to Networking > Segments. Click 'Add Segment'. Name: 'wld-vm-segment', VLAN: 1701 (or other non-overlapping ID), network: 192.168.1.0/24, gateway: 192.168.1.1. Attach to the workload Tier-0 gateway.

Segment created and shows in inventory with Up/Connected status to the Tier-0 gateway
Step 2

In vCenter, navigate to the workload domain cluster > Virtual Machines. Right-click > New Virtual Machine. Create a test VM: Name: 'wld-test-vm-01', Guest OS: Linux (CentOS 8 or Ubuntu 20.04 recommended), CPU: 2, Memory: 4 GB, Storage: place on workload domain vSAN datastore (not management domain).

VM creation wizard shows storage location is the workload datastore, not the management domain datastore
Step 3

Configure the VM's network interface: select the 'wld-vm-segment' created in Step 1. Do NOT use management domain network portgroups.

VM network interface assigned to the wld-vm-segment
Step 4

Complete VM creation. Power on the VM. In vCenter > VM console, wait for the guest OS to boot.

VM powers on and boots the guest OS successfully
Step 5

Verify VM networking: From the guest OS console, run 'ip addr' (Linux) or 'ipconfig /all' (Windows) to confirm the VM has an IP address on the 192.168.1.0/24 segment (e.g., 192.168.1.10). Try to ping the gateway (192.168.1.1) and a non-existent IP outside the segment to confirm isolation.

VM has IP in 192.168.1.0/24 range, can ping gateway, isolated from other segments
Step 6

Verify storage placement: In vCenter, right-click the VM > Summary. Confirm the VM .vmdk files are stored on the workload domain vSAN datastore (not management domain). You can verify by checking Storage > Datastores and seeing which datastore holds the VM folder.

VM storage shown as 'vSAN-WLD' or workload domain datastore, not 'vSAN-Mgmt'
Step 7

Verify DRS placement: In vCenter > cluster > Virtual Machines, the workload VM should be distributed across available hosts (if >1 host in cluster) or pinned to a single host with good consolidation. DRS should not attempt to migrate it to the management domain hosts.

DRS shows VM on a workload domain host, not a management domain host

Validation Gate

Check: vCenter: VM running on workload cluster, IP on 192.168.1.x segment, storage on workload vSAN. NSX: segment connected to Tier-0, VM can reach gateway. No management domain resource contention.

Expected: Test VM fully operational in workload domain with complete isolation from management domain

Common Errors

VM deployment fails with 'insufficient disk space' or 'disk format error'
Cause: Selected the wrong datastore (management domain vSAN instead of workload), or workload vSAN is full
Fix: In VM creation, explicitly select the workload datastore. Check vSAN capacity in vCenter > Cluster > Monitor > vSAN > Capacity. If workload vSAN is full, add a disk to a host in the workload cluster or delete old test VMs.
VM network interface shows 'Device is disconnected' after power-on
Cause: The wld-vm-segment is not connected to the Tier-0 gateway, or the edge cluster is not ready
Fix: In NSX Manager, verify the segment is connected to the Tier-0. Verify edge cluster is Up. Reconnect the segment to the Tier-0 if needed.
Guest OS boots but cannot reach 192.168.1.1 (gateway) — network unreachable
Cause: NSX routing between the workload Tier-0 and the segment is not established, or TEP network is broken
Fix: In NSX Manager, check the Tier-0 > Routing, verify the segment's logical router is attached. Ping the NSX edge cluster's TEP IP (10.1.14.x) from a workload host to confirm TEP connectivity.
VM is placed on a management domain host instead of workload domain host
Cause: VM was created on the wrong cluster (management instead of workload), or DRS settings are allowing cross-cluster affinity
Fix: Verify you are in the correct vCenter cluster (wld-cluster-1, not the management cluster). Recreate the VM ensuring it's on the workload cluster.

Task 5 Workload domain lifecycle operations — scale-out, scale-in, and cleanup

recoverability

Day 2 operations are where real VCDX design thinking emerges. Expanding and contracting a workload domain tests your understanding of cluster consistency, licensing, and the operational maturity required to manage multi-domain environments. Cleanup teaches failure recovery and lab reusability.

Step 1

Scale out: Add a third ESXi host to the workload domain cluster. In SDDC Manager, Inventory > Workload Domains > workload-domain-1 > Add Host. Commission a third host (esxi-wld-03 at 10.1.10.103) if not already done. The host will automatically join the existing workload cluster and vSAN storage.

Third host added, shows in vCenter cluster with 3/3 hosts connected, vSAN storage rebuilt to include third host
Step 2

Monitor vSAN rebalancing: In vCenter > Cluster > Monitor > vSAN > Capacity, watch as the new host's storage is integrated. vSAN will rebalance objects (VMs) across the now-3-node cluster to balance capacity.

vSAN status transitions from 'Rebalancing' to 'Healthy' within 5-15 minutes
Step 3

Scale in: Prepare to remove a workload host. In vCenter, select the workload cluster. Right-click a workload host > Enter Maintenance Mode. Confirm any VM migration warnings (your test VM will migrate to another host).

Host shows 'Maintenance Mode' status in vCenter. Test VM migrated to another host in the cluster.
Step 4

Once in maintenance mode, shut down the host gracefully. In SDDC Manager, you can now decommission it: Inventory > Hosts > (select the host) > Decommission. This removes it from SDDC Manager inventory.

Host decommissioned and removed from inventory. Workload domain now shows 2 hosts.
Step 5

Verify workload domain stability after scale-in: Check vSAN health in vCenter. With 2 hosts and FTT=1, the cluster still has 1 fault tolerance (if host-1 fails, host-2 survives with full data redundancy). This is the key benefit of proper storage architecture design.

vSAN health green with 2 hosts. DRS continues to manage the test VM effectively.
Step 6

Update/patch via SDDC Manager LCM (Lifecycle Management): Navigate to Administration > Lifecycle Management > Updates. Check for any available ESXi patches for the workload domain. In a real lab, you might schedule a cluster-wide patch. For now, just verify the UI shows the workload domain hosts in the LCM inventory.

Workload domain hosts listed under LCM, available patches shown (may be none in a lab with fresh deployment)
Step 7

Delete the workload domain for lab cleanup and reusability: In SDDC Manager, Workload Domains > workload-domain-1 > Delete. Confirm the deletion. This will: (a) Power off and delete the test VM, (b) Destroy the vSAN cluster, (c) Remove hosts from cluster membership, (d) Clean up NSX configuration. Full cleanup takes 10-20 minutes.

Workload domain deletion task progresses through: VM shutdown, vSAN cluster destruction, NSX cleanup, host decommissioning
Domain deletion is irreversible. Ensure you are ready to discard all workload domain data. After deletion, hosts return to 'Unassigned' status in SDDC Manager and can be recommissioned into a new domain or the management domain (not recommended).
Step 8

Verify cleanup: Once deletion completes, confirm in SDDC Manager Workload Domains that workload-domain-1 is gone. In vCenter, the workload cluster should no longer exist. In NSX, workload transport nodes should be in 'Detached' or 'Not Ready' state (edge cluster will be destroyed).

Workload domain removed, cluster gone from vCenter, NSX fabric cleaned up

Validation Gate

Check: After scale-out: 3 hosts in cluster, vSAN healthy. After scale-in: 2 hosts, vSAN still healthy with FTT=1. After deletion: workload domain absent from SDDC Manager, cluster removed from vCenter, hosts returned to Unassigned status.

Expected: All lifecycle operations complete successfully with data integrity and isolation maintained throughout

Common Errors

Scale-out hangs at 'Adding host to vSAN cluster' for >20 minutes
Cause: vSAN cluster is rebalancing from previous operations, or the new host's network/storage is not fully ready
Fix: Check vSAN health in vCenter. Verify the new host's vSAN VMkernel is on VLAN 1651 and can reach the vSAN network (ping 10.1.12.x from other hosts). Increase the operation timeout in SDDC Manager settings if needed.
Scale-in fails with 'Host has virtual machines that cannot be evacuated'
Cause: A VM is pinned to the host with affinity rules, or the VM is in an error state
Fix: In vCenter, check the test VM's affinity rules (right-click > VM Options > vSAN > VM Affinity). Remove any that pin it to the host being removed. Ensure the VM is in a healthy state (not grayed out).
Domain deletion fails partway through with 'Cannot delete NSX edge cluster'
Cause: An NSX service (e.g., Tier-0 router) is still dependent on the edge cluster, or the edge cluster hosts are not responsive
Fix: In NSX Manager, verify there are no additional Tier-0 or Tier-1 gateways using the workload edge cluster. If present, manually delete them before retrying the domain deletion in SDDC Manager.
After deletion, hosts show 'In Maintenance Mode' and cannot be recommissioned
Cause: The host is stuck in maintenance mode after workload domain teardown
Fix: In vCenter, exit maintenance mode for each host (right-click host > Exit Maintenance Mode). Then they can be decommissioned or recommissioned.

Final Validation

The VI workload domain is fully deployed, tested with live workloads, and successfully scaled and torn down. You have demonstrated the complete lifecycle: commissioning hosts, creating a domain with compute, storage, and NSX edge infrastructure, validating health across all components, deploying production-like workloads with isolation, and performing Day 2 operations (scale-out, scale-in, cleanup). This mirrors the real VCDX design decision: 'How do you architect a multi-domain VCF environment that maintains failure isolation while sharing management infrastructure?'

✓ Task 1: Workload hosts commissioned, network pool created → 6-7 total hosts in SDDC Manager, network pool with 4 ranges defined

✓ Task 2: Workload domain created and Active → SDDC Manager shows workload-domain-1 Active with 2-3 hosts, vSAN formed, NSX edge cluster operational

✓ Task 3: All health indicators green (vCenter, NSX, SDDC Manager) → Cluster connected, vSAN healthy, transport nodes Success/Up, edge cluster Up, Tier-0 active

✓ Task 4: Test VM deployed and isolated in workload domain → VM on workload cluster, IP on 192.168.1.0/24, storage on workload vSAN, separate from management domain

✓ Task 5: Lifecycle operations completed (scale-out, scale-in, delete) → 3-host scale successful, 2-host scale-in successful with data protection maintained, clean deletion with no orphaned resources

Cleanup / Restore

Snapshot: holodeck-05-complete

• Workload domain has been deleted as part of Task 5 — no further cleanup needed

• Workload hosts are returned to Unassigned status in SDDC Manager inventory (they can remain for future labs or be decommissioned)

• Network pool 'wld-netpool' remains in SDDC Manager (can be reused for subsequent workload domain labs)

Design Reflection (VCDX)

The VCDX panelist examining a multi-domain VCF architecture would probe these design decisions: (1) Failure isolation: How does a workload domain failure (e.g., vSAN loss) impact the management domain? Answer: zero impact — they are independent vSAN clusters, distinct NSX fabrics, separate vCenter clusters. (2) Resource sharing: The management domain provides SDDC Manager, vCenter, NSX Manager. Can workload domains share these? Answer: yes, by design — one management domain orchestrates multiple workload domains. This is the economy of the VCF architecture. (3) Why not 1 large cluster?

Single cluster is simpler but loses failure isolation. A single vSAN node failure could affect both management and user workloads. Multi-domain enforces the boundary. (4) NSX edge placement: Why must the workload domain have its own edge cluster? Because workload domain VMs require NSX-provided networking, and the edge cluster routes their traffic. Management domain edge cluster is for management traffic only. (5) Licensing: VCF licenses are per-domain. You need 1x management domain license + 1x WLD license per domain. This teaches the cost scaling of multi-domain.

(6) Nested architecture limits: In this lab, all domains share a single physical host. In production, you would have separate physical infrastructure per domain or at least per failure domain boundary. Acknowledge this limitation.

Requirements

  • Deploy a separate VI workload domain that does not impact the management domain if it fails
  • Workload domain must have compute (ESXi cluster), storage (vSAN), and networking (NSX segment with Tier-0 routing)
  • Workload VMs must be isolated from management infrastructure (different network, different storage, different cluster membership)
  • Workload domain must be scalable (add/remove hosts) and must support standard VM operations (vMotion, DRS, HA)

Constraints

  • Nested environment shares single physical host — limits total resource availability and creates resource contention between domains
  • Licensing: each domain requires its own evaluation license (Holodeck includes limited WLD licenses)
  • Network isolation in nested environment enforced by HoloRouter and VLAN tagging — relies on correct VLAN assignment (no cross-domain leakage)
  • vSAN FTT=1 on 2-3 node cluster means any host failure reduces capacity by 50% (teach the difference between resilience and recovery capacity)

Assumptions

  • Management domain is stable and healthy before workload domain deployment (prior lab success)
  • At least 2 uncommissioned ESXi hosts are available for workload domain (Holodeck 2.1.x default provides 4+ hosts)
  • Network pool with distinct VLAN ranges for workload domain is pre-created (prevents VLAN conflicts with management)
  • NSX version on management domain matches the requirements of the workload domain (VCF 9.0.2 bundles NSX 9.0.2 for all domains)

Risks

  • Resource exhaustion on physical host during workload domain deployment causes both management and new workload domain to fail — IMPACT: total lab loss, MITIGATION: monitor physical host CPU/memory during deployment, use snapshot-based rollback
  • Workload domain vSAN formation fails due to disk assignment conflict — IMPACT: wasted 20+ minutes, MITIGATION: pre-verify unformatted disks available on workload hosts via vSphere Client
  • NSX transport node installation times out on workload hosts due to network latency — IMPACT: deployment fails at 40% completion, MITIGATION: increase operation timeout, verify TEP network VLAN is routed and accessible
  • Test VM cannot reach gateway due to Tier-0 connectivity issue — IMPACT: false negative on workload domain validation, MITIGATION: manually verify NSX routing via API health checks, check edge cluster operational status
  • Workload domain deletion leaves orphaned VMs or NSX objects — IMPACT: subsequent domain deployments fail, MITIGATION: use SDDC Manager domain deletion (preferred) which cascades cleanup, avoid manual VM deletion

Self-Assessment Discussion Prompts

  1. If the workload domain vSAN goes offline, can the management domain continue to operate SDDC Manager, vCenter, NSX Manager? Why or why not?
  2. You have a VCF deployment with 1 management domain and 5 workload domains. A single physical ESXi host suffers a storage failure. How many domains are affected? How would you redesign to minimize blast radius?
  3. The lab uses a shared NSX Manager for both management and workload domains. Is this a best practice? What are the trade-offs of a per-domain NSX Manager vs shared?
  4. In production, would you deploy workload domains with vSAN or external storage (SAN/NAS)? What are the architectural differences and VCDX panelist talking points for each?
  5. A workload domain host needs emergency maintenance and must be powered off immediately (no evacuation). What data loss occurs with FTT=1 and 2 hosts? What if FTT=1 and 3 hosts?
  6. Explain the difference between scale-out (adding a host) and scale-in (removing a host) in the context of vSAN FTT and cluster sizing. What is the minimum cluster size you would run in production?

Extensions

Deploy a Second Workload Domain for Multi-Tenant Isolation

Repeat Tasks 1-4 to deploy a second workload domain (e.g., 'workload-domain-2') using the remaining uncommissioned hosts. Compare licensing, storage pool management, NSX edge cluster sprawl, and management overhead for 1x management + 2x workload domains. Document how SDDC Manager inventory, NSX fabric, and vCenter multisite capabilities differ with two workload domains.

same

Configure Cross-Workload-Domain vMotion and Test VM Migration

After deploying a second workload domain, configure vMotion network ranges that allow VM migration between workload domains (both domains share the same vMotion network range, e.g., 10.1.11.0/24). Attempt to vMotion the test VM from workload-domain-1 to workload-domain-2. Document any failures (expected due to storage differences) and the manual steps required to move a VM between domains (storage-level copy, re-IP). This exercises the constraint that vMotion within VCF is domain-local, not cross-domain.

harder

Implement Stretched Workload Domain Across Two Holodeck Pods (Dual-Site)

Configure two Holodeck instances (Site A and Site B) with a management domain each, then create a stretched workload domain spanning both sites. This requires: (a) dual-site HoloRouter networking (cross-site vMotion and vSAN network reachability), (b) stretched vSAN cluster (witness host on Site C or VM-based witness), (c) NSX L2 extension between sites. This is a VCDX advanced design topic and teaches failure domain boundaries, network latency tolerance, and split-brain scenarios. Reference: community thread 'Holodeck 9.0.2 dual-site IP config'.

much harder

Simulate Workload Domain Failure Recovery and Test Backup/Restore

Using the workload domain and test VM, intentionally corrupt the workload vSAN (e.g., mark a disk as failed, offline a host). Observe the failure cascade: (a) workload VMs transition to 'Degraded' state, (b) DRS may migrate them to remaining healthy hosts, (c) management domain remains unaffected. Then recover: restore the host online, re-add the disk to vSAN, verify VM recovery. Additionally, practice VM backup by cloning the test VM to the management domain, demonstrating cross-domain data replication for DR purposes.

harder

Measure Performance and Resource Utilization (Nested vs Production Baselines)

Use vCenter Performance Graphs and SDDC Manager reporting to establish baseline metrics for the workload domain: vSAN IOPS/throughput, vMotion migration time, DRS cluster balancing time, NSX edge cluster datapath throughput. Compare nested Holodeck baselines to production VCF SLAs (if available from Broadcom docs). This teaches the reality that nested architectures have overhead — typically 15-30% performance tax vs bare-metal. Document findings in a VCDX-style design assumption statement.

same

⚠ Known Pitfalls (from Community KB)

Unable to deploy additional Workload Domain RESOLVED
Problem: Workload domain deployment fails when trying to add a second domain after successful management domain deployment. User attempted Deploy-ManagementDomain instead of Deploy-WorkloadDomain. Related issues: host commissioning, network pool overlap, licensing limits.
Resolution: Use Deploy-WorkloadDomain command, not Deploy-ManagementDomain. Ensure network pools do not overlap with management domain ranges. Verify VCF workload domain licenses are available. Pre-validate host readiness and network connectivity before deployment.
Holodeck vcf 9 - Cannot select datastore RESOLVED
Problem: Workload domain host provisioning or cluster formation fails because the datastore is not accessible, full, or has naming conflicts. This is often the first deployment failure point.
Resolution: Verify physical datastore exists and is accessible from SDDC Manager. Ensure sufficient free space (>500 GB for workload cluster). Check that workload domain hosts have unformatted disks available for vSAN. Use vSphere Client to pre-inspect storage on target hosts.
Waiting for FRR service (HoloRouter BGP timeout during Prepare phase) LIKELY_RESOLVED
Problem: Prepare phase hangs waiting for HoloRouter FRR (BGP routing daemon) to become ready. Related to workload domain networking setup which requires HoloRouter to be stable.
Resolution: This issue is primarily in the Prepare phase (Task 1 of holodeck-02), but workload domain deployment depends on healthy HoloRouter. If FRR hangs, restart the service: SSH HoloRouter, systemctl restart frr.
Holodeck 9.0.1 VCF Deployment fails at NSX validation RESOLVED
Problem: NSX bundle version mismatch between management domain and workload domain. Workload domain deployment attempts to validate NSX but fails due to incompatible version.
Resolution: Ensure NSX 9.0.2.0 bundle is available before workload domain deployment. Update content manifest if using offline depot. Pre-validate NSX Manager is at correct version (9.0.2.x) before attempting WLD deployment.
VCF Installer is not ready yet (Cloud Builder initialization timeout) RESOLVED
Problem: Workload domain deployment may hang at NSX or bundle download steps if VCF Installer (in Cloud Builder) is not responding. Related to management domain health, not direct workload domain issue, but affects dependent operations.
Resolution: Verify management domain Cloud Builder is healthy. Check DNS resolution: nslookup sddc-manager.site-a.vcf.lab. Verify NTP sync on Cloud Builder. Increase operation timeout if deploying on resource-constrained physical host.
External Jumpserver Access — DNS and Network Configuration RESOLVED
Problem: Workload domain created but external access to management UIs (for validating WLD health) is blocked by DNS or network routing issues.
Resolution: Configure DNS forwarder to resolve *.site-a.vcf.lab to HoloRouter. Add static routes on workstation to reach nested networks. Use FQDNs (not IPs) for browser access to avoid certificate validation errors. This is often overlooked during multi-domain validation.

References

  • VMware Cloud Foundation 9.0 Workload Domain Planning and DeploymentTier 1 — Official
    Official Broadcom documentation — the authoritative reference for workload domain architecture, component sizing, and deployment workflow. Sections on cluster membership, vSAN design, NSX edge placement, and multi-domain resource sharing are critical.
  • VCF Holodeck Toolkit — Broadcom Community ForumTier 1 — Official
    Community forum with 63+ threads. Thread 52 'Unable to deploy additional Workload Domain' and related threads document host commissioning, licensing, and network pool configuration pitfalls. Threads 3, 5, 23, 37, 51 provide WLD-specific troubleshooting.
  • VCF Holodeck GitHub Repository — Issues and DiscussionsTier 1 — Official
    Official issue tracker. Issues related to workload domain deployment, scale-out, and lifecycle management are logged here. Useful for understanding edge cases and planned features.
  • Cormac Hogan — vSAN in VCF Workload DomainsTier 3 — Expert Blog
    Expert blog covering vSAN architecture in multi-domain deployments, stretched clusters, and performance tuning. Highly recommended for VCDX candidates preparing for storage-focused design questions.
  • William Lam — VCF Automation and Day 2 OperationsTier 3 — Expert Blog
    Deep dives on VCF lifecycle management, workload domain scaling, and automation. Covers host commissioning, NSX edge cluster deployment, and troubleshooting common gotchas in nested and production environments.
  • NSX 9.0 Design Guide — Multi-Domain ArchitectureTier 1 — Official
    NSX design considerations for VCF, including per-domain edge clusters, Tier-0 and Tier-1 hierarchy, and segment placement across domains. Critical for understanding NSX isolation in multi-domain.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.