Academy/VCP-VCF 9.0 Administrator (2V0-17.25)/Workload Domain Provisioning — Traditional vs Stretched
This lab targets VCF 9.0

Workload Domain Provisioning — Traditional vs Stretched

VCF 9.0Advancedadminarchitectvcdx⏱ 180 min

Workload domain provisioning architecture applies to VCF 9.0.0 through 9.0.2. Comparison with VCF 5.2 available in Extension 1. Focus is on traditional single-site WLD vs stretched cluster (multi-site/multi-zone) topology with witness hosts and fault domains.

Objectives

  • Plan a workload domain deployment: network pools, host commissioning prerequisites, licensing, and resource capacity analysis
  • Provision a traditional (single-site) workload domain using SDDC Manager and VCF Operations Manager, including NSX T0/T1 gateway configuration
  • Understand stretched cluster architecture: witness host deployment, fault domains, quorum model, and cross-site network requirements
  • Provision a stretched workload domain with multi-site replication and validate witness host quorum
  • Execute workload domain lifecycle operations: add hosts, expand cluster, shrink cluster, and validate placement policies
  • Compare traditional and stretched domain trade-offs (RTO, complexity, failure domain, cost)

Prerequisites

VCF 9.0.x management domain deployed and operational. SDDC Manager accessible. 8+ uncommissioned ESXi hosts available in two physical sites (4 hosts per site for stretched topology, or 4+ hosts in single site for traditional). NSX Manager and Cloud Builder operational. Lab account must have SDDC Manager Administrator role. Network infrastructure must support multiple VLAN segments for host management, vMotion, vSAN, and edge TEP networks.

Prior labs: holodeck-02

Required skills:

  • VCF management domain architecture (SDDC Manager, vCenter, NSX Manager)
  • ESXi host commissioning and network configuration (VMkernel, vMotion, vSAN)
  • NSX-T fundamentals: Tier-0 and Tier-1 gateways, segments, overlay networking
  • vSAN design: fault domains, stretched cluster quorum, witness host role
  • Network topology design: IP planning, VLAN allocation, BGP peering with NSX edges
  • VCDX-level understanding of failure domains, availability zones, and disaster recovery models

Lab Environment

VCF 9.0.x management domain with 8-12 commissioning-ready ESXi 8.0.x hosts. For traditional WLD lab: 4 hosts in single site. For stretched WLD lab (extension): 4 hosts in Site A, 3 hosts in Site B, 1 witness host. All hosts have: iLO/IPMI for out-of-band management, 10Gbps+ network connectivity, vSAN-capable storage (local NVMe), and configured for vMotion/vSAN trunk port groups.

Traditional WLD: graph TB
  SDDC[Management Domain]
  SDDC -->|vCenter Cluster 1| WLD-A[Workload Domain A: 4 hosts, vSAN]
  SDDC -->|Network| T0A[NSX T0 Gateway]
  T0A -->|Uplink| PAS[Physical Access Switch]

Stretched WLD: graph TB
  SDDC[Management Domain]
  SDDC -->|vCenter Stretched| WLD-S[Stretched WLD: Site A 4h + Site B 3h + Witness]
  SDDC -->|Network L3| Site-A[Site A Network]
  SDDC -->|Network L3| Site-B[Site B Network]
  Site-A -->|Witness Isolation| Witness[Witness Host]
  Site-A -->|vSAN Replication| Site-B

IP Addressing

NetworkPurposeVLAN
ManagementManagement domain (SDDC Manager, vCenter, NSX Manager)1644
Workload Host ManagementWLD ESXi host VMkernel (one VLAN per host group)1650-1659
vMotionvMotion VMkernel (for live migrations within WLD)1660-1669
vSANvSAN RDMA network (for disk group replication)1670-1679
NSX Host TEPHost Tunnel End Point (TEP) for NSX overlay1680-1689
NSX Edge TEPEdge node TEP for edge-based services1690-1699
Witness Isolation (Stretched only)Witness host isolated from data traffic — read-only vSAN participation1710

Credentials

SystemUsernamePassword
SDDC Manageradministrator@vsphere.localSet during VCF bring-up
ESXi Hosts (to be commissioned)rootSet during host onboarding
vCenter (new WLD vCenter if applicable)administrator@vsphere.localSet during WLD provisioning
NSX ManageradminSet during management domain bring-up

Tasks

Task 1 Plan and prepare for workload domain provisioning

manageability

Pre-deployment planning is where VCDX candidates demonstrate rigor. You must validate that the infrastructure is ready before initiating a 2+ hour provisioning operation. This mirrors production change management: plan, validate, execute, document.

Step 1

In SDDC Manager, navigate to Administration > Inventory. Review the list of all ESXi hosts that are NOT yet assigned to a workload domain. Verify each host status: should be 'Commissioned' and 'Connected'. You need 4+ uncommissioned hosts for this lab (or 8+ if doing stretched topology extension).

ESXi hosts list shows: esxi-05, esxi-06, esxi-07, esxi-08 with status Commissioned and Connected. (Or 8 hosts if using multiple sites.)
If hosts show 'Disconnected' or 'Commissioning in Progress', wait for commissioning to complete (or manually commission them: Administration > Inventory > Host Actions > Commission).
Step 2

Verify network infrastructure readiness for the WLD. For each proposed ESXi host in the WLD, validate: (a) Port group for host management (VMkernel) exists and is on dedicated VLAN, (b) Port group for vMotion exists, (c) Port group for vSAN exists, (d) Port group for NSX Host TEP exists, (e) All port groups are VLAN-enabled and do NOT have vLAN trunking misconfigured. Document the VLAN and IP range for each function.

Network readiness checklist completed. Example: 'Host Management VLAN 1650 (10.2.0.0/24), vMotion VLAN 1660 (10.2.1.0/24), vSAN VLAN 1670 (10.2.2.0/24), NSX TEP VLAN 1680 (10.2.3.0/24). All port groups configured on physical switches. VLAN trunking verified.'
VLAN misconfiguration is the #1 reason WLD provisioning fails at the host commissioning step. Verify this before starting.
Step 3

Create a workload domain specification document. Define: (a) WLD Name (e.g., 'Workload-Domain-01'), (b) Compute Cluster Name (e.g., 'WLD-Cluster-01'), (c) Host list (esxi-05, esxi-06, esxi-07, esxi-08), (d) vSAN configuration (Single-Fault Tolerance FTT=1, or Dual-Fault FTT=2?), (e) vSAN network pool (which segment for vSAN RDMA?), (f) vCenter credentials (will reuse management vCenter or deploy WLD-specific vCenter?), (g) NSX Tier-0 gateway (which uplink VLAN and BGP peers?), (h) License tier (Standard, Enterprise, or Ent Plus?).

Workload domain specification document (1-2 pages) with all parameters defined. Example snippet: 'WLD Name: Prod-Zone-A, Host Count: 4, vSAN FTT: 1 (3-2-1 layout on 4 nodes means 1 failure tolerated), License: VCF Enterprise for vSAN Analytics.'
Step 4

Validate licensing. In SDDC Manager, navigate to Administration > System Configuration > Licensing. Check current licenses: total seats, available seats, and per-workload-domain capacity. Ensure you have enough licenses for 4 hosts + overhead. Document any license constraints (e.g., 'License pool allows 8 hosts, planning 4 for WLD-01 and 4 for WLD-02 — at capacity').

Licensing status: 'Total VCF Enterprise licenses: 8 hosts. Available: 4 (after management domain: 4). Planning 4 hosts for WLD-01 — licenses: OK.'
Step 5

Validate infrastructure capacity. Calculate resource overhead: (a) vCenter Server VM (Small: 4 vCPU, 21 GB RAM) per WLD or shared with management, (b) NSX Edge cluster (typically 2-4 edge nodes × 4 vCPU, 8 GB RAM each), (c) vSAN witness appliance (if stretched, 2 vCPU, 8 GB RAM). Total the resources and verify against available physical capacity. In SDDC Manager, Inventory > System Capacity, check total host resources available.

Capacity calculation: 'WLD vCenter: 4C/16G, NSX Edges (optional): 8C/32G, Total workload capacity on 4 hosts: 128 vCPU, 384 GB RAM (gross). After SDDC Manager, vSAN, and overhead: ~90 vCPU, 250 GB RAM available for workload VMs. Document this for sizing discussions.'
Step 6

Create a provisioning checklist: (1) All 4 ESXi hosts status 'Commissioned' and 'Connected', (2) Network VLANs verified on physical switches and port groups created on each host, (3) Licensing capacity confirmed, (4) vSAN datastore pool planned (which disk groups per host?), (5) NSX T0 gateway uplink VLAN and BGP peering planned, (6) vCenter specifications (shared or WLD-specific?), (7) Risk mitigation: snapshot management domain before provisioning. Walk through each item. Check off as complete.

Provisioning checklist with 7 items, all verified and signed off. Example final item: 'Pre-provisioning snapshot taken: md-before-wld-01-provisioning. Rollback path clear.'

Validation Gate

Check: Planning documents complete: WLD specification, networking validation, licensing check, capacity calculation, provisioning checklist. All items signed off.

Expected: Pre-provisioning readiness confirmed. No blockers identified. Ready to proceed to Task 2 (traditional WLD provisioning).

Common Errors

VLAN mismatch during commissioning: ESXi-05 cannot reach vCenter on the planned management VLAN
Cause: Port group on the physical switch is not trunked for the planned VLAN, or VLAN ID is misconfigured
Fix: Before WLD provisioning, verify each VLAN is operational: ping from physical switch to the vCenter management IP on each VLAN. If VLAN is missing, add port group to physical switch and configure VLAN trunking.
Insufficient licenses: 'Cannot provision WLD — license capacity exceeded'
Cause: License pool is exhausted or licenses are expired
Fix: In SDDC Manager licensing, check expiration dates. If expired, request new licenses from Broadcom. If pool is exhausted, scale back WLD host count (use 3 hosts instead of 4) or defer WLD-02 provisioning.
Cannot commission hosts: 'Host esxi-05 network connectivity failed'
Cause: ESXi host does not have IP connectivity to SDDC Manager on the management network
Fix: SSH to the physical ESXi host. Verify: (1) Management VMkernel has an IP in the expected range (e.g., 10.2.0.5), (2) Default gateway is set and reachable, (3) Ping to SDDC Manager IP succeeds (ping 10.0.0.4). If not, reconfigure VMkernel networking on the host.

Task 2 Provision a traditional (single-site) workload domain

availability

Execute the WLD provisioning workflow using SDDC Manager. You'll orchestrate host commissioning, cluster creation, NSX deployment, and vCenter integration — all automated by SDDC Manager. Understand the sequence and where human intervention is required.

Step 1

In SDDC Manager, navigate to Workload Domains > Create New Workload Domain. Fill in: (a) Name: 'Workload-Domain-01' (or your chosen name), (b) Cluster Name: 'WLD-Cluster-01', (c) Type: 'VI Workload Domain' (traditional vSphere cluster), (d) Network Pool: select the pool with IP ranges for host management, vMotion, vSAN, and NSX TEP (created in Task 1).

WLD creation form displays with your specified values
Step 2

In the Hosts section, select the 4 ESXi hosts to be commissioned into this WLD (esxi-05, esxi-06, esxi-07, esxi-08). Verify each host shows status 'Ready for Commissioning'. Configure per-host properties: (a) Host name (FQDN), (b) Management VMkernel IP (will be auto-assigned from network pool), (c) vMotion IP, (d) vSAN IP, (e) NSX TEP IP.

Host list shows 4 hosts with auto-assigned IPs for each network. Example: 'esxi-05: Mgmt 10.2.0.5, vMotion 10.2.1.5, vSAN 10.2.2.5, TEP 10.2.3.5'
IP auto-assignment depends on the network pool configuration. If IPs look incorrect (e.g., all hosts get .5), the network pool VLAN ranges are misconfigured — go back and fix Task 1 network validation.
Step 3

Configure vSAN for the cluster. Select: (a) vSAN Fault Tolerance: 'FTT=1 (RAID-1)' for lab (or FTT=2 if larger cluster). (b) vSAN Network: select the vSAN VLAN/segment. (c) Disk Groups: specify which local disks on each host will be vSAN storage. In a Holodeck lab, each ESXi host may have 1 NVMe drive allocated for cache and 1 SSD for capacity. (d) Deduplication: disable for lab (use for larger production environments).

vSAN configuration: 'FTT=1, 4 hosts × 1 disk group = 4-node vSAN cluster, Disk layout: 1x cache (NVMe), 1x capacity (SSD) per host'
Step 4

Configure NSX T0 Gateway for the WLD. (a) Name: 'WLD-T0-01'. (b) Uplink Profile: select NSX-managed T0 or Edge-based T0 (Edge-based is more common for WLDs). (c) Uplink VLAN: which VLAN will edge nodes use to reach the external router/ToR? (d) BGP Peering: enable BGP and specify: BGP neighbor IP (physical router), BGP neighbor ASN, local ASN for this WLD. (e) Tier-1 segments: configure default T1 for tenant overlay traffic.

NSX T0 gateway configured: 'T0 Name: WLD-T0-01, Uplink VLAN 1700 (BGP 10.100.0.1 AS 65001), Local ASN 65002, Edge Transport Nodes: not yet deployed (optional for WLD)'
Step 5

Specify vCenter configuration for this WLD. Option 1: Reuse management vCenter (simplifies operations, vCenter scales to manage both management domain and WLD clusters). Option 2: Deploy WLD-specific vCenter (more isolation, higher licensing cost). For this lab, choose Option 1 (Reuse management vCenter). SDDC Manager will register the new cluster to the existing vCenter.

vCenter configuration: 'Reuse management vCenter (10.0.0.6). New cluster WLD-Cluster-01 will be registered to this vCenter.'
Step 6

Review the WLD provisioning summary. SDDC Manager displays the complete provisioning plan: Hosts, Networks, vSAN, NSX, vCenter. Verify all settings are correct. Click 'Validate' to run pre-flight checks. Watch for any validation failures (VLAN not reachable, IP conflicts, insufficient capacity). If validation passes, click 'Provision'.

Pre-flight validation passes. Provisioning begins. Progress indicator shows: (1) Host commissioning, (2) VMkernel network configuration, (3) vSAN cluster formation, (4) NSX fabric integration, (5) vCenter cluster registration. Estimated time: 45-75 minutes.
Step 7

Monitor provisioning progress in SDDC Manager. Navigate to Workload Domains > Workload-Domain-01 > Status. Watch the detailed task list. Key milestones: (a) 'Commissioning ESXi hosts' (5-10 min per host), (b) 'Creating vSAN cluster' (10-15 min), (c) 'Configuring NSX transport nodes' (5-10 min per host), (d) 'Registering cluster with vCenter' (2-3 min), (e) 'Final validation' (5 min).

Progress bar shows ~50% after 25 minutes, ~100% after 60 minutes. No errors in the task log.
If provisioning stalls at 'Commissioning ESXi hosts', check: (1) Network connectivity from SDDC Manager to the host (ping ESXi management IP), (2) SSH is enabled on the ESXi host, (3) Root credentials are correct. If stalled at 'Creating vSAN cluster', check: (1) vSAN licenses applied, (2) All hosts are healthy (no disk errors), (3) Network latency on vSAN VLAN is <1ms.

Validation Gate

Check: SDDC Manager: Workload Domains > Workload-Domain-01 status shows 'Active'. vCenter: navigate to https://10.0.0.6, login, verify new cluster 'WLD-Cluster-01' appears with 4 hosts connected. ESXi hosts status: 'Connected'. vSAN status: 'Healthy' (green). NSX: verify 4 transport nodes in Fabric > Nodes.

Expected: Traditional WLD fully provisioned and healthy. 4 ESXi hosts connected, vSAN cluster operational, NSX fabric integrated, vCenter managing the cluster.

Common Errors

Provisioning fails at 'Commissioning ESXi hosts': 'Cannot SSH to esxi-05'
Cause: Network connectivity issue or SSH keys not configured
Fix: Verify ESXi host is reachable from SDDC Manager. SSH manually from SDDC Manager: ssh root@10.2.0.5. If that works, verify SDDC Manager has the correct SSH key for the host. If SSH fails, check: (1) ESXi IP is on the correct management VLAN, (2) Firewall allows SSH port 22 from SDDC Manager network, (3) ESXi host is powered on.
vSAN cluster creation fails: 'Disk group initialization failed on esxi-06'
Cause: ESXi host disk is in use or not properly formatted
Fix: SSH to esxi-06. Check vSAN disk status: vsan cluster get. If a disk is in use, wipe it: vsan cluster leave --force. Then retry vSAN cluster formation from SDDC Manager.
NSX transport node installation fails: 'VIB installation timeout on esxi-07'
Cause: ESXi host is resource-starved or network latency is high
Fix: Check ESXi-07 resource usage in vCenter: CPU, memory, network utilization. If CPU >80%, defer commissioning other hosts. Check network latency: ping from SDDC Manager to esxi-07 management IP — latency should be <10ms. If >50ms, NSX VIB installation will timeout.

Task 3 Validate workload domain placement and network connectivity

manageability

A provisioned domain is not yet production-ready. You must validate that workloads can be placed, networks are segmented correctly, and edge nodes can route to external networks. This task covers the post-deployment checklist.

Step 1

In vCenter, navigate to Hosts and Clusters. Click on the new 'WLD-Cluster-01' cluster. Verify: (a) 4 hosts are connected and green, (b) No alarms on the cluster, (c) Resource summary shows total CPU and memory available.

Cluster summary: '4 hosts connected, 256 vCPU total, 768 GB RAM total (gross). vSAN: 16 GB cache, 400 GB capacity per host (example). DRS: Enabled. HA: Enabled with admission control.'
Step 2

Validate vSAN health. In vSAN, navigate to Monitor > vSAN > Cluster > Health. Check: (a) Object health: all objects Compliant (no degraded), (b) Disk health: all disk groups Healthy, (c) Network latency: <1ms between hosts (shown in latency matrix).

vSAN health status: 'Cluster: Healthy, 4/4 nodes, 16 components, 100% compliance. Latency: <1ms. Disk space: 1.2 TB capacity, 800 GB committed, 400 GB free.'
Step 3

Test workload placement. Deploy a test VM into the new cluster. In vCenter, create a new VM: (a) Name: 'wld-test-01', (b) Datastore: select vSAN datastore of WLD-Cluster-01, (c) Network: select an NSX segment (if Tier-1 was created during provisioning) or default VLAN, (d) Size: 2 vCPU, 8 GB RAM. Power on the VM.

VM is created, powered on, and receives an IP address. VM status in vCenter: 'Running' on one of the 4 WLD hosts.
Step 4

Validate NSX connectivity. In NSX Manager, navigate to Fabric > Nodes > Host Transport Nodes. Verify all 4 WLD hosts appear and show Configuration State: 'Success' and Transport Node Status: 'Up'. Click on each host to view: (a) TEP IP address (should be from the NSX TEP pool, e.g., 10.2.3.5), (b) VLAN associations (Management, vMotion, vSAN, TEP), (c) No connectivity warnings.

NSX fabric: '4 host transport nodes, all Success/Up. TEP IPs: 10.2.3.5-10.2.3.8. No warnings.'
Step 5

Validate NSX Tier-0 gateway. In NSX Manager, navigate to Networking > Tier-0 Gateways > WLD-T0-01. Verify: (a) Gateway status: 'Realized' (not Error), (b) Uplink segment: has an IP and is connected to physical network, (c) BGP status: neighbor relationship is 'Established' (if BGP is enabled), (d) Tier-1 gateways are attached and have routes.

NSX T0 gateway: 'Status: Realized. Uplink: 10.100.0.2 on VLAN 1700. BGP Neighbor 10.100.0.1 AS 65001: Established. 2 T1 gateways attached.'
Step 6

Test end-to-end connectivity. From the test VM (wld-test-01), ping an external IP (e.g., SDDC Manager 10.0.0.4 or a physical router 10.100.0.1). If the ping succeeds, NSX routing and the WLD network integration are working end-to-end.

Ping from wld-test-01 to 10.0.0.4: replies received. Round-trip latency <5ms. Network path: VM -> NSX segment -> T1 -> T0 -> physical uplink -> external network.

Validation Gate

Check: vCenter: WLD-Cluster-01 shows 4 hosts, resources available, no alarms. vSAN: all objects compliant, network latency <1ms. NSX: 4 transport nodes Success/Up, T0 gateway Realized, test VM connectivity working.

Expected: WLD is validated and production-ready for workload placement. All infrastructure components (compute, storage, network) verified.

Common Errors

Test VM cannot reach external network — ping to 10.100.0.1 times out
Cause: NSX T0 gateway is not connected to the physical uplink, or BGP neighbor is down
Fix: In NSX Manager, check T0 gateway: (1) Uplink segment status is 'Realized', (2) BGP neighbor is 'Established'. If neighbor is not established, check: (a) Physical router at 10.100.0.1 is reachable from NSX edge node, (b) BGP AS numbers match the configuration, (c) Edge node TEP IP is on correct VLAN.
vSAN cluster shows 1 host in 'degraded' state
Cause: One host lost network connectivity or has disk issues
Fix: Navigate to vSAN health and click the degraded host. Check: (1) Host is connected in vCenter, (2) Network latency to other hosts is normal, (3) Disk group health is green. If degraded, host may need to be removed and re-provisioned.

Task 4 Understand stretched cluster architecture and plan multi-site deployment

availability

Stretched clusters are the VCDX-level topology: multiple sites, witness quorum, fault domains, cross-site replication. This task is architecture and planning; the actual stretched WLD deployment is Extension 1. Understanding the design constraints is critical for VCDX panelists.

Step 1

Review stretched cluster quorum model. In your lab notes, document: (a) Stretched cluster components: Site A (4 ESXi hosts), Site B (3 ESXi hosts), Witness Host (1 host, isolated on separate VLAN). (b) Quorum calculation: (4 + 3 + 1) = 8 nodes. Quorum = 4 nodes (majority). (c) Failure scenarios: If Site A loses 2 hosts, it has 2 + 1 (witness) = 3 (< quorum) and becomes read-only. If Site B loses 2 hosts, it has 1 + 1 (witness) = 2 (< quorum) and becomes read-only. (d) Write winning site: the site with >50% of active nodes wins the quorum and can write. Document this in a table.

Quorum table: 'Site A: 4 nodes, Site B: 3 nodes, Witness: 1 node. Scenarios: (1) Site A down (network partition): Site B + Witness = 4 nodes (quorum, Site B wins), (2) Site B down: Site A + Witness = 5 nodes (quorum, Site A wins), (3) Both sites equally impaired (2+2): Witness breaks tie, winner determined by witness locality.'
Step 2

Understand witness host role and placement. The witness host does NOT participate in I/O — it only votes in quorum decisions. Placement: must be in a third location (third data center) or isolated network to avoid being partitioned along with one of the primary sites. In the lab, the witness can be on the same physical host as management, but on an isolated VLAN to simulate geographic separation. Document: (a) Witness host role: storage quorum voting, (b) Witness network isolation: separate VLAN with no data traffic, (c) Witness disk: typically a small disk (10-50 GB) used only for quorum metadata, not data storage.

Witness host architecture documented: 'Witness is a single vSAN-enabled ESXi host deployed on VLAN 1710 (isolated). It does not replicate data (no disk groups in the vSAN cluster). It votes on quorum only. If witness goes down, quorum is maintained as long as 4 of the 8 remaining nodes (4 Site A + 3 Site B) are up.'
Step 3

Design the stretched cluster network topology. For a 2-site stretched cluster: (a) Intra-site network latency: <1ms (within a data center, same network switch), (b) Inter-site network latency: <10ms (acceptable for vSAN replication), (c) Inter-site bandwidth: minimum 10 Gbps dedicated (25 Gbps recommended), RTT ≤5 ms (vSAN Stretched Cluster Guide). (d) Failure scenarios: if inter-site link fails, both sites enter read-only mode until the link is restored (network partition). Document the network design.

Network topology: 'Site A (10.2.0.0/24 management, 10.2.3.0/24 vSAN) <--L3 Gateway--> Site B (10.3.0.0/24 management, 10.3.3.0/24 vSAN). Inter-site links are redundant (2x 10 Gbps). Witness is on isolated VLAN 1710 at Site A. vSAN replication bandwidth: minimum 10 Gbps dedicated, RTT ≤5 ms.'
Step 4

Understand fault domains. In a stretched cluster, fault domains are used to ensure replicas are spread across sites. Configure: (a) Site A hosts assigned to Fault Domain A, (b) Site B hosts assigned to Fault Domain B, (c) Witness assigned to Fault Domain C. (d) Replication policy: vSAN enforces that a copy of data resides in each fault domain (RAID-1 across sites). Document the placement policy.

Fault domain configuration: 'WLD-Stretched-01: FD-A (4 hosts Site A), FD-B (3 hosts Site B), FD-C (Witness). vSAN replication policy: all objects have 1 copy in FD-A and 1 copy in FD-B (RAID-1 across sites). If FD-A entire site fails, FD-B still has all data.'
Step 5

Compare traditional vs stretched tradeoffs. Create a comparison table: Traditional (single-site WLD) vs Stretched (2-site WLD). Rows: RTO (recovery time objective if site fails), RPO (recovery point objective — data loss), complexity (operational overhead), cost (extra site, WAN bandwidth), failure scenarios.

Comparison table: Traditional WLD RTO: Re-provision from snapshot if site is destroyed (2-4 hours), RPO: depends on backup frequency (0-24 hours). Stretched WLD RTO: immediate failover to Site B (< 1 minute), RPO: zero (synchronous replication). Complexity: Traditional = low, Stretched = high (quorum management, fault domains, witness). Cost: Traditional = low, Stretched = high (2x infrastructure, WAN, witness). Document which use case each is appropriate for.
Step 6

Document the provisioning steps for stretched WLD (you will execute in Extension 1 if time permits). Steps: (1) Plan fault domains and network topology, (2) Commission hosts from both sites, (3) Configure inter-site network (L3 reachability, latency validate), (4) Deploy witness host on isolated VLAN, (5) Provision stretched WLD specifying fault domain assignments, (6) Configure vSAN stretch sync settings (e.g., change rebuild timeout from 30 min to 60 min for remote replication), (7) Validate quorum with network partition test (disable inter-site link, confirm read-only state, re-enable link, confirm quorum recovery).

Stretched WLD provisioning checklist with 7 steps documented and ready for execution

Validation Gate

Check: Quorum model explained with failure scenarios. Witness host role and placement documented. Network topology designed with latency/bandwidth specs. Fault domains defined. Traditional vs stretched comparison table complete. Stretched provisioning checklist ready.

Expected: Comprehensive understanding of stretched cluster architecture and design constraints. Ready to architect or deploy a stretched WLD in production.

Common Errors

Confusion: does the witness participate in data storage or just voting?
Cause: Witness role is often misunderstood — it's easy to think it replicates data
Fix: Remember: Witness = quorum voting only (no data). It has a minimal disk (metadata only). This is why it can be deployed anywhere without capacity concerns.
Network partition scenario unclear: what happens if the inter-site link fails?
Cause: The two sites become independent and can't agree on quorum
Fix: When inter-site link is down: Site A (4 nodes + witness) = 5 nodes, quorum met, can write. Site B (3 nodes) = 3 nodes, no quorum, becomes read-only. The witness is co-located with Site A, so Site A 'wins' the partition.

Task 5 Execute workload domain lifecycle operations: add hosts, expand cluster, and validate placement

manageability

Post-deployment, WLDs require day 2 operations: adding hosts to scale out, removing hosts for decommissioning, expanding storage. This task covers the operational commands and validation steps.

Step 1

Plan to add a 5th host to the traditional WLD (Workload-Domain-01). Ensure the host (esxi-09) is commissioned and ready. In SDDC Manager, navigate to Workload Domains > Workload-Domain-01 > Edit. In the Host List section, click 'Add Host'. Select esxi-09. Assign the same network IPs as the other hosts (management, vMotion, vSAN, TEP — auto-assigned from pool).

Host esxi-09 added to the WLD cluster specification. Status: 'Ready to Provision'.
Step 2

Click 'Save' to start the host addition operation. SDDC Manager will: (1) Commission esxi-09 (similar to initial commissioning), (2) Join it to the vSAN cluster, (3) Register transport node with NSX. Monitor progress in Workload Domains > Workload-Domain-01 > Status. Estimated time: 15-30 minutes.

Host addition completes. esxi-09 now appears in the cluster with status 'Active'. vSAN cluster resizes from 4 to 5 nodes.
While a new host is being added to vSAN, existing data is not automatically rebalanced. It remains on the original 4 hosts. To rebalance, you must explicitly request a vSAN rebalance (available in vCenter > vSAN > Rebalance). Rebalancing can take hours on large clusters.
Step 3

In vCenter, navigate to Hosts and Clusters > WLD-Cluster-01. Verify the 5th host (esxi-09) is now connected and green. Check vSAN health: it should show '5 nodes, 20 components (5 x 4 original)' — data is still on the original 4 nodes until rebalance is triggered.

vCenter cluster shows 5 hosts. vSAN: 5 nodes active, components still on original 4 hosts (unbalanced).
Step 4

Trigger a vSAN rebalance to distribute the data across all 5 hosts. In vCenter, navigate to vSAN > Monitor > Cluster > Rebalance. Click 'Start Rebalance'. Watch the progress. This operation: (1) Reads data from original 4 hosts, (2) Redistributes to the 5th host, (3) Maintains fault tolerance (FTT=1, so 1 failure tolerated at all times). Estimated time: 30-60 minutes depending on data volume.

Rebalance begins. Progress shows percentage of data migrated. Example: '45% migrated, estimated completion time: 25 minutes.'
Step 5

While rebalance is in progress, test workload placement policies. In vCenter, create a VM Storage Policy (in the Workload-Domain-01 vCenter): (a) Name: 'vSAN-FTT1-Balanced', (b) Fault Tolerance Objects: RAID-1 (FTT=1), (c) Storage Placement: prefer all 5 hosts equally. Apply this policy to the test VM (wld-test-01). vCenter will now manage VM placement to spread data evenly.

VM Storage Policy created and applied. When rebalance completes, VM data is distributed across all 5 hosts according to the policy.
Step 6

Validate the expanded cluster. In vSAN Monitor > Cluster > Health, confirm: (a) All 5 nodes are active, (b) No compliance violations, (c) Rebalance completed (100% migrated), (d) Data is spread evenly across 5 hosts (each host has 20% of data). Document the final state.

Final validation: 'WLD-Cluster-01 now has 5 hosts. vSAN health green. All objects distributed across 5 hosts. Capacity increased to 500 GB (5 x 100 GB per host example). Ready for additional workloads.'

Validation Gate

Check: SDDC Manager: Workload-Domain-01 shows 5 hosts. vCenter: 5 hosts connected in WLD-Cluster-01. vSAN: 5 nodes active, rebalance complete, data distributed evenly.

Expected: WLD scaled out successfully. Day 2 operations (add host, rebalance) demonstrated and validated.

Common Errors

Host addition fails: 'Cannot commission esxi-09 — network unreachable'
Cause: esxi-09 management IP is not on the expected VLAN or gateway is not configured
Fix: Verify esxi-09 VMkernel is on the same management VLAN as the other WLD hosts. Check IP address: esxcli network ip interface list. If wrong VLAN, reconfigure the VMkernel on the host before attempting to add to the WLD.
vSAN rebalance is stuck at 45% for 30+ minutes
Cause: Network congestion on vSAN VLAN or insufficient I/O bandwidth
Fix: Check vSAN VLAN network latency: ping between hosts should be <1ms. If >5ms, check for network congestion (other large transfers). Rebalance is I/O heavy — if cluster is running production workloads, they may be slowing rebalance. Consider running rebalance during off-hours.
After adding 5th host, vSAN compliance shows 'No compliance' for some VMs
Cause: VMs were created with older storage policy that specified exact host count (e.g., 'must be on hosts 1-4 only')
Fix: Update VM storage policies to include the new host (e.g., 'Preferred hosts: all'). Re-apply policy to non-compliant VMs. This will trigger a rebalance of those VMs.

Final Validation

You have planned and provisioned a traditional workload domain, validated network connectivity and placement, understood stretched cluster architecture including witness quorum and fault domains, and executed day 2 operations (host addition and rebalance). You can now architect workload domain topologies at the VCDX level, making data-driven decisions between traditional and stretched topologies based on RTO, RPO, cost, and complexity tradeoffs.

✓ Pre-provisioning Planning → Spec document, network validation, licensing check, capacity calculation, provisioning checklist — all complete

✓ Traditional WLD Provisioning → Workload-Domain-01 deployed with 4 hosts, vSAN healthy, NSX fabric integrated, vCenter managing cluster

✓ WLD Validation → vSAN cluster compliant, NSX transport nodes Success/Up, test VM deployed and connected, T0 gateway realized, end-to-end connectivity verified

✓ Stretched Cluster Architecture → Quorum model documented, witness host role explained, fault domains designed, traditional vs stretched comparison table completed

✓ Day 2 Operations → 5th host added, vSAN rebalance executed, data distributed evenly, storage policy applied, cluster validated at new size

Cleanup / Restore

Snapshot: vcp-admin-08-complete

• Power off test VM (wld-test-01): vCenter > VMs > right-click > Power Off > Remove from Inventory

• Document the WLD configuration: cluster name, host list, vSAN settings, NSX T0 configuration, IP pools used. Save to your lab notebook for future reference.

• Take a final snapshot: 'vcp-admin-08-complete — Traditional WLD fully provisioned and validated, 5 hosts after expansion, vSAN rebalanced, Day 2 operations completed.'

• If performing Extension 1 (stretched WLD), do NOT revert yet. Keep management domain and traditional WLD intact for stretched deployment.

Design Reflection (VCDX)

A VCDX panelist examining your WLD design would ask: Why choose traditional vs stretched topology for a given business case? What is the actual RTO/RPO if Site A fails (all 4 hosts down)? How does the witness host quorum work, and what if the witness host fails? Why is inter-site latency <10ms critical? Can you expand the cluster from 4 to 5 hosts without downtime? How does NSX scale with WLD cluster size? What is the blast radius of a vSAN failure? Be prepared to architect a multi-site WLD that balances availability, cost, and operational complexity for a real enterprise use case.

Requirements

  • Workload domain must provide dedicated compute for tenant workloads with performance and availability guarantees isolated from the management domain
  • Cluster must scale from initial size (4 hosts) to larger size (8-16 hosts) without downtime or data loss
  • Availability: single-site WLD must tolerate 1 host failure (FTT=1); stretched WLD must tolerate entire site failure
  • Network connectivity: WLD hosts must reach external networks via NSX T0 gateway with redundant uplinks

Constraints

  • vSAN requires <1ms intra-site latency and <10ms inter-site latency — cannot stretch WLD across distant data centers
  • Witness host for stretched cluster must be in third location (or isolated network) to avoid being partitioned with primary site
  • vSAN rebalance is I/O intensive and disruptive — must be scheduled during maintenance windows
  • NSX scale: Tier-0 gateway capacity depends on edge node resources; edge nodes compete for compute in the cluster
  • Licensing: each WLD requires proportional VCF licenses; stretched topology may require additional licensing (e.g., Site Recovery Manager for automated failover)

Assumptions

  • All ESXi hosts have sufficient local storage for vSAN disk groups (minimum 200 GB per host for lab, 1+ TB for production)
  • Network infrastructure supports dedicated VLANs for management, vMotion, vSAN, NSX TEP, and uplink traffic with adequate bandwidth (10 Gbps minimum per host for vSAN)
  • Operator has SDDC Manager Administrator role and understands vSphere, NSX, and vSAN administration
  • For stretched topology: inter-site network link is stable and redundant (2+ independent paths); witness host has guaranteed isolation from primary sites

Risks

  • vSAN cluster quorum loss if majority of nodes are down simultaneously — IMPACT: entire cluster becomes read-only, MITIGATION: over-provision hosts and monitor quorum health
  • Inter-site network partition on stretched cluster — IMPACT: both sites become read-only until link is restored, MITIGATION: design L3 redundancy, avoid single-point-of-failure link
  • vSAN rebalance triggers cascading failures if hosts are near capacity — IMPACT: cluster becomes unhealthy, MITIGATION: maintain 20-30% free space on vSAN datastore
  • Witness host failure in stretched cluster — IMPACT: quorum tilts to the physically co-located site (if witness is placed there), MITIGATION: place witness in true third location or use multiple witness nodes
  • NSX T0 uplink failure isolates WLD from external networks — IMPACT: workloads lose external connectivity, MITIGATION: use redundant edge nodes and multiple uplink connections

Self-Assessment Discussion Prompts

  1. You are designing a WLD for an e-commerce platform. Business requires <1 minute RTO if a single host fails and <1 hour RTO if entire data center fails. Which topology (traditional or stretched) would you recommend, and why?
  2. In a stretched WLD with 4+3+1 (witness) configuration, the inter-site link fails for 2 hours. What is the state of each site? Which site can still write? What happens to VMs running on the read-only site?
  3. You plan to add 8 new hosts to an existing 4-node WLD (doubling capacity). The vSAN rebalance is projected to take 8 hours. How would you execute this to minimize workload disruption?
  4. Compare NSX scale in a traditional 4-node WLD vs stretched 7-node WLD. How many edge nodes would you deploy in each case, and where would they run?
  5. A security audit requires all VMs in the WLD to be separated by fault domain (no two replicas on the same host, no two replicas in same failure domain for stretched). How would you design vSAN FTT and storage policies to meet this requirement?
  6. Stretched WLD witness host has a hardware failure and cannot be replaced for 24 hours. What is the state of the WLD? Can VMs still run? Can new VMs be created?

Extensions

Provision a Stretched Workload Domain (2-site topology)

Extend the lab to provision a stretched WLD using the architecture and planning from Task 4. Deploy: Site A with 4 ESXi hosts, Site B with 3 ESXi hosts, witness host on isolated VLAN. Configure fault domains and vSAN replication across sites. Execute a network partition test (disable inter-site link) to validate quorum behavior. This is the advanced topology that VCDX candidates must master.

harder

Design and Validate a Multi-Tier NSX Network for Workload Domain

Create a multi-tier NSX network architecture for the WLD: Tier-0 gateway for external routing, multiple Tier-1 gateways for tenant segmentation, overlay segments for different application tiers (web, app, db). Create workloads on different segments, enforce micro-segmentation policies, and validate east-west traffic. This demonstrates VCDX-level NSX design.

harder

Compare VCF 5.2 vs VCF 9.0 WLD Provisioning Workflows

If you have access to a VCF 5.2 lab, repeat the WLD provisioning (Task 2) in VCF 5.2. Document differences: SDDC Manager UI, provisioning automation maturity, NSX integration, vSAN configuration. Write a 2-page comparison addressing: what was improved in VCF 9.0, what was deprecated, and how the provisioning experience evolved.

same

Design and Implement WLD Placement Policies for Multi-Tenant Isolation

Design a placement policy strategy for a multi-tenant environment: Tenant A VMs must not run on the same physical host as Tenant B, each tenant has dedicated storage quota, each tenant has dedicated NSX segment. Use vCenter VM groups, host groups, and affinity rules to enforce these policies. Validate that workload placement complies with multi-tenant constraints.

harder

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.