Academy/VCAP — VCF Architect (3V0-12.26)/Lab: Multi-Domain VCF Physical Design & Sizing
This lab targets VCF 9.0

Lab: Multi-Domain VCF Physical Design & Sizing

VCF 9.0Advancedvcap-advancedvcdx⏱ 150 min

Physical design lab — host sizing, vSAN capacity calculations, network design, domain topology, bill of materials. VCF 9.0 with vSAN ESA.

Objectives

  • Size ESXi hosts for VCF management and workload domains using CPU, memory, and storage calculations
  • Calculate vSAN ESA usable capacity accounting for RAID overhead, slack space, and operations overhead
  • Design physical network topology for VCF including VLAN assignments, MTU, and uplink redundancy
  • Create a complete Bill of Materials (BoM) with hardware specifications and VCF licensing
  • Validate physical design against VCF HCL and configuration maximums

Prerequisites

No lab environment required — this is a design sizing exercise. Spreadsheet for capacity calculations. Access to VMware Compatibility Guide (VCG) for HCL validation.

Prior labs: vcap-architect-01

Required skills:

  • VCF domain architecture (management + workload domains)
  • vSAN ESA concepts (single-tier NVMe, storage policies)
  • Physical network design (VLANs, MTU, LAG)
  • Capacity planning fundamentals

Lab Environment

Design exercise using the healthcare scenario from Lab vcap-architect-01. Primary site (data center) + DR site. Target: 400 VMs across 2 workload domains.

Tasks

Task 1 Physical Design & Sizing for Multi-Domain VCF

Physical design errors are expensive — ordering wrong hardware, under-sizing storage, or miscounting network ports delays deployment by weeks. The capacity math must be defensible: show your work, document your ratios, and include growth projections. VCDX panelists will challenge your sizing assumptions.

Translate the logical design from Lab 01 into a physical implementation blueprint — specific server models, disk counts, network port assignments, VLAN IDs, and capacity calculations. This is the physical design layer that deployment engineers use to build the environment.

Step 1

Define workload profiles for sizing. From the scenario: (a) Clinical Domain: 200 VMs, avg profile = 4 vCPU, 16GB RAM, 200GB storage. Critical workloads: EHR system (8 VMs × 16 vCPU, 64GB, 500GB), database cluster (4 VMs × 8 vCPU, 32GB, 1TB). Total: ~1,000 vCPU, ~4TB RAM, ~60TB storage. (b) Development Domain: 200 VMs, avg profile = 2 vCPU, 8GB RAM, 100GB storage. Total: ~400 vCPU, ~1.6TB RAM, ~20TB storage. (c) Management Domain: VCF components (vCenter, NSX Manager ×3, VCF Operations ×3, SDDC Manager, Edge VMs ×4) ≈ 100 vCPU, 400GB RAM, 3TB storage. Document these profiles in a sizing spreadsheet.

Step 2

Size ESXi hosts for the management domain. Requirements: 100 vCPU, 400GB RAM, 3TB storage + HA overhead (N+1 = 25% for 4-host cluster). Target: 4 hosts. Per-host specification: CPU = 2× Intel Xeon Gold 6442Y (24 cores @ 2.6GHz) = 48 cores per host. Memory = 256GB per host (4 × 256GB = 1TB total, N+1 usable = 768GB, sufficient for 400GB requirement). Storage = 4× 3.84TB NVMe per host (vSAN ESA) = 15.36TB raw per host, 61.44TB raw total. RAID-5 usable = 46TB, 25% slack = 34.5TB effective — sufficient for 3TB requirement with room for growth. Validate against VCF HCL: confirm server model, NIC model, NVMe drive model are on VMware Compatibility Guide.

Step 3

Size ESXi hosts for the Clinical workload domain. Requirements: 1,000 vCPU (4:1 overcommit for general, 2:1 for EHR/DB), 4TB RAM, 60TB storage + HA (N+1). Over-commit calculation: EHR (128 vCPU at 2:1 = 64 cores dedicated) + DB (32 vCPU at 2:1 = 16 cores) + General (840 vCPU at 4:1 = 210 cores) = 290 cores needed. Target: 8 hosts with 2× 32-core CPUs (64 cores/host) = 512 cores total. N+1 usable = 448 cores — sufficient. Memory: 8× 512GB = 4TB total, N+1 usable = 3.5TB — requires 4TB, so increase to 8× 768GB = 6TB (N+1 = 5.25TB). Storage: 8× 8× 3.84TB NVMe = 245.76TB raw. RAID-1 for Clinical tier: 122.88TB. Subtract 25% slack: 92.16TB effective. Sufficient for 60TB.

Step 4

Size ESXi hosts for the Development workload domain. Requirements: 400 vCPU (6:1 overcommit acceptable), 1.6TB RAM, 20TB storage. Compute: 400 vCPU at 6:1 = 67 cores needed. Target: 4 hosts with 2× 24-core CPUs (48 cores/host) = 192 cores total. N+1 usable = 144 cores — sufficient. Memory: 4× 256GB = 1TB total, N+1 usable = 768GB — need 1.6TB, so increase to 4× 512GB = 2TB (N+1 = 1.5TB). Hmm, still short — either add a 5th host or increase memory to 4× 768GB. Decision: 5 hosts × 512GB = 2.5TB (N+1 = 2TB) — provides memory headroom. Storage: 5× 4× 3.84TB NVMe = 76.8TB raw. RAID-5: 57.6TB usable. 25% slack: 43.2TB effective — sufficient for 20TB.

Step 5

Design physical network topology. Per-host NIC configuration: 2× 25GbE dual-port NICs = 4× 25GbE ports per host. Port assignment: vmnic0 (25GbE): Management + vMotion (shared VDS uplink 1), vmnic1 (25GbE): Management + vMotion (shared VDS uplink 2), vmnic2 (25GbE): vSAN + TEP overlay (VDS uplink 3), vmnic3 (25GbE): vSAN + TEP overlay (VDS uplink 4). VLAN design: Management VLAN 10 (10.10.10.0/24), vMotion VLAN 20 (10.10.20.0/24), vSAN VLAN 30 (10.10.30.0/24), Host TEP VLAN 40 (10.10.40.0/24), Edge TEP VLAN 50 (10.10.50.0/24), Uplink VLAN 100 (192.168.100.0/24), VM Primary VLAN 110 (172.16.0.0/16). Configure MTU 9000 on VLANs 20, 30, 40, 50 (vMotion, vSAN, TEP). MTU 1600 minimum for GENEVE.

Step 6

Design Top-of-Rack (ToR) switch topology. Per-rack: 2× Cisco Nexus 93180YC-FX3 (48× 25GbE + 6× 100GbE uplinks). Configuration: MLAG/vPC between ToR pair for NIC redundancy, 2× 100GbE uplinks per ToR to spine switches (4× 100GbE total per rack to spine), Spanning Tree RSTP with BPDU guard on host ports, DHCP snooping disabled on management VLAN, MTU 9216 configured on all VLANs carrying overlay traffic. Rack layout: Management domain = 1 rack (4 hosts + 2 ToR switches), Clinical domain = 2 racks (4 hosts per rack + 2 ToR per rack), Development domain = 1 rack (5 hosts + 2 ToR). Total: 4 racks, 8 ToR switches.

Step 7

NSX Edge physical design. Deploy 4 Edge VMs (2 for T0, 2 for T1/services) — Large form factor: 8 vCPU, 32GB RAM. Place Edge VMs on management domain hosts (not workload hosts) for lifecycle independence. Edge VM NIC mapping: vNIC0 = Management, vNIC1 = Edge TEP VLAN 50 (overlay connectivity), vNIC2 = T0 Uplink VLAN 100 (first uplink to physical router), vNIC3 = T0 Uplink VLAN 101 (second uplink for ECMP). BGP configuration: Edge AS 65001, Physical router AS 65000, ECMP across both Edge pairs. Anti-affinity rules: Edge-1 and Edge-2 must not colocate on same host.

Step 8

Build the Bill of Materials (BoM). Create a table with: Item | Quantity | Unit Price | Total. Servers: Management hosts (4× Dell R760, 48-core, 256GB, 4×NVMe) ≈ $25K each = $100K. Clinical hosts (8× Dell R760, 64-core, 768GB, 8×NVMe) ≈ $45K each = $360K. Development hosts (5× Dell R760, 48-core, 512GB, 4×NVMe) ≈ $30K each = $150K. Networking: ToR switches (8× Nexus 93180YC-FX3) ≈ $15K each = $120K. Spine switches (2× Nexus 9336C-FX2) ≈ $25K each = $50K. Licensing: VCF subscription (17 hosts × $8K/socket × 2 sockets) ≈ $272K. Total hardware + licensing ≈ $1.052M against $2M budget. Remaining: DR site ($700K), implementation services ($248K).

Step 9

Validate against VCF configuration maximums. Check: (a) Hosts per cluster: max 64 — Clinical domain has 8 (OK). (b) VMs per host: max 1024 — Clinical 200/8 = 25 per host (OK). (c) vSAN ESA requirements: min 4 hosts per cluster (OK for all domains), min 25GbE for vSAN (OK), min 128GB RAM (OK — management has 256GB). (d) NSX Edge cluster: max 10 Edge VMs per cluster — using 4 (OK). (e) Total hosts per VCF Instance: max 192 — using 17 (OK). Document any limits that constrain future growth and the growth capacity remaining.

Step 10

Create capacity growth projections. Assuming 20% VM growth per year for 3 years: Year 1: 400 VMs (current). Year 2: 480 VMs (+80). Year 3: 576 VMs (+96). Assess runway: Clinical domain can grow to ~350 VMs before needing additional hosts (from 200, runway = 2 years). Development domain can grow to ~400 VMs (from 200, runway = 3+ years). Management domain: fixed overhead, minimal growth. Plan: add 4 hosts to Clinical domain in Year 2 ($180K — within remaining budget). Document growth triggers: 'When Clinical domain CPU utilization exceeds 70% sustained or vSAN utilization exceeds 75%, add hosts.'

Validation Gate

Check: Complete physical design with sizing, BoM, and capacity projections

Expected: Host specifications for 3 domains (17 hosts total), vSAN capacity calculations with RAID overhead and slack space, network topology with VLAN assignments and MTU, BoM within $2M budget, configuration maximum validation, 3-year growth projection

Common Errors

vSAN capacity calculation ignores RAID overhead and slack space
Fix: Raw capacity ≠ usable capacity. Apply: Raw → RAID overhead (RAID-1 = 50%, RAID-5 = 25%, RAID-6 = 33%) → Slack space (25% for vSAN operations) → Operations overhead (vSAN metadata ~1-2%). Document the full waterfall.
Memory sizing doesn't account for vSAN ESA overhead
Fix: vSAN ESA requires significant memory for buffer cache — at least 128GB RAM per host before workload allocation. Subtract vSAN overhead from available workload memory. For a 512GB host: ~128GB vSAN + ~384GB workload.
Network design uses MTU 1500 for TEP and vSAN VLANs
Fix: vSAN requires MTU 9000 for optimal performance. TEP (GENEVE) needs MTU 1600 minimum but 9000 is recommended. Failure to configure jumbo frames causes fragmentation and performance degradation.
BoM doesn't include licensing costs
Fix: VCF licensing is a significant portion of total cost. VCF is now subscription-based (per-core or per-socket). Include: VCF subscription, support/subscription renewal, and any add-on licenses (vDefend, VCF Automation, etc.).

Final Validation

Physical design blueprint ready for procurement and deployment

✓ Host sizing → All domains sized with CPU, memory, storage meeting workload requirements + HA overhead

✓ vSAN capacity → Capacity waterfall documented: raw → RAID → slack → usable for each domain

✓ Network design → VLAN assignments, MTU, ToR topology, uplink redundancy documented

✓ Bill of Materials → Complete BoM within $2M budget including hardware, networking, and licensing

✓ Configuration maximums → All VCF limits validated with growth headroom documented

✓ Growth projection → 3-year capacity projection with expansion triggers

Cleanup / Restore

• Save physical design documentation for use in subsequent labs

Design Reflection (VCDX)

Physical design demonstrates your ability to translate logical architecture into implementable specifications. VCDX panelists will challenge your sizing: 'Why 8 hosts instead of 6?' — you must show the math. They'll probe vSAN capacity: 'What's your effective capacity after RAID, slack, and overhead?' They'll question network: 'Why 4 NICs per host? What if the customer only has 2-port NICs?' Have alternatives ready.

Requirements

  • Host 400 VMs across 3 domains with specified resource profiles
  • Meet $2M budget including DR site
  • Support 20% annual growth for 3 years

Constraints

  • Budget: $2M total including DR
  • vSAN ESA minimum: 4 hosts, 25GbE, 128GB RAM
  • Existing Cisco Nexus switches must be reused

Assumptions

  • Workload profiles (vCPU, RAM, storage per VM) are representative of actual deployment
  • VCF subscription pricing is $8K/socket/year
  • Hardware delivery lead time is 6-8 weeks

Risks

  • Workload profiles underestimate actual resource consumption — VM right-sizing needed post-deployment
  • Hardware price increases between design and procurement
  • vSAN ESA NVMe drive failures exceed spare capacity during warranty period

Self-Assessment Discussion Prompts

  1. How do you handle a scenario where the BoM exceeds the budget?
  2. When should you use RAID-1 vs RAID-5 for vSAN storage policies?
  3. How does HCI Mesh change your sizing approach for compute-heavy workloads?
  4. What's your approach to sizing Edge VMs — when do you go from Medium to Large to XLarge?

Extensions

Add DR site physical design — mirror primary site at 50% capacity for failover

Design a phased deployment plan: management domain first, then Clinical, then Development

Create a physical design for a VCF stretched cluster between the primary and DR sites

Build a capacity planning spreadsheet template that auto-calculates vSAN usable capacity from raw inputs

⚠ Known Pitfalls (from Community KB)

Sizing to exact requirements without HA overhead — N+1 for HA means you lose one host's capacity from usable totals
Forgetting vSAN ESA memory overhead — 128GB+ per host goes to vSAN buffer cache, not workloads
Using vSAN RAID-1 everywhere — RAID-1 uses 2x capacity; RAID-5 is 1.33x and appropriate for non-critical workloads
Not validating HCL compatibility — ordering hardware not on the VMware Compatibility Guide causes deployment delays and support issues

References

  • VCF 9.0 Planning and Preparation Guide: techdocs.broadcom.com
  • VMware Compatibility Guide (VCG): vmware.com/resources/compatibility
  • vSAN ESA Design and Sizing Guide: techdocs.broadcom.com
  • VCF Configuration Maximums: techdocs.broadcom.com
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.