Compute Cluster Design with HA Admission Control
Objectives
- Size compute clusters for workload requirements with growth projection
- Configure HA admission control policies (slot-based, percentage, dedicated failover)
- Design DRS affinity and anti-affinity rules for availability and performance
- Evaluate EVC mode requirements for heterogeneous hardware environments
- Justify cluster boundary decisions based on failure domain isolation
Prerequisites
Access to VCF 9.0 documentation and vSphere cluster architecture references
Prior labs: vcp-architect-01
Required skills:
- Understanding of vSphere HA and DRS concepts
- Basic capacity planning arithmetic
- VCF workload domain architecture
Lab Environment
Design exercise using VCF 9.0 reference architecture — can be validated on Holodeck with nested ESXi hosts if available
Tasks
Task 1 Workload Profiling & Cluster Sizing
Translate business workload requirements into concrete compute cluster sizing with growth headroom and failure capacity.
Define workload profiles for a 600-VM environment:
Tier-1 (Mission Critical, 80 VMs):
- Average: 4 vCPU, 16 GB RAM per VM
- Peak CPU utilization: 70%, Memory: 85%
- Total: 320 vCPU, 1,280 GB RAM
- Overcommit policy: CPU 3:1 max, Memory 1.25:1 max
Tier-2 (Business Applications, 300 VMs):
- Average: 2 vCPU, 8 GB RAM per VM
- Peak CPU utilization: 50%, Memory: 60%
- Total: 600 vCPU, 2,400 GB RAM
- Overcommit policy: CPU 5:1, Memory 1.5:1
Tier-3 (Dev/Test, 220 VMs):
- Average: 2 vCPU, 4 GB RAM per VM
- Peak CPU utilization: 30%, Memory: 40%
- Total: 440 vCPU, 880 GB RAM
- Overcommit policy: CPU 8:1, Memory 2:1
Calculate physical host requirements:
Assume host specification: Dual-socket, 32 cores per socket (64 cores total), 512 GB RAM
VCF 9.0 licensing: per-core (minimum 16 cores per CPU socket)
Tier-1 cluster:
- CPU: 320 vCPU / 3:1 overcommit = 107 physical cores needed
- Memory: 1,280 GB / 1.25 = 1,024 GB physical RAM needed
- Hosts: ceil(107/64) = 2 (CPU), ceil(1,024/512) = 2 (RAM) → 2 hosts minimum - With N+1 HA: 3 hosts minimum - With 3-year growth (20% p.a.): 2 × 1.2³ = 3.46 → 4 hosts + 1 HA = 5 hosts
Repeat calculation for Tier-2 and Tier-3.
Document the governing constraint (CPU-bound vs. memory-bound) for each tier.
Apply VCF 9.0 cluster minimums and maximums:
- Management domain: minimum 4 hosts (VCF requirement for vSAN and management VMs)
- Workload domain: minimum 3 hosts for vSAN, but 4 recommended for HA + maintenance
- Maximum: 64 hosts per cluster (vSphere 8.x/9.0 limit), 96 hosts per vSAN cluster
- Consider VCF-specific overhead: SDDC Manager, NSX Manager cluster (3 nodes), vCenter, Aria Suite
- Management domain typically consumes 80-120 GB RAM across 3-4 management VMs
Document final cluster topology:
- Cluster A (Management): X hosts
- Cluster B (Tier-1 Workload): Y hosts
- Cluster C (Tier-2/3 Workload): Z hosts — justify consolidation or separation decision
Create a capacity summary table:
| Cluster | Hosts | Cores/Host | Total Cores | RAM/Host | Total RAM | VMs | CPU Overcommit | Mem Overcommit | HA Reserve | Growth Headroom |
|---|---|---|---|---|---|---|---|---|---|---|
| Mgmt | ||||||||||
| Tier-1 | ||||||||||
| Tier-2/3 |
Validate: Total host count × 64 cores × per-core license cost = total VCF licensing cost
Note: VCF 9.0 uses per-core licensing — every physical core on every socket counts.
Validation Gate
Check: Capacity model shows full sizing math with overcommit ratios, growth factors, and HA overhead for all tiers
Expected: Each cluster has documented host count, resource totals, overcommit ratios within policy, N+1 (or N+2 for critical) HA capacity, and 3-year growth headroom
Common Errors
Task 2 HA Admission Control Policy Design
Select and configure the optimal HA admission control policy for each cluster based on workload criticality and cluster size.
Evaluate three HA admission control methods for the Tier-1 cluster:
Option A — Slot-based (Host Failures Cluster Tolerates):
- Slot size = largest VM reservation (4 vCPU, 16 GB RAM)
- Slots per host: min(64 cores/4, 512 GB/16 GB) = min(16, 32) = 16 slots
- With 5 hosts: 80 total slots, tolerating 1 failure = 64 usable slots
- Drawback: Large 'monster VMs' inflate slot size, wasting capacity
Option B — Percentage-based:
- Reserve 25% of cluster CPU and memory (equivalent to 1 host in a 4-host cluster)
- More flexible than slot-based for heterogeneous VM sizes
- Risk: Doesn't guarantee a specific number of host failures tolerated
Option C — Dedicated Failover Hosts:
- Designate 1 specific host as failover target
- Pros: Predictable, that host is always available
- Cons: Wastes resources (dedicated host sits idle), doesn't support DRS load balancing
Document decision: Select percentage-based (25%) for Tier-1 with justification.
Configure HA settings for each cluster:
Management Cluster:
- Admission control: Percentage-based, 25% (tolerate 1 host failure in 4-host cluster)
- VM restart priority: Highest for SDDC Manager, NSX Manager, vCenter; Medium for Aria
- Host isolation response: Leave Powered On (prevent false positive isolation events)
- Heartbeat datastore: 2 datastores minimum (vSAN + one additional if available)
- VM Component Protection (VMCP): Enabled with 'Power off and restart VMs' for APD after 140s
Tier-1 Workload Cluster:
- Admission control: Percentage-based, 25%
- VM restart priority: Highest for database VMs, Medium for app servers
- Proactive HA: Enabled (if supported by hardware — requires IPMI/iLO integration)
Tier-2/3 Workload Cluster:
- Admission control: Percentage-based, 20% (can tolerate slightly longer recovery)
- VM restart priority: Medium for all VMs
- Proactive HA: Optional
Design VM restart priority and orchestrated restart:
- Define restart priority groups for the Management cluster:
- Group 1 (Highest): vCenter Server, SDDC Manager
- Group 2 (High): NSX Manager nodes (stagger with 120s delay between nodes)
- Group 3 (Medium): Aria Suite components
- Group 4 (Low): Monitoring, logging agents
- Configure VM-VM dependency rules:
- NSX Manager depends on vCenter availability
- Aria Suite depends on NSX Manager availability
- Use 'Orchestrated Restart' (vSphere 8+) to define startup order
- Set VM Monitoring (application-level HA):
- Sensitivity: Medium (120s failure window)
- Maximum per-VM resets: 3 within 1 hour
- Maximum resets window: Reset count resets after 24 hours
Validate HA design against failure scenarios:
Scenario 1 — Single host failure in Tier-1 cluster:
- Calculate: Can remaining 4 hosts accommodate all 80 Tier-1 VMs?
- Memory check: 80 × 16 GB = 1,280 GB needed; 4 × 512 GB = 2,048 GB available → Pass - CPU check: 80 × 4 vCPU = 320 vCPU; 4 × 64 cores × 3:1 = 768 vCPU capacity → Pass
Scenario 2 — Host failure during maintenance (N-2 condition):
- Can 3 remaining hosts handle all VMs? Recalculate.
- If not: Document maintenance window policy (no concurrent maintenance + failure tolerance)
Scenario 3 — Network partition (split-brain):
- Which host is master? (Host with most datastores mounted, then highest MOID)
- What happens to VMs on isolated hosts? (Host isolation response = Leave Powered On)
- How does vSAN handle split-brain? (Witness vote determines data availability)
Document each scenario with pass/fail and mitigation actions.
Validation Gate
Check: HA admission control configured with documented policy choice, restart priorities, and failure scenario analysis
Expected: Each cluster has admission control method selected with alternatives considered, VM restart priorities defined, and 3+ failure scenarios validated with capacity math
Common Errors
Task 3 DRS Affinity Rules & EVC Mode Design
Design DRS rules that enforce availability boundaries and optimize performance, and evaluate EVC mode requirements for hardware compatibility.
Design DRS anti-affinity rules (VM-VM):
Rule 1 — Management HA Separation (MUST):
- NSX Manager node 1, 2, 3: MUST run on separate hosts
- Type: Separate Virtual Machines (anti-affinity)
- Priority: Required (not preferred) — DRS will not violate
Rule 2 — vCenter + SDDC Manager Separation (MUST):
- vCenter and SDDC Manager must not share a host
- Reason: Both are management plane critical — co-location creates dual SPOF
Rule 3 — Database Cluster Separation (SHOULD):
- Database primary and replica VMs: SHOULD run on separate hosts
- Type: Preferred (DRS may violate during resource contention to avoid worse outcome)
Rule 4 — App Tier Load Distribution (SHOULD):
- Web server VMs spread across minimum 3 hosts
- Type: Preferred anti-affinity
Design DRS affinity rules (VM-Host):
Rule 1 — GPU Workloads (if applicable):
- VMs requiring vGPU pinned to GPU-equipped hosts
- Type: Must run on hosts in group 'GPU-Hosts'
Rule 2 — Licensed Software Locality:
- Oracle/SQL Server VMs restricted to specific hosts (license compliance)
- Type: Must run on hosts in group 'Licensed-Hosts'
- VCDX consideration: Document licensing cost impact of VM mobility restrictions
Rule 3 — Latency-Sensitive Workloads:
- Trading/real-time VMs on hosts with Latency Sensitivity = High
- Configure: vSphere Latency Sensitivity setting, disable C-states, reserve CPU/memory
Document the trade-off: Affinity rules reduce DRS flexibility and can prevent HA restart if target hosts are unavailable.
Configure DRS automation and migration thresholds:
- DRS Automation Level per cluster:
- Management: Partially Automated (conservative — manual approval for management VM moves)
- Tier-1: Fully Automated, threshold = 3 (moderate — balance load without excessive vMotion)
- Tier-2/3: Fully Automated, threshold = 4 (aggressive — maximize utilization)
- vMotion rate-limiting:
- Maximum concurrent vMotions per host: 4 (for 25 GbE network)
- Maximum concurrent vMotions per datastore: 8
- Resource pools (if used):
- Design warning: Nested resource pools create 'DRS tax' — only use when mandatory for organizational isolation
- If used: Set shares proportionally (High:Normal:Low = 4:2:1)
- Recommended alternative: VM-level reservations for critical VMs instead of resource pools
Evaluate and configure EVC mode:
Scenario: Mixed CPU generations in the environment:
- Generation A: Intel Sapphire Rapids (6th Gen Xeon Scalable)
- Generation B: Intel Emerald Rapids (5th Gen Xeon Scalable)
- Generation C: Intel Ice Lake (3rd Gen Xeon Scalable)
- Determine lowest common EVC baseline:
- EVC mode: 'Intel Ice Lake Generation' (lowest generation)
- Impact: Sapphire Rapids and Emerald Rapids features masked (AVX-512 differences, AMX)
- VMs created under this EVC mode can vMotion across all three generations
- Document EVC constraints:
- Cannot raise EVC mode with VMs powered on
- Cannot lower EVC mode without powering off all VMs
- Some workloads (AI/ML using AMX instructions) may need a separate cluster without EVC restrictions
- Design decision: Separate clusters by CPU generation vs. unified EVC cluster?
- If performance-sensitive: Separate clusters, accept reduced vMotion mobility
- If operational simplicity preferred: Unified EVC at lowest baseline
- Document trade-off and recommendation
Validation Gate
Check: DRS rules documented with type (required/preferred), EVC mode evaluated with performance impact analysis
Expected: Minimum 4 DRS rules defined with justification, EVC mode decision documented with alternatives, DRS automation levels set per cluster with rationale
Common Errors
Task 4 Cluster Boundary Design & VCDX Defense
Make and defend the decision of how many clusters to deploy, what goes in each, and why the boundaries are where they are.
Document cluster boundary decision matrix:
| Factor | Consolidate (fewer, larger clusters) | Isolate (more, smaller clusters) |
|---|---|---|
| HA blast radius | Larger (more VMs affected) | Smaller (fewer VMs per failure domain) |
| Resource efficiency | Higher utilization | Lower (overhead per cluster) |
| Operational complexity | Simpler (fewer objects) | Complex (more vCenters, more DRS configs) |
| Upgrade risk | Higher (more VMs at risk during upgrade) | Lower (rolling upgrade by cluster) |
| Licensing cost | Lower (fewer management hosts) | Higher (minimum host count per cluster) |
| Compliance isolation | Weaker | Stronger (PCI/SOC 2 isolation) |
Map each factor to a specific requirement or constraint from the RCAR framework.
Design the final cluster topology and document justification:
Recommended topology for 600-VM environment:
- Management Domain (1 cluster, 4 hosts):
- Contains: vCenter, SDDC Manager, NSX Manager cluster, Aria Suite
- Justification: VCF mandates dedicated management domain; 4 hosts = N+1 with 25% HA
- Workload Domain — Tier-1 (1 cluster, 5 hosts):
- Contains: Mission-critical production VMs (80 VMs)
- Justification: Separate failure domain for SOC 2 compliance; conservative overcommit ratios
- Workload Domain — Tier-2/3 (1 cluster, 8 hosts):
- Contains: Business apps + dev/test (520 VMs)
- Justification: Consolidation acceptable for lower-tier workloads; higher overcommit saves licensing cost
- Alternative considered: Separate Tier-2 and Tier-3 clusters → Rejected due to minimum host count overhead (3+3 vs 8 hosts)
Total: 17 hosts, 3 clusters, 2 workload domains
Prepare VCDX defense responses for cluster design challenges:
Challenge 1: 'Why not consolidate everything into one large cluster?'
Response: SOC 2 requires isolation of PCI-scoped workloads. Single cluster means one upgrade window affects all tiers. HA blast radius of 600 VMs is unacceptable for Tier-1 RTO of 4 hours.
Challenge 2: 'Why not separate Tier-2 and Tier-3 into individual clusters?'
Response: 3-host minimum per vSAN cluster × 2 = 6 hosts vs 8 consolidated. Net savings of 2 hosts × 64 cores × per-core license = significant cost reduction. Tier-2 and Tier-3 have similar overcommit tolerance.
Challenge 3: 'What if your Tier-1 cluster needs to scale beyond 5 hosts?'
Response: VCF allows host addition to existing clusters via SDDC Manager. If exceeding 8 hosts, evaluate splitting into two Tier-1 clusters with different workload profiles (e.g., database vs application).
Challenge 4: 'How do you handle a host failure in the management domain during an upgrade?'
Response: Management domain upgrade is always done first per VCF upgrade sequence (KB 390634). With 4 hosts and 25% HA, one host in maintenance + one failure = N-2, which is tolerable for management workload sizing. If not: pause upgrade, remediate, then continue.
Create a design decision register entry for the cluster design:
Decision D-005 — Cluster Topology:
Decision: 3 clusters across 2 workload domains (Management + Tier-1 + Tier-2/3)
Alternatives:
A) Single consolidated cluster (17 hosts) — Rejected: violates SOC 2 isolation, unacceptable blast radius
B) 4 separate clusters (Mgmt + T1 + T2 + T3) — Rejected: 4-host minimum overhead, excessive licensing cost
C) 2 clusters (Mgmt + single workload) — Rejected: cannot enforce different overcommit policies per tier
Requirement traceability: R-001 (isolation), R-003 (cost), C-004 (SOC 2), A-002 (growth)
Risk: R-003 (cluster at capacity sooner than projected if growth exceeds 20% p.a.)
Mitigation: Quarterly capacity reviews; SDDC Manager host addition workflow documented in runbook
Validation Gate
Check: Cluster topology documented with decision matrix, design decision register entry, and 4+ defense responses prepared
Expected: Final topology shows host count per cluster, workload placement rationale, licensing impact, and traceable justification for every boundary decision
Common Errors
Final Validation
Complete compute cluster design with sizing, HA, DRS, EVC, and boundary decisions documented and defensible
✓ Capacity model covers all tiers with growth projection → Sizing math shows host count derivation with overcommit ratios and HA overhead
✓ HA admission control configured per cluster → Policy method selected with alternatives evaluated; failure scenarios validated
✓ DRS rules enforce availability boundaries → Minimum 4 rules defined with type (required/preferred) and justification
✓ Cluster boundaries justified with design decisions → Decision register entry links to RCAR requirements; defense responses prepared
Cleanup / Restore
• Save all design documents and decision registers
• If using Holodeck: revert nested ESXi hosts to base snapshot
Design Reflection (VCDX)
Compute cluster design is one of the most heavily challenged areas in VCDX defense. Panelists probe three things: (1) Can you show the math? Every host count must trace to workload sizing. (2) Do you understand HA nuances? Slot-based vs percentage, host isolation response, VMCP — these are differentiators. (3) Can you defend boundaries? Why this number of clusters, why these workloads together? The answer must always connect back to requirements.
Requirements
- 600 VMs across 3 tiers with different SLAs and overcommit policies
- N+1 HA for all clusters, N+2 for management domain during maintenance windows
- 3-year lifecycle with 20% annual VM growth
- SOC 2 compliance requires tier isolation
Constraints
- VCF 9.0 per-core licensing — every additional host adds license cost
- Management domain minimum 4 hosts (VCF hard requirement)
- Maximum 64 hosts per cluster (vSphere limit)
- Mixed CPU generations in hardware inventory
Assumptions
- Host hardware is dual-socket, 64 cores, 512 GB RAM (on VCF HCL)
- 25 GbE network available for vMotion and vSAN
- Growth rate of 20% p.a. is consistent across all tiers
- No GPU or FPGA requirements in initial deployment
Risks
- Actual growth exceeds 20% — cluster runs out of headroom before 3-year lifecycle
- Monster VM deployed that inflates HA slot size and reduces usable capacity
- EVC mode prevents workload migration if CPU generation gap widens
- DRS anti-affinity rule prevents HA restart when insufficient hosts in target group
Self-Assessment Discussion Prompts
- What changes if the customer switches from Intel to AMD processors mid-lifecycle — does EVC mode break?
- How would you redesign if the RTO for Tier-1 dropped from 4 hours to 15 minutes?
- If you could only have 2 clusters total (management + one workload), how would you enforce tier isolation within a single cluster?
- What is the licensing cost impact of adding a 6th host to the Tier-1 cluster vs. increasing overcommit ratio?
Extensions
Proactive HA with Hardware Health Integration
Configure Proactive HA with Dell iDRAC or HPE iLO health providers. When a host reports degraded health (fan failure, memory ECC errors), DRS proactively migrates VMs before a hard failure. Document the health provider configuration, quarantine mode behavior, and impact on DRS automation level.
Predictive DRS with Aria Operations Integration
Enable Predictive DRS using Aria Operations workload forecasting. DRS pre-positions VMs based on predicted demand spikes rather than reactive balancing. Evaluate the data retention period needed for accurate predictions (minimum 2 weeks of metrics) and the impact on vMotion rates.
vSphere Lifecycle Manager (vLCM) Image-Based Cluster Management
Design a vLCM image composition for each cluster with ESXi base image, vendor add-on (Dell/HPE), firmware/driver bundles, and component specifications. Document how image compliance checking integrates with maintenance mode workflows and VCF upgrade sequencing.
⚠ Known Pitfalls (from Community KB)
References
- vSphere 9.0 Availability Guide — HA Admission Control section
- vSphere 9.0 Resource Management Guide — DRS chapter
- VMware VCF 9.0 Planning and Preparation Guide — Management Domain sizing
- VMware KB 1003212 — Understanding HA Slot Size Calculation
- VMware KB 2108935 — EVC Processor Support matrix