HA Admission Control Validation
Objectives
- Understand HA admission control policies (slot-based and percentage-based)
- Calculate slot sizes for different VM profiles in a 4-host cluster
- Evaluate admission control scenarios and policy recommendations
- Design N+1 HA policy for VCF management domain
Prerequisites
VCF 9.0 lab environment deployed with 4-host management cluster
Prior labs: vcp-architect-01: Cluster Basics
Required skills:
- Understanding vSphere HA concepts
- Basic cluster architecture knowledge
Lab Environment
4-host VCF management cluster (each host: 96GB RAM, 16 vCPU cores, 2.5GHz)
Tasks
Task 1 Understand HA Admission Control Policies
Establish foundational knowledge of slot-based and percentage-based admission control mechanisms and their impact on cluster capacity planning.
Review the 4-host VCF management cluster configuration. Document: total cluster RAM (384GB), total vCPU count (64 cores), reserved memory for hypervisors per host (8GB), and available memory per host for VMs (88GB per host = 352GB total).
Define three VM profiles: Small (2vCPU, 4GB RAM), Medium (4vCPU, 16GB RAM), Large (8vCPU, 64GB RAM). For each profile, calculate the slot size in bytes using the formula: Slot Size = MAX(CPU requirement, Memory requirement in bytes / total available memory per host).
Slot size calculations for each VM profile:
- Small VM slot: MAX(2/16 cores, 4GB/88GB) = MAX(0.125, 0.045) = 0.125 slots
- Medium VM slot: MAX(4/16 cores, 16GB/88GB) = MAX(0.25, 0.182) = 0.25 slots
- Large VM slot: MAX(8/16 cores, 64GB/88GB) = MAX(0.5, 0.727) = 0.727 slots
Determine the maximum slot size (most resource-intensive VM that could run in this cluster). Using the 4-host cluster, calculate maximum slots available: if maximum slot size = 0.727 (the Large VM requirement), then total available slots = 4 hosts / 0.727 = 5.5 slots (round down to 5 usable slots).
Calculate available slots for admission control with N+1 failover (one host failure allowed). With N+1, usable capacity = capacity of (n-1) hosts = 3 hosts. Recalculate available slots for admission control: 3 / 0.727 = 4.1 slots, rounded down to 4 slots. This ensures one failed host still leaves remaining hosts to run all admitted VMs.
Validation Gate
Check: Verify admission control calculations and N+1 capacity planning
Expected: Documented slot sizes for all VM profiles, total and N+1-compliant slot counts
Common Errors
Task 2 Evaluate Admission Control Scenarios
Apply admission control calculations to real-world scenarios and compare policy approaches (slot-based vs. percentage-based) for cluster design decisions.
Scenario 1: A new 96GB VM is requested for the management cluster. Using slot-based calculation from Task 1, determine if the cluster can admit this VM while maintaining N+1 HA. Slot size for 96GB VM = MAX(1/16, 96/88) = MAX(0.0625, 1.09) = 1.09 slots. Available slots with N+1 = 4. Can the cluster admit this VM? Calculate: 1.09 slots > 4 available slots? No—this VM cannot be admitted.
Calculate the impact of losing 1 host on remaining cluster capacity. Current state: 4 hosts, 4 slots available for Large VMs (N+1 compliant). After host failure: 3 hosts remain, but all admitted VMs must still run on these 3 hosts. If 4 Large VMs are currently admitted (4 slots × 0.727 = 2.9 VMs worth of capacity, let's say 4 Large VMs requiring 4 × 0.727 = 2.9 host equivalents), can all 4 Large VMs run on 3 remaining hosts? Remaining capacity after failover: 3 / 0.727 = 4.1 slots. Yes, 4 admitted Large VMs (using 2.9 host equivalents) fit in 3 remaining hosts.
Compare slot-based vs. percentage-based admission control for this cluster. Slot-based: Uses predefined slot size (0.727 for Large VMs); hard limit prevents oversubscription; predictable. Percentage-based: Reserves a percentage (e.g., 25%) of cluster resources for failover; allows heterogeneous workloads without recalculation. For this cluster: Slot-based reserves (1 - 3/4) = 25% of cluster capacity for failover. Percentage-based with 25% reserve achieves same result but with flexibility. Recommendation: Compare predictability vs. flexibility trade-off.
Comparison table:
- Slot-based: Predictable, hard limits, requires slot recalculation for new VM types, prevents oversubscription
- Percentage-based: Flexible, accommodates heterogeneous workloads, simpler to understand, less granular control
For this VCF management cluster with homogeneous VMs (vCenter, NSX, SDDC Manager, Ops), slot-based is more precise.
Formulate a policy recommendation for this VCF management cluster. Decision: Which admission control method and why? Justification: 'Recommend slot-based admission control for this cluster because: (1) VCF management VMs have well-defined resource profiles (Small=4GB, Medium=16GB, Large=64GB); (2) slot-based provides deterministic capacity planning; (3) prevents accidental oversubscription of large VMs; (4) VCF design already specifies resource tiers.' Document the policy choice with rationale.
Validation Gate
Check: Verify scenario calculations and policy recommendation
Expected: Documented capacity impact analysis, policy comparison, and justified recommendation
Common Errors
Task 3 Design HA Policy for VCF Management Domain
Create a comprehensive HA design document for the VCF management domain that demonstrates N+1 availability architecture and applies admission control concepts to production VCF components.
Identify all management VMs and their resource requirements in the VCF management domain. List: vCenter Server (8vCPU, 32GB RAM), NSX Manager cluster nodes (8vCPU each, 16GB RAM each, 3 nodes = 24vCPU/48GB), SDDC Manager (8vCPU, 32GB RAM), VCF Ops (formerly vRealize Operations) (16vCPU, 64GB RAM). Total: (8+24+8+16) = 56 vCPU and (32+48+32+64) = 176GB RAM across management VMs.
Documented inventory of management domain VMs with resource specifications. Example table:
| VM | vCPU | RAM GB | Notes |
| vCenter | 8 | 32 | HA required |
| NSX Manager-1 | 8 | 16 | Part of 3-node cluster |
| NSX Manager-2 | 8 | 16 | Part of 3-node cluster |
| NSX Manager-3 | 8 | 16 | Part of 3-node cluster |
| SDDC Manager | 8 | 32 | HA required |
| VCF Ops | 16 | 64 | Shared infrastructure service |
Calculate the minimum cluster size to maintain N+1 availability for all management VMs. Total management VM capacity needed: 176GB RAM + 56vCPU. With N+1 reserve (reserve 1 host): If each host = 96GB/16vCPU, then minimum usable capacity = 176GB RAM. Divide by host capacity: 176GB / 96GB = 1.83 host equivalents needed. Add N+1 reserve: 1.83 + 1 = 2.83, round up to 3 hosts minimum. Verify vCPU: 56vCPU / 16vCPU per host = 3.5 host equivalents; 3.5 + 1 reserve = 4.5, round up to 5 hosts. Thus, minimum cluster size = 5 hosts to accommodate all management VMs + N+1 failover reserve.
Document the HA design decision (DD format). Create a formal design decision document with sections: (1) Requirement: 'Maintain N+1 availability for VCF management domain to support 99.95% uptime SLA (max 22 minutes downtime/month)'; (2) Constraints: 'Cluster must survive single host failure without VM restart cascades'; (3) Assumptions: 'VMs will be distributed across all 5 hosts (no single-host concentration)'; (4) Design Decision: 'Deploy 5-host management cluster with slot-based HA admission control. Max slot size = 64GB (VCF Ops VM). Reserve 1 host capacity for failover. Admit only management VMs to this cluster (no user workloads)'; (5) Justification: 'Sizing calculation: 176GB management workload + 64GB + 96GB reserve = 336GB; 5 hosts × 96GB = 480GB available. N+1 compliant with 144GB spare capacity for growth.'
Formal HA Design Decision document:
VCF Management Domain HA Design (v9.0)
Requirement: N+1 availability for VCF management domain
Current State: 4-host cluster (insufficient for management workload)
Proposed Design: 5-host cluster with slot-based HA admission control
Rationale: Management workload totals 176GB + largest VM (VCF Ops, 64GB) + 1-host reserve = 336GB required; 5 hosts × 96GB = 480GB available. Provides N+1 failover coverage.
Impact: Requires upgrade from 4-host to 5-host cluster.
Test design with what-if scenario: What if 2 hosts fail simultaneously (N+2)? Current design assumes N+1 (1 host failure). With 2 host failures: 5 hosts - 2 = 3 hosts remain. Remaining capacity = 3 × 96GB = 288GB. Management VMs = 176GB. Remaining capacity (288GB - 176GB = 112GB) is sufficient. However, this exceeds N+1 design goal. Document: 'Current HA design supports N+1 (single host failure). N+2 survivability is not a requirement per business SLA but is a positive secondary outcome. If N+2 requirement emerges, cluster would need to grow to 6 hosts.'
Validation Gate
Check: Verify HA design document and what-if scenario analysis
Expected: Formal design decision document with justification and what-if scenario analysis
Common Errors
Final Validation
HA admission control lab completed with slot-based analysis, scenario evaluation, and VCF management cluster design
✓ Admission control calculations (Task 1, Steps 1-4) → Documented slot sizes for all VM profiles, N+1 capacity calculations
✓ Scenario evaluation and policy recommendation (Task 2, Steps 1-4) → 96GB VM admissibility determination, failover impact analysis, slot vs. percentage comparison, and justified policy choice
✓ VCF management design decision (Task 3, Steps 1-4) → Management VM inventory, minimum cluster sizing, formal DD document, N+2 what-if scenario
Cleanup / Restore
• Revert to initial snapshot if any lab VMs were created
• Document all design decisions and calculations for VCDX study
Design Reflection (VCDX)
During a VCDX defense, panelists will probe your HA design with questions like: 'You sized the cluster for 5 hosts. What happens if a host fails during a maintenance window when a 3rd host is offline?' Or: 'How would you handle a scenario where NSX Managers are concentrated on 2 hosts and one host fails?' Be prepared to explain admission control trade-offs, justify why you chose slot-based over percentage-based, and articulate what would change if requirements shifted to N+2 or if workload profiles expanded. Panelists expect designers to show flexibility—how would you adjust the design if vCF Ops expanded from 64GB to 128GB?
Requirements
- VCF management domain must maintain N+1 availability for 99.95% uptime SLA
- Cluster must support heterogeneous management VMs (vCenter, NSX Managers, SDDC Manager, VCF Ops)
- Admission control must prevent oversized VM admissions that would violate HA during failover
- Design must be documented in formal DD format with sizing calculations and assumptions
- Must address minimum cluster size, slot calculation methodology, and what-if failure scenarios
Constraints
- Each host in the management cluster has fixed capacity (96GB RAM, 16 vCPU cores)
- NSX Manager requires 3-node cluster for quorum and HA
- VCF management domain must not share hosts with user workloads (isolation requirement)
- Admission control policy chosen must apply consistently across all admitted VMs
- One full host capacity must always be reserved for failover (N+1 requirement)
Assumptions
- All management VMs will be distributed across available hosts (no single-host concentration)
- Hypervisor reserves ~8GB per host, leaving 88GB available for VMs
- vCPU allocation is 1:1 (no oversubscription) for deterministic performance
- Management VM resource requirements are stable (no dynamic scaling)
- Slot size is calculated once and remains stable (no mid-cluster addition of different-sized VMs)
- Network and storage are not the limiting factors (CPU and memory are the constraints)
- Single host failure is the design-basis failure scenario; N+2 is not required
Risks
- Risk: Cluster growth adds larger VMs without recalculating slot size. Mitigation: Document when slot size becomes invalid and requires cluster redesign.
- Risk: VCF Ops or another management VM is unexpectedly resized upward. Mitigation: Establish change control requiring HA impact analysis before VM resize in management domain.
- Risk: Operational team over-admits VMs by disabling HA or ignoring admission control. Mitigation: Use vSphere admission control policies that are enforced and non-bypassable without explicit approval.
- Risk: Network or storage becomes the limiting factor instead of CPU/memory. Mitigation: Extend this analysis to validate network (NSX bandwidth) and storage (IOPS, capacity) separately.
- Risk: A third host fails within 15 minutes of the first host failure (before first host VMs have fully restarted). Mitigation: Monitor first host failure recovery time; if recovery takes >15 min and concurrent failures are possible, increase cluster size or add recovery time SLA.
Self-Assessment Discussion Prompts
- You proposed a 5-host management cluster to maintain N+1 availability. A business stakeholder asks: 'Why do we need 5 hosts? We currently run on 4, and nothing has failed in 2 years.' How would you justify the upgrade? What specific risks does the 4-host cluster face during a single host failure?
- Compare slot-based vs. percentage-based admission control. In a cluster with heterogeneous VM sizes (2GB to 64GB), which method provides better capacity utilization and why? What are the trade-offs?
- VCF Ops is a 64GB VM (the largest in your cluster). A second ops tool is proposed that also requires 64GB. How would your HA design need to change? What's the new minimum cluster size?
- Your cluster is sized for N+1 but a compliance requirement emerges: 'Infrastructure must survive simultaneous failure of any 2 hosts.' How do you redesign the cluster? What's the new minimum size and cost implications?
- During a post-mortem after a host failure, you discover that all NSX Manager VMs restarted on a single surviving host, overloading it momentarily. What admission control policy would have prevented this? Should you separate management VMs across multiple failure zones?
- You're asked to design HA for a smaller regional VCF instance with only vCenter and SDDC Manager (no NSX, no Ops). How does this change your admission control calculations? Could a 3-host cluster suffice?
Extensions
Extend to Multi-Region HA Design
Design HA for a VCF environment spanning 2 geographic regions with asynchronous replication. How would you adapt admission control policies? What is minimum cluster size per region?
Integrate Storage Admission Control
Extend the CPU/memory analysis to include storage constraints. If each management VM requires 500GB storage and the cluster has 10TB total, does storage change the minimum cluster size calculation?
Dynamic Workload Scenario
Model a scenario where VCF Ops is expanded from 64GB to 128GB. Recalculate cluster sizing. At what point does slot size change? What remediation steps would you take during the upgrade?
⚠ Known Pitfalls (from Community KB)
References
- vSphere 7.0 High Availability AdministrationTier 1 — Official
Core reference for vSphere HA concepts, admission control, and slot calculation - VCF 9.0 Architecture and SizingTier 1 — Official
Defines minimum management VM profiles and cluster sizing guidelines - vCenter Server 8.0 Hardware RequirementsTier 1 — Official
Reference for vCenter sizing and NSX Manager cluster requirements