Academy/VCP-VCF 9.0 Architect (2V0-13.25)/HA Admission Control Validation
This lab targets VCF 9.0

HA Admission Control Validation

VCF 9.0Intermediatevcp-foundation⏱ 75 min

Objectives

  • Understand HA admission control policies (slot-based and percentage-based)
  • Calculate slot sizes for different VM profiles in a 4-host cluster
  • Evaluate admission control scenarios and policy recommendations
  • Design N+1 HA policy for VCF management domain

Prerequisites

VCF 9.0 lab environment deployed with 4-host management cluster

Prior labs: vcp-architect-01: Cluster Basics

Required skills:

  • Understanding vSphere HA concepts
  • Basic cluster architecture knowledge

Lab Environment

4-host VCF management cluster (each host: 96GB RAM, 16 vCPU cores, 2.5GHz)

Tasks

Task 1 Understand HA Admission Control Policies

Demonstrates ability to analyze admission control mechanisms and understand trade-offs between predictability (slot-based) and flexibility (percentage-based).

Establish foundational knowledge of slot-based and percentage-based admission control mechanisms and their impact on cluster capacity planning.

Step 1

Review the 4-host VCF management cluster configuration. Document: total cluster RAM (384GB), total vCPU count (64 cores), reserved memory for hypervisors per host (8GB), and available memory per host for VMs (88GB per host = 352GB total).

Cluster capacity summary documenting 352GB total available VM memory and 64 vCPU cores available for VMs.
Do not confuse hypervisor reserved memory with VM memory reservations—they are calculated separately.
vSphere reserves approximately 8GB of RAM per host for ESXi hypervisor and management. Always subtract this from total host memory.
Step 2

Define three VM profiles: Small (2vCPU, 4GB RAM), Medium (4vCPU, 16GB RAM), Large (8vCPU, 64GB RAM). For each profile, calculate the slot size in bytes using the formula: Slot Size = MAX(CPU requirement, Memory requirement in bytes / total available memory per host).

Expected

Slot size calculations for each VM profile:

  • Small VM slot: MAX(2/16 cores, 4GB/88GB) = MAX(0.125, 0.045) = 0.125 slots
  • Medium VM slot: MAX(4/16 cores, 16GB/88GB) = MAX(0.25, 0.182) = 0.25 slots
  • Large VM slot: MAX(8/16 cores, 64GB/88GB) = MAX(0.5, 0.727) = 0.727 slots
Using the same slot size for all VMs ignores heterogeneous workload requirements and can lead to inefficient capacity utilization.
The slot size is determined by whichever constraint (CPU or memory) is more restrictive. In this cluster, memory typically becomes the limiting factor for large VMs.
Step 3

Determine the maximum slot size (most resource-intensive VM that could run in this cluster). Using the 4-host cluster, calculate maximum slots available: if maximum slot size = 0.727 (the Large VM requirement), then total available slots = 4 hosts / 0.727 = 5.5 slots (round down to 5 usable slots).

Maximum slot size: 0.727 (for 8vCPU/64GB Large VMs). Total available slots in 4-host cluster: 5 slots. This means the cluster can accommodate 5 Large VMs before hitting admission control threshold.
Confusing total slots with maximum simultaneously admissible VMs is a common error. N+1 failover reduces this number further.
The number of available slots determines absolute cluster capacity in slot-based admission control. A 4-host cluster can hold 5 Large VMs but more Small VMs.
Step 4

Calculate available slots for admission control with N+1 failover (one host failure allowed). With N+1, usable capacity = capacity of (n-1) hosts = 3 hosts. Recalculate available slots for admission control: 3 / 0.727 = 4.1 slots, rounded down to 4 slots. This ensures one failed host still leaves remaining hosts to run all admitted VMs.

With N+1 failover requirement, available slots for admission control = 4 slots for Large VMs. Smaller VMs would allow more slots: 3 hosts / 0.125 slot size (Small VM) = 24 slots available.
Ignoring N+1 requirements leads to HA violations during actual host failures. This is a critical production readiness failure.
N+1 availability is the golden standard for production clusters. Always calculate admission control with one host failure already assumed.

Validation Gate

Check: Verify admission control calculations and N+1 capacity planning

Expected: Documented slot sizes for all VM profiles, total and N+1-compliant slot counts

Common Errors

Calculated slot sizes that don't account for the maximum of CPU and memory constraints
Cause: Misunderstanding that slot size = MAX(CPU%, Memory%), not addition or separate calculations
Fix: Review admission control mechanism: use whichever constraint is larger (CPU cores as % of host, or Memory as % of host). The result is expressed as a fraction of one host.
Calculating available slots as if all 4 hosts are usable for admitted VMs
Cause: Not reserving one entire host's capacity for failover capacity
Fix: Always divide by (n-1) hosts for admission control calculations: 3 hosts / slot size = available slots. The 4th host is reserved for failover.
Stating 'cluster supports 4 slots' but thinking this means 4 VMs total
Cause: Not understanding that slot is a unit of resource allocation, not a VM count. A Small VM takes 0.125 slots.
Fix: Remember: if slot size = 0.125, then 4 slots / 0.125 = 32 Small VMs can be admitted. Slots are units of capacity, not VM instances.

Task 2 Evaluate Admission Control Scenarios

Demonstrates ability to evaluate trade-offs between admission control methods and make data-driven policy recommendations backed by concrete calculations.

Apply admission control calculations to real-world scenarios and compare policy approaches (slot-based vs. percentage-based) for cluster design decisions.

Step 1

Scenario 1: A new 96GB VM is requested for the management cluster. Using slot-based calculation from Task 1, determine if the cluster can admit this VM while maintaining N+1 HA. Slot size for 96GB VM = MAX(1/16, 96/88) = MAX(0.0625, 1.09) = 1.09 slots. Available slots with N+1 = 4. Can the cluster admit this VM? Calculate: 1.09 slots > 4 available slots? No—this VM cannot be admitted.

Determination: 96GB VM cannot be admitted to the 4-host cluster while maintaining N+1 HA. Required slots: 1.09; Available slots: 4. If admitted, a single host failure would leave insufficient capacity.
Attempting to force-admit an oversized VM by disabling HA is a critical production risk. The VM would be unprotected.
VMs larger than the max slot-sized VM in the cluster may be impossible to admit under slot-based admission control, even if raw capacity exists.
Step 2

Calculate the impact of losing 1 host on remaining cluster capacity. Current state: 4 hosts, 4 slots available for Large VMs (N+1 compliant). After host failure: 3 hosts remain, but all admitted VMs must still run on these 3 hosts. If 4 Large VMs are currently admitted (4 slots × 0.727 = 2.9 VMs worth of capacity, let's say 4 Large VMs requiring 4 × 0.727 = 2.9 host equivalents), can all 4 Large VMs run on 3 remaining hosts? Remaining capacity after failover: 3 / 0.727 = 4.1 slots. Yes, 4 admitted Large VMs (using 2.9 host equivalents) fit in 3 remaining hosts.

With one host failure: 3 remaining hosts provide 4.1 slots of capacity. Cluster can support 4 admitted Large VMs. Calculate actual VM count: 4 Large VMs use 4 × 0.727 = 2.9 host equivalents; 3 remaining hosts have 2.18 host equivalents capacity after removing HA reserve.
Oversubscribing during normal operation (admitting 5 Large VMs when only 4 slots available) causes cascading failures when a host fails—excess VMs cannot restart.
During failover, the failed host's VMs are restarted on remaining hosts. Verify that admitted VMs fit in (n-1) hosts, not n hosts.
Step 3

Compare slot-based vs. percentage-based admission control for this cluster. Slot-based: Uses predefined slot size (0.727 for Large VMs); hard limit prevents oversubscription; predictable. Percentage-based: Reserves a percentage (e.g., 25%) of cluster resources for failover; allows heterogeneous workloads without recalculation. For this cluster: Slot-based reserves (1 - 3/4) = 25% of cluster capacity for failover. Percentage-based with 25% reserve achieves same result but with flexibility. Recommendation: Compare predictability vs. flexibility trade-off.

Expected

Comparison table:

  • Slot-based: Predictable, hard limits, requires slot recalculation for new VM types, prevents oversubscription
  • Percentage-based: Flexible, accommodates heterogeneous workloads, simpler to understand, less granular control

For this VCF management cluster with homogeneous VMs (vCenter, NSX, SDDC Manager, Ops), slot-based is more precise.

Percentage-based with insufficient reserve percentage (e.g., 10%) may fail to protect oversized VMs during failover.
Percentage-based is often easier to communicate to stakeholders: 'We reserve 25% capacity for failover' is simpler than explaining slots.
Step 4

Formulate a policy recommendation for this VCF management cluster. Decision: Which admission control method and why? Justification: 'Recommend slot-based admission control for this cluster because: (1) VCF management VMs have well-defined resource profiles (Small=4GB, Medium=16GB, Large=64GB); (2) slot-based provides deterministic capacity planning; (3) prevents accidental oversubscription of large VMs; (4) VCF design already specifies resource tiers.' Document the policy choice with rationale.

Written recommendation: 'Implement slot-based HA admission control with slot size 0.727 for Large VMs (8vCPU/64GB). Reserve 25% of cluster capacity for N+1 failover. This allows 4 Large VMs to be admitted while guaranteeing single-host-failure survivability.'
Never recommend a policy without understanding the specific workload. 'Use percentage-based for flexibility' without knowing the workload requirements is insufficient justification.
Always justify policy recommendations with specific cluster characteristics, workload profiles, and business requirements (e.g., 'VCF management domain requires N+1 availability for 99.9% uptime SLA').

Validation Gate

Check: Verify scenario calculations and policy recommendation

Expected: Documented capacity impact analysis, policy comparison, and justified recommendation

Common Errors

Calculating capacity as if the cluster has 5 hosts worth of resources (4 hosts + 1 extra)
Cause: Misinterpreting N+1 as adding a host rather than reserving one host's worth of capacity
Fix: N+1 means capacity of (n-1) hosts is available for VMs; 1 host worth of capacity is reserved for failover. For 4 hosts: 3 hosts = available capacity, 1 host = reserved.
Recommending percentage-based for a cluster with homogeneous VM sizes, losing precision
Cause: Not analyzing the actual VMs in the cluster before choosing admission control method
Fix: Document the VM profiles first. If VMs are similar sizes (vCenter 2GB, NSX 4GB, Ops 8GB), slot-based is more precise. If highly heterogeneous, percentage-based is more practical.
Recommending a policy without recalculating how many VMs it actually allows
Cause: Separating the 'which policy' decision from the 'how many VMs' consequence
Fix: Always end policy recommendations with concrete capacity statements: 'This policy allows 4 Large VMs' or 'This policy allows up to 100GB of mixed workloads'.

Task 3 Design HA Policy for VCF Management Domain

Produces a professional design decision document (DD format) that specifies HA policy, justifies cluster sizing, documents reservation requirements, and addresses what-if failure scenarios.

Create a comprehensive HA design document for the VCF management domain that demonstrates N+1 availability architecture and applies admission control concepts to production VCF components.

Step 1

Identify all management VMs and their resource requirements in the VCF management domain. List: vCenter Server (8vCPU, 32GB RAM), NSX Manager cluster nodes (8vCPU each, 16GB RAM each, 3 nodes = 24vCPU/48GB), SDDC Manager (8vCPU, 32GB RAM), VCF Ops (formerly vRealize Operations) (16vCPU, 64GB RAM). Total: (8+24+8+16) = 56 vCPU and (32+48+32+64) = 176GB RAM across management VMs.

Expected

Documented inventory of management domain VMs with resource specifications. Example table:
| VM | vCPU | RAM GB | Notes |
| vCenter | 8 | 32 | HA required |
| NSX Manager-1 | 8 | 16 | Part of 3-node cluster |
| NSX Manager-2 | 8 | 16 | Part of 3-node cluster |
| NSX Manager-3 | 8 | 16 | Part of 3-node cluster |
| SDDC Manager | 8 | 32 | HA required |
| VCF Ops | 16 | 64 | Shared infrastructure service |

Underestimating management VM resource requirements can cause uncontrolled restarts during host failure when there's insufficient spare capacity.
VCF NSX Managers require quorum (N+1 minimum = 3 nodes). If any single NSX Manager fails, remaining 2 are still quorum-capable.
Step 2

Calculate the minimum cluster size to maintain N+1 availability for all management VMs. Total management VM capacity needed: 176GB RAM + 56vCPU. With N+1 reserve (reserve 1 host): If each host = 96GB/16vCPU, then minimum usable capacity = 176GB RAM. Divide by host capacity: 176GB / 96GB = 1.83 host equivalents needed. Add N+1 reserve: 1.83 + 1 = 2.83, round up to 3 hosts minimum. Verify vCPU: 56vCPU / 16vCPU per host = 3.5 host equivalents; 3.5 + 1 reserve = 4.5, round up to 5 hosts. Thus, minimum cluster size = 5 hosts to accommodate all management VMs + N+1 failover reserve.

Minimum cluster sizing calculation: Management workload = 176GB RAM (equivalent to 1.83 hosts). With N+1 reserve = 2.83 host equivalents. vCPU constraint = 56 vCPU (3.5 hosts); with N+1 = 4.5 host equivalents. Determine limiting factor: vCPU (need 5 hosts) is more restrictive than memory (need 3 hosts). Minimum cluster size = 5 hosts.
Calculating minimum size for only one resource type (e.g., memory) and ignoring CPU can lead to CPU oversubscription during failover.
Always calculate for both memory and vCPU; use the larger result. Often vCPU is the limiting factor for compute-intensive workloads like NSX Managers.
Step 3

Document the HA design decision (DD format). Create a formal design decision document with sections: (1) Requirement: 'Maintain N+1 availability for VCF management domain to support 99.95% uptime SLA (max 22 minutes downtime/month)'; (2) Constraints: 'Cluster must survive single host failure without VM restart cascades'; (3) Assumptions: 'VMs will be distributed across all 5 hosts (no single-host concentration)'; (4) Design Decision: 'Deploy 5-host management cluster with slot-based HA admission control. Max slot size = 64GB (VCF Ops VM). Reserve 1 host capacity for failover. Admit only management VMs to this cluster (no user workloads)'; (5) Justification: 'Sizing calculation: 176GB management workload + 64GB + 96GB reserve = 336GB; 5 hosts × 96GB = 480GB available. N+1 compliant with 144GB spare capacity for growth.'

Expected

Formal HA Design Decision document:
VCF Management Domain HA Design (v9.0)
Requirement: N+1 availability for VCF management domain
Current State: 4-host cluster (insufficient for management workload)
Proposed Design: 5-host cluster with slot-based HA admission control
Rationale: Management workload totals 176GB + largest VM (VCF Ops, 64GB) + 1-host reserve = 336GB required; 5 hosts × 96GB = 480GB available. Provides N+1 failover coverage.
Impact: Requires upgrade from 4-host to 5-host cluster.

Vague designs like 'Add more hosts' without sizing calculations fail VCDX review. Always provide concrete numbers and justifications.
Design decision documents are expected in VCDX defense. Use formal structure: Requirement, Constraint, Assumption, Design Decision, Justification, Risk/Mitigation. Panelists probe how you arrived at your sizing and whether you considered failure scenarios.
Step 4

Test design with what-if scenario: What if 2 hosts fail simultaneously (N+2)? Current design assumes N+1 (1 host failure). With 2 host failures: 5 hosts - 2 = 3 hosts remain. Remaining capacity = 3 × 96GB = 288GB. Management VMs = 176GB. Remaining capacity (288GB - 176GB = 112GB) is sufficient. However, this exceeds N+1 design goal. Document: 'Current HA design supports N+1 (single host failure). N+2 survivability is not a requirement per business SLA but is a positive secondary outcome. If N+2 requirement emerges, cluster would need to grow to 6 hosts.'

What-if analysis: N+2 scenario. With 2 host failures: 3 remaining hosts = 288GB capacity. Management workload = 176GB. Conclusion: Design incidentally supports N+2, but is not validated for N+2 SLA. Document as 'secondary resilience, not primary design goal.' If customer requires N+2, update design to 6 hosts or larger.
Do not claim N+2 support as a design goal if you did not explicitly plan for it. Accurately represent your design as N+1 compliant, with incidental N+2 tolerance.
VCDX panelists appreciate when you analyze what-if scenarios proactively. Showing that your design handles N+2 as a bonus is valuable.

Validation Gate

Check: Verify HA design document and what-if scenario analysis

Expected: Formal design decision document with justification and what-if scenario analysis

Common Errors

Calculating minimum cluster size = 2 hosts (enough for 176GB management VMs) without adding N+1 reserve
Cause: Not including the (n-1) calculation; forgetting that 1 host must be reserved
Fix: Always add 1 full host to your sizing: If workload fits in 1.83 hosts, you need 2.83 hosts minimum. Round up to 3, then add reserve if not already accounted for. Use formula: Workload host equivalents + 1 = minimum sizing.
Proposing a 3-host management cluster, which can support NSX if only 1-2 NSX Managers fail
Cause: Not understanding that NSX Manager needs 3 nodes for HA and the cluster itself needs reserve capacity separate from NSX HA
Fix: NSX Manager HA (3-node cluster) is separate from vSphere cluster HA (N+1 failover reserve). Size the vSphere cluster to hold 3 NSX Managers PLUS reserve 1 host. Total: 3 NSX nodes + other management VMs + 1 reserve host.
Providing a cluster size (5 hosts) without explaining assumptions about VM distribution, workload consolidation, or resource allocation
Cause: Jumping to conclusion without showing decision logic
Fix: Always document assumptions explicitly: 'Assume all management VMs distributed across 5 hosts, no single-host concentration.' 'Assume no over-provisioning of vCPU (1:1 allocation).' This allows panelists to probe your reasoning and ask 'what if we over-provision?'

Final Validation

HA admission control lab completed with slot-based analysis, scenario evaluation, and VCF management cluster design

✓ Admission control calculations (Task 1, Steps 1-4) → Documented slot sizes for all VM profiles, N+1 capacity calculations

✓ Scenario evaluation and policy recommendation (Task 2, Steps 1-4) → 96GB VM admissibility determination, failover impact analysis, slot vs. percentage comparison, and justified policy choice

✓ VCF management design decision (Task 3, Steps 1-4) → Management VM inventory, minimum cluster sizing, formal DD document, N+2 what-if scenario

Cleanup / Restore

• Revert to initial snapshot if any lab VMs were created

• Document all design decisions and calculations for VCDX study

Design Reflection (VCDX)

During a VCDX defense, panelists will probe your HA design with questions like: 'You sized the cluster for 5 hosts. What happens if a host fails during a maintenance window when a 3rd host is offline?' Or: 'How would you handle a scenario where NSX Managers are concentrated on 2 hosts and one host fails?' Be prepared to explain admission control trade-offs, justify why you chose slot-based over percentage-based, and articulate what would change if requirements shifted to N+2 or if workload profiles expanded. Panelists expect designers to show flexibility—how would you adjust the design if vCF Ops expanded from 64GB to 128GB?

Requirements

  • VCF management domain must maintain N+1 availability for 99.95% uptime SLA
  • Cluster must support heterogeneous management VMs (vCenter, NSX Managers, SDDC Manager, VCF Ops)
  • Admission control must prevent oversized VM admissions that would violate HA during failover
  • Design must be documented in formal DD format with sizing calculations and assumptions
  • Must address minimum cluster size, slot calculation methodology, and what-if failure scenarios

Constraints

  • Each host in the management cluster has fixed capacity (96GB RAM, 16 vCPU cores)
  • NSX Manager requires 3-node cluster for quorum and HA
  • VCF management domain must not share hosts with user workloads (isolation requirement)
  • Admission control policy chosen must apply consistently across all admitted VMs
  • One full host capacity must always be reserved for failover (N+1 requirement)

Assumptions

  • All management VMs will be distributed across available hosts (no single-host concentration)
  • Hypervisor reserves ~8GB per host, leaving 88GB available for VMs
  • vCPU allocation is 1:1 (no oversubscription) for deterministic performance
  • Management VM resource requirements are stable (no dynamic scaling)
  • Slot size is calculated once and remains stable (no mid-cluster addition of different-sized VMs)
  • Network and storage are not the limiting factors (CPU and memory are the constraints)
  • Single host failure is the design-basis failure scenario; N+2 is not required

Risks

  • Risk: Cluster growth adds larger VMs without recalculating slot size. Mitigation: Document when slot size becomes invalid and requires cluster redesign.
  • Risk: VCF Ops or another management VM is unexpectedly resized upward. Mitigation: Establish change control requiring HA impact analysis before VM resize in management domain.
  • Risk: Operational team over-admits VMs by disabling HA or ignoring admission control. Mitigation: Use vSphere admission control policies that are enforced and non-bypassable without explicit approval.
  • Risk: Network or storage becomes the limiting factor instead of CPU/memory. Mitigation: Extend this analysis to validate network (NSX bandwidth) and storage (IOPS, capacity) separately.
  • Risk: A third host fails within 15 minutes of the first host failure (before first host VMs have fully restarted). Mitigation: Monitor first host failure recovery time; if recovery takes >15 min and concurrent failures are possible, increase cluster size or add recovery time SLA.

Self-Assessment Discussion Prompts

  1. You proposed a 5-host management cluster to maintain N+1 availability. A business stakeholder asks: 'Why do we need 5 hosts? We currently run on 4, and nothing has failed in 2 years.' How would you justify the upgrade? What specific risks does the 4-host cluster face during a single host failure?
  2. Compare slot-based vs. percentage-based admission control. In a cluster with heterogeneous VM sizes (2GB to 64GB), which method provides better capacity utilization and why? What are the trade-offs?
  3. VCF Ops is a 64GB VM (the largest in your cluster). A second ops tool is proposed that also requires 64GB. How would your HA design need to change? What's the new minimum cluster size?
  4. Your cluster is sized for N+1 but a compliance requirement emerges: 'Infrastructure must survive simultaneous failure of any 2 hosts.' How do you redesign the cluster? What's the new minimum size and cost implications?
  5. During a post-mortem after a host failure, you discover that all NSX Manager VMs restarted on a single surviving host, overloading it momentarily. What admission control policy would have prevented this? Should you separate management VMs across multiple failure zones?
  6. You're asked to design HA for a smaller regional VCF instance with only vCenter and SDDC Manager (no NSX, no Ops). How does this change your admission control calculations? Could a 3-host cluster suffice?

Extensions

Extend to Multi-Region HA Design

Design HA for a VCF environment spanning 2 geographic regions with asynchronous replication. How would you adapt admission control policies? What is minimum cluster size per region?

Integrate Storage Admission Control

Extend the CPU/memory analysis to include storage constraints. If each management VM requires 500GB storage and the cluster has 10TB total, does storage change the minimum cluster size calculation?

Dynamic Workload Scenario

Model a scenario where VCF Ops is expanded from 64GB to 128GB. Recalculate cluster sizing. At what point does slot size change? What remediation steps would you take during the upgrade?

⚠ Known Pitfalls (from Community KB)

Confusing admission control reserve with spare capacity. Reserve capacity (N+1) is not spare; it's guaranteed to be held unused in steady state to support failover.
Problem: Teams provisioning 'extra' hosts for growth while failing to reserve 1 host for HA end up with oversubscribed clusters that fail under single host failure.
Resolution: Establish cluster sizing policy: 'Total capacity = Workload + 1 host reserve.' Growth requires adding hosts in pairs or clusters of 3+ to maintain reserve.
Assuming all VMs fit in one slot and using a single cluster-wide slot size. If VMs vary from 2GB to 64GB, using the 64GB slot size wastes capacity for small VMs.
Problem: Capacity is underutilized. A cluster that could admit 10 small VMs may be designed to admit only 4 large VMs, wasting 60% of available capacity.
Resolution: Use different admission control thresholds for different VM classes, or use percentage-based admission control for heterogeneous workloads.
Not communicating HA design decisions to the operations team. If ops doesn't understand why the cluster is sized for N+1, they may disable HA or over-admit VMs to 'get more capacity.'
Problem: Operational teams accidentally reduce HA compliance, violating SLAs unintentionally.
Resolution: Provide training and documentation. Create alerts if admission control is disabled or if admitted VM count exceeds design threshold.
Designing HA for current workload only, without growth buffer. If management workload grows 20% in year 2, the cluster is suddenly oversubscribed.
Problem: Cluster becomes non-compliant with N+1 SLA mid-year, requiring emergency host addition.
Resolution: Size clusters with 20-30% growth headroom: 'Design for 176GB current workload + 50GB forecast growth = 226GB design basis.' Establish annual review cadence to reassess sizing.

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.