Academy/VCAP — VCF Architect (3V0-12.26)/Lab: AMPRS Design — HA, DR, Security Architecture
This lab targets VCF 9.0

Lab: AMPRS Design — HA, DR, Security Architecture

VCF 9.0Advancedvcap-advancedvcdx⏱ 120 min

AMPRS deep-dive lab — vSphere HA/DRS design, DR with VMware Live Recovery, NSX security architecture, encryption strategy. VCF 9.0.

Objectives

  • Design vSphere HA admission control policies for different availability SLA tiers
  • Architect DR topology using VMware Live Recovery with RPO/RTO mapping to protection mechanisms
  • Design NSX micro-segmentation architecture with trust zones and DFW rule categories
  • Create an encryption strategy covering VM encryption, vSAN encryption, and network encryption
  • Document AMPRS trade-offs for each design decision

Prerequisites

Holodeck VCF 9.0 lab with management + workload domain deployed. NSX operational with T0/T1 gateways. VMs deployed for HA/DRS testing.

Prior labs: vcap-architect-01, vcap-architect-02

Required skills:

  • vSphere HA/DRS concepts
  • NSX DFW and gateway firewall
  • DR concepts (RPO, RTO, replication)
  • Encryption fundamentals (at-rest, in-transit)

Lab Environment

VCF 9.0 management domain + 1 workload domain. 4+ hosts in workload domain for HA testing. NSX overlay networking operational. VMs deployed across hosts.

Tasks

Task 1 Design and Validate AMPRS Architecture

AMPRS design is where architecture decisions have the most visible impact on operations. Over-designing availability (e.g., FT for all VMs) wastes resources; under-designing (no HA admission control) risks downtime during host failure. The architect must balance each AMPRS dimension against cost and complexity.

Design the availability, security, and recoverability architecture for a VCF deployment — configuring HA admission control, DRS automation, NSX security zones, encryption layers, and DR topology. This lab combines hands-on configuration with design documentation to practice both implementation and defense skills.

Step 1
Design HA Admission Control for the Clinical workload domain. In vSphere Client → Cluster → Configure → vSphere Availability. HA Admission Control policy options: (a) Host Failures Cluster Tolerates: set to 1 (N+1) for 8-host cluster — reserves 12.5% capacity. (b) Percentage of Cluster Resources Reserved: CPU 15%, Memory 15% (more precise control). (c) Dedicated Failover Hosts: designate 1 specific host as standby (deterministic but wastes capacity). Design Decision: choose 'Host Failures Cluster Tolerates = 1' — it's the recommended approach for VCF, automatically calculates reservation based on largest host. Document: justification, trade-off (simplicity vs precision of percentage-based), AMPRS impact (Availability ↑, Performance → capacity reserved reduces available resources).
Step 2

Configure HA advanced settings for tiered protection. (a) VM Monitoring: enable — restarts VMs that stop sending VMware Tools heartbeats (application-level HA). Set sensitivity to 'High' for EHR VMs (30-second failure detection), 'Low' for dev VMs (120-second). (b) Proactive HA: enable — pre-migrates VMs from hosts reporting hardware warnings via IPMI/BMC. Set 'Automated' for EHR VMs (migrate immediately on warning), 'Manual' for development VMs. (c) VM Component Protection (VMCP): enable for APD (All Paths Down) and PDL (Permanent Device Loss). APD response: 'Power off and restart VMs' after 140-second timeout. Document each configuration as a design decision with AMPRS impact.

Step 3

Design DRS resource allocation. (a) DRS Automation Level: 'Fully Automated' for the workload cluster — DRS places and balances VMs without admin intervention. Migration threshold: set to 3 (moderate aggressiveness — balances load distribution vs. migration overhead). (b) Resource Pools: create 'RP-Clinical-Tier1' (shares: High, reservation: 50% of cluster CPU/memory for EHR/DB VMs), 'RP-Clinical-General' (shares: Normal, no reservation), 'RP-Development' (shares: Low, expandable reservation). (c) VM-Host Affinity: create 'Should' rule keeping EHR VMs in rack 1 fault domain for data locality (not 'Must' — HA must be able to override). Document DRS design decisions.

Step 4
Design DR topology with VMware Live Recovery. DR Architecture: Primary site (active) ← asynchronous replication → DR site (standby). Component-level DR: vCenter: File-Based Backup to NFS share replicated to DR site (manual restore, RTO ~30 min). NSX: configuration backup replicated to DR site (manual redeploy, RTO ~1h). VCF Operations: Cassandra backup replicated (manual restore, RTO ~1h). Workload DR: Tier-1 (EHR): vSphere Replication with RPO=15 min to DR site → Live Recovery plan for automated failover. Tier-2 (Admin): vSphere Replication with RPO=4h → Live Recovery plan. Tier-3 (Dev): backup-only, no replication (RPO=24h, restore from backup). Document per-tier protection with RPO/RTO justification.
Step 5
Design Live Recovery plans. Create recovery plan structure: (a) Plan: 'DR-Clinical-EHR' — recovery priority groups: Group 1 (DNS, AD, DHCP — infrastructure VMs, boot first), Group 2 (database servers — boot after infrastructure, wait for DB service start), Group 3 (application servers — boot after DB, verify connectivity), Group 4 (web servers — boot last, verify external access). (b) IP customization: DR site uses different subnet — configure per-VM IP mapping (production 172.16.1.x → DR 10.20.1.x). (c) Pre-power-on script: verify storage is accessible. Post-power-on script: validate application health endpoint. (d) Test recovery: schedule monthly non-disruptive DR test — Live Recovery creates isolated network bubble for testing without affecting production replication.
Step 6
Design NSX security architecture with trust zones. Define trust zones: (a) Zone-Untrusted: external traffic, internet-facing (T0 gateway firewall). (b) Zone-DMZ: web tier VMs (limited inbound from untrusted, outbound to app tier only). (c) Zone-Application: application servers (inbound from DMZ only, outbound to DB tier only). (d) Zone-Data: database servers (inbound from application tier only, no external access). (e) Zone-Management: VCF management components (restricted access, admin-only). Map zones to NSX DFW categories: Emergency → Infrastructure → Environment → Application → Default. Design rule: Emergency rules for quarantine (block compromised VMs), Infrastructure rules for management access, Environment rules for zone isolation, Application rules for intra-zone communication.
Step 7
Design DFW rule sets per zone. (a) Infrastructure category: Allow vCenter/NSX/VCF Ops management traffic from Zone-Management to all zones (required for operations). (b) Environment category: 'Zone-Isolation-DMZ-to-App': Source=Zone-DMZ, Dest=Zone-Application, Service=HTTPS(443), Action=Allow. 'Zone-Isolation-App-to-Data': Source=Zone-Application, Dest=Zone-Data, Service=MySQL(3306)/PostgreSQL(5432), Action=Allow. 'Zone-Default-Deny': Source=Any, Dest=Any, Action=Drop, Applied To=All workload VMs. (c) Application category: intra-zone rules (e.g., web server cluster communication on port 8080). Design Decision: default-deny posture with explicit allow rules — zero-trust model. Document: AMPRS impact (Security ↑, Manageability ↓ — every new application requires rule creation).
Step 8

Design encryption strategy. Layers: (a) Data at rest: vSAN encryption enabled — uses KMS (Key Management Server) for key management. Design Decision: use vSAN native encryption (AES-256-XTS) rather than VM-level encryption to avoid double encryption overhead. KMS: deploy KMIP-compliant KMS cluster (2 nodes for HA) in management domain. Key rotation: annual, automated via KMS policy. (b) Data in transit: vSAN traffic encryption (enabled per cluster), vMotion encryption (required — enabled by default in VCF 9.0), management traffic via TLS 1.2+ (enforced by VCF). (c) VM-level encryption: reserve for Tier-1 VMs requiring per-VM key management (HIPAA PHI workloads). Uses vTPM for key sealing. Note: VM encryption prevents vSAN deduplication/compression — storage efficiency trade-off.

Step 9
Document AMPRS trade-off matrix. Create a table: Design Decision | Availability Impact | Manageability Impact | Performance Impact | Recoverability Impact | Security Impact. Fill for each major decision: (a) HA Admission Control N+1: A↑ M→ P↓(capacity reserved) R→ S→. (b) DRS Fully Automated: A→ M↑(less manual) P↑(balanced load) R→ S→. (c) vSAN Encryption: A→ M↓(KMS dependency) P↓(~5% CPU overhead) R→(backup must handle encrypted data) S↑. (d) NSX Zero-Trust DFW: A→ M↓(rule management) P→ R→ S↑↑. (e) Live Recovery RPO=15min: A↑ M↓(replication monitoring) P↓(replication I/O overhead) R↑↑ S→. Identify which AMPRS dimensions conflict most in your design.

Validation Gate

Check: AMPRS architecture documented with HA, DR, security, and encryption design decisions

Expected: HA admission control configured with tiered VM protection policies. DR topology designed with per-tier RPO/RTO mapping. NSX security zones defined with DFW rule categories. Encryption strategy covering at-rest, in-transit, and per-VM layers. AMPRS trade-off matrix completed.

Common Errors

HA admission control disabled to maximize usable capacity
Fix: Without admission control, HA may not have sufficient resources to restart all VMs after a host failure. This is the most common HA misconfiguration. Always reserve capacity — the 12.5% 'cost' of N+1 in an 8-host cluster is insurance against downtime.
Using 'Must' affinity rules that prevent HA failover
Fix: 'Must' rules are absolute — HA cannot override them. If the target host group is down, VMs cannot restart. Use 'Should' rules for recommendations that HA can override during failure. Reserve 'Must' only for licensing or regulatory requirements.
VM encryption enabled alongside vSAN encryption (double encryption)
Fix: Double encryption wastes CPU cycles without additional security benefit. Choose one layer: vSAN encryption for blanket coverage (all VMs on the datastore), VM encryption only for specific VMs requiring per-VM key management. Document the decision.
DR plan missing IP customization for subnet change
Fix: If the DR site uses different IP subnets (common), VMs must be re-IPed during failover. Live Recovery supports per-VM IP customization — configure this in the recovery plan. Missing IP customization = VMs boot at DR site with wrong IPs = no connectivity.

Final Validation

Complete AMPRS architecture with HA, DR, security, and encryption design

✓ HA design → Admission control configured, tiered VM monitoring, Proactive HA enabled

✓ DRS design → Fully automated with resource pools and soft affinity rules

✓ DR topology → Per-tier RPO/RTO mapping with Live Recovery plans and test schedule

✓ NSX security → Trust zones defined, DFW categories mapped, zero-trust default-deny posture

✓ Encryption strategy → At-rest (vSAN), in-transit (vMotion/vSAN), per-VM (HIPAA) with KMS architecture

✓ AMPRS trade-off matrix → Impact analysis for each major design decision across all 5 dimensions

Cleanup / Restore

• Revert HA/DRS settings if modified in lab environment

• Remove test DFW rules

• Save design documentation

Design Reflection (VCDX)

AMPRS design is the most heavily scrutinized part of any VCDX defense. Panelists will probe: 'What happens when your KMS cluster is unavailable — can VMs still boot?' (answer: vSAN encryption key cache survives KMS outage for 30 days, but new VMs cannot be encrypted). They'll challenge: 'Why RPO=15 min for Tier-1 instead of RPO=0 with stretched cluster?' Be ready with cost/complexity trade-off analysis.

Requirements

  • 99.99% availability for clinical applications
  • HIPAA compliance for PHI workloads
  • DR with RPO ≤ 15 min / RTO ≤ 1h for Tier-1

Constraints

  • KMS infrastructure required for encryption — adds 2 VMs to management domain
  • vSphere Replication RPO minimum is 5 minutes (cannot achieve RPO < 5 min without stretched cluster)

Assumptions

  • Network bandwidth sufficient for replication traffic at RPO=15 min
  • KMS vendor supports KMIP 1.1+ for vSAN integration

Risks

  • KMS outage prevents new VM encryption operations — mitigate with HA KMS cluster
  • DFW rule complexity grows unmanageable — mitigate with NSX Intelligence for flow-based rule optimization

Self-Assessment Discussion Prompts

  1. When should you choose vSAN stretched cluster (RPO=0) vs vSphere Replication (RPO ≥ 5 min)?
  2. How do you handle DFW rule lifecycle — who creates rules, who reviews, who decommissions?
  3. What is the performance impact of vSAN encryption and how do you measure it?
  4. How does Proactive HA interact with DRS and anti-affinity rules during a pre-failure migration?

Extensions

Design a vSAN stretched cluster alternative for RPO=0 and compare with the replication-based approach

Implement NSX IDS/IPS (vDefend) for the DMZ trust zone with signature-based threat detection

Build an automated DR test runbook that validates failover monthly without admin intervention

Design certificate-based mutual TLS between application tiers using NSX service insertion

⚠ Known Pitfalls (from Community KB)

Disabling HA admission control to 'maximize capacity' — guarantees VM restart failure during host outage
Using 'Must' affinity rules without understanding HA override limitations — can prevent VM failover
Enabling both vSAN and VM encryption on same datastore — double encryption wastes CPU with no security benefit
Designing DR without testing — untested DR plans have high failure rates; schedule monthly non-disruptive tests

References

  • vSphere HA Design Guide: techdocs.broadcom.com
  • VMware Live Recovery Administration Guide: techdocs.broadcom.com
  • NSX Security Design Guide: techdocs.broadcom.com
  • vSAN Encryption and KMS Integration: techdocs.broadcom.com
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.