Lab: AMPRS Design — HA, DR, Security Architecture
Objectives
- Design vSphere HA admission control policies for different availability SLA tiers
- Architect DR topology using VMware Live Recovery with RPO/RTO mapping to protection mechanisms
- Design NSX micro-segmentation architecture with trust zones and DFW rule categories
- Create an encryption strategy covering VM encryption, vSAN encryption, and network encryption
- Document AMPRS trade-offs for each design decision
Prerequisites
Holodeck VCF 9.0 lab with management + workload domain deployed. NSX operational with T0/T1 gateways. VMs deployed for HA/DRS testing.
Prior labs: vcap-architect-01, vcap-architect-02
Required skills:
- vSphere HA/DRS concepts
- NSX DFW and gateway firewall
- DR concepts (RPO, RTO, replication)
- Encryption fundamentals (at-rest, in-transit)
Lab Environment
VCF 9.0 management domain + 1 workload domain. 4+ hosts in workload domain for HA testing. NSX overlay networking operational. VMs deployed across hosts.
Tasks
Task 1 Design and Validate AMPRS Architecture
Design the availability, security, and recoverability architecture for a VCF deployment — configuring HA admission control, DRS automation, NSX security zones, encryption layers, and DR topology. This lab combines hands-on configuration with design documentation to practice both implementation and defense skills.
Design HA Admission Control for the Clinical workload domain. In vSphere Client → Cluster → Configure → vSphere Availability. HA Admission Control policy options: (a) Host Failures Cluster Tolerates: set to 1 (N+1) for 8-host cluster — reserves 12.5% capacity. (b) Percentage of Cluster Resources Reserved: CPU 15%, Memory 15% (more precise control). (c) Dedicated Failover Hosts: designate 1 specific host as standby (deterministic but wastes capacity). Design Decision: choose 'Host Failures Cluster Tolerates = 1' — it's the recommended approach for VCF, automatically calculates reservation based on largest host. Document: justification, trade-off (simplicity vs precision of percentage-based), AMPRS impact (Availability ↑, Performance → capacity reserved reduces available resources).
Configure HA advanced settings for tiered protection. (a) VM Monitoring: enable — restarts VMs that stop sending VMware Tools heartbeats (application-level HA). Set sensitivity to 'High' for EHR VMs (30-second failure detection), 'Low' for dev VMs (120-second). (b) Proactive HA: enable — pre-migrates VMs from hosts reporting hardware warnings via IPMI/BMC. Set 'Automated' for EHR VMs (migrate immediately on warning), 'Manual' for development VMs. (c) VM Component Protection (VMCP): enable for APD (All Paths Down) and PDL (Permanent Device Loss). APD response: 'Power off and restart VMs' after 140-second timeout. Document each configuration as a design decision with AMPRS impact.
Design DRS resource allocation. (a) DRS Automation Level: 'Fully Automated' for the workload cluster — DRS places and balances VMs without admin intervention. Migration threshold: set to 3 (moderate aggressiveness — balances load distribution vs. migration overhead). (b) Resource Pools: create 'RP-Clinical-Tier1' (shares: High, reservation: 50% of cluster CPU/memory for EHR/DB VMs), 'RP-Clinical-General' (shares: Normal, no reservation), 'RP-Development' (shares: Low, expandable reservation). (c) VM-Host Affinity: create 'Should' rule keeping EHR VMs in rack 1 fault domain for data locality (not 'Must' — HA must be able to override). Document DRS design decisions.
Design DR topology with VMware Live Recovery. DR Architecture: Primary site (active) ← asynchronous replication → DR site (standby). Component-level DR: vCenter: File-Based Backup to NFS share replicated to DR site (manual restore, RTO ~30 min). NSX: configuration backup replicated to DR site (manual redeploy, RTO ~1h). VCF Operations: Cassandra backup replicated (manual restore, RTO ~1h). Workload DR: Tier-1 (EHR): vSphere Replication with RPO=15 min to DR site → Live Recovery plan for automated failover. Tier-2 (Admin): vSphere Replication with RPO=4h → Live Recovery plan. Tier-3 (Dev): backup-only, no replication (RPO=24h, restore from backup). Document per-tier protection with RPO/RTO justification.
Design Live Recovery plans. Create recovery plan structure: (a) Plan: 'DR-Clinical-EHR' — recovery priority groups: Group 1 (DNS, AD, DHCP — infrastructure VMs, boot first), Group 2 (database servers — boot after infrastructure, wait for DB service start), Group 3 (application servers — boot after DB, verify connectivity), Group 4 (web servers — boot last, verify external access). (b) IP customization: DR site uses different subnet — configure per-VM IP mapping (production 172.16.1.x → DR 10.20.1.x). (c) Pre-power-on script: verify storage is accessible. Post-power-on script: validate application health endpoint. (d) Test recovery: schedule monthly non-disruptive DR test — Live Recovery creates isolated network bubble for testing without affecting production replication.
Design NSX security architecture with trust zones. Define trust zones: (a) Zone-Untrusted: external traffic, internet-facing (T0 gateway firewall). (b) Zone-DMZ: web tier VMs (limited inbound from untrusted, outbound to app tier only). (c) Zone-Application: application servers (inbound from DMZ only, outbound to DB tier only). (d) Zone-Data: database servers (inbound from application tier only, no external access). (e) Zone-Management: VCF management components (restricted access, admin-only). Map zones to NSX DFW categories: Emergency → Infrastructure → Environment → Application → Default. Design rule: Emergency rules for quarantine (block compromised VMs), Infrastructure rules for management access, Environment rules for zone isolation, Application rules for intra-zone communication.
Design DFW rule sets per zone. (a) Infrastructure category: Allow vCenter/NSX/VCF Ops management traffic from Zone-Management to all zones (required for operations). (b) Environment category: 'Zone-Isolation-DMZ-to-App': Source=Zone-DMZ, Dest=Zone-Application, Service=HTTPS(443), Action=Allow. 'Zone-Isolation-App-to-Data': Source=Zone-Application, Dest=Zone-Data, Service=MySQL(3306)/PostgreSQL(5432), Action=Allow. 'Zone-Default-Deny': Source=Any, Dest=Any, Action=Drop, Applied To=All workload VMs. (c) Application category: intra-zone rules (e.g., web server cluster communication on port 8080). Design Decision: default-deny posture with explicit allow rules — zero-trust model. Document: AMPRS impact (Security ↑, Manageability ↓ — every new application requires rule creation).
Design encryption strategy. Layers: (a) Data at rest: vSAN encryption enabled — uses KMS (Key Management Server) for key management. Design Decision: use vSAN native encryption (AES-256-XTS) rather than VM-level encryption to avoid double encryption overhead. KMS: deploy KMIP-compliant KMS cluster (2 nodes for HA) in management domain. Key rotation: annual, automated via KMS policy. (b) Data in transit: vSAN traffic encryption (enabled per cluster), vMotion encryption (required — enabled by default in VCF 9.0), management traffic via TLS 1.2+ (enforced by VCF). (c) VM-level encryption: reserve for Tier-1 VMs requiring per-VM key management (HIPAA PHI workloads). Uses vTPM for key sealing. Note: VM encryption prevents vSAN deduplication/compression — storage efficiency trade-off.
Document AMPRS trade-off matrix. Create a table: Design Decision | Availability Impact | Manageability Impact | Performance Impact | Recoverability Impact | Security Impact. Fill for each major decision: (a) HA Admission Control N+1: A↑ M→ P↓(capacity reserved) R→ S→. (b) DRS Fully Automated: A→ M↑(less manual) P↑(balanced load) R→ S→. (c) vSAN Encryption: A→ M↓(KMS dependency) P↓(~5% CPU overhead) R→(backup must handle encrypted data) S↑. (d) NSX Zero-Trust DFW: A→ M↓(rule management) P→ R→ S↑↑. (e) Live Recovery RPO=15min: A↑ M↓(replication monitoring) P↓(replication I/O overhead) R↑↑ S→. Identify which AMPRS dimensions conflict most in your design.
Validation Gate
Check: AMPRS architecture documented with HA, DR, security, and encryption design decisions
Expected: HA admission control configured with tiered VM protection policies. DR topology designed with per-tier RPO/RTO mapping. NSX security zones defined with DFW rule categories. Encryption strategy covering at-rest, in-transit, and per-VM layers. AMPRS trade-off matrix completed.
Common Errors
Final Validation
Complete AMPRS architecture with HA, DR, security, and encryption design
✓ HA design → Admission control configured, tiered VM monitoring, Proactive HA enabled
✓ DRS design → Fully automated with resource pools and soft affinity rules
✓ DR topology → Per-tier RPO/RTO mapping with Live Recovery plans and test schedule
✓ NSX security → Trust zones defined, DFW categories mapped, zero-trust default-deny posture
✓ Encryption strategy → At-rest (vSAN), in-transit (vMotion/vSAN), per-VM (HIPAA) with KMS architecture
✓ AMPRS trade-off matrix → Impact analysis for each major design decision across all 5 dimensions
Cleanup / Restore
• Revert HA/DRS settings if modified in lab environment
• Remove test DFW rules
• Save design documentation
Design Reflection (VCDX)
AMPRS design is the most heavily scrutinized part of any VCDX defense. Panelists will probe: 'What happens when your KMS cluster is unavailable — can VMs still boot?' (answer: vSAN encryption key cache survives KMS outage for 30 days, but new VMs cannot be encrypted). They'll challenge: 'Why RPO=15 min for Tier-1 instead of RPO=0 with stretched cluster?' Be ready with cost/complexity trade-off analysis.
Requirements
- 99.99% availability for clinical applications
- HIPAA compliance for PHI workloads
- DR with RPO ≤ 15 min / RTO ≤ 1h for Tier-1
Constraints
- KMS infrastructure required for encryption — adds 2 VMs to management domain
- vSphere Replication RPO minimum is 5 minutes (cannot achieve RPO < 5 min without stretched cluster)
Assumptions
- Network bandwidth sufficient for replication traffic at RPO=15 min
- KMS vendor supports KMIP 1.1+ for vSAN integration
Risks
- KMS outage prevents new VM encryption operations — mitigate with HA KMS cluster
- DFW rule complexity grows unmanageable — mitigate with NSX Intelligence for flow-based rule optimization
Self-Assessment Discussion Prompts
- When should you choose vSAN stretched cluster (RPO=0) vs vSphere Replication (RPO ≥ 5 min)?
- How do you handle DFW rule lifecycle — who creates rules, who reviews, who decommissions?
- What is the performance impact of vSAN encryption and how do you measure it?
- How does Proactive HA interact with DRS and anti-affinity rules during a pre-failure migration?
Extensions
Design a vSAN stretched cluster alternative for RPO=0 and compare with the replication-based approach
Implement NSX IDS/IPS (vDefend) for the DMZ trust zone with signature-based threat detection
Build an automated DR test runbook that validates failover monthly without admin intervention
Design certificate-based mutual TLS between application tiers using NSX service insertion
⚠ Known Pitfalls (from Community KB)
References
- vSphere HA Design Guide: techdocs.broadcom.com
- VMware Live Recovery Administration Guide: techdocs.broadcom.com
- NSX Security Design Guide: techdocs.broadcom.com
- vSAN Encryption and KMS Integration: techdocs.broadcom.com