End-to-End Architecture Design Scenario
Objectives
- Design a complete VCF 9.0 architecture from requirements to implementation plan
- Create conceptual, logical, and physical architecture diagrams
- Produce detailed sizing calculations for compute, storage, network, and IOPS
- Design comprehensive DFW rule set with healthcare compliance (HIPAA/HITRUST)
- Build encryption design with key management hierarchy and compliance mapping
- Create HA/DR design with stretched cluster, witness, and failover runbook
- Compile design decision register with 10+ decisions and complete risk register
Prerequisites
All prior VCF Architect labs completed (vcp-architect-01 through -12)
Prior labs: vcp-architect-01, vcp-architect-02, vcp-architect-03, vcp-architect-04, vcp-architect-05, vcp-architect-06, vcp-architect-10, vcp-architect-11, vcp-architect-12
Required skills:
- All VCF architecture design skills from prior labs
- Healthcare IT regulatory knowledge (HIPAA basics)
- End-to-end design methodology
- VCDX defense preparation
Lab Environment
Full design exercise — produces a complete VCDX-quality architecture for a healthcare scenario
Tasks
Task 1 Scenario Brief & Requirements Analysis
Analyze a complex real-world scenario, extract requirements, and establish the RCAR framework for a healthcare VCF deployment.
Read the scenario brief:
MedCity Regional Health System operates 3 hospitals and 12 outpatient clinics across a metropolitan area. They are migrating from a legacy vSphere 7.0 environment (no VCF) to VMware Cloud Foundation 9.0. The environment includes:
- 450 existing VMs (mix of Windows and Linux)
- Electronic Health Record (EHR) system: Epic, 40 production VMs, database-intensive
- Picture Archiving (PACS/DICOM): 30 VMs, storage-intensive (150 TB medical images)
- Clinical applications: 120 VMs (lab systems, pharmacy, radiology viewers)
- Business applications: 100 VMs (HR, finance, email, collaboration)
- Research computing: 60 VMs (genomics analysis, clinical trials data)
- Dev/Test/Training: 100 VMs (Epic sandbox, clinical training environments)
Business requirements:
- HIPAA compliance (Security Rule and Privacy Rule)
- HITRUST CSF certification target (within 18 months)
- EHR downtime tolerance: <15 minutes per year (99.997% availability)
- RPO=0 for EHR and PACS, RPO<1 hour for clinical apps
- Two data centers: 15km apart, dark fiber connected
- Planned growth: 15% annually for clinical, 30% for research computing
- Budget: $4.5M capex for infrastructure (3-year lifecycle), $800K/year opex
Constraints:
- Must maintain HIPAA-compliant audit trail for all PHI access
- Air-gapped environment: No internet connectivity for clinical systems
- Radiologists require low-latency PACS access (<5ms to storage)
- Research data must be isolated from clinical data (IRB requirement)
- 24/7 operations — no maintenance windows longer than 4 hours
Extract and document requirements:
Business Requirements:
- BR-001: HIPAA Security Rule compliance (encryption, access controls, audit)
- BR-002: HITRUST CSF certification readiness within 18 months
- BR-003: EHR availability ≥99.997% (≤15 min downtime/year)
- BR-004: Zero data loss for EHR and PACS on site failure
- BR-005: Research data isolation from clinical data
- BR-006: 3-year infrastructure lifecycle with 15-30% annual growth
- BR-007: Budget: $4.5M capex, $800K/year opex
Technical Requirements:
- TR-001: VCF 9.0 with vSAN ESA, NSX micro-segmentation
- TR-002: Stretched cluster for EHR/PACS (RPO=0, RTO<5 min)
- TR-003: Separate workload domain for research (IRB isolation)
- TR-004: PACS storage: 150 TB usable with <5ms latency
- TR-005: Encryption at rest and in transit for all PHI
- TR-006: Air-gapped depot for SDDC Manager (no internet)
- TR-007: DFW micro-segmentation isolating patient/admin/research zones
Document constraints, assumptions, and risks using the RCAR framework from Lab 01.
Classify workloads into tiers:
| Tier | Workload | VMs | RPO | RTO | Storage | Sensitivity | DR Pattern |
|---|---|---|---|---|---|---|---|
| 1-Critical | EHR (Epic) | 40 | 0 | <5 min | 20 TB | PHI | Stretched Cluster |
| 1-Critical | PACS/DICOM | 30 | 0 | <5 min | 150 TB | PHI | Stretched Cluster |
| 2-Clinical | Clinical Apps | 120 | <1 hour | <30 min | 15 TB | PHI | SRM |
| 3-Business | Business Apps | 100 | <4 hours | <2 hours | 12 TB | PII (non-PHI) | SRM |
| 4-Research | Research Computing | 60 | <24 hours | <4 hours | 30 TB | De-identified | Backup |
| 5-DevTest | Dev/Test/Training | 100 | None | N/A | 8 TB | Synthetic data | None |
| Total | 450 | 235 TB |
Note: PACS is the storage-dominant workload (150 TB = 64% of total storage). This drives the storage architecture.
Identify design challenges unique to healthcare:
- PACS storage challenge:
- 150 TB of medical images with <5ms access latency
- DICOM images are large (10-500 MB each) and rarely modified after creation
- Dedup ratio: Very low (~1.1:1) — medical images are unique
- Growth: 15% p.a. → 230 TB in 3 years
- Design implication: vSAN ESA with additional NVMe capacity, or NFS external storage for archive tier
- Air-gapped environment:
- No internet connectivity for clinical VMs
- SDDC Manager cannot download bundles from VMware depot
- Solution: Air-gapped depot (local HTTP server with manually transferred bundles)
- Skyline Health: Use Skyline Collector appliance (on-premises, forwards data via proxy)
- HIPAA audit requirements:
- All access to PHI must be logged and auditable
- Log retention: Minimum 6 years (HIPAA requirement)
- DFW logs, vCenter tasks, NSX audit logs → centralized SIEM
- Storage for 6 years of logs: Significant (plan log archival to external storage)
- Research data isolation:
- IRB (Institutional Review Board) requires complete isolation of research from clinical
- Separate workload domain (not just separate cluster)
- Network: No connectivity between research and clinical segments
- DFW: Explicit block rules between research and clinical groups
- Exception: De-identified data pipeline (one-way, audited)
Validation Gate
Check: Scenario analyzed with requirements extracted, workloads classified, and design challenges identified
Expected: 7+ business requirements, 7+ technical requirements, 6 workload tiers with RPO/RTO/storage, 4+ healthcare-specific design challenges documented
Common Errors
Task 2 Architecture Design — Compute, Storage, Network
Design the complete VCF 9.0 architecture including cluster topology, vSAN storage, and NSX networking for the healthcare scenario.
Design cluster topology and sizing:
Domain architecture:
- Management Domain (4 hosts):
- vCenter, SDDC Manager, NSX Manager (3-node), Aria Suite
- Sizing: 4 × 64 cores, 4 × 512 GB RAM
- Storage: 15 TB vSAN ESA, FTT=1 RAID-5
- Workload Domain 1 — Clinical (production):
- Cluster A: EHR + PACS (stretched cluster, 8+8 hosts across 2 sites)
- 70 VMs (40 EHR + 30 PACS)
- Storage challenge: 170 TB usable (20 TB EHR + 150 TB PACS)
- Host spec: 64 cores, 768 GB RAM, 8× 7.68 TB NVMe (for PACS storage density)
- Raw per host: 61.44 TB; 16 hosts total = 983 TB raw
- After FTT=1 site mirror + 30% slack + overhead: ~330 TB usable → sufficient
- Cluster B: Clinical Applications (4 hosts)
- 120 VMs, 15 TB storage
- Standard host spec: 64 cores, 512 GB RAM, 4× 3.84 TB NVMe
- Workload Domain 2 — Business (4 hosts):
- 100 business VMs, 12 TB storage
- Standard host spec
- Workload Domain 3 — Research (separate domain, 4 hosts):
- 60 research VMs, 30 TB storage
- Isolated: Separate vCenter, separate NSX transport zone
- Higher compute density: GPU-equipped hosts for genomics analysis
- Dev/Test cluster (within WD-2, 4 hosts):
- 100 VMs, 8 TB storage
- Lower-spec hosts acceptable
Total: 4 + 16 + 4 + 4 + 4 + 4 = 36 hosts (primary site + DR combined)
Design storage architecture:
Challenge: PACS drives 64% of total storage (150 TB) with unique characteristics:
- Large sequential writes (DICOM image storage)
- Random reads (radiologist viewing)
- Very low dedup ratio (~1.1:1 for medical images)
- Latency requirement: <5ms (radiologist workflow SLA)
Storage design:
- EHR + PACS cluster (stretched, 16 hosts):
- Hosts: 8× 7.68 TB NVMe per host = 61.44 TB raw per host
- Per-site raw: 8 × 61.44 = 491.52 TB
- FTT=1 site mirror: Data mirrored between sites (effective 2× overhead)
- After mirror: 491.52 / 2 = 245.76 TB per-site usable before slack
- After 30% slack: 245.76 × 0.70 = 172 TB usable → sufficient for 170 TB
- Headroom: ~1% — tight. Plan expansion at 70% utilization alert.
- Storage policy design:
| Policy | FTT | Method | Dedup/Comp | Encryption | IOPS Limit | Workload |
|---|---|---|---|---|---|---|
| EHR-Critical | 2 (within-site) | RAID-6 | Yes | Yes (AES-256) | None | Epic DB |
| PACS-Archive | 1 (site mirror) | RAID-5 | Compression only | Yes | 2,000/VMDK | DICOM images |
| Clinical-Standard | 1 | RAID-5 | Yes | Yes | 1,500/VMDK | Clinical apps |
| Business-Standard | 1 | RAID-5 | Yes | Yes | 1,000/VMDK | Business apps |
| Research-HighPerf | 1 | RAID-5 | No (unique data) | Yes | None | Genomics |
| DevTest-Economy | 1 | RAID-5 | Yes | Optional | 500/VMDK | Dev/test |
Note: PACS uses compression-only (not dedup) because medical images have near-zero dedup ratio.
Design NSX network with healthcare segmentation:
Segment design (healthcare-specific zones):
| Zone | Segment | CIDR | Tier-1 | DFW Zone | PHI Level |
|---|---|---|---|---|---|
| Patient Care | EHR-Prod | 172.16.1.0/24 | T1-Clinical | PHI-Clinical | High |
| Patient Care | PACS-Prod | 172.16.2.0/24 | T1-Clinical | PHI-Clinical | High |
| Patient Care | Clinical-Apps | 172.16.3.0/23 | T1-Clinical | PHI-Clinical | Medium |
| Administration | Business-HR | 172.16.10.0/24 | T1-Business | PII-Admin | Medium |
| Administration | Business-Finance | 172.16.11.0/24 | T1-Business | PII-Admin | Medium |
| Research | Research-Compute | 172.16.20.0/23 | T1-Research | Research-Isolated | De-identified |
| Research | Research-Data | 172.16.22.0/24 | T1-Research | Research-Isolated | De-identified |
| Dev/Test | EHR-Sandbox | 172.16.100.0/24 | T1-DevTest | NonProd | Synthetic |
| Dev/Test | Training | 172.16.101.0/24 | T1-DevTest | NonProd | Synthetic |
| Infrastructure | Mgmt | 172.16.200.0/24 | T1-Mgmt | Mgmt | N/A |
Critical DFW rules:
- BLOCK: Research ↔ Clinical (any direction, any port) — IRB requirement - BLOCK: Dev/Test → Clinical (any direction) — prevent test data contamination - ALLOW: Clinical-Apps → EHR (TCP/1433 for Epic Caché DB) - ALLOW: PACS → Clinical (DICOM port TCP/104, HTTPS/443)
- LOG ALL: Any traffic touching PHI-Clinical zone (HIPAA audit requirement)
Design encryption and HIPAA compliance controls:
- Encryption architecture:
- Data at rest: vSAN encryption (AES-256) — ALL workload domains
- Data in transit: vSAN in-transit encryption between hosts
- VM-level: vTPM on all clinical VMs (required for HIPAA)
- Key management: External KMS (Thales CipherTrust) for production, NKP for dev
- Key hierarchy: KEK (KMS) → DEK (per-host) → standard two-tier model
- HIPAA Security Rule mapping:
| HIPAA § | Requirement | VCF Control | Evidence |
|---|---|---|---|
| 164.312(a)(1) | Access control | NSX DFW + RBAC + MFA | DFW rule export, vCenter audit log |
| 164.312(a)(2)(iv) | Encryption | vSAN encryption, vTPM | KMS audit log, encryption status report |
| 164.312(b) | Audit controls | DFW logging, vCenter events | Aria Ops for Logs, 6-year retention |
| 164.312(c)(1) | Integrity | vSAN checksums, DFW | vSAN health, DFW integrity checks |
| 164.312(d) | Authentication | AD + MFA, SSO | AD logs, MFA enrollment report |
| 164.312(e)(1) | Transmission security | vSAN in-transit encryption, TLS | Network capture verification |
- Audit log architecture:
- Sources: vCenter, NSX DFW, SDDC Manager, ESXi, vSAN
- Collector: Aria Operations for Logs (primary), external SIEM (secondary)
- Retention: 90 days hot (Aria Ops for Logs), 6 years cold (archive to NFS)
- Archive storage: ~2 TB/year of compressed logs → 12 TB for 6-year retention
- Integrity: Log checksums to prevent tampering (HIPAA requirement)
Validation Gate
Check: Complete architecture designed for healthcare scenario with compute, storage, network, and compliance
Expected: 36-host topology across 4 domains, storage design handling 150 TB PACS, NSX segmentation with healthcare zones, HIPAA compliance mapping with evidence
Common Errors
Task 3 HA/DR Design & Failover Runbook
Design the complete HA and DR architecture for the healthcare scenario, including stretched cluster for EHR/PACS and SRM for clinical applications.
Design HA for EHR/PACS (stretched cluster):
- Stretched cluster configuration:
- Site A (Primary Hospital DC): 8 hosts
- Site B (DR DC): 8 hosts
- Witness: Medium OVA at administrative office (Site C)
- Inter-site: 100 Gbps dark fiber, 0.6ms RTT (15km)
- Preferred site: Site A
- EHR-specific HA:
- Epic Caché database VMs: Highest restart priority
- Epic Hyperspace (client) VMs: High priority
- Epic Interconnect (integration engine): High priority
- Anti-affinity: Caché primary and shadow never on same host
- VM restart order: Caché DB → Interconnect (120s delay) → Hyperspace (60s delay)
- PACS-specific HA:
- PACS archive server: Highest priority (must be available for image retrieval)
- PACS web viewer: Medium priority (radiologists need within 5 minutes)
- DICOM gateway: High priority (incoming studies from imaging equipment)
- Site affinity: PACS VMs prefer Site A (where imaging equipment connects)
- Availability calculation:
- Target: 99.997% = ≤15 min downtime/year
- HA detection: ~35 seconds (heartbeat miss + declaration)
- VM restart: ~2-3 minutes (boot + Epic service start)
- Total RTO: ~3.5 minutes per event
- Allowed events: 15 min / 3.5 min = ~4 events per year
- vSAN stretched cluster + HA + anti-affinity achieves this target
Design DR for clinical and business applications (SRM):
- SRM configuration for clinical apps (120 VMs):
- RPO: 1 hour (vSphere Replication)
- RTO: 30 minutes (SRM recovery plan)
- DR hosts: 4 hosts at Site B (separate from stretched cluster)
- Over-subscription: 120 VMs at DR on 4 hosts → acceptable with degraded performance
- Normal: 4 vCPU avg, 3:1 overcommit → 4 hosts handle 768 vCPU
- DR mode: 120 VMs × 4 vCPU = 480 vCPU → within capacity- Recovery plan for clinical apps:
- Priority Group 1: Laboratory Information System (10 VMs)
- Priority Group 2: Pharmacy system (8 VMs)
- Priority Group 3: Radiology viewers (15 VMs)
- Priority Group 4: Remaining clinical apps (87 VMs)
- Startup delays: 120s between groups
- Post-recovery: Verify HL7 interface connectivity to EHR
- Business application DR:
- RPO: 4 hours
- RTO: 2 hours
- DR hosts: Shared with clinical DR hosts (lower priority in recovery plan)
- Recovery plan: Executed AFTER clinical apps are confirmed running
- Research DR:
- RPO: 24 hours (daily backup)
- RTO: 4 hours
- No dedicated DR hosts — restore to available capacity at Site B
- Lowest priority in any multi-tier recovery scenario
Write the failover runbook:
Runbook: EHR/PACS Stretched Cluster — Site A Failure
Step 1: Detection (0-5 minutes)
- [ ] vSAN Health alert received: Site A hosts unreachable
- [ ] Verify: Ping Site A management IPs from Site B → No response - [ ] Verify: Check IPMI/iLO at Site A → No response (confirms site failure, not network partition)
- [ ] Escalation: Page on-call infrastructure lead AND clinical informatics officer
Step 2: Automatic Failover (5-10 minutes)
- [ ] Confirm: HA has restarted EHR VMs on Site B (check vCenter events)
- [ ] Confirm: Caché database shadow promoted to primary (Epic-specific verification)
- [ ] Confirm: PACS archive server accessible from radiology workstations
- [ ] Monitor: vSAN health → verify all objects accessible at Site B
Step 3: Clinical Verification (10-20 minutes)
- [ ] Clinical test: Open Epic Hyperspace client → verify patient record retrieval - [ ] Clinical test: Open PACS viewer → verify recent study retrieval - [ ] Clinical test: Send test HL7 message through Interconnect → verify integration
- [ ] Notify: Clinical informatics that EHR is operational on DR site
- [ ] Notify: Radiology department that PACS is operational
Step 4: Secondary Recovery (20-60 minutes)
- [ ] Assess: Are clinical app VMs (SRM) also affected?
- If yes: Execute SRM recovery plan for clinical apps
- If no: Clinical apps still running on Site A infrastructure (separate failure domain)
- [ ] Monitor: Aria Operations dashboards for performance degradation
- [ ] Document: Timeline of events for incident report
Step 5: Communication (ongoing)
- [ ] Update: Hospital incident command (if activated)
- [ ] Update: VMware Support SR (Severity 1)
- [ ] Update: Executive stakeholders every 30 minutes until stable
Step 6: Post-Incident (24-72 hours)
- [ ] Root cause analysis
- [ ] Site A recovery and vSAN resync
- [ ] DRS failback to restore site affinity
- [ ] Update risk register with lessons learned
Design DR testing for healthcare:
- Testing constraints:
- Clinical systems cannot have extended downtime for testing
- SRM test recovery uses isolated network (no production impact)
- Stretched cluster planned failover during lowest-census period (Sunday 2-6 AM)
- Joint Commission and HIPAA require documented DR testing
- Annual DR test calendar:
| Quarter | Test | Scope | Duration | Approval |
|---|---|---|---|---|
| Q1 | SRM test recovery | Clinical apps (120 VMs) | 4 hours | Clinical Ops + IT Director |
| Q2 | Stretched cluster planned failover | EHR + PACS (70 VMs) | 6 hours | CMO + CIO + IT Director |
| Q3 | Backup restore validation | 20 random VMs + 1 EHR test | 3 hours | IT Director |
| Q4 | Full multi-tier DR exercise | All tiers | 8 hours | CMO + CFO + CIO |
- Success criteria for each test:
- EHR functional: Patient search, order entry, result review within 5 minutes
- PACS functional: Study retrieval, image display within 10 seconds
- Clinical apps: HL7 message processing within 60 seconds
- RTO measured: Stopwatch from failover initiation to clinical verification
- Documentation: Test report filed within 5 business days
- HIPAA DR documentation:
- HIPAA § 164.308(a)(7)(ii)(D): Testing and revision procedures
- Requirement: Regular testing of contingency plan
- Evidence: DR test reports with timestamps, success/failure, corrective actions
- Retention: 6 years (align with other HIPAA records)
Validation Gate
Check: Complete HA/DR design with stretched cluster, SRM, failover runbook, and healthcare-specific testing
Expected: Stretched cluster for EHR/PACS with 99.997% availability math, SRM for clinical with recovery plan, step-by-step failover runbook, annual DR test calendar with HIPAA compliance
Common Errors
Task 4 Design Decision Register & Final Presentation
Compile the complete design decision register for the healthcare scenario and prepare the VCDX defense presentation.
Compile healthcare-specific design decisions:
| ID | Decision | Key Justification | HIPAA Mapping |
|---|---|---|---|
| H-001 | 4 workload domains (Clinical + Business + Research + Dev) | IRB isolation for research; HIPAA PHI scope reduction | 164.312(a)(1) |
| H-002 | Stretched cluster for EHR + PACS | RPO=0 for patient data; 99.997% availability target | 164.308(a)(7) |
| H-003 | External KMS (Thales) for all production domains | HIPAA encryption mandate; key separation from data | 164.312(a)(2)(iv) |
| H-004 | PACS on dedicated cluster with high-density NVMe | 150 TB with <5ms latency; low dedup ratio (1.1:1) | 164.312(c)(1) |
| H-005 | Air-gapped SDDC Manager depot | Clinical network no-internet policy | N/A (operational) |
| H-006 | 6-year log retention to NFS archive | HIPAA audit trail requirement | 164.312(b) |
| H-007 | DFW: Block research ↔ clinical (absolute deny) | IRB data isolation requirement | 164.312(a)(1) |
| H-008 | Separate transport zone per workload domain | Prevent network-level cross-domain leakage | 164.312(e)(1) |
| H-009 | vTPM on all clinical VMs | HIPAA integrity; platform attestation | 164.312(c)(1) |
| H-010 | Quarterly DR testing with clinical verification | HIPAA contingency plan testing requirement | 164.308(a)(7)(ii)(D) |
Create the risk register with healthcare-specific risks:
| ID | Risk | P | I | Score | Mitigation | HIPAA Impact |
|---|---|---|---|---|---|---|
| HR-001 | EHR downtime >15 min during site failure | L(2) | C(5) | 10 | Stretched cluster + HA + anti-affinity + clinical verification runbook | Patient safety |
| HR-002 | PHI data breach via network lateral movement | M(3) | C(5) | 15 | DFW zero-trust + logging + SIEM correlation + incident response plan | HIPAA violation, OCR notification |
| HR-003 | Ransomware encrypts clinical VMs | M(3) | C(5) | 15 | Immutable backups + air-gapped recovery environment + 4-hour RTO restore | Patient safety + HIPAA breach |
| HR-004 | KMS failure during host reboot cycle | L(2) | H(4) | 8 | KMS HA cluster + quarterly recovery test + documented break-glass | Data inaccessibility |
| HR-005 | PACS storage capacity exhaustion | M(3) | H(4) | 12 | 70% alert, compression analysis, expansion procurement at 12-week lead time | Inability to store new studies |
| HR-006 | Research data leaks to clinical network | L(2) | C(5) | 10 | Separate WD + separate TZ + DFW absolute deny + network monitoring | IRB violation, research shutdown |
| HR-007 | Audit log loss or tampering | L(2) | H(4) | 8 | Log integrity checksums + dual-destination logging + NFS write-once | HIPAA audit failure |
| HR-008 | VCF upgrade fails mid-sequence | M(3) | H(4) | 12 | Pre-upgrade health checks + snapshots + VMware SR pre-opened + rollback plan | Extended downtime risk |
Top risks (score ≥12): HR-002 (data breach), HR-003 (ransomware), HR-005 (PACS capacity), HR-008 (upgrade failure)
Prepare VCDX defense for healthcare challenges:
Challenge 1: 'Your PACS cluster is at 98% capacity after 3 years of growth. What do you do?'
Response: The design includes a 70% capacity alert triggering procurement with 12-week lead time. At 98%, emergency options: (1) Enable HCI Mesh from underutilized business cluster for temporary overflow, (2) Archive older studies (>5 years) to external NFS with DICOM viewer redirect, (3) Add hosts to PACS cluster via SDDC Manager. Long-term: Evaluate tiered storage with cold tier for archived images.
Challenge 2: 'How do you handle a ransomware attack that encrypts both data sites?'
Response: The design includes immutable backups with an air-gapped recovery environment. Recovery procedure: (1) Isolate affected segments via DFW emergency quarantine rule, (2) Stand up clean vCenter from air-gapped backup, (3) Restore VMs from immutable backup to isolated recovery network, (4) Forensic analysis of infection vector, (5) Validate clean state before reconnecting to production. RTO: 4-8 hours for EHR, 24 hours for full environment. Immutable backups use WORM storage that cannot be encrypted by ransomware.
Challenge 3: 'A researcher needs access to clinical data for a study. How does your design handle this?'
Response: The design enforces complete isolation between research and clinical domains (separate workload domain, separate transport zone, DFW absolute deny). For approved studies: (1) Clinical data team de-identifies patient data per HIPAA Safe Harbor or Expert Determination method, (2) De-identified dataset is exported to a staging area, (3) Staging area is scanned for residual PHI, (4) Approved dataset is transferred to research domain via audited one-way data pipeline, (5) Full audit trail logged for IRB compliance.
Challenge 4: 'Why not use a single workload domain with DFW rules instead of 4 separate domains?'
Response: HIPAA requires demonstrable isolation — DFW rules can be misconfigured, but separate workload domains provide domain-level isolation that survives DFW errors. Separate domains also enable independent upgrade schedules (clinical cannot wait for research domain upgrades), separate vSAN datastores (research workloads cannot impact clinical storage performance), and cleaner audit scope (HIPAA audit covers clinical domain only, not research).
Create the final design summary and quality checklist:
- Architecture summary:
- 4 workload domains + 1 management domain = 5 domains
- 36 hosts total (including DR)
- 2 sites + 1 witness
- 235 TB usable storage (150 TB PACS dominant)
- VCF 9.0 with vSAN ESA, NSX, Aria Suite
- Compliance summary:
- HIPAA Security Rule: 6 § mapped to VCF controls
- HITRUST CSF: Architecture supports certification timeline
- Audit: 6-year log retention, quarterly DR testing, annual access review
- Cost summary:
- 36 hosts × 64 cores × 2 sockets = 4,608 licensed cores
- VCF per-core licensing + external KMS + Aria Suite
- Target: $4.5M capex (verify against host + storage + networking BOM)
- Opex: $800K/year (licensing subscription + support + operations staff)
- Final quality checklist:
- [ ] Every HIPAA § has a corresponding VCF control
- [ ] Research isolation verified at domain + network + DFW + storage levels
- [ ] PACS storage sized for 3-year growth with alerts
- [ ] Failover runbook includes clinical-specific verification steps
- [ ] Air-gapped depot documented for SDDC Manager
- [ ] 6-year log archival sized and designed
- [ ] All 10 design decisions trace to requirements
- [ ] All 8 risks have specific mitigations
- [ ] VCDX defense responses prepared for 4+ healthcare challenges
- [ ] Design internally consistent — no contradictions between sections
Validation Gate
Check: Complete design decision register, risk register, defense preparation, and quality checklist for healthcare scenario
Expected: 10 healthcare-specific design decisions with HIPAA mapping, 8 risks with scores and mitigations, 4+ defense responses, comprehensive quality checklist
Common Errors
Final Validation
Complete end-to-end VCF 9.0 architecture design for healthcare, production-ready for VCDX submission
✓ Scenario requirements fully analyzed → 7+ business requirements, 7+ technical requirements, 6 workload tiers classified
✓ Architecture covers all domains → 36-host topology, 5 domains, 150 TB PACS addressed, NSX healthcare segmentation
✓ HA/DR design meets healthcare SLAs → 99.997% availability for EHR, failover runbook with clinical verification, annual DR testing
✓ HIPAA compliance demonstrated → 6 HIPAA sections mapped, 6-year log retention, encryption, audit controls
✓ Design defensible in VCDX context → 10 design decisions, 8 risks, 4+ defense responses, quality checklist complete
Cleanup / Restore
• Save the complete healthcare design document as the capstone artifact
• Archive all supporting worksheets, diagrams, and calculations
• This design serves as a portfolio piece for VCDX preparation
Design Reflection (VCDX)
This capstone lab is the closest simulation of a real VCDX design exercise. Healthcare adds layers of complexity that generic designs do not face: regulatory compliance with specific statute references, clinical workflow integration (not just IT SLAs), data classification with IRB isolation, and life-safety implications of downtime. VCDX panelists who review healthcare designs look for: (1) Do you understand that EHR downtime = patient safety risk? The design must reflect this urgency. (2) Can you explain the data isolation model? Research/clinical/admin boundaries must be absolute. (3) Is your DR testing realistic? Healthcare cannot tolerate extended test windows. (4) Do you know HIPAA well enough to map controls to specific sections? Generic 'we comply with HIPAA' is insufficient.
Requirements
- HIPAA Security Rule compliance with 6 specific section mappings
- 99.997% availability for EHR (≤15 min downtime/year)
- 150 TB PACS storage with <5ms read latency for radiologists
- Complete isolation between research and clinical data (IRB mandate)
- 6-year audit log retention with integrity verification
Constraints
- Air-gapped clinical network — no internet access
- 24/7 operations — maintenance windows ≤4 hours
- $4.5M capex budget for 3-year lifecycle
- PACS images have ~1.1:1 dedup ratio (cannot rely on dedup for capacity savings)
- 15km between sites — stretched cluster viable (0.6ms RTT)
Assumptions
- Epic and PACS vendors support VMware VCF 9.0 (vendor certification obtained)
- Dark fiber available between hospital data centers
- Active Directory with LDAPS available for authentication
- De-identification pipeline exists for research data access
- Hospital leadership has approved RPO/RTO targets per tier
Risks
- PHI data breach via lateral movement (DFW misconfiguration or compromise)
- Ransomware attack encrypting both data sites simultaneously
- PACS storage capacity exhaustion before procurement completes
- EHR downtime exceeding 15-minute target during complex failure scenario
- HIPAA audit finding due to incomplete log retention or access controls
Self-Assessment Discussion Prompts
- If a new HIPAA regulation requires end-to-end encryption for all data in motion, what changes in your design?
- How would you handle the scenario where the Epic vendor does not support vSAN ESA — only external storage?
- If the hospital acquires a second health system with 300 additional VMs, how does the architecture scale?
- What is the business case for spending $4.5M on VCF when a public cloud (AWS/Azure) option exists?
Extensions
Telehealth Integration Architecture
Extend the design to support a telehealth platform with video conferencing VMs, patient portal, and external-facing web services. Design the DMZ architecture with NSX, the DFW rules for external access, and the HIPAA-compliant video/audio encryption requirements.
Medical IoT Device Integration
Design the network architecture for integrating medical IoT devices (patient monitors, infusion pumps, imaging equipment) with the VCF environment. Address device segmentation, traffic inspection, and the unique security challenges of medical devices that cannot be patched.
Multi-Hospital VCF Federation
Extend the single-hospital design to a 3-hospital federation with shared EHR but independent PACS and clinical systems. Design the cross-site Enhanced Linked Mode, shared Content Library, federated monitoring, and the data sovereignty implications of multi-site PHI storage.
⚠ Known Pitfalls (from Community KB)
References
- HIPAA Security Rule (45 CFR Part 164, Subpart C)
- HITRUST CSF v11 — Control Category Reference
- VMware VCF 9.0 Healthcare Reference Architecture
- Epic Systems — VMware Support Matrix and Best Practices
- NIST SP 800-66 — Implementing the HIPAA Security Rule