Academy/VCP-VCF 9.0 Architect (2V0-13.25)/End-to-End Architecture Design Scenario
This lab targets VCF 9.0

End-to-End Architecture Design Scenario

VCF 9.0Advancedvcp-foundation⏱ 180 min

Capstone design exercise — complete VCF 9.0 architecture for a healthcare organization with HIPAA requirements

Objectives

  • Design a complete VCF 9.0 architecture from requirements to implementation plan
  • Create conceptual, logical, and physical architecture diagrams
  • Produce detailed sizing calculations for compute, storage, network, and IOPS
  • Design comprehensive DFW rule set with healthcare compliance (HIPAA/HITRUST)
  • Build encryption design with key management hierarchy and compliance mapping
  • Create HA/DR design with stretched cluster, witness, and failover runbook
  • Compile design decision register with 10+ decisions and complete risk register

Prerequisites

All prior VCF Architect labs completed (vcp-architect-01 through -12)

Prior labs: vcp-architect-01, vcp-architect-02, vcp-architect-03, vcp-architect-04, vcp-architect-05, vcp-architect-06, vcp-architect-10, vcp-architect-11, vcp-architect-12

Required skills:

  • All VCF architecture design skills from prior labs
  • Healthcare IT regulatory knowledge (HIPAA basics)
  • End-to-end design methodology
  • VCDX defense preparation

Lab Environment

Full design exercise — produces a complete VCDX-quality architecture for a healthcare scenario

Tasks

Task 1 Scenario Brief & Requirements Analysis

This is a VCDX-level capstone exercise. The scenario intentionally includes ambiguity — you must make assumptions, document them, and be prepared to defend them. Real-world designs always have incomplete information.

Analyze a complex real-world scenario, extract requirements, and establish the RCAR framework for a healthcare VCF deployment.

Step 1

Read the scenario brief:

MedCity Regional Health System operates 3 hospitals and 12 outpatient clinics across a metropolitan area. They are migrating from a legacy vSphere 7.0 environment (no VCF) to VMware Cloud Foundation 9.0. The environment includes:

  • 450 existing VMs (mix of Windows and Linux)
  • Electronic Health Record (EHR) system: Epic, 40 production VMs, database-intensive
  • Picture Archiving (PACS/DICOM): 30 VMs, storage-intensive (150 TB medical images)
  • Clinical applications: 120 VMs (lab systems, pharmacy, radiology viewers)
  • Business applications: 100 VMs (HR, finance, email, collaboration)
  • Research computing: 60 VMs (genomics analysis, clinical trials data)
  • Dev/Test/Training: 100 VMs (Epic sandbox, clinical training environments)

Business requirements:

  • HIPAA compliance (Security Rule and Privacy Rule)
  • HITRUST CSF certification target (within 18 months)
  • EHR downtime tolerance: <15 minutes per year (99.997% availability)
  • RPO=0 for EHR and PACS, RPO<1 hour for clinical apps
  • Two data centers: 15km apart, dark fiber connected
  • Planned growth: 15% annually for clinical, 30% for research computing
  • Budget: $4.5M capex for infrastructure (3-year lifecycle), $800K/year opex

Constraints:

  • Must maintain HIPAA-compliant audit trail for all PHI access
  • Air-gapped environment: No internet connectivity for clinical systems
  • Radiologists require low-latency PACS access (<5ms to storage)
  • Research data must be isolated from clinical data (IRB requirement)
  • 24/7 operations — no maintenance windows longer than 4 hours
Step 2

Extract and document requirements:

Business Requirements:

  • BR-001: HIPAA Security Rule compliance (encryption, access controls, audit)
  • BR-002: HITRUST CSF certification readiness within 18 months
  • BR-003: EHR availability ≥99.997% (≤15 min downtime/year)
  • BR-004: Zero data loss for EHR and PACS on site failure
  • BR-005: Research data isolation from clinical data
  • BR-006: 3-year infrastructure lifecycle with 15-30% annual growth
  • BR-007: Budget: $4.5M capex, $800K/year opex

Technical Requirements:

  • TR-001: VCF 9.0 with vSAN ESA, NSX micro-segmentation
  • TR-002: Stretched cluster for EHR/PACS (RPO=0, RTO<5 min)
  • TR-003: Separate workload domain for research (IRB isolation)
  • TR-004: PACS storage: 150 TB usable with <5ms latency
  • TR-005: Encryption at rest and in transit for all PHI
  • TR-006: Air-gapped depot for SDDC Manager (no internet)
  • TR-007: DFW micro-segmentation isolating patient/admin/research zones

Document constraints, assumptions, and risks using the RCAR framework from Lab 01.

Step 3

Classify workloads into tiers:

TierWorkloadVMsRPORTOStorageSensitivityDR Pattern
1-CriticalEHR (Epic)400<5 min20 TBPHIStretched Cluster
1-CriticalPACS/DICOM300<5 min150 TBPHIStretched Cluster
2-ClinicalClinical Apps120<1 hour<30 min15 TBPHISRM
3-BusinessBusiness Apps100<4 hours<2 hours12 TBPII (non-PHI)SRM
4-ResearchResearch Computing60<24 hours<4 hours30 TBDe-identifiedBackup
5-DevTestDev/Test/Training100NoneN/A8 TBSynthetic dataNone
Total450235 TB

Note: PACS is the storage-dominant workload (150 TB = 64% of total storage). This drives the storage architecture.

Step 4

Identify design challenges unique to healthcare:

  1. PACS storage challenge:
  • 150 TB of medical images with <5ms access latency
  • DICOM images are large (10-500 MB each) and rarely modified after creation
  • Dedup ratio: Very low (~1.1:1) — medical images are unique
   - Growth: 15% p.a. → 230 TB in 3 years
  • Design implication: vSAN ESA with additional NVMe capacity, or NFS external storage for archive tier
  1. Air-gapped environment:
  • No internet connectivity for clinical VMs
  • SDDC Manager cannot download bundles from VMware depot
  • Solution: Air-gapped depot (local HTTP server with manually transferred bundles)
  • Skyline Health: Use Skyline Collector appliance (on-premises, forwards data via proxy)
  1. HIPAA audit requirements:
  • All access to PHI must be logged and auditable
  • Log retention: Minimum 6 years (HIPAA requirement)
   - DFW logs, vCenter tasks, NSX audit logs → centralized SIEM
  • Storage for 6 years of logs: Significant (plan log archival to external storage)
  1. Research data isolation:
  • IRB (Institutional Review Board) requires complete isolation of research from clinical
  • Separate workload domain (not just separate cluster)
  • Network: No connectivity between research and clinical segments
  • DFW: Explicit block rules between research and clinical groups
  • Exception: De-identified data pipeline (one-way, audited)

Validation Gate

Check: Scenario analyzed with requirements extracted, workloads classified, and design challenges identified

Expected: 7+ business requirements, 7+ technical requirements, 6 workload tiers with RPO/RTO/storage, 4+ healthcare-specific design challenges documented

Common Errors

Treating all healthcare data the same — EHR, PACS, research, and business have very different requirements
Forgetting air-gapped constraints — many healthcare environments restrict internet access for clinical systems
Underestimating PACS storage — medical imaging dominates storage requirements and has very low dedup ratios
Not separating research from clinical — IRB requirements mandate isolation, not just network segmentation

Task 2 Architecture Design — Compute, Storage, Network

This task synthesizes all prior labs into a single coherent design. Every number must be justified, every decision must trace to a requirement.

Design the complete VCF 9.0 architecture including cluster topology, vSAN storage, and NSX networking for the healthcare scenario.

Step 1

Design cluster topology and sizing:

Domain architecture:

  1. Management Domain (4 hosts):
  • vCenter, SDDC Manager, NSX Manager (3-node), Aria Suite
  • Sizing: 4 × 64 cores, 4 × 512 GB RAM
  • Storage: 15 TB vSAN ESA, FTT=1 RAID-5
  1. Workload Domain 1 — Clinical (production):
  • Cluster A: EHR + PACS (stretched cluster, 8+8 hosts across 2 sites)
    • 70 VMs (40 EHR + 30 PACS)
    • Storage challenge: 170 TB usable (20 TB EHR + 150 TB PACS)
    • Host spec: 64 cores, 768 GB RAM, 8× 7.68 TB NVMe (for PACS storage density)
    • Raw per host: 61.44 TB; 16 hosts total = 983 TB raw
     - After FTT=1 site mirror + 30% slack + overhead: ~330 TB usable → sufficient
  • Cluster B: Clinical Applications (4 hosts)
    • 120 VMs, 15 TB storage
    • Standard host spec: 64 cores, 512 GB RAM, 4× 3.84 TB NVMe
  1. Workload Domain 2 — Business (4 hosts):
  • 100 business VMs, 12 TB storage
  • Standard host spec
  1. Workload Domain 3 — Research (separate domain, 4 hosts):
  • 60 research VMs, 30 TB storage
  • Isolated: Separate vCenter, separate NSX transport zone
  • Higher compute density: GPU-equipped hosts for genomics analysis
  1. Dev/Test cluster (within WD-2, 4 hosts):
  • 100 VMs, 8 TB storage
  • Lower-spec hosts acceptable

Total: 4 + 16 + 4 + 4 + 4 + 4 = 36 hosts (primary site + DR combined)

Step 2

Design storage architecture:

Challenge: PACS drives 64% of total storage (150 TB) with unique characteristics:

  • Large sequential writes (DICOM image storage)
  • Random reads (radiologist viewing)
  • Very low dedup ratio (~1.1:1 for medical images)
  • Latency requirement: <5ms (radiologist workflow SLA)

Storage design:

  1. EHR + PACS cluster (stretched, 16 hosts):
  • Hosts: 8× 7.68 TB NVMe per host = 61.44 TB raw per host
  • Per-site raw: 8 × 61.44 = 491.52 TB
  • FTT=1 site mirror: Data mirrored between sites (effective 2× overhead)
  • After mirror: 491.52 / 2 = 245.76 TB per-site usable before slack
   - After 30% slack: 245.76 × 0.70 = 172 TB usable → sufficient for 170 TB
  • Headroom: ~1% — tight. Plan expansion at 70% utilization alert.
  1. Storage policy design:
PolicyFTTMethodDedup/CompEncryptionIOPS LimitWorkload
EHR-Critical2 (within-site)RAID-6YesYes (AES-256)NoneEpic DB
PACS-Archive1 (site mirror)RAID-5Compression onlyYes2,000/VMDKDICOM images
Clinical-Standard1RAID-5YesYes1,500/VMDKClinical apps
Business-Standard1RAID-5YesYes1,000/VMDKBusiness apps
Research-HighPerf1RAID-5No (unique data)YesNoneGenomics
DevTest-Economy1RAID-5YesOptional500/VMDKDev/test

Note: PACS uses compression-only (not dedup) because medical images have near-zero dedup ratio.

Step 3

Design NSX network with healthcare segmentation:

Segment design (healthcare-specific zones):

ZoneSegmentCIDRTier-1DFW ZonePHI Level
Patient CareEHR-Prod172.16.1.0/24T1-ClinicalPHI-ClinicalHigh
Patient CarePACS-Prod172.16.2.0/24T1-ClinicalPHI-ClinicalHigh
Patient CareClinical-Apps172.16.3.0/23T1-ClinicalPHI-ClinicalMedium
AdministrationBusiness-HR172.16.10.0/24T1-BusinessPII-AdminMedium
AdministrationBusiness-Finance172.16.11.0/24T1-BusinessPII-AdminMedium
ResearchResearch-Compute172.16.20.0/23T1-ResearchResearch-IsolatedDe-identified
ResearchResearch-Data172.16.22.0/24T1-ResearchResearch-IsolatedDe-identified
Dev/TestEHR-Sandbox172.16.100.0/24T1-DevTestNonProdSynthetic
Dev/TestTraining172.16.101.0/24T1-DevTestNonProdSynthetic
InfrastructureMgmt172.16.200.0/24T1-MgmtMgmtN/A

Critical DFW rules:

- BLOCK: Research ↔ Clinical (any direction, any port) — IRB requirement
- BLOCK: Dev/Test → Clinical (any direction) — prevent test data contamination  
- ALLOW: Clinical-Apps → EHR (TCP/1433 for Epic Caché DB)
- ALLOW: PACS → Clinical (DICOM port TCP/104, HTTPS/443)
  • LOG ALL: Any traffic touching PHI-Clinical zone (HIPAA audit requirement)
Step 4

Design encryption and HIPAA compliance controls:

  1. Encryption architecture:
  • Data at rest: vSAN encryption (AES-256) — ALL workload domains
  • Data in transit: vSAN in-transit encryption between hosts
  • VM-level: vTPM on all clinical VMs (required for HIPAA)
  • Key management: External KMS (Thales CipherTrust) for production, NKP for dev
   - Key hierarchy: KEK (KMS) → DEK (per-host) → standard two-tier model
  1. HIPAA Security Rule mapping:
HIPAA §RequirementVCF ControlEvidence
164.312(a)(1)Access controlNSX DFW + RBAC + MFADFW rule export, vCenter audit log
164.312(a)(2)(iv)EncryptionvSAN encryption, vTPMKMS audit log, encryption status report
164.312(b)Audit controlsDFW logging, vCenter eventsAria Ops for Logs, 6-year retention
164.312(c)(1)IntegrityvSAN checksums, DFWvSAN health, DFW integrity checks
164.312(d)AuthenticationAD + MFA, SSOAD logs, MFA enrollment report
164.312(e)(1)Transmission securityvSAN in-transit encryption, TLSNetwork capture verification
  1. Audit log architecture:
  • Sources: vCenter, NSX DFW, SDDC Manager, ESXi, vSAN
  • Collector: Aria Operations for Logs (primary), external SIEM (secondary)
  • Retention: 90 days hot (Aria Ops for Logs), 6 years cold (archive to NFS)
   - Archive storage: ~2 TB/year of compressed logs → 12 TB for 6-year retention
  • Integrity: Log checksums to prevent tampering (HIPAA requirement)

Validation Gate

Check: Complete architecture designed for healthcare scenario with compute, storage, network, and compliance

Expected: 36-host topology across 4 domains, storage design handling 150 TB PACS, NSX segmentation with healthcare zones, HIPAA compliance mapping with evidence

Common Errors

Using standard vSAN dedup ratios for PACS — medical images have ~1.1:1 dedup ratio; compression-only is the correct approach
Not isolating research in a separate workload domain — IRB requirements mandate domain-level isolation, not just DFW rules
Forgetting 6-year log retention — HIPAA requires long-term audit log preservation; plan the archive storage
Undersizing the PACS cluster — 150 TB with stretched cluster effectively requires ~350 TB raw per site with slack space

Task 3 HA/DR Design & Failover Runbook

Healthcare DR design has no room for error. EHR downtime directly affects patient care. The failover runbook must be specific enough that an on-call engineer can execute it at 3 AM without architectural guidance.

Design the complete HA and DR architecture for the healthcare scenario, including stretched cluster for EHR/PACS and SRM for clinical applications.

Step 1

Design HA for EHR/PACS (stretched cluster):

  1. Stretched cluster configuration:
  • Site A (Primary Hospital DC): 8 hosts
  • Site B (DR DC): 8 hosts
  • Witness: Medium OVA at administrative office (Site C)
  • Inter-site: 100 Gbps dark fiber, 0.6ms RTT (15km)
  • Preferred site: Site A
  1. EHR-specific HA:
  • Epic Caché database VMs: Highest restart priority
  • Epic Hyperspace (client) VMs: High priority
  • Epic Interconnect (integration engine): High priority
  • Anti-affinity: Caché primary and shadow never on same host
   - VM restart order: Caché DB → Interconnect (120s delay) → Hyperspace (60s delay)
  1. PACS-specific HA:
  • PACS archive server: Highest priority (must be available for image retrieval)
  • PACS web viewer: Medium priority (radiologists need within 5 minutes)
  • DICOM gateway: High priority (incoming studies from imaging equipment)
  • Site affinity: PACS VMs prefer Site A (where imaging equipment connects)
  1. Availability calculation:
  • Target: 99.997% = ≤15 min downtime/year
  • HA detection: ~35 seconds (heartbeat miss + declaration)
  • VM restart: ~2-3 minutes (boot + Epic service start)
  • Total RTO: ~3.5 minutes per event
  • Allowed events: 15 min / 3.5 min = ~4 events per year
  • vSAN stretched cluster + HA + anti-affinity achieves this target
Step 2

Design DR for clinical and business applications (SRM):

  1. SRM configuration for clinical apps (120 VMs):
  • RPO: 1 hour (vSphere Replication)
  • RTO: 30 minutes (SRM recovery plan)
  • DR hosts: 4 hosts at Site B (separate from stretched cluster)
   - Over-subscription: 120 VMs at DR on 4 hosts → acceptable with degraded performance
     - Normal: 4 vCPU avg, 3:1 overcommit → 4 hosts handle 768 vCPU
     - DR mode: 120 VMs × 4 vCPU = 480 vCPU → within capacity
  1. Recovery plan for clinical apps:
  • Priority Group 1: Laboratory Information System (10 VMs)
  • Priority Group 2: Pharmacy system (8 VMs)
  • Priority Group 3: Radiology viewers (15 VMs)
  • Priority Group 4: Remaining clinical apps (87 VMs)
  • Startup delays: 120s between groups
  • Post-recovery: Verify HL7 interface connectivity to EHR
  1. Business application DR:
  • RPO: 4 hours
  • RTO: 2 hours
  • DR hosts: Shared with clinical DR hosts (lower priority in recovery plan)
  • Recovery plan: Executed AFTER clinical apps are confirmed running
  1. Research DR:
  • RPO: 24 hours (daily backup)
  • RTO: 4 hours
  • No dedicated DR hosts — restore to available capacity at Site B
  • Lowest priority in any multi-tier recovery scenario
Step 3

Write the failover runbook:

Runbook: EHR/PACS Stretched Cluster — Site A Failure

Step 1: Detection (0-5 minutes)

  • [ ] vSAN Health alert received: Site A hosts unreachable
- [ ] Verify: Ping Site A management IPs from Site B → No response
- [ ] Verify: Check IPMI/iLO at Site A → No response (confirms site failure, not network partition)
  • [ ] Escalation: Page on-call infrastructure lead AND clinical informatics officer

Step 2: Automatic Failover (5-10 minutes)

  • [ ] Confirm: HA has restarted EHR VMs on Site B (check vCenter events)
  • [ ] Confirm: Caché database shadow promoted to primary (Epic-specific verification)
  • [ ] Confirm: PACS archive server accessible from radiology workstations
- [ ] Monitor: vSAN health → verify all objects accessible at Site B

Step 3: Clinical Verification (10-20 minutes)

- [ ] Clinical test: Open Epic Hyperspace client → verify patient record retrieval
- [ ] Clinical test: Open PACS viewer → verify recent study retrieval
- [ ] Clinical test: Send test HL7 message through Interconnect → verify integration
  • [ ] Notify: Clinical informatics that EHR is operational on DR site
  • [ ] Notify: Radiology department that PACS is operational

Step 4: Secondary Recovery (20-60 minutes)

  • [ ] Assess: Are clinical app VMs (SRM) also affected?
    • If yes: Execute SRM recovery plan for clinical apps
    • If no: Clinical apps still running on Site A infrastructure (separate failure domain)
  • [ ] Monitor: Aria Operations dashboards for performance degradation
  • [ ] Document: Timeline of events for incident report

Step 5: Communication (ongoing)

  • [ ] Update: Hospital incident command (if activated)
  • [ ] Update: VMware Support SR (Severity 1)
  • [ ] Update: Executive stakeholders every 30 minutes until stable

Step 6: Post-Incident (24-72 hours)

  • [ ] Root cause analysis
  • [ ] Site A recovery and vSAN resync
  • [ ] DRS failback to restore site affinity
  • [ ] Update risk register with lessons learned
Step 4

Design DR testing for healthcare:

  1. Testing constraints:
  • Clinical systems cannot have extended downtime for testing
  • SRM test recovery uses isolated network (no production impact)
  • Stretched cluster planned failover during lowest-census period (Sunday 2-6 AM)
  • Joint Commission and HIPAA require documented DR testing
  1. Annual DR test calendar:
QuarterTestScopeDurationApproval
Q1SRM test recoveryClinical apps (120 VMs)4 hoursClinical Ops + IT Director
Q2Stretched cluster planned failoverEHR + PACS (70 VMs)6 hoursCMO + CIO + IT Director
Q3Backup restore validation20 random VMs + 1 EHR test3 hoursIT Director
Q4Full multi-tier DR exerciseAll tiers8 hoursCMO + CFO + CIO
  1. Success criteria for each test:
  • EHR functional: Patient search, order entry, result review within 5 minutes
  • PACS functional: Study retrieval, image display within 10 seconds
  • Clinical apps: HL7 message processing within 60 seconds
  • RTO measured: Stopwatch from failover initiation to clinical verification
  • Documentation: Test report filed within 5 business days
  1. HIPAA DR documentation:
  • HIPAA § 164.308(a)(7)(ii)(D): Testing and revision procedures
  • Requirement: Regular testing of contingency plan
  • Evidence: DR test reports with timestamps, success/failure, corrective actions
  • Retention: 6 years (align with other HIPAA records)

Validation Gate

Check: Complete HA/DR design with stretched cluster, SRM, failover runbook, and healthcare-specific testing

Expected: Stretched cluster for EHR/PACS with 99.997% availability math, SRM for clinical with recovery plan, step-by-step failover runbook, annual DR test calendar with HIPAA compliance

Common Errors

Not including clinical verification in the failover runbook — VMs running ≠ EHR functional; Epic/PACS-specific checks are required
Scheduling DR tests without clinical leadership approval — healthcare DR affects patient care; CMO/clinical ops must approve
Forgetting HIPAA documentation requirements — every DR test must be documented and retained for 6 years
Not calculating availability math — '99.997%' must be supported by specific RTO × max events per year calculation

Task 4 Design Decision Register & Final Presentation

This is the capstone. The design decision register must be comprehensive, internally consistent, and defensible. Every decision must trace to a healthcare-specific requirement.

Compile the complete design decision register for the healthcare scenario and prepare the VCDX defense presentation.

Step 1

Compile healthcare-specific design decisions:

IDDecisionKey JustificationHIPAA Mapping
H-0014 workload domains (Clinical + Business + Research + Dev)IRB isolation for research; HIPAA PHI scope reduction164.312(a)(1)
H-002Stretched cluster for EHR + PACSRPO=0 for patient data; 99.997% availability target164.308(a)(7)
H-003External KMS (Thales) for all production domainsHIPAA encryption mandate; key separation from data164.312(a)(2)(iv)
H-004PACS on dedicated cluster with high-density NVMe150 TB with <5ms latency; low dedup ratio (1.1:1)164.312(c)(1)
H-005Air-gapped SDDC Manager depotClinical network no-internet policyN/A (operational)
H-0066-year log retention to NFS archiveHIPAA audit trail requirement164.312(b)
H-007DFW: Block research ↔ clinical (absolute deny)IRB data isolation requirement164.312(a)(1)
H-008Separate transport zone per workload domainPrevent network-level cross-domain leakage164.312(e)(1)
H-009vTPM on all clinical VMsHIPAA integrity; platform attestation164.312(c)(1)
H-010Quarterly DR testing with clinical verificationHIPAA contingency plan testing requirement164.308(a)(7)(ii)(D)
Step 2

Create the risk register with healthcare-specific risks:

IDRiskPIScoreMitigationHIPAA Impact
HR-001EHR downtime >15 min during site failureL(2)C(5)10Stretched cluster + HA + anti-affinity + clinical verification runbookPatient safety
HR-002PHI data breach via network lateral movementM(3)C(5)15DFW zero-trust + logging + SIEM correlation + incident response planHIPAA violation, OCR notification
HR-003Ransomware encrypts clinical VMsM(3)C(5)15Immutable backups + air-gapped recovery environment + 4-hour RTO restorePatient safety + HIPAA breach
HR-004KMS failure during host reboot cycleL(2)H(4)8KMS HA cluster + quarterly recovery test + documented break-glassData inaccessibility
HR-005PACS storage capacity exhaustionM(3)H(4)1270% alert, compression analysis, expansion procurement at 12-week lead timeInability to store new studies
HR-006Research data leaks to clinical networkL(2)C(5)10Separate WD + separate TZ + DFW absolute deny + network monitoringIRB violation, research shutdown
HR-007Audit log loss or tamperingL(2)H(4)8Log integrity checksums + dual-destination logging + NFS write-onceHIPAA audit failure
HR-008VCF upgrade fails mid-sequenceM(3)H(4)12Pre-upgrade health checks + snapshots + VMware SR pre-opened + rollback planExtended downtime risk

Top risks (score ≥12): HR-002 (data breach), HR-003 (ransomware), HR-005 (PACS capacity), HR-008 (upgrade failure)

Step 3

Prepare VCDX defense for healthcare challenges:

Challenge 1: 'Your PACS cluster is at 98% capacity after 3 years of growth. What do you do?'
Response: The design includes a 70% capacity alert triggering procurement with 12-week lead time. At 98%, emergency options: (1) Enable HCI Mesh from underutilized business cluster for temporary overflow, (2) Archive older studies (>5 years) to external NFS with DICOM viewer redirect, (3) Add hosts to PACS cluster via SDDC Manager. Long-term: Evaluate tiered storage with cold tier for archived images.

Challenge 2: 'How do you handle a ransomware attack that encrypts both data sites?'
Response: The design includes immutable backups with an air-gapped recovery environment. Recovery procedure: (1) Isolate affected segments via DFW emergency quarantine rule, (2) Stand up clean vCenter from air-gapped backup, (3) Restore VMs from immutable backup to isolated recovery network, (4) Forensic analysis of infection vector, (5) Validate clean state before reconnecting to production. RTO: 4-8 hours for EHR, 24 hours for full environment. Immutable backups use WORM storage that cannot be encrypted by ransomware.

Challenge 3: 'A researcher needs access to clinical data for a study. How does your design handle this?'
Response: The design enforces complete isolation between research and clinical domains (separate workload domain, separate transport zone, DFW absolute deny). For approved studies: (1) Clinical data team de-identifies patient data per HIPAA Safe Harbor or Expert Determination method, (2) De-identified dataset is exported to a staging area, (3) Staging area is scanned for residual PHI, (4) Approved dataset is transferred to research domain via audited one-way data pipeline, (5) Full audit trail logged for IRB compliance.

Challenge 4: 'Why not use a single workload domain with DFW rules instead of 4 separate domains?'
Response: HIPAA requires demonstrable isolation — DFW rules can be misconfigured, but separate workload domains provide domain-level isolation that survives DFW errors. Separate domains also enable independent upgrade schedules (clinical cannot wait for research domain upgrades), separate vSAN datastores (research workloads cannot impact clinical storage performance), and cleaner audit scope (HIPAA audit covers clinical domain only, not research).

Step 4

Create the final design summary and quality checklist:

  1. Architecture summary:
  • 4 workload domains + 1 management domain = 5 domains
  • 36 hosts total (including DR)
  • 2 sites + 1 witness
  • 235 TB usable storage (150 TB PACS dominant)
  • VCF 9.0 with vSAN ESA, NSX, Aria Suite
  1. Compliance summary:
  • HIPAA Security Rule: 6 § mapped to VCF controls
  • HITRUST CSF: Architecture supports certification timeline
  • Audit: 6-year log retention, quarterly DR testing, annual access review
  1. Cost summary:
  • 36 hosts × 64 cores × 2 sockets = 4,608 licensed cores
  • VCF per-core licensing + external KMS + Aria Suite
  • Target: $4.5M capex (verify against host + storage + networking BOM)
  • Opex: $800K/year (licensing subscription + support + operations staff)
  1. Final quality checklist:
  • [ ] Every HIPAA § has a corresponding VCF control
  • [ ] Research isolation verified at domain + network + DFW + storage levels
  • [ ] PACS storage sized for 3-year growth with alerts
  • [ ] Failover runbook includes clinical-specific verification steps
  • [ ] Air-gapped depot documented for SDDC Manager
  • [ ] 6-year log archival sized and designed
  • [ ] All 10 design decisions trace to requirements
  • [ ] All 8 risks have specific mitigations
  • [ ] VCDX defense responses prepared for 4+ healthcare challenges
  • [ ] Design internally consistent — no contradictions between sections

Validation Gate

Check: Complete design decision register, risk register, defense preparation, and quality checklist for healthcare scenario

Expected: 10 healthcare-specific design decisions with HIPAA mapping, 8 risks with scores and mitigations, 4+ defense responses, comprehensive quality checklist

Common Errors

Not mapping every design decision to a HIPAA section — healthcare auditors require this traceability
Ignoring ransomware as a risk — it is the #1 threat to healthcare IT; the design must include immutable backups and recovery procedure
Not addressing the data de-identification pipeline — researchers will need clinical data; the design must show how this happens securely
Presenting cost estimates without a Bill of Materials — $4.5M must be broken down into hosts, storage, licensing, networking, and professional services

Final Validation

Complete end-to-end VCF 9.0 architecture design for healthcare, production-ready for VCDX submission

✓ Scenario requirements fully analyzed → 7+ business requirements, 7+ technical requirements, 6 workload tiers classified

✓ Architecture covers all domains → 36-host topology, 5 domains, 150 TB PACS addressed, NSX healthcare segmentation

✓ HA/DR design meets healthcare SLAs → 99.997% availability for EHR, failover runbook with clinical verification, annual DR testing

✓ HIPAA compliance demonstrated → 6 HIPAA sections mapped, 6-year log retention, encryption, audit controls

✓ Design defensible in VCDX context → 10 design decisions, 8 risks, 4+ defense responses, quality checklist complete

Cleanup / Restore

• Save the complete healthcare design document as the capstone artifact

• Archive all supporting worksheets, diagrams, and calculations

• This design serves as a portfolio piece for VCDX preparation

Design Reflection (VCDX)

This capstone lab is the closest simulation of a real VCDX design exercise. Healthcare adds layers of complexity that generic designs do not face: regulatory compliance with specific statute references, clinical workflow integration (not just IT SLAs), data classification with IRB isolation, and life-safety implications of downtime. VCDX panelists who review healthcare designs look for: (1) Do you understand that EHR downtime = patient safety risk? The design must reflect this urgency. (2) Can you explain the data isolation model? Research/clinical/admin boundaries must be absolute. (3) Is your DR testing realistic? Healthcare cannot tolerate extended test windows. (4) Do you know HIPAA well enough to map controls to specific sections? Generic 'we comply with HIPAA' is insufficient.

Requirements

  • HIPAA Security Rule compliance with 6 specific section mappings
  • 99.997% availability for EHR (≤15 min downtime/year)
  • 150 TB PACS storage with <5ms read latency for radiologists
  • Complete isolation between research and clinical data (IRB mandate)
  • 6-year audit log retention with integrity verification

Constraints

  • Air-gapped clinical network — no internet access
  • 24/7 operations — maintenance windows ≤4 hours
  • $4.5M capex budget for 3-year lifecycle
  • PACS images have ~1.1:1 dedup ratio (cannot rely on dedup for capacity savings)
  • 15km between sites — stretched cluster viable (0.6ms RTT)

Assumptions

  • Epic and PACS vendors support VMware VCF 9.0 (vendor certification obtained)
  • Dark fiber available between hospital data centers
  • Active Directory with LDAPS available for authentication
  • De-identification pipeline exists for research data access
  • Hospital leadership has approved RPO/RTO targets per tier

Risks

  • PHI data breach via lateral movement (DFW misconfiguration or compromise)
  • Ransomware attack encrypting both data sites simultaneously
  • PACS storage capacity exhaustion before procurement completes
  • EHR downtime exceeding 15-minute target during complex failure scenario
  • HIPAA audit finding due to incomplete log retention or access controls

Self-Assessment Discussion Prompts

  1. If a new HIPAA regulation requires end-to-end encryption for all data in motion, what changes in your design?
  2. How would you handle the scenario where the Epic vendor does not support vSAN ESA — only external storage?
  3. If the hospital acquires a second health system with 300 additional VMs, how does the architecture scale?
  4. What is the business case for spending $4.5M on VCF when a public cloud (AWS/Azure) option exists?

Extensions

Telehealth Integration Architecture

Extend the design to support a telehealth platform with video conferencing VMs, patient portal, and external-facing web services. Design the DMZ architecture with NSX, the DFW rules for external access, and the HIPAA-compliant video/audio encryption requirements.

Medical IoT Device Integration

Design the network architecture for integrating medical IoT devices (patient monitors, infusion pumps, imaging equipment) with the VCF environment. Address device segmentation, traffic inspection, and the unique security challenges of medical devices that cannot be patched.

Multi-Hospital VCF Federation

Extend the single-hospital design to a 3-hospital federation with shared EHR but independent PACS and clinical systems. Design the cross-site Enhanced Linked Mode, shared Content Library, federated monitoring, and the data sovereignty implications of multi-site PHI storage.

⚠ Known Pitfalls (from Community KB)

Treating healthcare VCF like any other enterprise deployment — healthcare has unique constraints (HIPAA, IRB, clinical workflows, 24/7 operations, air-gapped networks) that fundamentally change design decisions
Using standard dedup ratios for PACS storage sizing — medical images (DICOM) are unique per-patient and achieve only ~1.1:1 dedup; relying on vendor-quoted 3:1+ ratios will result in capacity shortfall
Not including clinical-specific verification in DR runbooks — 'VMs are running' is insufficient in healthcare; the runbook must include Epic login test, PACS image retrieval, and HL7 interface verification
Failing to address ransomware recovery — healthcare is the #1 ransomware target; the design must include immutable backups, air-gapped recovery environment, and a documented clean-room restore procedure

References

  • HIPAA Security Rule (45 CFR Part 164, Subpart C)
  • HITRUST CSF v11 — Control Category Reference
  • VMware VCF 9.0 Healthcare Reference Architecture
  • Epic Systems — VMware Support Matrix and Best Practices
  • NIST SP 800-66 — Implementing the HIPAA Security Rule
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.