VCAP — VCF Architect (3V0-12.26)
Study guide for the VCAP-Architect exam (3V0-12.26) — VMware Certified Advanced Professional VMware Cloud Foundation Architect. Covers advanced VCF design including multi-region and multi-organization private cloud architecture, conceptual/logical/physical design, AMPRS (Availability, Manageability, Performance, Recoverability, Security), consumption strategy, monitoring strategy, and end-to-end solution design.
Version Evolution
New VCAP exam introduced in 2026, part of the restructured VCF 9.0 certification path. Replaces legacy VCAP-DCV Design track. Based on VCF 9.0 architecture including VKS, NSX VPC, VCF Automation, and VCF Operations.
Learning Outcomes
- Apply RCAR methodology to decompose business requirements into technical design decisions with full traceability
- Design for all AMPRS dimensions with explicit trade-off analysis — balancing availability, manageability, performance, recoverability, and security
- Differentiate between VCF architecture options (consolidated vs standard vs multi-instance) and select the right topology for given requirements
- Create conceptual, logical, and physical VCF designs that progress from business-language vision through technology choices to implementation specifications
- Design identity management, consumption strategy, and monitoring architecture using VCF 9.0 services (SSO/OIDC, VCF Automation, VCF Operations)
- Defend design decisions under scrutiny using the VCDX defense methodology — justifying from requirements, discussing alternatives, and handling failure scenario challenges
Design Methodology: Requirements, Constraints, Assumptions & Risks (RCAR)#
The VCAP Architect exam and VCDX defense both center on a structured design methodology. Every VCF design begins with gathering business requirements and translating them into technical decisions — documented through RCAR (Requirements, Constraints, Assumptions, Risks) and justified against AMPRS (Availability, Manageability, Performance, Recoverability, Security).
Business vs Technical Requirements:
---
Business Requirements: what the organization needs in business terms
Examples:
- "99.99% application uptime for customer-facing services"
- "Support 200 concurrent developers across 3 geographic regions"
- "Achieve 30% infrastructure cost reduction within 18 months"
- "Meet HIPAA compliance for all workloads processing PHI"
Technical Requirements: derived from business requirements, expressed in infrastructure terms
Examples (mapped from above):
- Business: 99.99% uptime → Technical: vSphere HA with zero-downtime maintenance, stretched cluster or SRM for site failover, N+2 host capacity
- Business: 200 developers, 3 regions → Technical: 3 VCF Instances, VCF Automation multi-cloud catalog, NSX Federation for stretched networking
- Business: 30% cost reduction → Technical: VCF Operations rightsizing, vSAN ESA (fewer disk groups, lower licensing), DRS resource optimization
- Business: HIPAA compliance → Technical: VM Encryption, vSAN encryption at rest, NSX micro-segmentation for PHI workload isolation, audit loggingRCAR Framework:
---
Requirements (R):
- Functional: what the system must do (workload hosting, DR capability, multi-tenancy)
- Non-functional: how well it must do it (performance targets, availability SLAs, compliance mandates)
- Capacity: scale targets (VM count, storage capacity, network throughput)
- Rule: every requirement must be measurable and traceable to a business objective
Constraints (C):
- Budget: hardware/licensing spend limits
- Technical: existing infrastructure that must be reused (e.g., "must integrate with existing Cisco Nexus 9000 switches")
- Organizational: staffing, skills, operational model
- Regulatory: data sovereignty, compliance frameworks
- Time: deployment timeline
- Rule: constraints limit design options. A constraint is not negotiable — if it were, it would be a requirement.
Assumptions (A):
- Statements believed to be true but not yet validated
- Example: "WAN bandwidth between Site A and Site B is sufficient for vSAN stretched cluster witness traffic"
- Example: "All ESXi hosts will have TPM 2.0 modules installed"
- Rule: every assumption carries risk. If an assumption proves false, the design may need revision.
- Best practice: rank assumptions by impact — which assumptions, if wrong, would invalidate the design?
Risks (R):
- Potential events that could negatively impact the design
- Each risk needs: probability assessment, impact severity, and mitigation strategy
- Example: Risk: "vSAN stretched cluster inter-site latency exceeds 5ms during peak hours"
Probability: Medium. Impact: High (vSAN I/O performance degradation).
Mitigation: dedicated dark fiber between sites, QoS marking for vSAN traffic, pre-deployment latency testing.
- Rule: risks are the inverse of assumptions — they describe what could go wrong.
Design Decisions:
---
Every design choice must be documented as a formal design decision:
Structure:
Decision ID: DESIGN-001
Decision Statement: "Deploy a 4-node vSAN stretched cluster for the management domain"
Justification: Meets the 99.99% availability requirement by providing site-level fault tolerance. vSAN stretched cluster provides synchronous replication with RPO=0 across sites.
Implications: Requires <5ms RTT between sites, dedicated witness host at third site, additional licensing for vSAN stretched cluster feature.
Alternatives Considered:
- SRM with array-based replication (rejected: RPO > 0, more complex failover)
- Active-passive with shared storage (rejected: constraint — no SAN infrastructure available)
AMPRS Impact: Availability ↑ (site failover), Performance ↓ (write latency increases by ~1ms due to synchronous replication), Manageability → (vSAN health monitoring covers both sites)
Risk: If inter-site latency exceeds 5ms, vSAN performance degrades. Mitigation: dedicated link with SLA.
Design Validation:
---
A validation strategy ensures the design meets requirements before implementation:
Validation Methods:
- Requirements Traceability Matrix: map every requirement to design decisions that address it
- Proof of Concept: deploy in lab (Holodeck Toolkit) to validate key design assumptions
- Peer Review: architecture review board or VCDX panelist feedback
- Failure Mode Analysis: for each AMPRS dimension, enumerate failure scenarios and verify the design handles them
- Capacity Model Validation: confirm sizing meets requirements with growth projections
VCDX Defense Context:
- Panelists will challenge your design decisions — expect "why not X instead of Y?"
- Every decision must trace back to a requirement or constraint
- "It's best practice" is not a valid justification — justify from YOUR requirements
- Be prepared to discuss trade-offs: what did you sacrifice and why was it acceptable?
Key Takeaways
- Business requirements describe WHAT the org needs; technical requirements describe HOW infrastructure delivers it. Every technical requirement must trace to a business driver.
- RCAR: Requirements are measurable needs, Constraints are non-negotiable limits, Assumptions are unvalidated beliefs that carry risk, Risks are potential negative events with probability and mitigation.
- Design decisions must follow a structured format: statement, justification (from requirements/constraints), implications, alternatives considered, and AMPRS impact analysis.
- Validation strategy includes: requirements traceability matrix, proof-of-concept (Holodeck), peer review, failure mode analysis, and capacity model validation.
AMPRS Design Framework: Availability, Manageability, Performance, Recoverability, Security#
AMPRS is the core design quality framework for VCF architecture. Each dimension represents a design quality that must be explicitly addressed, balanced, and justified in trade-off analysis.
Availability Design:
---
Availability = minimizing unplanned downtime. VCF provides multiple availability layers:
Compute Availability:
vSphere HA: restarts VMs on surviving hosts after host failure
- Admission Control: reserves capacity for failover (e.g., 1 host failure tolerated = N+1)
- Proactive HA: migrates VMs away from hosts reporting hardware warnings (IPMI/iLO sensors)
- VM Component Protection (VMCP): responds to datastore accessibility failures (APD/PDL)
- Application-level HA: vSphere HA monitors VMware Tools heartbeat — restarts unresponsive VMs
vSphere FT (Fault Tolerance): zero-downtime protection for mission-critical single VMs
- Creates live shadow instance on separate host
- Consumes 2x resources — use only for VMs where even HA restart time is unacceptable
- VCF 9.0: FT supports up to 8 vCPUs per protected VM
Design patterns by SLA:
99.9% (8.7h/year downtime): vSphere HA, N+1 capacity, single site
99.99% (52.6min/year): vSphere HA + Proactive HA, N+2 capacity, maintenance without impact
99.999% (5.3min/year): Stretched cluster or active-active multi-site, application-level clustering
Storage Availability:
vSAN Fault Domains: distribute hosts across racks/power circuits to survive rack failure
- Minimum 3 fault domains for FTT=1 (RAID-1 mirroring)
- 4+ fault domains for OSA RAID-5 3+1 (FTT=1, 1.33x vs 2x for RAID-1); ESA RAID-5 is adaptive - 2+1 (1.5x) below 6 hosts, 4+1 (1.25x) at 6+
vSAN Stretched Cluster: synchronous replication across 2 sites + witness at 3rd site
- RPO=0, RTO=minutes (HA restart at surviving site)
- Requirements: <5ms RTT between data sites, <200ms RTT to witness- vSAN ESA stretched cluster improvements: single-tier storage pool simplifies stretched operations
Network Availability:
NSX Edge HA: active-standby Edge pairs for T0 gateways
- NSX Edge Cluster: multi-node Edge cluster with ECMP for north-south load distribution
- Dual TEP uplinks: redundant tunnel endpoints on each host
- Physical network: dual ToR switches, LAG/LACP, diverse uplink paths
Manageability Design:
---
Manageability = operational complexity, day-2 effort, and skill requirements.
VCF Manageability Features:
VCF Lifecycle Management (LCM): automated stack upgrades (NSX → vCenter → ESXi → vSAN)
VCF Operations: centralized monitoring, alerting, capacity planning
VCF Automation: self-service catalog, infrastructure-as-code templates
VCF Configuration Drift: desired-state configuration baselines, automated compliance scanning
Design Trade-offs:
More availability layers → more complexity → harder to manage
Example: stretched cluster (high availability) requires cross-site network management, witness monitoring, split-brain handling — more operational burden than single-site HA
Standardization reduces manageability burden:
- Standard host profiles across all domains
- Consistent VLAN assignments per domain type
- Uniform VM templates with approved OS versions
- Configuration-as-code via VCF Automation
Performance Design:
---
Performance = responsiveness and throughput under expected workload.
Compute Performance:
CPU: cores per socket, clock speed, NUMA awareness
- NUMA alignment: VM vCPU count ≤ physical cores per socket to avoid cross-NUMA memory access (30-40% latency penalty)
- Over-commitment ratios: depends on workload type
- Production databases: 1:1 to 2:1 vCPU-to-pCPU
- General servers: 4:1 to 6:1
- VDI: 6:1 to 10:1
Memory: ECC, bandwidth, NUMA-local allocation
- vSAN ESA minimum: 128GB RAM per host (vSAN buffer cache requires significant memory)
Storage Performance:
- vSAN ESA performance tiers:
- All-NVMe (required for ESA): mixed workload IOPS at sub-100μs latency
- Storage policy based management (SPBM): per-VM performance via storage policies
- FTT setting affects performance: RAID-1 (mirror) = better read performance, RAID-5/6 (erasure coding) = better capacity efficiency but slightly higher write latency
Network Performance:
NSX overlay overhead: ~50 bytes GENEVE encapsulation per packet
- MTU 9000 (jumbo frames): required for TEP interfaces to avoid fragmentation
- 25GbE minimum for vSAN ESA (10GbE insufficient for NVMe-tier storage traffic)
- Edge sizing: DPDK-accelerated data path — Large/XLarge VMs for >10Gbps throughput
Recoverability Design:
---
Recoverability = ability to restore service after data loss or site-level failure.
Backup Strategy:
VCF component backups: SDDC Manager, vCenter, NSX Manager, VCF Operations
- File-based backup (FBB) for vCenter: scheduled, stored on remote datastore or NFS
- NSX Manager backup: configuration + distributed firewall rules
- VCF Operations: Cassandra database backup
VM-level backup: third-party (Veeam, Cohesity, Dell Avamar) or vSphere API for Data Protection (VADP)
RPO/RTO Targets by Tier:
Tier 1 (mission-critical): RPO ≤ 15 min, RTO ≤ 1 hour → vSAN stretched cluster or SRM with synchronous replication Tier 2 (business-important): RPO ≤ 4 hours, RTO ≤ 4 hours → SRM with asynchronous replication Tier 3 (non-critical): RPO ≤ 24 hours, RTO ≤ 24 hours → backup/restore
VCF Site Recovery:
VMware Live Recovery (formerly SRM): automated DR orchestration
- Recovery plans: ordered VM startup, IP customization, runbook integration
- Test recovery: non-disruptive DR testing with network isolation
- Reprotect: reverse replication after failover
Security Design:
---
Security = protecting infrastructure, data, and access at all layers.
Defense in Depth for VCF:
- Network: NSX micro-segmentation (DFW), gateway firewall, IDS/IPS (vDefend)
- Compute: VM Encryption (vSphere 7.0+), vTPM, Secure Boot (host + VM)
- Storage: vSAN encryption at rest (data + cache tiers), encryption in transit
- Identity: RBAC, SSO federation, MFA via OIDC, least-privilege roles
- Compliance: STIGs, CIS Benchmarks, configuration drift baselines
- Audit: syslog forwarding, VCF Operations audit trails, NSX IPFIX flow logging
Key Takeaways
- Availability: layer vSphere HA, Proactive HA, vSAN fault domains, stretched cluster, and NSX Edge HA according to SLA tiers (99.9% → 99.999%). Each layer adds complexity — justify with requirements.
- Performance: NUMA alignment (vCPU ≤ cores/socket), vSAN ESA requires 25GbE and 128GB RAM minimum, MTU 9000 for TEP interfaces, Edge sizing matches throughput needs.
- Recoverability: map RPO/RTO targets to protection tiers — stretched cluster for RPO=0, SRM async for RPO ≤ 4h, backup for RPO ≤ 24h. Include VCF component backups (vCenter FBB, NSX config, VCF Ops Cassandra).
- Security: defense in depth across all layers — NSX DFW (network), VM Encryption + vTPM (compute), vSAN encryption (storage), RBAC + MFA (identity), STIGs + drift baselines (compliance).
- AMPRS trade-offs are the heart of design justification: more availability = more complexity (manageability cost), more security = more operational overhead, higher performance = higher cost. Document trade-offs explicitly.
VCF Architecture Options, Topology & Advanced Services#
VCF supports multiple architecture options depending on scale, workload requirements, and operational model. The VCAP Architect must differentiate between recommended and supported configurations, and select the right topology for each use case.
VCF Deployment Models:
---
Consolidated Architecture (Single Domain):
- Management + workload VMs in the same domain (management domain hosts workloads)
- Minimum footprint: 4 hosts (vSAN ESA minimum)
- Use case: small deployments, edge sites, development/test environments
- Constraint: no workload isolation from management plane
- Recommendation: supported but NOT recommended for production — management domain resource contention can impact VCF component health
Standard Architecture (Multi-Domain):
- Dedicated management domain + one or more workload domains
- Management domain: VCF components only (vCenter, NSX, VCF Operations, SDDC Manager)
- Workload domains: application workloads, each with dedicated vSAN/networking
- Minimum: 4 hosts management + 4 hosts per workload domain = 8 hosts minimum for standard
- Use case: production environments, multi-tenant deployments
- Recommendation: RECOMMENDED for production. Clear separation of management and workload fault domains.
Multi-Instance Architecture:
- Multiple independent VCF Instances, managed as a Fleet
- Each Instance has its own management domain + workload domains
- Fleet-level management via VCF Operations
- Use case: multi-region deployments, large enterprises, managed service providers
- Example: 3 VCF Instances (Americas, EMEA, APAC), each with management + workload domains, all in one Fleet
VCF Domain Topologies:
---
Workload Domain Types:
vSphere + vSAN + NSX: standard VCF workload domain (most common)
vSphere + External Storage + NSX: for workloads requiring NFS/FC/iSCSI storage (not vSAN)
vSphere + vSAN + NSX + VKS: Kubernetes workload domain (VCF 9.0) — deploys Tanzu/VKS for container workloads
vSphere + vSAN ReadyNode + NSX: uses HCI Mesh for disaggregated storage (compute-only hosts accessing vSAN storage from other clusters)
Cluster Configurations:
Standard Cluster: all hosts identical (balanced compute + storage)
Stretched Cluster: hosts split across 2 fault domains (sites) with witness at 3rd
- vSAN stretched cluster configuration: site affinity, preferred/secondary site
- NSX spanning across sites: cross-site segments, cross-site DFW
HCI Mesh: compute-only hosts mounting remote vSAN datastores from storage-contributing hosts
- Decouples compute scaling from storage scaling
- Constraint: network latency between compute and storage clusters < 0.5ms RTT
Recommended vs Supported:
---
VMware classifies configurations as:
Recommended: tested, validated, fully supported, design guidance provided
Supported: works, is supported, but may have limitations or require extra validation
Key distinctions for VCAP:
Recommended: Standard architecture with dedicated management domain
- Supported: Consolidated architecture (small footprint, but not for production)
- Recommended: vSAN ESA (single-tier NVMe, new default for VCF 9.0)
- Supported: vSAN OSA (legacy dual-tier with cache + capacity, for brownfield)
- Recommended: NSX overlay networking for all workload domains
- Supported: VLAN-backed networking (limited — no micro-segmentation, no NSX DFW)
Recommended: VCF Operations for monitoring/capacity
Not Supported: third-party monitoring as replacement for VCF Operations health checks in LCM
VCF Advanced Services:
---
VCF 9.0 includes advanced services beyond core compute/storage/network:
VCF Automation (formerly vRealize Automation / Aria Automation):
- Infrastructure-as-Code: Cloud Templates (YAML) define VMs, networks, load balancers
- Service Catalog: self-service portal for tenants to request infrastructure
- Extensibility: ABX (Action Based Extensibility) for custom workflows, event-driven automation
- Multi-cloud: provision to vSphere, AWS, Azure, GCP from single catalog
- Day-2 operations: resize, snapshot, reconfigure through catalog actions
- Use case: any organization with >50 VM deployment requests per month, DevOps teams requiring IaC
VCF Operations (formerly vRealize Operations / Aria Operations):
- Monitoring: real-time and historical metrics, alerts, dashboards
- Capacity Planning: time remaining, what-if analysis, rightsizing
- Cost Management: chargeback/showback, cost optimization
- Compliance: configuration drift, STIG/CIS baselines
- Multi-cloud: monitor AWS, Azure, GCP alongside vSphere
- Use case: ALL VCF deployments (required for LCM health checks)
VCF Operations for Networks (formerly vRealize Network Insight / Aria Operations for Networks):
- Network visibility: flow analytics, dependency mapping
- Security planning: DFW 1-2-3-4 workflow — discover, segment, monitor, enforce
- Troubleshooting: path analysis, latency detection
- Use case: NSX micro-segmentation planning, compliance auditing, network troubleshooting
VMware Kubernetes Service (VKS) — VCF 9.0:
- Integrated Kubernetes: deploy Tanzu Kubernetes clusters on VCF infrastructure
- vSphere Pod Service: run containers directly on ESXi (no worker node VM overhead)
- Supervisor Cluster: Kubernetes control plane integrated with vSphere
- Use case: organizations running containerized workloads alongside VMs, Kubernetes-native DevOps
VMware Live Recovery (formerly Site Recovery Manager):
- DR orchestration: automated failover/failback with recovery plans
- Continuous replication: vSphere Replication for VM-level RPO
- Test recovery: non-disruptive DR testing
- Use case: regulated industries requiring documented DR testing, multi-site deployments
vDefend (formerly NSX Advanced Threat Prevention):
- Distributed IDS/IPS: signature-based threat detection at every vNIC
- Network Detection and Response (NDR): behavioral analytics, campaign tracking
- Malware Prevention: sandboxing, file analysis
- Use case: zero-trust security model, regulated workloads requiring perimeter + east-west threat detection
Key Takeaways
- Consolidated architecture (single domain) is supported but NOT recommended for production. Standard architecture (management + workload domains) is the production recommendation.
- Recommended vs Supported distinction is critical for VCAP: vSAN ESA (recommended) vs OSA (supported/brownfield), NSX overlay (recommended) vs VLAN-backed (supported/limited).
- VCF Advanced Services: Automation (IaC + self-service), Operations (monitoring + capacity), Ops for Networks (flow analytics + DFW planning), VKS (Kubernetes), Live Recovery (DR), vDefend (IDS/IPS + NDR).
- HCI Mesh decouples compute from storage scaling — compute-only hosts mount remote vSAN datastores. Requires <0.5ms RTT between clusters.
- VCF 9.0 new capabilities: VKS for Kubernetes, NSX VPC for multi-tenancy, VCF Operations as primary management plane (SDDC Manager UI deprecated).
Conceptual, Logical & Physical Design Progression#
VCF design follows a three-layer progression: Conceptual → Logical → Physical. Each layer adds implementation detail while maintaining traceability to business requirements. This progression is the backbone of VCDX documentation.
Conceptual Design:
---
Purpose: communicate the high-level vision to business stakeholders (non-technical audience)
Content: what the solution does, not how it does it
Conceptual Design Elements:
- Solution Overview: one-page description of the VCF private cloud and its business purpose
- Stakeholder Mapping: who sponsors, consumes, operates, and governs the solution
- Use Case Diagram: application workload types and their relationship to infrastructure
- Service Level Definitions: availability tiers, performance expectations, recovery objectives
- Scope & Boundaries: what is included/excluded from the design
Example Conceptual Statement:
"The VCF Private Cloud will provide a self-service, multi-tenant infrastructure platform hosting 500 virtual machines across 3 workload domains, with 99.99% availability for Tier-1 applications, integrated disaster recovery with RPO ≤ 15 minutes, and compliance with PCI-DSS and HIPAA frameworks."
Key Principles:
- No vendor-specific technology names (say "software-defined networking" not "NSX")
- No IP addresses, VLAN IDs, or host counts
- Focus on business outcomes: cost savings, agility, compliance, risk reduction
- Diagram: simple boxes showing workload tiers, connectivity, and management
Logical Design:
---
Purpose: define the architecture without specifying physical implementation (what components, how they connect, what services they provide)
Content: technology choices, component relationships, network topology, storage architecture
Logical Design Elements:
Compute:
- vSphere cluster architecture: HA/DRS configuration, resource pool strategy
- VM sizing methodology: T-shirt sizes (S/M/L/XL), resource allocation policies
- NUMA awareness requirements: maximum vCPU-per-VM guidelines
Storage:
- vSAN architecture: ESA vs OSA, FTT policy, storage policies per workload tier
- Storage Policy Based Management (SPBM): policy definitions mapped to service levels
- Data services: deduplication, compression, encryption, snapshots
Network:
- NSX overlay architecture: T0/T1 gateway hierarchy, segment design
- VPC model: Project → VPC → Subnet hierarchy for multi-tenancy
- Physical network requirements: VLAN topology, MTU, uplink redundancy
- Security zones: trust zones mapped to NSX DFW categories
Management:
- VCF Operations architecture: node sizing, retention policies
- VCF Automation: cloud template strategy, approval workflows
- Monitoring & Alerting: alert escalation paths, notification channels
Identity & Security:
- Identity sources: SSO, AD/LDAP, OIDC federation
- Role mapping: personas (infra admin, tenant admin, developer) → RBAC roles
- Certificate strategy: VMCA subordinate CA, rotation schedule
DR & Backup:
- Recovery architecture: primary-secondary site topology
- RPO/RTO mapping to recovery technology (stretched cluster, SRM, backup)
- Recovery plan structure: order, dependencies, validation tests
Logical Design Diagrams:
- Logical network topology (T0/T1/segments, not physical switches)
- Logical storage architecture (vSAN clusters, storage policies, not disk models)
- Logical management plane (component relationships, not IP addresses)
- Data flow diagrams: how workload traffic traverses the architecture
Physical Design:
---
Purpose: specify exact implementation details — product models, quantities, IP addresses, configurations
Content: bill of materials, rack layouts, cabling, IP schemas, ESXi configurations
Physical Design Elements:
Hardware:
- Server specifications: Dell PowerEdge R760, 2x Intel Xeon Gold 6448Y (32-core), 1TB RAM, 8x 3.84TB NVMe
- Host count per cluster: management domain (4 hosts), workload domain A (8 hosts), workload domain B (6 hosts)
- Network infrastructure: 2x Cisco Nexus 93180YC-FX3 per rack (ToR), MTU 9216
- Edge appliance sizing: Large (8 vCPU, 32GB) for T0 gateway, Medium (4 vCPU, 8GB) for T1
Network:
- VLAN assignments: Management (VLAN 10), vMotion (VLAN 20), vSAN (VLAN 30), TEP (VLAN 40), Uplink (VLAN 100-110)
- IP addressing: Management subnet 10.10.10.0/24, vMotion 10.10.20.0/24, vSAN 10.10.30.0/24
- BGP peering: T0 uplinks to physical router (AS 65001), ECMP with 4 Edge nodes
- DNS/NTP/Syslog: specific server IPs, forwarder configuration
Storage:
- vSAN ESA configuration: single storage pool, RAID-5 erasure coding (FTT=1)
- Usable capacity calculation: Raw (8 hosts × 8 disks × 3.84TB = 245.76TB) → RAID-5 (184TB usable) → Slack space 25% (138TB effective)
- Storage policies: Gold (RAID-1, thick), Silver (RAID-5, thin), Bronze (RAID-6, thin + compression)
Sizing & Capacity:
- CPU: 8 hosts × 64 cores = 512 cores total. HA reservation (N+1) = 448 usable cores. Target 4:1 ratio = 1,792 vCPU capacity.
- Memory: 8 hosts × 1TB = 8TB total. HA reservation = 7TB. vSAN overhead (128GB × 8) = 1TB. Workload available = 6TB.
- Network: 2x 25GbE per host. vSAN (1x 25GbE), vMotion (1x 25GbE via sharing), management + TEP (remaining bandwidth).
Physical Design Diagrams:
- Rack elevation diagrams: server placement, switch placement, cabling
- Physical network topology: switch-to-host cabling, port assignments
- Physical storage layout: disk placement, NVMe slot assignments
- IP address allocation tables: every IP assigned to a specific component
Design Decision Documentation Throughout:
---
Each layer produces design decisions:
Conceptual: "The solution will provide three tiers of service with differentiated SLAs"
Logical: "Tier-1 workloads will use vSAN RAID-1 storage policy with FTT=1 for performance"
Physical: "vSAN ESA RAID-1 on Dell PowerEdge R760 with 8x 3.84TB NVMe per host provides 92TB usable for Tier-1"
Traceability: Physical → Logical → Conceptual → Business Requirement
Every physical configuration choice should trace back through the logical design to a conceptual decision to a business requirement.
Key Takeaways
- Conceptual design: business-language, no vendor terms, stakeholder mapping, service level definitions, scope boundaries. Audience: business stakeholders and sponsors.
- Logical design: technology choices without implementation specifics — vSAN ESA with RAID-5 (not 'Dell R760 with 8 NVMe disks'), NSX T0/T1 hierarchy (not 'VLAN 40, IP 10.10.40.0/24'). Audience: technical architects and reviewers.
- Physical design: exact specifications — server models, core counts, IP addresses, VLAN IDs, rack elevations. This is the implementation blueprint. Audience: deployment engineers.
- Traceability is the golden thread: every physical choice → logical design decision → conceptual requirement → business objective. VCDX panelists will trace this chain.
- vSAN capacity calculation: Raw capacity → RAID overhead (OSA RAID-5 3+1 = 1.33x; ESA adaptive RAID-5 = 1.25x at 6+ hosts or 1.5x below; RAID-6 4+2 = 1.5x; RAID-1 = 2x) → Reserved Capacity (Operations Reserve + Host Rebuild Reserve since vSAN 7 U1, NOT a flat 25% slack) → Operations overhead → Effective usable capacity.
Identity, Consumption Strategy & Monitoring Design#
Three critical design domains complete the VCF architecture: identity management (who accesses what), consumption strategy (how users request and use infrastructure), and monitoring strategy (how the platform is observed and maintained).
Identity Management Design:
---
VCF identity architecture must address four personas with distinct access requirements:
Persona-to-Role Mapping:
Infrastructure Administrator:
- VCF: Super Admin (all VCF operations)
- vCenter: Administrator (vsphere.local\Administrators group)
- NSX: Enterprise Admin (full NSX access)
- VCF Operations: Content Administrator
- Scope: full platform access for Day 0/Day 1 operations
Tenant/Domain Administrator:
- VCF: Domain Admin (scoped to assigned workload domain)
- vCenter: Resource Pool/Folder Admin (delegated, scoped)
- NSX: Project Admin (VPC management within assigned Project)
- VCF Operations: Dashboard/Report consumer
- Scope: self-service within assigned boundary, no infrastructure modification
Developer/Consumer:
- VCF Automation: Catalog Consumer (request VMs, VPCs, Kubernetes clusters)
- vCenter: Read-only or VM User (power operations on owned VMs)
- NSX: VPC User (manage subnets/security within assigned VPC)
- Scope: consume services, no infrastructure visibility
Auditor/Compliance:
- VCF Operations: Read-only access to compliance dashboards
- vCenter: Read-only (no modification rights)
- NSX: Auditor role (view-only for DFW rules and flow logs)
- Scope: verify compliance without operational access
Identity Federation Architecture:
Single-Site:
Active Directory → vCenter SSO (LDAP identity source) → all VCF components
Multi-Site:
Corporate AD (global catalog) → vCenter SSO at each site → NSX + VCF Ops at each site
OR: OIDC provider (Okta/Azure AD) → vCenter SSO federation → centralized MFA + governanceDesign Decision: OIDC vs AD/LDAP:
OIDC: modern, MFA-native, cloud-compatible, but requires OIDC provider infrastructure
AD/LDAP: mature, widely deployed, Kerberos-capable, but MFA requires additional integration
Recommendation: OIDC for new VCF 9.0 deployments, AD/LDAP where AD is already the enterprise standard
Consumption Strategy Design:
---
Consumption strategy defines how users request, provision, and manage infrastructure resources.
Consumption Models:
- Ticket-Based (traditional):
- User submits ticket → infra team provisions manually → handover
- Pros: full control, human review
- Cons: slow (days/weeks), error-prone, doesn't scale
- When to use: regulated environments requiring human approval for every change
- Self-Service Catalog (VCF Automation):
- Admin publishes Cloud Templates (VM, VPC, Kubernetes cluster)
- Users request from catalog → automated provisioning (minutes)
- Approval workflows: optional, configurable per template
- Day-2 actions: resize, snapshot, reconfigure through catalog
- When to use: DevOps teams, development environments, any scale >50 VMs
- Infrastructure-as-Code (IaC):
- Terraform provider for vSphere, NSX, VCF
- Git-based workflows: PR review → merge → automated deployment
- When to use: mature DevOps organizations, CI/CD integration, GitOps model
- Kubernetes-Native (VKS):
- VCF 9.0 VMware Kubernetes Service
- Developers deploy containers; VKS manages infrastructure
- When to use: cloud-native applications, microservices architectures
Cloud Template Design:
- T-Shirt Sizing: Small (2 vCPU, 4GB), Medium (4 vCPU, 8GB), Large (8 vCPU, 16GB), XLarge (16 vCPU, 32GB)
- Naming convention: enforce via VCF Automation custom naming (prefix-env-app-index)
- Network assignment: Cloud Templates auto-select segment/VPC based on project mapping
- Storage policy: mapped to service tier (Gold/Silver/Bronze → vSAN storage policy)
- Lifecycle: templates versioned in Git, tested in dev before publishing to production catalog
Quota & Governance:
- Resource Quotas: limit CPU, memory, storage per project/tenant
- Lease Policies: auto-expire VMs after configurable period (30/60/90 days)
- Approval Workflows: multi-level approval for large resource requests
- Naming Policies: enforce consistent naming across all provisioned resources
- Cost Transparency: showback/chargeback dashboards per project/team
Monitoring Strategy Design:
---
Monitoring design ensures the VCF platform is observable, alertable, and actionable.
Monitoring Architecture Layers:
- Infrastructure Monitoring (VCF Operations):
- Metrics: CPU, memory, disk, network for hosts, VMs, datastores
- Collection: 5-minute default interval (configurable per object type)
- Retention: real-time (1 day, 5-min granularity) → hourly (30 days) → daily (6 months) → weekly (2 years)
- Super Metrics: custom calculated metrics for business-level KPIs (cost per VM, efficiency scores)
- Network Monitoring (VCF Operations for Networks):
- Flow collection: IPFIX from NSX DFW, NetFlow from physical switches
- Flow analytics: application dependency mapping, micro-segmentation planning
- Path analysis: trace packet path through overlay and underlay
- Log Monitoring:
- Syslog forwarding: all VCF components → central syslog (vRealize Log Insight or Splunk/ELK)
- Log retention: compliance-driven (PCI-DSS: 1 year, HIPAA: 6 years)
- Structured logging: ESXi, vCenter, NSX, vSAN emit structured syslog events
- Availability Monitoring:
- Synthetic monitoring: periodic health checks for management components
- SoS utility: on-demand health checks across the VCF stack
- LCM integration: VCF Operations health check required before upgrade operations
Alert Strategy:
Severity Levels:
Critical: immediate action required (host failure, datastore full, component down)
Warning: attention needed within business hours (capacity threshold, certificate expiry, drift detected)
Info: awareness only (VM migration completed, backup successful)
Alert Routing:
Critical → PagerDuty/ServiceNow + email + Slack → on-call engineer
Warning → email + Slack → platform team queue
Info → dashboard only (no notification)Alert Hygiene:
- Tune thresholds to reduce false positives (dynamic thresholds preferred over static)
- Correlation rules: suppress child alerts when parent is known (host down → suppress VM-level alerts)
- Maintenance windows: suppress alerts during planned maintenance
Dashboard Strategy:
- Executive Dashboard: availability SLA, capacity utilization, cost trends, compliance score
- Operations Dashboard: real-time health, active alerts, recent changes, top consumers
- Capacity Dashboard: time remaining, growth trends, what-if scenarios
- Tenant Dashboard: per-project resource usage, cost allocation, VM inventory
Key Takeaways
- Four identity personas: Infrastructure Admin (full access), Tenant Admin (domain-scoped), Developer (catalog consumer), Auditor (read-only compliance). Map each to specific VCF/vCenter/NSX/VCF Ops roles.
- Consumption models: Ticket-based (regulated, slow), Self-service catalog (VCF Automation, minutes), IaC (Terraform/Git, DevOps), Kubernetes-native (VKS, cloud-native). Select based on organizational maturity and workload type.
- Governance essentials: resource quotas, lease policies (auto-expire), approval workflows, naming conventions, cost transparency (showback/chargeback). Without governance, self-service leads to sprawl.
- Monitoring layers: Infrastructure metrics (VCF Operations), Network flows (Ops for Networks), Logs (syslog/SIEM), Availability (SoS/synthetic). Alert routing by severity to appropriate channels with maintenance window suppression.
- Dashboard hierarchy: Executive (SLA + cost), Operations (health + alerts), Capacity (time remaining + what-if), Tenant (per-project usage + cost). Each audience sees what they need to act on.
End-to-End Solution Design & VCDX Defense Preparation#
The VCDX defense evaluates your ability to create, justify, and defend a complete VCF solution design under scrutiny. This theory synthesizes all design domains into a cohesive approach and prepares you for the defense experience.
End-to-End Design Document Structure:
---
A complete VCF design document follows this structure:
- Executive Summary (1-2 pages):
- Business context and strategic objectives
- Solution overview in business language
- Key outcomes: cost reduction, agility, compliance, risk mitigation
- Scope and timeline
- Requirements & Constraints (5-10 pages):
- Business requirements (traced from stakeholder interviews)
- Technical requirements (derived from business requirements)
- Constraints (budget, technical, organizational, regulatory, timeline)
- Assumptions (each with impact assessment if invalidated)
- Risks (probability, impact, mitigation for each)
- Requirements Traceability Matrix: every requirement → design decision → validation method
- Conceptual Design (5-10 pages):
- Solution vision and use case model
- Service tier definitions (Tier-1/2/3 with SLAs)
- Stakeholder and RACI matrix
- Conceptual architecture diagram (no vendor terms)
- Logical Design (20-30 pages):
- Compute: cluster architecture, HA/DRS, resource allocation
- Storage: vSAN architecture, storage policies, data services
- Network: NSX topology (T0/T1/segments), security zones, VPC model
- Management: VCF Operations, Automation, monitoring strategy
- Identity & Security: federation, RBAC, encryption, compliance
- DR & Recoverability: site topology, RPO/RTO mapping, recovery plans
- Design decisions: formal documentation for each major choice
- Physical Design (20-30 pages):
- Bill of Materials: hardware specifications, quantities, costs
- Rack layouts and cabling diagrams
- IP addressing and VLAN assignments
- ESXi host configuration: NTP, DNS, syslog, firewall rules
- vSAN configuration: disk group layout, storage policies, capacity calculations
- NSX configuration: Edge sizing, ECMP, TEP pools, DFW rule sets
- VCF Operations: node sizing, retention, integration configuration
- Validation & Testing (5-10 pages):
- Validation strategy: how each requirement will be verified
- Test plan: functional tests, performance tests, failover tests
- Acceptance criteria: measurable outcomes for sign-off
- Operational Model (5-10 pages):
- Day-2 operations: monitoring, patching, capacity management
- Escalation paths and support model
- Change management process
- Documentation and training plan
VCDX Defense Format:
---
The VCDX defense consists of three components:
Part 1: Design Presentation (75 minutes):
- You present YOUR design to a panel of VCDX holders
- Format: slides + design documentation
- Focus: justify design decisions, demonstrate RCAR analysis, show AMPRS trade-offs
- Tip: lead with the problem (business requirements), then show how your design solves it
- Tip: anticipate challenges — where panelists will push back
Part 2: Design Defense (75 minutes):
- Panelists ask probing questions about YOUR design
- They will challenge assumptions, constraints, and decisions
- They will propose failure scenarios and ask how your design responds
- They will suggest alternative approaches and ask why you didn't choose them
Common Defense Questions:
"What happens when [component X] fails?" → describe HA mechanism, RTO, blast radius
"Why did you choose [option A] over [option B]?" → trace to requirement or constraint
"What if [assumption Y] is wrong?" → describe risk mitigation, fallback design
"How does this scale to [2x/5x/10x] the current requirement?" → capacity model, growth plan
"Show me the requirement that drives this decision" → requirements traceability
"What would you change if the budget were halved?" → prioritization, trade-off analysisPart 3: Design Scenario (50 minutes):
- Panelists present a NEW scenario you haven't prepared for
- You must design a solution live using VCF components
- Tests your depth of knowledge and ability to think architecturally under pressure
- Approach: immediately gather RCAR, sketch conceptual design, then work through logical design
- Tip: think aloud — panelists evaluate your process as much as your answer
Defense Anti-Patterns:
---
✗ "It's VMware best practice" → not a justification. WHY is it best practice for YOUR design? ✗ "We've always done it this way" → not a design rationale ✗ Avoiding trade-off discussion → panelists will press harder ✗ Over-engineering without requirements backing → every feature must trace to a need ✗ Dismissing alternatives without analysis → "we didn't consider that" is worse than "we considered X but chose Y because..." ✗ Panicking on unknowns → "I would need to validate that assumption" is a valid answer
Defense Best Practices:
---
✓ Every decision traces to RCAR → requirement, constraint, assumption, or risk mitigation ✓ AMPRS trade-offs are explicit → "this improves availability but adds manageability complexity" ✓ Failure scenarios are anticipated → you've done failure mode analysis for each AMPRS dimension ✓ Alternatives are documented → for each major decision, you considered 2-3 alternatives ✓ Capacity model is defensible → show the math, include growth projections ✓ You know your constraints → "we chose this because the customer requires X" ✓ You own your design → "I would change this in retrospect" shows maturity ✓ Quantified outcomes → "this design delivers 99.99% availability and 30% cost reduction vs. the previous architecture"
Design Scenario Methodology:
---
When facing an unfamiliar scenario in Part 3, follow this structured approach:
Step 1 (5 min): Gather Requirements
- Ask clarifying questions: "What are the availability requirements? How many users? What compliance frameworks?"
- Identify the top 3 business drivers
Step 2 (5 min): Identify Constraints & Assumptions
- Budget, timeline, existing infrastructure
- Skills and operational model
- Document assumptions explicitly: "I'm assuming inter-site latency < 5ms"
Step 3 (10 min): Conceptual Architecture
- Sketch the high-level solution on a whiteboard
- Show site topology, workload tiers, management plane
- Validate with panelists: "Does this align with the requirements you described?"
Step 4 (15 min): Logical Design
- Walk through compute, storage, network, management, security
- Make design decisions verbally: "I'm choosing vSAN stretched cluster because the RPO=0 requirement eliminates async replication"
- Call out AMPRS implications for each decision
Step 5 (10 min): Physical Considerations & Risks
- Sizing: "For 200 VMs at 4:1 ratio, I need approximately 200 cores = 3 hosts with 64 cores"
- Network: "I would use NSX overlay with T0 ECMP for this scale"
- Risks: "The main risk is inter-site latency — I would validate with a PoC before committing to stretched cluster"
Step 6 (5 min): Summary
- Restate the requirements and how the design addresses each one
- Acknowledge what you would need to validate further
- Open for questions
Key Takeaways
- Design document structure: Executive Summary → Requirements & Constraints (RCAR) → Conceptual → Logical → Physical → Validation → Operational Model. Each layer traces to the one above.
- VCDX defense has 3 parts: Design Presentation (75 min — your design), Design Defense (75 min — Q&A on your design), Design Scenario (50 min — new problem, design live). Total ~200 minutes.
- Defense golden rule: every decision traces to a requirement or constraint. 'Best practice' and 'we've always done it' are not valid justifications.
- Design Scenario methodology: Gather requirements (5 min) → Constraints/assumptions (5 min) → Conceptual sketch (10 min) → Logical design walkthrough (15 min) → Physical/risks (10 min) → Summary (5 min).
- Anti-patterns to avoid: over-engineering without backing requirements, dismissing alternatives without analysis, avoiding trade-off discussion, panicking on unknowns (saying 'I would validate that' is acceptable).
Exam Mapping: 3V0-12.26 — VCAP — VMware Cloud Foundation Architect
- 1.1 Differentiate between business and technical requirements
- 1.2 Differentiate between requirements, assumptions, constraints and risks
- 1.3 Differentiate between AMPRS (availability, manageability, performance, recoverability, security)
- 1.4 Develop design decisions
- 1.5 Develop a design validation strategy
- 2.1 Differentiate between VCF architecture options
- 2.2 Evaluate recommended vs supported design choices
- 2.3 Identify use case for VCF Advanced Services
- 3.1 Gather and analyze business objectives and requirements
- 3.2 Create a conceptual model
- 3.3 Create VCF logical designs
- 3.4 Create VCF physical designs
- 3.5-3.9 Design for Availability, Manageability, Performance, Recoverability, Security
- 3.10 Design for Identity Management
- 3.11 Design a consumption strategy for VCF
- 3.12 Design a monitoring strategy for VCF