VCP-VCF 9.0 Architect (2V0-13.25)
Comprehensive study guide for VCP-VCF 9.0 Architect.
Exam Blueprint Weights
Version Evolution
VCF architect skills evolved from vSphere-centric design (VCF 3.x) to multi-domain architecture (5.x) to unified platform design (9.0). Key shifts: NSX integration changed network architecture fundamentals, vSAN became the default storage model, and VCF Operations replaced multiple monitoring tools. Architect exam now covers workload domain design, management domain sizing, and cross-domain networking.
Learning Outcomes
- VCDX Insight:During design review, candidates must articulate how each design decision traces back to a specific requirement. Avoid over-engineering (e.g., building 99.999% availability when 99.9% is
- Design Pattern:For stretched cluster across two datacenters, place witness at a third location (cloud, co-location facility, even a small edge appliance). If forced to choose between sites, place witn
- DFW Rule Performance:More rules = higher CPU overhead. Consolidate rules using groups and tags. Aim for <100 rules for typical environment. Use rule hit counters to identify unused rules monthly.
- Compliance Audit Preparation:Maintain a control matrix (Requirement ID → Design Component → Evidence) updated quarterly. Assign ownership (e.g., Security team = DFW rules, Ops team = backup verificati
- VCDX Insight:A complete design decision register is non-negotiable. Panelists expect you to explain "why this shape, not that shape?" for every choice. Document assumptions (e.g., "CPU contention acce
- Obj 1.1/3.1: Gather business objectives, differentiate business vs technical requirements, create requirements traceability matrix
- Obj 2.1: Evaluate VCF architecture options — standard, consolidated, multi-site, multi-AZ — and justify selection
- Obj 3.10/3.11: Design workload migration strategy (wave planning, HCX), VCF Automation tenant model, self-service governance, and modern application consumption
VCF Design Methodology — Objectives 1.1 through 1.7#
The VCF Architect exam Section 1 tests design methodology fundamentals. Every objective here applies across all VCF design decisions.
Objective 1.1 — Differentiate Business and Technical Requirements
Business requirements come from stakeholders and define WHAT the solution must achieve: RTO/RPO targets (e.g., RTO < 1 hour, RPO < 15 minutes), SLA uptime commitments (99.9% = 8.76 hours downtime/year, 99.99% = 52.6 minutes/year, 99.999% = 5.26 minutes/year), budget constraints, compliance mandates (HIPAA, PCI-DSS, GDPR, SOC2), geographic distribution needs, growth projections (3-5 year capacity plan), and time-to-delivery expectations.
Technical requirements translate business needs into HOW: workload characterization (CPU overcommit ratios, memory reservation policies, storage IOPS profiles), VM density targets, network segmentation model, storage tier performance (vSAN ESA NVMe latency < 200 microseconds vs HDD-based OSA), encryption requirements (vSAN data-at-rest encryption, NSX in-transit encryption), identity integration (Active Directory, LDAP, SAML), and automation maturity level.
Requirements Traceability: Every technical requirement must trace to a business requirement. Example: Business says "zero data loss for financial transactions" → Technical translates to "synchronous replication with RPO=0, requiring vSAN stretched cluster with < 5ms RTT between sites." A requirements traceability matrix (RTM) maps each business requirement to one or more technical requirements, design decisions, and validation tests.
Objective 1.2 — Differentiate Conceptual, Logical, and Physical Design
Conceptual Design: High-level architecture answering "what are we building?" — defines scope, boundaries, and major components without specifying technology. Example: "Two-site active-active private cloud with self-service consumption and automated disaster recovery." Audience: business stakeholders and CIO. Contains no vendor-specific technology references.
Logical Design: Technology-specific architecture answering "how will we build it?" — defines VCF components, relationships, network topology, security zones, storage tiers, and automation workflows WITHOUT specifying exact hardware models, IP addresses, or host counts. Example: "VCF management domain with 4 ESXi hosts, vSAN ESA all-NVMe storage, NSX overlay networking with Tier-0/Tier-1 gateway hierarchy, VCF Operations for monitoring." Audience: solution architects and technical leads.
Physical Design: Implementation-ready specification answering "what exactly do we deploy?" — defines specific server models (Dell VxRail E660F), NIC configurations (dual 25GbE for management, dual 100GbE for vSAN/vMotion), IP addressing schemes, VLAN assignments, rack layouts, power/cooling calculations, and cable maps. Example: "4x Dell PowerEdge R760, each with 2x Intel Xeon 6430 (32 cores), 1TB RAM, 8x 3.84TB NVMe TLC, 2x Mellanox ConnectX-6 100GbE." Audience: implementation engineers and procurement.
Key Distinction: Logical design is technology-specific but environment-independent (portable across sites). Physical design is environment-specific (tied to a particular datacenter, rack, and IP scheme).
Objective 1.3 — Differentiate Requirements, Constraints, Assumptions, Risks (RCAR)
Requirements: What the solution MUST do. Non-negotiable functional and non-functional needs. Example: "Must support 500 VMs with 99.99% availability."
Constraints: Limitations that CANNOT be changed. External boundaries the design must work within. Example: "Maximum budget of $2M," "Cannot add a third datacenter site," "Must use existing Cisco Nexus 9000 switches," "Only 8 rack units available."
Assumptions: Beliefs about the environment that are EXPECTED to be true but not guaranteed. Must be validated during implementation. Example: "Network latency between sites is < 5ms RTT," "All ESXi hosts will have identical hardware," "10GbE network available for vMotion." CRITICAL: Every assumption is a potential risk. If an assumption proves false, the design may fail.
Risks: Potential problems that COULD impact the design. Each risk needs probability, impact, and mitigation. Example: "WAN link congestion could increase latency beyond 5ms (probability: medium, impact: high). Mitigation: dedicated MPLS circuit with QoS for vSAN traffic, monitoring with threshold alerts."
VCDX Defense: Panelists will test whether you can distinguish these four categories. A common trap: stating "we assume the budget is $2M" — budget is a CONSTRAINT (fixed), not an assumption (uncertain). Another trap: "risk is that we run out of storage" — this is too vague. State the specific condition, probability, impact, and mitigation.
Objective 1.4 — Differentiate AMPRS Design Qualities
Every design decision impacts one or more of the five design qualities. The architect must understand trade-offs between them.
Availability: Can the system survive component failures? Measured in uptime percentage. Key VCF mechanisms: vSphere HA (host failure), vSAN data redundancy (disk/host failure), NSX Edge HA (gateway failure), stretched cluster (site failure). Trade-off: higher availability = more redundant components = higher cost and complexity.
Manageability: Can the system be operated, maintained, and evolved efficiently? Key VCF mechanisms: SDDC Manager lifecycle management, VCF Operations monitoring, VCF Automation self-service, fleet management. Trade-off: more automation = less operational burden but higher initial setup complexity.
Performance: Does the system meet latency, throughput, and IOPS requirements? Key VCF mechanisms: vSAN ESA single-tier NVMe (< 200 microsecond latency), NSX distributed routing (East-West without hairpinning), DRS load balancing. Trade-off: higher performance = more expensive hardware, potentially less consolidation.
Recoverability: Can the system recover from failures and disasters? Measured in RTO/RPO. Key VCF mechanisms: vSAN native replication, SRM orchestrated failover, vSphere Replication, backup integration (Veeam, Cohesity). Trade-off: lower RTO/RPO = synchronous replication = higher bandwidth costs and latency sensitivity.
Security: Is the system protected against threats? Key VCF mechanisms: NSX DFW micro-segmentation, vSAN encryption (at-rest and in-transit), certificate management, Identity Broker SSO, CIS compliance benchmarks. Trade-off: stronger security = more operational overhead (certificate rotation, rule management, audit logging).
VCDX Principle: No design quality exists in isolation. Improving availability (stretched cluster) impacts performance (cross-site latency), cost (duplicate infrastructure), and manageability (more complex operations). The architect must articulate these trade-offs explicitly.
Objective 1.5 — Develop and Document Risk Mitigation Strategy
Risk Register Structure: For each identified risk, document: Risk ID, Description, Category (technical/operational/financial/compliance), Probability (High/Medium/Low), Impact (High/Medium/Low), Risk Score (Probability x Impact), Mitigation Strategy, Residual Risk (after mitigation), Owner, Review Date.
Mitigation Strategies: Accept (risk is low and cost of mitigation exceeds impact), Mitigate (reduce probability or impact), Transfer (insurance, SLA penalties, vendor responsibility), Avoid (change design to eliminate the risk entirely).
VCF-Specific Risk Examples:
1. Single management domain failure takes down all fleet management → Mitigate: vCenter HA, management domain on dedicated hardware, regular backups, documented manual recovery procedure. 2. vSAN rebuild storm after multiple disk failures → Mitigate: configure rebuild rate throttling, maintain 25% free capacity buffer, proactive disk health monitoring via VCF Operations. 3. NSX control plane split-brain during network partition → Mitigate: NSX Manager cluster across fault domains, anti-affinity rules, network path redundancy.
Objective 1.6 — Document Design Decisions
Every design choice must be formally documented with this structure:
- Decision ID: Unique identifier for traceability (e.g., ARCH-COMPUTE-001)
- Decision: Specific technical choice made
- Justification: Why this choice was selected (trace to requirement)
- Implication: Consequences of this choice (positive and negative)
- Alternative Considered: Other options evaluated and why rejected
- Risk: What could go wrong with this choice
- Dependency: Other decisions this depends on or enables
Example:
Decision ID: ARCH-STORAGE-003
Decision: Deploy vSAN ESA with RAID-5 erasure coding (4+1) instead of RAID-1 mirroring
Justification: Business requirement BR-007 specifies 80% storage utilization efficiency. RAID-5 provides 80% efficiency vs RAID-1 at 50%.
Implication: 4-host minimum per fault domain. Write performance ~15% lower than RAID-1. Single host failure tolerated without data loss.
Alternative: RAID-1 mirroring (rejected: only 50% efficiency, doesn't meet BR-007); RAID-6 (4+2) (rejected: requires 6 hosts minimum, exceeds budget)
Risk: RAID-5 rebuild on large NVMe drives takes longer than RAID-1 resync. If a second failure occurs during rebuild, data loss is possible. Mitigation: maintain hot spare capacity, monitor rebuild progress.
Dependency: Depends on ARCH-COMPUTE-001 (minimum 4 hosts per cluster)
Objective 1.7 — Develop a Design Validation Strategy
A design validation strategy ensures every requirement is testable and every design decision is verifiable before production deployment.
Validation Categories:
- Functional Validation: Does the design meet requirements? Test: deploy workloads matching the design spec, verify SLA compliance, measure actual RTO/RPO via failover testing.
- Performance Validation: Does the design meet performance targets? Test: run HCIBench or vdbench with production-representative I/O profiles, measure latency/IOPS/throughput under load.
- Availability Validation: Does the design survive failures? Test: simulate host failure (HA failover), disk failure (vSAN rebuild), site failure (stretched cluster switchover), NSX edge failure (gateway HA).
- Security Validation: Does the design meet security requirements? Test: run CIS benchmark compliance scan, verify micro-segmentation rules block unauthorized traffic via Traceflow, test certificate rotation procedure.
- Operational Validation: Can the design be operated? Test: perform lifecycle upgrade in lab, execute backup/restore procedure, verify monitoring alerts fire correctly.
Validation Traceability: Each validation test maps to a specific requirement and design decision. The traceability chain: Business Requirement → Technical Requirement → Design Decision → Validation Test → Pass/Fail Result.
Holodeck Lab Strategy: Use Holodeck Toolkit to build a representative VCF environment and execute all validation tests before production deployment. This is exactly what VCDX candidates should do — validate design decisions with evidence from a working lab.
Key Takeaways
- Obj 1.1: Business requirements define WHAT (SLA, budget, compliance); technical requirements define HOW (CPU, storage, network specs). Every technical requirement must trace back to a business requirement.
- Obj 1.2: Conceptual = scope/boundaries (no technology), Logical = technology-specific but environment-independent, Physical = implementation-ready with exact hardware/IP/VLAN specs.
- Obj 1.3: RCAR — Requirements are non-negotiable, Constraints are fixed boundaries, Assumptions are uncertain beliefs (each is a potential risk), Risks need probability + impact + mitigation.
- Obj 1.4: AMPRS qualities (Availability, Manageability, Performance, Recoverability, Security) are interconnected — improving one often impacts others. The architect must articulate trade-offs.
- Obj 1.5-1.7: Risk register with Accept/Mitigate/Transfer/Avoid strategies. Design decisions documented with ID/Decision/Justification/Implication/Alternative/Risk. Validation strategy maps tests to requirements.
Conceptual Model Development — Objective 3.2#
A conceptual model translates gathered business objectives into a high-level architecture vision. It is technology-agnostic and stakeholder-facing — no vendor names, no product versions, no IP addresses.
Purpose of the Conceptual Model
The conceptual model answers three questions: (1) What business capabilities does the solution provide? (2) What are the major architectural boundaries and relationships? (3) What quality attributes (AMPRS) does the design prioritize?
Building a VCF Conceptual Model — Step by Step
Step 1 — Define Business Capabilities: List the capabilities the private cloud must deliver. Example capabilities: "Self-service VM provisioning for 10 development teams," "Automated disaster recovery with < 1 hour RTO," "Centralized compliance monitoring across all workloads," "Network isolation between business units."
Step 2 — Identify Architectural Boundaries: Define the major zones/domains: Management zone (platform control plane), Workload zones (tenant compute), Edge zone (external connectivity), Operations zone (monitoring/logging). These become the logical separation points in the design.
Step 3 — Map Quality Attributes to Capabilities: For each capability, identify which AMPRS quality is primary. Self-service provisioning → Manageability. Disaster recovery → Recoverability. Compliance monitoring → Security. Network isolation → Security + Availability.
Step 4 — Define External Interfaces: Identify integration points: corporate Active Directory, external DNS, NTP, syslog/SIEM, backup infrastructure, WAN/internet connectivity, cloud provider interconnects. These become the physical design boundary conditions.
Step 5 — Document Constraints and Assumptions: Capture all constraints (budget, site limitations, existing hardware reuse) and assumptions (network bandwidth, vendor support availability, team skill level) that will shape the logical and physical designs.
Conceptual Model Artifacts
The conceptual model produces three artifacts:
- Conceptual Architecture Diagram: Boxes-and-lines showing major zones, relationships, and external interfaces. No technology names — just functional labels ("Compute Platform," "Network Fabric," "Monitoring System").
- Quality Attribute Priority Matrix: Ranked list of AMPRS qualities with justification. Example: "Security > Availability > Performance > Manageability > Recoverability" with rationale for the ordering.
- Constraints and Assumptions Register: Formal list of all constraints and assumptions with ownership and validation plan.
VCDX Defense Application: The conceptual model is your first slide in a VCDX defense. Panelists want to see that you understood the business problem BEFORE jumping to technology. A common failure: presenting a physical design (server models, IP addresses) without first establishing WHY those choices were made. The conceptual model provides the WHY.
Key Takeaways
- Obj 3.2: Conceptual model is technology-agnostic — no vendor names, no IP addresses, no product versions. It maps business capabilities to architectural boundaries.
- Three artifacts: Conceptual Architecture Diagram (zones/relationships), Quality Attribute Priority Matrix (AMPRS ranking), Constraints/Assumptions Register.
- VCDX Defense: The conceptual model is your opening argument — it establishes WHY before WHAT. Panelists will ask 'what business problem does this solve?' before examining any technical detail.
VCF Logical Design Decisions — Objective 3.3 (All Sub-Objectives)#
Logical design translates the conceptual model into technology-specific architecture. It defines WHAT VCF components are used and HOW they relate — without specifying exact hardware, IP addresses, or host counts. The 2V0-13.25 blueprint tests logical design for 8 VCF areas.
Obj 3.3.1 — Logical Design: VCF Prerequisites
DNS Architecture: Forward and reverse DNS for all VCF management components (vCenter, NSX Manager, SDDC Manager, ESXi hosts). Dedicated DNS zone (e.g., vcf.lab.local) recommended. DNS must be available BEFORE VCF bring-up — circular dependency if DNS runs on VCF itself.
NTP Architecture: All VCF components must synchronize to the same NTP source. Maximum acceptable drift: 5 seconds (vSAN CMMDS requirement). Stratum-2 or better NTP servers, minimum 2 for redundancy.
Network Prerequisites: Management network (routable, all components), vMotion network (dedicated, jumbo frames optional), vSAN network (dedicated, MTU 9000 mandatory), NSX overlay network (TEP, MTU 9000 mandatory for GENEVE encapsulation), uplink networks for north-south traffic.
Certificate Authority: VCF supports Microsoft CA, OpenSSL CA, or VMCA as intermediate CA. Decision: use enterprise CA for production (compliance), VMCA for lab/dev. Certificate lifecycle: auto-renewal for internal certs, manual process for external-facing certs.
Active Directory/LDAP: Identity source for SSO. VCF Identity Broker provides single sign-on across vCenter, NSX, SDDC Manager. Decision: dedicated service account per VCF component vs shared service account (dedicated is more secure, shared is simpler).
Obj 3.3.2 — Logical Design: Fleet Topologies
Single-Instance Fleet: One SDDC Manager managing one management domain and multiple workload domains. Suitable for: single datacenter, < 1000 VMs, single operations team. Simplest to deploy and manage.
Multi-Instance Fleet: Multiple SDDC Manager instances, each managing its own VCF instance. Suitable for: multi-region deployments, organizational isolation requirements, separate management domains per business unit. VCF Fleet Management provides single-pane visibility across instances.
Consolidated Architecture: Management and workload VMs share the same cluster (minimum 4 hosts). Suitable for: small deployments, edge/ROBO locations, cost-constrained environments. Trade-off: management workload contention possible, but reduced hardware footprint.
Standard Architecture: Dedicated management domain cluster (4 hosts) separate from workload domain clusters. Suitable for: production enterprise deployments, compliance requirements demanding management isolation. Trade-off: more hardware required, but clean separation of concerns.
Multi-AZ (Availability Zone): Stretched clusters spanning two sites with witness at a third site. Each site is an AZ. Suitable for: active-active datacenter designs, metro-distance deployments (< 5ms RTT (no geographic distance restriction — only latency matters)). Trade-off: 2x infrastructure cost, complex networking, WAN dependency.
Obj 3.3.3 — Logical Design: Network Infrastructure
Underlay Network: Physical switching fabric that carries all VCF traffic types. Design decisions: L2 vs L3 spine-leaf topology (L3 preferred for scale), VLAN allocation strategy (separate VLANs per traffic type), MTU 9000 for vSAN and NSX overlay, redundant uplinks with LACP or active-standby.
Overlay Network: NSX GENEVE-encapsulated virtual networks running on top of the physical underlay. Design decisions: transport zone scope (per-cluster vs shared across clusters), TEP IP pool sizing, host TEP vs edge TEP separation.
Traffic Flow Design: Define North-South (external to/from VCF) and East-West (VM-to-VM within VCF) traffic patterns. North-South flows through NSX Tier-0 gateway → physical router. East-West stays on NSX distributed router (DFW enforces segmentation). Design decision: centralized vs distributed services (NAT, load balancing, firewall).
Obj 3.3.4 — Logical Design: Management Domain
Management domain hosts the VCF control plane: vCenter Server, NSX Manager cluster (3 nodes), SDDC Manager, VCF Operations, VCF Automation (optional). Logical design decisions:
Cluster sizing: Minimum 4 hosts for vSAN (FTT=1 with RAID-5). Recommended: 4 hosts for small (< 500 workload VMs), 6-8 hosts for large (> 1000 workload VMs) to handle management workload growth.
Resource reservation: Management VMs should have CPU and memory reservations to prevent workload contention from affecting control plane. vCenter: 4 vCPU, 24GB reserved. NSX Manager: 6 vCPU, 24GB reserved per node. SDDC Manager: 4 vCPU, 16GB reserved.
Anti-affinity: NSX Manager nodes must run on separate hosts (anti-affinity rules). vCenter HA nodes (active/passive/witness) must run on separate hosts.
Storage policy: Management domain vSAN policy: FTT=1, RAID-5 erasure coding for space efficiency (management VMs are not IOPS-intensive). Alternatively FTT=1, RAID-1 for simpler operations.
Obj 3.3.5 — Logical Design: Workload Domain
Workload domains host tenant/application VMs. Each workload domain has its own vCenter instance (since VCF 5.0+) and dedicated clusters. Logical design decisions:
Domain boundary strategy: One workload domain per business unit (strong isolation), one per environment (dev/staging/prod), or one per application tier (web/app/DB). Decision depends on: security isolation requirements, lifecycle independence, operational team structure.
Cluster composition: All hosts in a cluster must be identical (CPU, memory, storage, NIC). Mixed clusters cause DRS imbalance and vSAN degradation. Design decision: cluster size (minimum 4 for vSAN, recommended 8-16 for production).
Storage tier: vSAN ESA (NVMe-only, single-tier, higher performance) vs vSAN OSA (HDD + SSD cache tier, lower cost). Decision based on workload IOPS requirements and budget. Can mix storage tiers across different workload domains.
Network isolation: Each workload domain can have independent NSX transport zones for network isolation, or share transport zones for cross-domain communication. Decision depends on security requirements.
Obj 3.3.6 — Logical Design: VCF Networking (NSX)
NSX Manager cluster: 3-node cluster for production (quorum-based). Shared across management and workload domains (single NSX Manager instance manages all domains in a VCF instance). Decision: single vs multi-site NSX federation (multi-site for cross-datacenter consistent policy).
Tier-0 Gateway: Provides north-south connectivity. Design decisions: active-active (ECMP) for high throughput or active-standby for simplicity. BGP peering with physical routers. Route redistribution strategy (which routes advertise to physical network).
Tier-1 Gateways: Provides routing for tenant segments. Design decisions: one Tier-1 per workload domain (simple), one per application (granular control), or one per tenant (multi-tenancy). Tier-1 connects upstream to Tier-0 for external access.
DFW Micro-segmentation: Logical security policy model. Design decisions: rule category strategy (Emergency/Infrastructure/Environment/Application/Default), Applied To scope (limit rule processing per VM), tag-based grouping (application-centric) vs IP-based grouping (network-centric).
Edge Cluster: Hosts Tier-0 and Tier-1 service routers. Design decisions: dedicated edge cluster (recommended) vs shared with compute. Edge VM sizing: Large (8 vCPU, 32GB) for production. Minimum 2 Edge VMs for HA.
Obj 3.3.7 — Logical Design: VCF Automation
Deployment model: Single-node (< 500 managed VMs, dev/lab), clustered 3-node (production, HA). Decision based on scale and availability requirements.
Cloud account architecture: One vCenter cloud account per VCF instance. One NSX cloud account associated with the vCenter account. Cloud accounts define the infrastructure VCF Automation can provision to.
Cloud zone strategy: One cloud zone per workload domain cluster (maps compute boundaries). Capability tags on cloud zones enable constraint-based placement. Decision: fine-grained zones (per cluster) vs coarse zones (per domain).
Project and tenancy model: One project per team/business unit. Projects define resource boundaries, naming conventions, lease policies, and content sharing. Multi-organizational tenancy for service provider scenarios (separate organizations with independent identity and governance).
Content strategy: Shared template library (organization-level) for standard VM patterns. Team-specific templates for specialized workloads. Template versioning with release workflow. Service Broker catalog as the self-service storefront.
Obj 3.3.8 — Logical Design: VCF Operations
Deployment model: Single-node (< 2000 VMs, dev/lab), 2-node HA (production, < 10000 VMs), multi-node cluster (> 10000 VMs). Remote collector nodes for distributed VCF instances.
Adapter architecture: vCenter adapter (compute/storage metrics), NSX adapter (network metrics), vSAN adapter (storage health), SDDC Manager adapter (lifecycle status). Each adapter provides specific metric domains.
Dashboard and alerting strategy: Role-based dashboards — executive summary, ops team operational view, security team compliance view. Alert definitions with notification channels (email, webhook, SNMP). Escalation hierarchy: warning → critical → emergency.
Content pack strategy: CIS compliance benchmarks, VMware hardening guide, custom VCF operational benchmarks. Decision: which compliance frameworks to enable based on regulatory requirements.
Integration with VCF Automation: VCF Operations can trigger Automation workflows based on alert conditions (e.g., auto-scale workload domain when capacity threshold reached). This creates a closed-loop operations model.
Key Takeaways
- Obj 3.3: Logical design is technology-specific (VCF, NSX, vSAN) but environment-independent (no IP addresses, no specific hardware models). It defines component relationships and architectural patterns.
- Fleet topology decision: Single-instance (simple) vs Multi-instance (isolation) vs Consolidated (cost) vs Standard (separation) vs Multi-AZ (availability). Each has clear trade-offs in cost, complexity, and availability.
- Management domain logical design: 4-8 hosts, resource reservations for control plane VMs, anti-affinity rules for NSX Manager and vCenter HA, vSAN RAID-5 for space efficiency.
- Workload domain boundary strategy: per business unit (isolation), per environment (lifecycle), or per application tier (performance). Decision depends on security, operational, and compliance requirements.
- NSX logical design: Tier-0 (north-south, BGP peering) → Tier-1 (per-tenant routing) → Segments (VM connectivity). DFW provides micro-segmentation with 5 rule categories.
VCF Physical Design Decisions — Objective 3.4 (All Sub-Objectives)#
Physical design translates logical architecture into implementation-ready specifications. It defines EXACT hardware models, IP addresses, VLAN IDs, cable maps, and rack layouts. The 2V0-13.25 blueprint tests physical design for 8 VCF areas.
Obj 3.4.1 — Physical Design: VCF Prerequisites
DNS Physical Design: Specific DNS server IPs (e.g., 10.0.0.11, 10.0.0.12), zone configuration files, forward/reverse zone entries for every VCF component FQDN. DNS must be deployed and validated BEFORE VCF bring-up.
NTP Physical Design: NTP server IPs, stratum level, chrony vs ntpd configuration on ESXi hosts. Example: ntp1.corp.local (10.0.0.21, stratum-2), ntp2.corp.local (10.0.0.22, stratum-2).
Network Physical Design: Switch model and firmware (e.g., Cisco Nexus 93180YC-FX3S running NX-OS 10.3), port allocation map, VLAN table:
- VLAN 10: Management (10.10.10.0/24, MTU 1500)
- VLAN 20: vMotion (10.10.20.0/24, MTU 9000)
- VLAN 30: vSAN (10.10.30.0/24, MTU 9000)
- VLAN 40: NSX TEP (10.10.40.0/24, MTU 9000)
- VLAN 50: Uplink/External (public or routable subnet)
Cabling: Each host needs minimum 4 physical NICs — 2x 25GbE for management/vMotion (active-standby or LACP), 2x 25GbE or 100GbE for vSAN/NSX overlay (LACP recommended for throughput). ToR switch redundancy: dual switches per rack.
Obj 3.4.2 — Physical Design: Fleet Topology
For a standard architecture single-instance fleet:
- Management Domain: 1 rack, 4 hosts, 2 ToR switches
- Workload Domain 1: 1-2 racks, 8-16 hosts, 2 ToR switches
- Workload Domain 2: 1-2 racks, 8-16 hosts, 2 ToR switches
- Edge rack: 2 bare-metal Edge nodes (or Edge VMs on compute hosts)
Rack layout considerations: Power distribution (A+B feeds per rack), cooling (front-to-back airflow), cable management (structured cabling with patch panels), physical security (locked racks, access logging).
For multi-AZ stretched cluster:
- Site A: 2 hosts (management) + 4 hosts (workload) + 1 Edge
- Site B: 2 hosts (management) + 4 hosts (workload) + 1 Edge
- Witness Site: 1 lightweight host or VM for vSAN witness + ESXi witness
- WAN: Dedicated dark fiber or MPLS circuit, < 5ms RTT, minimum 10Gbps for vSAN replication
Obj 3.4.3 — Physical Design: Network Infrastructure
Spine-Leaf Physical Topology (recommended for > 2 racks):
- 2 Spine switches (e.g., Cisco Nexus 9336C-FX2) providing inter-rack L3 routing
- 2 Leaf switches per rack (e.g., Cisco Nexus 93180YC-FX3S) providing host connectivity
- BGP EVPN/VXLAN fabric for scalable L2/L3 overlay (or simple OSPF for smaller deployments)
Port Allocation Per Host (4-NIC design):
- vmnic0 → Leaf-A port (management/vMotion uplink-1) - vmnic1 → Leaf-B port (management/vMotion uplink-2) - vmnic2 → Leaf-A port (vSAN/overlay uplink-1) - vmnic3 → Leaf-B port (vSAN/overlay uplink-2)
VDS Configuration:
- Management VDS: vmnic0 + vmnic1, active-standby teaming, VLAN 10 (management) + VLAN 20 (vMotion)
- Overlay VDS: vmnic2 + vmnic3, LACP teaming, VLAN 30 (vSAN) + VLAN 40 (NSX TEP)
Physical router integration: BGP peering between NSX Tier-0 Edge and physical router. Advertised routes: workload subnets. Received routes: default route or specific external routes. BFD for fast failure detection (< 1 second).
Obj 3.4.4 — Physical Design: Management Domain
Host Specification (example for 500-VM environment):
- 4x Dell PowerEdge R760 (or HPE ProLiant DL380 Gen11, Lenovo ThinkSystem SR650 V3)
- CPU: 2x Intel Xeon Gold 6430 (32 cores each, 64 cores total per host) — meets VCF 9.0 minimum 16 cores per socket
- RAM: 1024 GB DDR5-4800 (management VMs need ~200GB, remainder for growth)
- Storage: 8x 3.84TB NVMe TLC SSD (vSAN ESA, ~30TB raw per host, ~92TB raw total)
- NIC: 2x Mellanox ConnectX-6 Dx dual-port 25GbE (4 ports total per host)
- Boot: 2x M.2 SATA SSD in RAID-1 (ESXi boot device)
IP Addressing (management domain):
- vCenter: 10.10.10.10
- SDDC Manager: 10.10.10.11
- NSX Manager 1/2/3: 10.10.10.12-14, VIP: 10.10.10.15
- ESXi hosts: 10.10.10.101-104
- vSAN VMkernel: 10.10.30.101-104
- vMotion VMkernel: 10.10.20.101-104
- TEP Pool: 10.10.40.101-116 (4 TEPs per host x 4 hosts)
VCF 9.0 Licensing: Per-core licensing model. 4 hosts x 64 cores = 256 cores. Minimum purchase: 16 cores per socket = 128 cores minimum. Actual: 256 VCF cores licensed.
Obj 3.4.5 — Physical Design: Workload Domain
Host Specification (example for performance tier):
- 8x Dell PowerEdge R760 (identical to management domain hosts for operational simplicity, or different spec for cost optimization)
- CPU: 2x Intel Xeon Gold 6438N (32 cores, optimized for VMs)
- RAM: 2048 GB DDR5 (high density for VM consolidation)
- Storage: 12x 7.68TB NVMe TLC SSD (vSAN ESA, ~92TB raw per host)
- NIC: 2x Mellanox ConnectX-6 Dx dual-port 100GbE (for high-bandwidth workloads)
Capacity Planning:
- Usable vSAN capacity: 8 hosts x 92TB raw x 0.75 (RAID-5 overhead) x 0.70 (slack space) = ~386TB usable
- Compute capacity: 8 hosts x 64 cores x 4:1 overcommit = 2048 vCPUs available
- Memory capacity: 8 hosts x 2048GB x 0.85 (reservation buffer) = ~13.9TB available
Obj 3.4.6 — Physical Design: VCF Networking (NSX)
Edge Node Physical Design:
- 2x Edge VMs (Large: 8 vCPU, 32GB RAM, 200GB disk) on management domain hosts
- Or 2x bare-metal Edge nodes (dedicated hardware for high-throughput north-south traffic)
- Edge uplink: dedicated VLAN for BGP peering with physical router
NSX TEP IP Pool Sizing:
- Management domain: 4 hosts x 4 TEPs = 16 IPs + 4 spare = /27 subnet (10.10.40.0/27)
- Workload domain 1: 8 hosts x 4 TEPs = 32 IPs + 8 spare = /26 subnet (10.10.40.64/26)
- Edge TEPs: 4 TEPs per edge x 2 edges = 8 IPs + 4 spare = /28 (10.10.40.128/28)
Physical Firewall Integration: North-south traffic flow: VM → DFW → NSX Tier-1 → Tier-0 → physical firewall (Palo Alto/Fortinet) → internet. Decision: insert physical firewall inline (transparent bridge mode) or as next-hop (routed mode).
Obj 3.4.7 — Physical Design: VCF Automation
Appliance Sizing:
- Single-node: 1 VM — 12 vCPU, 48GB RAM, 200GB disk (< 500 managed VMs)
- Clustered: 3 VMs — each 12 vCPU, 48GB RAM, 200GB disk (production, HA)
- External PostgreSQL: separate DB VM for scale (> 5000 managed VMs)
Network Requirements: Management network access to vCenter and NSX Manager APIs. If multi-site: dedicated WAN link between Automation instances for content synchronization.
Integration Points:
- vCenter cloud account: API endpoint, service account credentials
- NSX cloud account: API endpoint, associated with vCenter
- Active Directory: identity source for project membership
- IPAM (optional): Infoblox or Bluecat for enterprise IP management
- CMDB (optional): ServiceNow for asset registration via ABX extensibility
Obj 3.4.8 — Physical Design: VCF Operations
Appliance Sizing:
- Small (< 2000 VMs): 1 node — 8 vCPU, 32GB RAM, 500GB disk
- Medium (< 8000 VMs): 2 nodes (HA) — each 16 vCPU, 48GB RAM, 1TB disk
- Large (< 25000 VMs): 4+ nodes (cluster) — each 16 vCPU, 48GB RAM, 2TB disk
- Remote Collector: 2 vCPU, 4GB RAM per remote VCF instance
Storage Retention: Metric data retention affects disk sizing. Default: 6 months real-time, 3 years rolled-up. For compliance: extend to 5-7 years (increases disk requirements significantly).
Network Requirements: Collector-to-adapter traffic: port 443 (HTTPS) to vCenter, NSX Manager, SDDC Manager. Outbound: port 443 to Broadcom for content pack downloads (or offline import for air-gapped environments).
Integration with Logging: VCF Operations for Logs (formerly vRealize Log Insight): separate appliance for syslog aggregation. Sizing: 1 node per 5000 events/second. Integration: log events correlate with metric alerts for root cause analysis.
Key Takeaways
- Obj 3.4: Physical design is implementation-ready — specific server models, CPU/RAM/storage specs, IP addresses, VLAN IDs, port assignments, rack layouts, and cable maps.
- Management domain physical: 4 hosts minimum, 64+ cores per host (VCF 9.0 per-core licensing), 1TB+ RAM, 8+ NVMe drives per host for vSAN ESA, 4 NICs per host (2 pairs for redundancy).
- Network physical: Spine-leaf topology for > 2 racks. Per-host 4-NIC design: 2 for management/vMotion (active-standby), 2 for vSAN/overlay (LACP). MTU 9000 for vSAN and NSX TEP VLANs.
- Capacity planning formula: Usable vSAN = hosts x raw x RAID-overhead x slack-factor. Compute = hosts x cores x overcommit-ratio. Memory = hosts x RAM x reservation-buffer.
- NSX TEP pool sizing: 4 TEPs per host. Size subnet to 1.5x current host count for growth. Edge TEPs on separate subnet from host TEPs.
Compute Design#
ESXi Host Sizing & Capacity Planning
Host sizing begins with workload characterization, not with purchasing decisions. Understand your VM mix before selecting hardware.
CPU Sizing Methodology:
vCPU allocation:
Sum all VM vCPU requirements + 10-15% headroom for management services + vSAN operations
Overcommit ratios:
CPU is highly virtualized; typical ratios 4:1 (bursty workloads) to 8:1 (stable baseline workloads). Advanced users understand the difference between entitlement and allocation.
pCPU selection:
Calculate required pCPUs = total vCPU demand / overcommit ratio. Example: 1000 vCPUs with 4:1 ratio = 250 pCPUs = 8 x 32-core hosts (with overhead).
Memory overcommit:
Unlike CPU, memory is not aggressively overcommitted. Guideline: 95% of host RAM allocated to VMs, with remainder reserved for hypervisor, management, ballooning overhead. vSAN dedups/compresses reduce effective memory requirements.
vSAN-specific overhead:
Reserve 32GB minimum for vSAN control plane, 1GB per disk in disk groups for iSCSI stack, additional memory for dedup/compression data structures.
Sizing Example:
Workload: 800 vCPUs, 3200GB VM memory, 4:1 CPU overcommit
Host Spec: 32 cores (64 logical threads), 512GB RAM
CPU Calculation:
Required pCPUs = 800 / 4 = 200 pCPUs
Threads per host = 64
Hosts needed = 200 / 64 = 3.125 → 4 hosts (with 12.5% spare)
Memory Calculation:
VM Memory = 3200GB
Hypervisor overhead per host ≈ 12GB (kernel, drivers, management)
vSAN overhead ≈ 16GB (disk group metadata, dedup/compression)
Per-host allocation target = 512 - 12 - 16 = 484GB for VMs
Memory needed = 3200 / 484 = 6.6 → 7 hosts required
Limiting factor: Memory (7 hosts), not CPU (4 hosts)
Final design: 7-host cluster with capacity headroom
Cluster Design & Scalability
vSphere Cluster Maximums (vSAN environment):
Maximum hosts per cluster: 64 (vSphere 7.0+) with all-flash vSAN
- Practical maximum for vSAN stretched cluster: 16 hosts per site (32 total)
- Minimum hosts: 3 (2-node clusters not supported for availability)
- Odd host numbers preferred for quorum (3, 5, 7, etc.)
DRS (Distributed Resource Scheduler) Configuration:
Automation level:
Manual (explicit approval required), Partially Automated (only correct imbalances), or Fully Automated (immediate migrations)
Affinity rules:
VM-to-VM affinity (keep together for low latency), VM-to-host affinity (pin to specific hardware)
Anti-affinity rules:
Separate VMs to different hosts (high availability pattern), separate VMs from host groups (licensing constraints)
VM groups:
Label VMs logically (e.g., "database tier", "web tier") for policy application
Host groups:
Designate hardware groups (e.g., GPU hosts, high-memory hosts) and apply rules to place specific workloads there
Cluster Architecture Patterns:
General-purpose cluster:
Homogeneous hardware, mixed workload types, standard sizing
Specialized clusters:
Separate clusters for CPU-intensive, memory-intensive, GPU-required, or I/O-bound workloads with right-sized hardware
Multi-tier architecture:
Separate clusters for different application tiers (presentation, application, database) enabling tier-specific policies
HA Design & Admission Control
vSphere HA protects against host failures through VM restart mechanisms. Design HA capacity carefully to ensure failover capability.
HA Admission Control Policies:
Slot-based policy:
Calculates a "slot" size based on largest VM CPU/memory requirements. Cluster can host N-1 failures if spare capacity ≥ 1 slot. Sensitive to large VM outliers.
Percentage-based policy:
Reserve percentage of cluster resources for failover (e.g., 25% means cluster can handle 1 failure in 4-host cluster). Simpler but less precise.
Dedicated failover hosts policy:
Explicitly reserve one host's capacity for failovers. Predictable but ties up resources; suitable for small clusters.
Disabled admission control:
Not recommended; allows VM placement that violates HA guarantees.
HA Pit:
If admission control is set to percentage-based (25%) on a 4-host cluster, only 3 hosts worth of workload can be running. The 4th host stays empty. Design cluster size to balance capacity utilization and failover headroom.
Enhanced vMotion Compatibility (EVC)
EVC allows mixed-generation CPU hardware in a single cluster by masking advanced CPU features. Critical for multi-site designs or phased upgrades.
EVC Mode Selection:
Match the lowest-generation CPU in the cluster (e.g., if mixing Intel Skylake and Cascade Lake, use Skylake EVC mode)
VM Compatibility:
VMs booted in lower EVC mode retain that capability; cannot hot-migrate to higher EVC mode without power cycle
Performance implication:
VMs lose access to newer CPU features (e.g., AVX2, TSX); measure impact for floating-point workloads
Multi-site scenario:
If replicating VMs between vSAN stretched cluster sites with different CPU generations, EVC must match across sites
Key Takeaways
- Compute density determines cluster performance: CPU overcommit ratio (typically 4:1) affects vMotion feasibility and CPU ready time.
- Host affinity and anti-affinity rules enforce compliance but must be designed for resiliency—avoid creating single points of failure.
- NUMA topology impacts performance: VMs with large vCPU counts (>16) require NUMA awareness for optimal memory latency.
Storage Design#
vSAN Architecture & Capacity Planning
vSAN is a hyperconverged storage platform that pools local storage from ESXi hosts. Architecture choices directly impact performance, capacity, and operational complexity.
- vSAN Object Storage Architecture (OSA)
- the traditional vSAN model:
Disk groups:
Logical grouping of 1 cache tier + 1-7 capacity tiers per host. Host can have multiple disk groups for scaling.
Cache (read/write buffer):
All-flash tier (NVMe or SSD) holds hot data. Dual-controller RAID 1 by default; RAID 5/6 not recommended for cache.
Capacity tier:
HDD or SSD backing store. Single or multiple devices per disk group.
Cache-to-capacity ratio:
Industry guideline 1:10 to 1:30. Example: 400GB NVMe to 4-12TB HDD. If ratio is too high, cache fills unnecessarily; too low, high misses and destaging overhead.
- vSAN Enterprise Storage Architecture (ESA)
- introduced in vSAN 8.x+:
Distributed pool architecture:
No separate cache/capacity tiers. All-flash (NVMe) pool, with auto-tiering based on access patterns.
Advantages:
Higher performance (no destaging bottleneck), simplified capacity planning, better dedup/compression efficiency, supports encryption natively.
Minimum requirement:
All-NVMe TLC drives (no HDDs in ESA). Minimum 1 NVMe TLC device per storage pool (official requirement). ReadyNode configurations typically require 4+ drives for production performance. Requires 128 GB minimum host memory.
When to choose ESA:
Performance-critical workloads, cloud-native (high churn), capacity-constrained environments where dedup/compression matter.
Fault Tolerance & Capacity Math
Failures to Tolerate (FTT):
vSAN replicates objects to survive host/disk failures. FTT=1 (mirrored) survives 1 failure; FTT=2 (striped) survives 2 failures.
Capacity Calculation with FTT:
Object Size = 100GB (single VM disk)
FTT=1 (Mirroring - RAID 1):
Replicas = 2 (primary + 1 copy)
Effective capacity consumed = 100GB × 2 = 200GB
Usable capacity = 50% of raw
FTT=2 with Erasure Coding (RAID 6):
- Primary object = 100GB, split into 4 stripes (25GB each) + 2 parity stripes (25GB each)
- Total 6 stripes across 6 hosts = 150GB consumed
- Usable capacity = 66.7% of raw (RAID 6 = n/(n+2) for n data stripes)
Cluster 10 hosts × 1.6TB = 16TB raw capacity
With FTT=2 RAID 6: 16TB × 0.667 = 10.67TB usable
With FTT=1 mirroring: 16TB × 0.5 = 8TB usable
Difference: 2.67TB (25% more capacity with erasure coding)Erasure Coding vs Mirroring Trade-offs:
Mirroring (FTT=1):
Simple, low latency, 50% usable capacity, faster rebuild, better for small clusters (<6 hosts)
Erasure Coding (FTT=2):Complex, higher latency for random writes, ~67% usable capacity, slower rebuild, better for large clusters (>6 hosts), requires even object placement across hosts
Hybrid approach:
Use mirroring for critical VMs, erasure coding for bulk storage (e.g., archives, backups)
vSAN Stretched Cluster Design
A stretched cluster spans two sites with a third witness site for quorum. Enables active-active replication with sub-second RPO.
Architecture Requirements:
Network latency:
<5ms RTT between primary sites (stretched cluster); <200ms to witness (asynch commit possible). Measure with
ping -c 100 remote-site
and examine standard deviation; high variance causes problems.
Host count:
For standard stretched clusters: minimum 3 hosts per site (6 total) for quorum. 2-node stretched clusters (1 host per site + witness) are also supported for smaller deployments. Witness is deployed as a dedicated ESXi appliance (OVA). Sizing tiers: Tiny (≤750 components/≤10 VMs), Medium (≤21,833 components/≤500 VMs), Large (≤45,000 components), Extra Large (≤64,000 components). vSAN ESA does not support Tiny witness.
Partition handling:
If sites partition, the site with majority of hosts continues; minority partition stops VM execution to prevent split-brain.
Witness placement:
MUST be at third location. If witness is co-located with one site, it's not truly a tiebreaker (dependent on that site's availability).
Witness VM Sizing:
Deployed using the Tiny or Medium OVA option depending on cluster size. Can be hosted on external vCenter or local cluster. Must use a different datastore than the vSAN stretched cluster datastore.
External Storage Integration
VCF supports FC, iSCSI, and NFS for non-vSAN deployments or hybrid configurations.
FC (Fibre Channel):
High performance (10/16/32Gbps), stable, best for database tier. Requires dedicated fabric.
iSCSI:
IP-based, easier to deploy (standard Ethernet), sufficient for most workloads. Software iSCSI on ESXi handles well; hardware iSCSI NICs reduce CPU overhead.
NFS v4.1:
Modern, supports Kerberos authentication, good for file-based workloads (VDI, file services).
Multi-site consideration:
FC requires stretching fabric (expensive); iSCSI/NFS stretch easily over IP WAN.
Datastore Sizing & Storage Policy Alignment
Datastores must map to SLA requirements through storage policies (vSAN policy objects or external storage class definitions).
Storage Tier to Policy Mapping Example:
Tier 1 - Premium (Databases, Real-time):
vSAN FTT=2, RAID 6, no dedup/compression, all-NVMe pool (ESA)
Policy: NumberOfFailuresToTolerate=2, CacheReservation=25%
Tier 2 - Standard (Applications, Mid-priority):
vSAN FTT=1, mirroring, enable dedup/compression
Policy: NumberOfFailuresToTolerate=1, CacheReservation=0%
Tier 3 - Archive (Backups, Logs):
External NFS, RAID 6, dedup on storage array
Policy: external array policy, lower priority QoS
SLA Mapping:
Database: 99.99% uptime → Premium t
Key Takeaways
- Design Pattern:For stretched cluster across two datacenters, place witness at a third location (cloud, co-location facility, even a small edge appliance). If forced to choose between sites, place witness on lower-cost infrastructure.
Network Design#
NSX Design Fundamentals
NSX (Network Security) abstracts the physical network through virtualization, enabling micro-segmentation and zero-trust security at scale.
NSX Architecture Layers:
Management Plane:
NSX Manager cluster (3-node), handles policy definition, audit logging, REST API
Control Plane:
Central Control Cluster (CCP), distributed edge agents, maintains BGP neighbor state, manages segment membership
Data Plane:
Edge Transport Nodes (Edge VMs or hardware), Kernel Modules on ESXi hosts (KNI), forward traffic per policy
Tier-0 & Tier-1 Gateway Design
Tier-0 Gateway:
Top-level logical router, performs North-South routing (external connectivity). Typically placed at network edge.
HA Mode - Active-Standby:
One edge processes traffic, backup takes over on failure. Simple, predictable traffic path, lower CPU on backup node.
HA Mode - Active-Active:
Both edges process traffic simultaneously using ECMP (Equal Cost Multi-Path), requires stateful traffic pinning for flows. Higher throughput (scales linearly with edge count), complex BGP design.
Edge cluster sizing:
Minimum 2 edges (HA pair), typically 2-4 for medium deployments, 4-8 for large. Each edge = dedicated VM (8vCPU, 16GB) or hardware appliance.
Tier-1 Gateway:
Logical router below Tier-0, performs East-West routing within segments. One Tier-1 per tenant/workload domain is common. Connects to Tier-0 for external access.
Placement:
Distributed (running on each ESXi host) or centralized (on edge nodes). Distributed scales better, centralized is simpler to troubleshoot.
Failover:
Tier-1 typically advertises routes to Tier-0; if one edge fails, with BFD enabled (default in NSX), failover detection is sub-second (~1.5s with default 500ms interval × 3 multiplier); without BFD, BGP hold timer default is 180 seconds.
Physical Network Design Integration
NSX requires a stable underlying physical network. Design physical topology to support NSX overlay.
Physical Network Topology:
Spine-Leaf (CLOS):
Recommended for modern data centers. Spine switches connect all leaf switches (no leaf-to-leaf links). Any host can reach any other in 2 hops. Supports ECMP load balancing.
3-Tier (Traditional):
Access layer (ToR), aggregation layer, core. More expensive, more oversubscription if not carefully designed.
Underlay Network Requirements for NSX:
MTU:
Minimum 1600 (overlay adds encapsulation headers: Geneve encapsulation adds ~54 bytes (outer Ethernet + outer IP + UDP + Geneve header); 1700 MTU recommended for future header expansion). Set physical NICs to 9000 (jumbo frames) to handle overlays safely.
Bandwidth planning:
East-West overlay traffic adds ~10-15% overhead. Dimension uplinks to carry combined North-South + East-West.
VLAN design:
Separate VLANs for management, vMotion, vSAN, NSX controllers, NSX edge uplinks. Avoid trunking all VLANs to all hosts (security, VLAN scalability).
QoS:
Tag NSX-managed traffic (DSCP or 802.1p) to ensure control plane (low latency) and data plane (high throughput) are not starved.
NSX Segmentation Design
Segments are logical Layer 2 domains, replacing VLANs. Design segments around security zones and application tiers.
Segment Design Example:
Application Stack Segmentation:
Segment: Production-Web
- Connected to: Tier-1-Production
- Addresses: 10.10.1.0/24
- QoS: High priority (web-facing)
Segment: Production-App
- Connected to: Tier-1-Production
- Addresses: 10.10.2.0/24
- QoS: Medium priority
Segment: Production-Database
- Connected to: Tier-1-Production
- Addresses: 10.10.3.0/24
- QoS: High priority (latency-sensitive)
Multi-Tenant Segmentation:
Tenant-A-Segment (isolated from Tenant-B-Segment via Tier-1 policies)
Uses same IP schema (10.10.0.0/16 internally, isolated at DFW)
Distributed Firewall (DFW) Design
DFW enforces micro-segmentation at the kernel level on each host. Design firewall rules to achieve zero-trust security.
Rule organization:
Rules are processed top-to-bottom; first match wins. Organize by application tier (web → app → database).
Rule type:
Allow-listing (deny-by-default, explicit allows) is more secure than block-listing.
Rule scope:
Can target VMs, segments, security groups, tags, or combinations.
Stateful rules:
DFW maintains connection state; reverse traffic automatically allowed if forward allowed.
Zero-Trust DFW Policy Example:
Rule Set: Production-Zero-Trust
- Rule 1: Allow inbound HTTPS to web servers
- Source: External (any)
- Destination: Segment:Production-Web, VM Tag:web-server
- Service: TCP/443
- Direction: Inbound
Rule 2: Allow web→app tier
Source: Segment:Production-Web, VM Tag:web-server
- Destination: Segment:Production-App, VM Tag:app-server
- Service: TCP/8080
- Direction: Inbound (app-server-side)
Rule 3: Allow app→database
Source: Segment:Production-App, VM Tag:app-server
- Destination: Segment:Production-Database, VM Tag:database
- Service: TCP/5432
- Direction: Inbound
- Rule 4: Deny all (implicit)
- Source: Any
- Destination: Any
- Service: Any
- Action: Drop
Result: Lateral movement is prevented; only east-west flows matching rules are allowed
BGP Design for NSX Edge Routing
NSX edges peer with physical routers via BGP to exchange routes.
BGP neighbor d
Key Takeaways
- DFW Rule Performance:More rules = higher CPU overhead. Consolidate rules using groups and tags. Aim for <100 rules for typical environment. Use rule hit counters to identify unused rules monthly.
Management Domain & Workload Domain Design#
Management Domain Architecture
The management domain (formerly Platform Services Controller + vCenter in vSphere 7.0, now unified as vCenter 8.x) is the control plane of VCF. Its availability directly impacts the ability to manage infrastructure.
Management Cluster Sizing:
Minimum size:
3 hosts for HA quorum (vCenter can restart on any host if one fails).
Typical size:
5-7 hosts to allow 1-2 host failures without performance impact on remaining workload domains.
CPU/Memory:
Management cluster does not run user VMs; reserve resources entirely for management services.
vCenter VM sizing:
8vCPU, 32GB minimum; up to 16vCPU/64GB for 10,000+ VMs.
Consolidated vs Standard Architecture:
Consolidated:
Management domain is physically separate (e.g., on-premises). Workload domains are where user VMs run. Easier to manage, allows independent scaling.
Standard (legacy):
No separation; management services and user VMs co-exist. Not recommended for production; difficult to patch management independently.
Platform Services Controller (PSC) in vCenter 8.x
vCenter 8.x has embedded Single Sign-On (SSO) and embedded PSC. Separate PSC is no longer supported.
Embedded SSO:
Manages authentication and authorization for vCenter, NSX, and other VCF components.
vCenter Replication:
In a multi-vCenter environment (e.g., management at HQ, remote vCenter at branch office), vCenter can be configured with replication topology (e.g., linear, star), but these are separate vCenter instances, not replicas.
Backup & Disaster Recovery for Management
Management domain failure = entire VCF is unmanageable. Design robust backup strategy.
vCenter backup:
Native backup (HTTPS export), or file-level snapshots to external storage. Test restore procedure quarterly.
NSX Manager backup:
3-node cluster, auto-replicated. Backup one node separately for offsite storage.
SDDC Manager backup:
Includes VCF configuration, cluster info, bundle cache. Critical for day-2 operations.
Backup frequency:
Nightly for vCenter/NSX/SDDC Manager; weekly full backup to offsite storage.
RTO for management:
If management is down, RTO for recovery should be <4 hours. Design backup to support this.
Workload Domain Design
A workload domain is a logical grouping of compute resources (clusters) managed by a single vCenter instance. Design workload domains around organizational or operational boundaries.
VI Workload Domains:
Traditional vSphere clusters for VM workloads.
Single vs multiple domains:
Single domain (all clusters under one vCenter) is simpler. Multiple domains (one vCenter per domain) allows independent scaling, separate security policies, organizational alignment.
When to split into multiple domains:
>3000 VMs per vCenter (performance limits), separate business units (cost allocation), different SLAs (upgrade windows), geographic distribution (latency concerns).
Resource boundaries:
Each domain has own storage pools, network segments, security policies. Limited cross-domain sharing.
VKS Workload Domains:
Kubernetes clusters managed by VCF (vSphere with Kubernetes).
K8s workload domain:
Separate from VI workload domains. One cluster can have both VI and VKS capability enabled.
Supervisor cluster:
The Kubernetes control plane running on vSphere nodes. Manages container workloads directly on ESXi (no VMs required).
Tanzu Kubernetes cluster (TKC):
User-deployed K8s clusters on top of supervisor, for microservices.
Multi-AZ (Availability Zone) Design
Multi-AZ architecture deploys workload domains across geographic sites for disaster resilience.
Active-active multi-AZ:
Both sites run production workload, replication is synchronous (RPO=0). Requires <5ms latency, orchestration complexity.
Active-passive multi-AZ:
Primary site active, secondary on standby (via SRM or manual failover). Simpler operationally, higher RPO.
Stretched cluster multi-AZ:
vSAN stretched cluster spans sites; VMs can run on either site transparently. Requires very low latency (<5ms), witness site for quorum.
Key Takeaways
- Management domain sizing: minimum 3 hosts for quorum, maximum 16 hosts per cluster for management overhead control.
- Workload domain isolation requires separate vSAN clusters per tier (dev/prod/sensitive data) for failure domain segmentation.
- Cross-domain networking: NSX Tier-0 must be shared management component; Tier-1 routers per workload domain for traffic isolation.
Security Design#
Zero-Trust Architecture with NSX
Zero-trust security assumes no trust by default; every flow must be explicitly authorized. NSX enables zero-trust at the infrastructure layer.
Zero-Trust Design Pillars:
Identity-based access:
VMs authenticated by tags/groups, not IP addresses. Enables workload mobility without firewall reconfiguration.
Least privilege:
Every flow (even east-west) requires explicit allow rule. Deny-by-default posture.
Microsegmentation:
Segment network into small zones; lateral movement is contained.
Continuous monitoring:
DFW flow logs and NSX Intelligence (behavioral analytics) detect anomalies.
DFW Rule Design Methodology
DFW rules are the enforcement mechanism for zero-trust policy. Design rules systematically to avoid gaps.
Rule Design Steps:
Map application architecture (tiers, inter-tier flows).
For each flow, identify source and destination (use VM tags, not IPs).
Determine service (protocol/port) for each flow.
Write explicit allow rule.
Test rule with traceflow before enforcing.
Enable logging on rules for audit compliance.
Review rule hit counters monthly; remove unused rules.
Rule Template:
- Rule Name: Allow-WebServer-to-AppServer
- Source: VM Tag: tier=web
- Destination: VM Tag: tier=app
- Service: TCP/8080
- Action: Allow
- Direction: Both (initiator and responder)
- Logging: On
- Applied To: All Edges and Hosts
NSX Intelligence & Behavioral Analytics
NSX Intelligence (optional module) uses ML to detect anomalous traffic flows and auto-generate recommended firewall rules.
Behavior learning:
Observes flows in "monitor mode" for 2-4 weeks, builds baseline of normal behavior.
Anomaly detection:
Flags flows that deviate from baseline (e.g., web server suddenly making external API calls at 3am).
Rule recommendations:
Auto-suggests DFW rules to allow baseline, or deny anomalies.
Threat intelligence:
Integrates with threat feeds (VMware, third-party) to block known malicious IPs/domains.
Certificate Management Design
NSX components (Manager, Controllers, Edges) require TLS certificates. Design certificate lifecycle to avoid expiry outages.
Certificate authority (CA):
Use internal CA (e.g., Microsoft CA) or public CA. Internal is common for lab/non-public NSX deployments.
Certificate lifecycle:
1-3 year validity. Rotate 30-60 days before expiry. Automate with monitoring.
Multi-node NSX Manager:
All 3 managers need valid certs. If one expires unnoticed, cluster health degrades.
Renewal strategy:
Use certificate monitoring tool (Prometheus, vRealize Operations) to alert 90 days before expiry.
Identity Federation Design
vCenter and NSX can integrate with corporate identity provider (Okta, Azure AD, Ping Identity) for SSO.
vCenter SSO:
Embedded PSC in vCenter 8.x; configure to trust external IdP (OpenLDAP, AD, OIDC).
NSX authentication:
NSX Manager can use local users, LDAP, or OIDC. LDAP is common for AD-integrated environments.
Design consideration:
Keep backup local accounts (vCenter: administrator@vsphere.local, NSX: admin) in case IdP is unavailable.
Role-Based Access Control (RBAC):
Map corporate groups (e.g., "Cloud-Admins", "Database-Ops") to vCenter/NSX roles for delegation.
Compliance Mapping: PCI-DSS, HIPAA, SOC2
Design security controls to satisfy regulatory requirements. Document mapping of controls to design decisions.
PCI-DSS Compliance (Payment Card Industry):
Requirement 1 (Firewall):
DFW rules enforce network segmentation (cardholder data must be isolated).
Requirement 2 (Defaults):
Change default vCenter/NSX passwords, enable RBAC with least privilege.
Requirement 6 (Code security):
VM patching automation, container image scanning (if using K8s).
Requirement 8 (Authentication):
SSO integration, MFA for admin accounts, local account disablement policy.
Requirement 10 (Logging):
vCenter audit logs, DFW flow logs, NSX manager logs sent to SIEM (minimum 1 year retention).
Requirement 11 (Testing):
Quarterly vulnerability scans (ESXi, vCenter), annual penetration test of cardholder network.
HIPAA Compliance (Healthcare):
Encryption at rest:
vSAN encryption, LUKS on Linux VMs, BitLocker on Windows.
Encryption in transit:
vMotion encryption enabled, DPM (Distributed Port Mirror) not allowed (could expose data).
Access controls:
Role-based (clinician vs IT vs admin), audit trail of who accessed what PHI (Protected Health Info).
Logging/audit:
vCenter event database (minimum 1 year), application logs, access logs to all healthcare data.
Breach notification:
Monitoring for data exfiltration (anomalous large data transfers); incident response plan.
SOC2 Type II Compliance (Service Organizations):
Availability:
Demonstrated by SLA uptime, redundancy design (HA, multi-AZ), disaster recovery testing.
Processing integrity:
Configuration management (Terraform, Ansible), change control process (all infrastructure changes logged).
Confidentiality:
Encryption, RBAC, data classification policies.
Security:
DFW micro-segmentation, vulnerability management, threat monitoring.
Priva
Key Takeaways
- Compliance Audit Preparation:Maintain a control matrix (Requirement ID → Design Component → Evidence) updated quarterly. Assign ownership (e.g., Security team = DFW rules, Ops team = backup verification).
Availability & Recoverability Design#
RPO/RTO Mapping to Technology
Recovery objectives directly drive architectural choices. Understand the cost/complexity of each recovery strategy.
RPO vs RTO Matrix:
Recovery Objective → Technology Mapping:
Scenario 1: RTO=4hrs, RPO=1hr (Standard SLA 99.9%)
- Technology: Snapshot-based backup, off-site vaults
- Tool: Veeam, Nakivo, or native vCenter snapshots
- Recovery: 3-4 hours to restore, test, validate VMs
- Cost: Moderate backup infrastructure, tape library optional
- Example: Non-critical applications, bulk processing
Scenario 2: RTO=30min, RPO=15min (Advanced SLA 99.95%)
- Technology: vSAN stretched cluster OR SRM with continuous replication
- Tool: vSAN async replication + SRM orchestration
- Recovery: Automated failover <5min, automated networking failover via NSX
- Cost: 2 full datacenters, WAN bandwidth for 15min RPO
- Example: Production applications, customer-facing services
Scenario 3: RTO=<1min, RPO=0 (Mission-critical 99.99%)
- Technology: vSAN synchronous stretched cluster (active-active)
- Tool: vSAN replication (sync), NSX for unified networking
- Recovery: Transparent to apps (same site-to-site replication)
- Cost: 3 datacenters (two active + witness), low-latency WAN <5ms
- Example: Databases, financial trading, real-time systems
SRM (Site Recovery Manager) Design
SRM is VMware's disaster recovery orchestration tool. It coordinates failover of entire VM groups and ensures network reconfiguration.
Protected site:
Primary production environment with VMs and storage replication (vSAN, SRA plugin to external array, vSphere Replication).
Recovery site:
Secondary location with SRM server, placeholder VMs (not running), storage for replicas.
SRM server placement:
One SRM instance per site, connected via secure tunnel. SRM at primary drives failover; SRM at secondary can failback.
Recovery plan:
Logical grouping of VMs to failover together (e.g., "App Tier", "Database Tier"). Controls startup order, IP address reconfiguration, network configuration.
Test capability:
SRM can test failover non-disruptively on recovery site (doesn't affect production).
SRM + vSAN Integration:
Storage replication:
vSAN provides native replication (async or sync). SRM controls RPO/RTO, retry logic, cleanup after failover.
Network failover:
SRM customization scripts can trigger NSX segment failover (change next-hop, BGP re-peer), or manually reconfigure routes at recovery site.
Application consistency:
SRM can trigger VSS (Volume Shadow Copy) on Windows VMs before failover to ensure app-level consistency.
Backup Architecture Design
Comprehensive backup strategy combines multiple methods for defense-in-depth.
Backup Tiers:
Tier 1: Continuous snapshots:
vSAN snapshots (on-array), retained for 7-14 days. Fast restore (<1min), protects against accidental deletion/corruption. RPO = seconds.
Tier 2: Daily incremental backups:
Veeam/Nakivo backup appliance, incremental after full backup. Retained 30 days on disk. Restores from full+incrementals take 5-30min depending on size. RPO = 24 hours.
Tier 3: Weekly full offsite backup:
Full backup copied to geographically distant vault (cloud, co-location). Retained 1 year+. RTO for full restore = hours. RPO = 7 days.
Tier 4: Immutable offsite copy:
Monthly full backup to immutable object storage (S3 WORM, Azure Immutable Blobs). Protection against ransomware (cannot be deleted even by admin). RTO = days. RPO = 30 days.
Backup for Management Domain:
vCenter native backup:
HTTPS export snapshot, weekly full, retained 4 weeks locally.
NSX Manager backup:
Cluster backup (all 3 nodes), weekly, offsite copy.
SDDC Manager backup:
Configuration backup, daily, offsite copy.
Target RPO:
If management is lost, recovery from backup should be <2 hours (RTO).
Multi-Site Design Patterns
Different multi-site topologies serve different RTO/RPO targets and operational models.
Active-Passive (Traditional DR):
Primary site runs all production; secondary is dormant (cold standby) or warm standby (replicas kept updated).
RPO: High (depends on replication frequency, typically 1-4 hours for async replication).
RTO: 30min - 4 hours (failover manual or via SRM, network reconfiguration needed).
Cost: Secondary site is underutilized (infrastructure sitting idle).
Use case: Non-critical workloads, cost-sensitive organizations, regulatory requirement for geographic separation.
Active-Active (Load-Distributed):
Both sites run production workload simultaneously. VMs can fail over to other site transparently.
RPO: Low (sync replication, ~100ms latency acceptable).
RTO: <5min (automatic failover via monitoring/orchestration tools).
Cost: Higher (both sites must be sized for full workload), network latency overhead, complex DNS/load balancing.
Use case: High-availability critical apps, elastic scalability, geographic load distribution for latency optimization.
Stretched Cluster (Transparent Failover):
vSAN stretched cluster spans 2 sites with witness at 3rd loca
Key Takeaways
- RPO/RTO targets drive replication strategy: synchronous (0 RPO, high latency) vs asynchronous (data loss risk, better latency).
- Backup frequency and retention policy depend on compliance needs—ISO 27001 requires 3-year retention minimum.
- Site failover decision tree: >100ms network latency favors stretched cluster with witness; <100ms favors active-passive replication.
Design Process Masterclass — From RFP to Implementation#
The VMware Design Framework & Methodology
VCF 9.0 architectures begin with rigorous discovery and structured documentation. The VMware Design Framework (evolved from VVD principles) consists of five phases: Discovery → Requirements → Design → Documentation → Handoff.
Discovery Workshops: The Foundation
Effective discovery captures the true business drivers, not just technology lists.
Requirements Categorization
Organize findings into four requirement classes:
- Category
- Examples
- Design Impact
- Business
- RTO: 4h, RPO: 1h; opex target -30%; platform consolidation on VCF
- Drives multi-site vs single-site, backup strategy, automation scope
- Technical
- Performance SLA: <10ms latency for tier-1 app; vSAN stretched cluster; GPU acceleration
- Cluster layout, vMotion constraints, edge node requirements
Compliance
HIPAA: encryption at rest + in transit; audit logging; workload isolation; PCI-DSS: DFW segmentation
Encryption strategy, vDefend rules, namespace design, vTPM requirements
Functional
Kubernetes automation; self-service catalog; cost showback; capacity planning integration
VCF Ops dashboard design, API integration, policy framework, telemetry ingestion
Architecture Artefacts & Their Purpose
- Conceptual Architecture
High-level block diagram showing business domains (e.g., "Development", "Production", "Edge") and their interconnections. No technical depth yet. Audience: CIO, CFO, business stakeholders.
- Logical Architecture
Maps conceptual domains to VCF constructs: which domains become which workload domains (VI, VKS, Kubernetes)? NSX Federation topology (shared Tier-0 vs local T0 per domain)? Security zones (DMZ tier, app tier, database tier, admin tier)? Data flow (east-west, north-south, client-to-service). Audience: architects, security.
- Physical Architecture
Real hardware, IP ranges, vlan/vxlan mappings, vSAN disk layouts, NSX node placement, edge location, failure domain boundaries, redundancy paths. Audience: implementation team, NetOps.
Design Decision Register Template
Track every major choice and its rationale:
| Decision ID | Area | Decision | Rationale | Alternatives | Risk Mitigation |
|---|---|---|---|---|---|
| D-001 | Cluster Design | 4-node management + 8-node VI (split) | Isolation for stability; independent scaling | Consolidated | Monitor resource trending |
| D-002 | Storage | vSAN ESA 8+2 FTT2 (20 hosts) | Performance for latency-sensitive; cost per IOPS | OSA 4+2 | Planned ESA->OSA migration |
| D-003 | DR Strategy | Pilot-light site with auto-scale | Cost-optimal 2h RTO; data replication via native | Active-active | Test DR quarterly |
| D-004 | NSX Federation | Shared T0 per VCF instance | Simplified BGP (1 neighbor per site) | Federated T0 | Monitor T0 HA failover |
| D-005 | Encryption | Native KMS on primary, external KMIP 1.1 | Simplicity primary; operational independence DR | All external | Failover runbook testing |
Sizing Tool Outputs → Design Validation
VMware Capacity Planner output feeds into design. Cross-check against:
CPU oversubscription ratio:
Tool says 4:1 for general compute; your design assumes 4:1 → validated. Memory overhead: Tool reports 8 GB per host (vmkernel, services); your 512 GB host with 50 VMs gets 10.2 GB avg = 8 overhead, so ~502 GB usable for VMs. vSAN capacity: Tool output: "16 TB raw × 0.8 (efficiency) × 0.67 (FTT2 → 2/3 usable) = 8.5 TB per host; 20 hosts = 170 TB usable" — matches your BoM?
Network bandwidth:
Tool flags if East-West traffic (VM-to-VM) exceeds uplink capacity; validate with DFW rules in logical design.
Bill of Materials & Statement of Work Handoff
BoM must be granular and traceable to design decisions:
Management Cluster (4 nodes):
- CPU: 2 × 20-core (40 cores/node) = 160 cores total
- RAM: 384 GB/node (1.5 TB total)
- Storage: 2 × 1.6 TB NVMe (local) + 4 × 10 TB SAS (shared vSAN, mirrored, FTT1)
VI Workload Domain (8 nodes):
- CPU: 2 × 28-core = 56 cores/node × 8 = 448 cores
- RAM: 512 GB/node × 8 = 4.096 TB
- vSAN: 2 × 3.2 TB NVMe (cache) + 8 × 14 TB SAS (capacity) per node
NSX (Shared):
- 2 × Edge nodes (4 cores, 32 GB each, 2× 500 GB SSD)
- 3 × Segments (2× prod T1, 1× mgmt T1)
Network:
- 2× 25 Gbps for vsan replication
- 4× 10 Gbps for workload traffic (bonded)
Transition to Implementation
Final design doc hands off to Implementation with:
Deployment sequencing:
Rack → cable → BIOS/firmware → ESXi install → vCenter → vSAN formation → NSX bootstrap → workload domain creation → application deployment.
Success criteria:
All health checks green; performance baseline estab
Key Takeaways
- VCDX Insight:A complete design decision register is non-negotiable. Panelists expect you to explain "why this shape, not that shape?" for every choice. Document assumptions (e.g., "CPU contention acceptable up to 25% RDY") and trade-offs explicitly.
Mathematical Sizing — Show Your Work#
CPU Sizing: Subscription Ratios by Workload Class
Oversubscription is acceptable only when utilization data supports it. The formula:
Total vCPU = (Sum of vCPU requests across all VMs) / Subscription Ratio
Number of Hosts = Total vCPU / (pCPU/host × EVC de-rating × Overhead reservation)
Subscription ratios by workload class (based on utilization patterns):
- Workload Class
- Ratio
- Justification
- Peak CPU %RDY Tolerance
- Latency-Sensitive (DB, real-time)
- 1:1
- No contention; strict performance SLA
- <5%
- Mixed Tier-1 (ERP, CRM)
- 2:1
- Moderate burstiness; tolerate occasional spikes
- <15%
- General Purpose (web, app servers)
- 3:1 to 4:1
- Load-balanced; average utilization 20–30%
- <25%
- VDI / Batch
- 4:1 to 6:1
- Highly variable; idle periods; DRS consolidation
- <30%
CPU Sizing Example: Mixed Tier-1 Workload
Requirements:
120 VMs averaging 6 vCPU each = 720 total vCPU demand
Workload class: mixed tier-1 (ERP + CRM) → 2:1 subscription Hardware: dual-socket, 28-core per socket = 56 pCPU per host EVC baseline: Broadwell (legacy requirement) → de-rate to 90% = ~50 effective pCPU
Overhead reservation: 5% for vmkernel/services = 2.5 pCPU reserved
Available pCPU for VMs per host: 50 × 0.95 = 47.5 pCPU
Step 1: Apply subscription ratio
vCPU to allocate = 720 vCPU / 2 = 360 pCPU required
Step 2: Number of hosts
Hosts = 360 pCPU / 47.5 pCPU-per-host = 7.58 → round up to 8 hosts
Step 3: Validate headroom
Actual allocation: 8 × 47.5 = 380 pCPU available
Utilization: 360 / 380 = 94.7% → tight, recommend 10 hosts (76% utilization)
Result:
10-host cluster for sustainable 2:1 oversubscription with 24% headroom for growth.
Memory Sizing: The Multi-Layer Math
Memory is the hardest component to right-size because it includes hidden overhead:
Host Memory Budget:
- = Σ(VM memory allocations)
- + vmkernel overhead (~8–12 GB per host, varies by ESXi version)
- + vSAN cache requirement (5–15% of raw capacity, typically 10%)
- + NSX buffer (if TEP on host: ~2 GB for overlay)
- + HA reservation (either dedicated standby host OR percentage model)
- + 20% headroom for transient spikes
Example for 50 VMs on single host:
VM memory sum: 50 VMs × 10 GB avg = 500 GB
vmkernel overhead: 10 GB
vSAN cache budget: 200 GB raw SSD → 20 GB cache = 20 GB reservation
NSX buffer: 2 GB
- HA reservation: (500 + 10 + 20 + 2) × 25% = 133 GB (if 25% HA policy)
- Headroom: 20% of above = 133 GB
- Total needed: 500 + 10 + 20 + 2 + 133 + 133 = 798 GB
Physical host size: 1 TB (1,024 GB) ✓
Usable for new VMs: 1024 - 798 = 226 GB headroom
Transparent Page Sharing (TPS) & Memory Compression Impact
Modern ESXi (9.x) uses TPS primarily for security isolation (not performance). Memory compression is more aggressive:
TPS savings:
Typically 5–10% on homogeneous workloads (e.g., 100 identical VDI desktops), negligible on heterogeneous environments. Do NOT rely on TPS for capacity planning.
Compression (in-guest):
zswap + balloon driver can save 10–30% on bursty workloads but introduces CPU tax (~5% CPU per 10% compression rate). Monitor esxtop CMPRS/s and UNCOMP/s.
vSAN Read Cache:
Reduces DRAM pressure for read-heavy workloads. If vSAN hosts 80% hot data in cache, cold data spill to disk doesn't consume host memory.
vSAN Capacity Sizing: Raw → Usable Math
Formula (vSAN OSA traditional):
- Usable Capacity = Raw Capacity × Efficiency × Fault Tolerance × Slack Reserve
- where:
- Efficiency = 0.80–0.85 (RAID-1/5/6 overhead)
Fault Tolerance (FTT):
- FTT1 (mirrored): 0.5 (you keep half)
- FTT2 (3 copies): 0.33
- FTT3 (4 copies): 0.25
RAID-6 (striped): ~0.667
NOTE ON SLACK: the flat 25-30% 'slack space' rule applied only BEFORE vSAN 7 U1. From 7 U1 it is Reserved Capacity = Operations Reserve (hardware-dependent, documented examples 6-17%) + Host Rebuild Reserve (~1/N of the cluster: 25% at 4 hosts, ~8% at 12). Combined can be under 10%. With fault domains configured the toggles are unavailable - fall back to ~25% free. In VCF 9.1 both reserves are removed in favour of the Auto-RAID Effective Capacity view. Size with the vSAN ReadyNode Sizer, not a flat multiplier. The 0.85 factors below are exam-style arithmetic, not Broadcom guidance.
Slack Reserve = 0.85 (exam-style shorthand; leave 15% unfilled for resync headroom)
- Example: 20-host cluster, 10 TB raw capacity per host, FTT2, RAID-1 mirrored
- Raw: 20 hosts × 10 TB = 200 TB
- Usable = 200 × 0.8 × 0.33 × 0.85 = 45.1 TB
vSAN ESA (9.x):
ESA is single-tier (NVMe only) and uses RAID-5/6 per default, higher efficiency:
- Example: 8-host ESA cluster, 8 TB NVMe per host, 8+2 RAID (stripe width 8)
- Raw: 8 × 8 TB = 64 TB
- RAID efficiency: 8/(8+2) = 0.8
- Fault tolerance: can lose 2 devices
- Slack: 0.85
- Usable = 64 × 0.8 × 0.85 = 43.52 TB
Per-host share: 43.52 / 8 = 5.44 TB usable per node
vSAN Max (9.0.2+, disaggregated):
Compute and storage clusters are separate. Storage cluster can be 4–32 nodes. Eliminates co-location penalty.
NSX Edge Throughput Sizing
Edge throughput is bottlenecked by:
Throughput line-rate:
25 Gbps (standard T0/T1 edge container) or 100 Gbps (large edge container). Shared across all services (NAT, FW, LB, IPSEC).
Stateful connection limit:
~4M connections per large edge node (smaller for small containers).
DFW impact:
Each rule adds ~5% overhead. 200 rules = ~1 Gbps throughput penalty.
NAT overhead:
~15% CPU/throughput cost per NAT operation.
Service insertion (vDefend/Avi):
Minimum 1 Gbps guaranteed throughput even with chaining.
Edge Sizing Example:
Workload: 3 production Tier-1s, 5 Gbps per T1 (peak)
Base demand: 15 Gbps
Overhead factors:
-
Key Takeaways
- VCDX Insight:Panelists expect you to present the full math, not just the answer. Show your assumptions (subscription ratio, FTT strategy, NVMe model), intermediate steps, and sensitivity analysis (e.g., "if actual utilization is 20% higher, we need X hosts instead of Y"). Defend non-standard choices
Cluster Design Patterns#
Management Cluster: The Foundation
Minimum sizing: 4 nodes
3-node fails VSAN quorum (PSA required).
4-node: can tolerate 1 node failure + maintain vSAN quorum + HA admission control space.
Larger: 5–7 nodes for large environments (>500 workload VMs) to reduce management overhead impact during failures.
Consolidated vs. Standard Model:
- Model
- Configuration
- Pros
- Cons
- Size Threshold
- Consolidated
- 1 cluster: 4 management nodes + VMs coexist
- Low capex; reduced VM sprawl
- Management node failure impacts VMs; harder capacity planning
- <100 workload VMs
- Standard
- 4-node management + separate 4+ node VI cluster
- Isolated; management failure doesn't affect workloads
- Higher capex; dual infrastructure
- >100 workload VMs, tier-1 SLA
EVC (Enhanced vMotion Compatibility) requirement:
VI Workload Domain: 4 to 96 Hosts
Sizing guidance:
4–8 hosts:
Single rack, single rack PDU pair. HA admission control via percentage model (25% reserved). DRS fully effective.
8–32 hosts:
Multi-rack; split PDU redundancy across rows. HA admission control: slot-based (requires careful CPU/RAM profile homogeneity) or percentage. DRS with rack affinity rules.
32–96 hosts:
Potential for multiple vSAN racks. Fault domain boundaries via rack/PDU group. DRS with NUMA-aware scheduling.
Stretched VI Cluster (vSAN stretched, vMotion <10 ms RTT):
Typically 6–12 nodes per site (3 per site minimum for vSAN quorum).
Witness appliance 3rd site (or dedicated witness node in primary site).
Active-active HA: each site can tolerate 1 node failure independently.
RPO/RTO: ~0 RPO (synchronous replication), <1 min RTO (automated failover via HA).
VKS-Dedicated Workload Domain
If containerized workloads are a primary use case, isolate in dedicated WLD:
Cluster size:
4–12 nodes minimum (Kubernetes masters require quorum, 3 minimum; recommend 5).
vSAN design:
OSA ESA for performance; smaller disk footprints per VM (Kubernetes workers are lean: 4 vCPU, 8 GB RAM typical).
Network isolation:
Dedicated NSX segments for Kubernetes: kube-control, pod-overlay, ingress, service-lb.
Storage class:
StorageClass → vSAN SPBM + vasa provider ensures Kubernetes PVC maps to correct vSAN datastore.
Namespace isolation:
vDefend micro-segmentation per namespace; DFW rules block cross-namespace traffic by default.
Storage-Dedicated Cluster (vSAN Max)
Disaggregated architecture for extreme scale (1000+ VMs, 50+ PB capacity):
Cluster composition:
4–32 dedicated storage nodes; compute nodes separate. NO VM workloads on storage cluster.
vSAN Max specifics:
Requires vSAN 9.0.2+; RAID-EC(6) for capacity efficiency; hierarchical object model (containers → services → objects).
Network:
Dedicated vSAN network (10× 25 Gbps bonded recommended) between compute and storage clusters.
Advantage:
Decouple capacity from compute; scale storage independently.
GPU-Enabled Cluster Design
NVIDIA GRID / vGPU for VDI:
Cluster requirement:
All nodes homogeneous GPU model (e.g., all A100-80GB or all L4).
Licensing:
Per-VM vGPU license (tied to vSphere host UUID) + GRID Manager.
Host sizing:
1 GPU per 4–12 VMs (varies by profile: M10 quad-desktop, M60 8-desktop, A100 single high-performance).
vMotion:
Requires EVC CPU profile match + GPU BIOS compatibility; vMotion latency penalty ~500 ms for GPU memory transfer.
HA consideration:
Slot calculation must account for GPU count. A 4-GPU host failure → reserve 4-GPU capacity on another host for HA slots.
AI Enterprise Stack (e.g., inference):
Typically baremetal or container-based (VKS); less common in traditional VMs.
If VM-based: require dedicated cluster to avoid noisy-neighbor contention (GPU-hungry inference can starve other VMs).
Monitor GPU memory with nvidia-smi; set GPU memory reservation in VM hardware profile.
Edge Cluster Sizing
Single Edge Node (small deployment, <5 Gbps throughput):
Typically T0 SPOF; T1s stretched to management site for HA.
CPU: 4 vCPU, RAM: 32 GB (minimum for stateful services); connections: ~1M.
Edge HA Pair (production edge, active-active T0 BGP):
2 large edge nodes; BGP neighbor count per edge: limit ~64 upstream neighbors.
CPU: 16 vCPU, RAM: 128 GB per edge; connections: ~4M per node.
Active-active BGP: Upstream route learning (BGP ECMP) balances ingress traffic; each edge learns full routing table and advertises local T1s.
Failover: sub-second BGP reconvergence; stateful connections may be lost (no NSX Federation backup path).
BGP neighbor calculation:
Edge HA pair, autonomous system: 65000
Upstream routers (core): 4 routers, each in AS 65001
Leaf switches: 8 VLANs, each VLAN has 1 leaf (8 uplinks total)
BGP session count per edge:
- 4 upstream routers (BGP EVPN eBGP) = 4 sessions
- 8 leaf switches (BGP IP routing iBGP) = 8 sessions
Total: 12 sessions per edge well below 64-session limit ✓
Multi-AZ vs. Cross-Cluster Stretched: Decision Framework
- Pattern
- Topology
- Latency RTT
- RPO
- Capex
- Complexity
- Stretched Cluster
- 1 vSAN, 2 sites, 6 nodes (3 per site)
- <5 ms
- 0 (sync replication)
- Higher (2 sites, uniform
Key Takeaways
- VCDX Insight:Choose stretched cluster ONLY if latency <5 ms AND synchronized replication RPO is a hard requirement. Multi-AZ is more common for geo-distributed deployments; SRM for cost-optimized DR where RTO/RPO can be 1–4 hours. Panelists will challenge your topology if it doesn't match the busine
Availability Design — Aligning to SLAs#
Three 9s, Four 9s, Five 9s: Architecture Mapping
- SLA Tier
- Uptime %
- Downtime/year
- Architecture Minimum
- Key Design Points
- 3-9s
- 99.9%
- ~8.75 hours
- Single site, HA within cluster
- 4-node cluster; HA admission control (25% reserved); DRS (avoid consolidation to <3 nodes)
- 4-9s
- 99.99%
- ~52 minutes
Stretched cluster OR primary + hot DR
Stretched vSAN + HA, OR separate cluster + NSX Federation + Avi GSLB; zero data loss not required
5-9s
99.999%
~26 seconds
Multi-site active-active (rare)
Stretched vSAN (synchronous replication) + HA + active-active HA across sites + multi-master NSX Federation; mission-critical only
HA Admission Control Strategies
Goal: reserve capacity on remaining hosts to guarantee failed-VM restart space.
- Dedicated Host Model
How it works:
Designate 1 host as failover standby; no VMs run on it normally.
Pros:
Simple math; guaranteed capacity; no oversubscription risk.
Cons:
Wastes ~25% capex (1 of 4 hosts idle); poor resource utilization.
When to use:
Small 4-node clusters; mission-critical single-app workloads; cost no concern.
- Percentage-Based Model
How it works:
Reserve X% of cluster CPU/Memory capacity for HA slots (default 25%).
Pros:
Allows hosts to run VMs; dynamic scaling (remove host → reduce reserve). Cons: Assumes homogeneous workload; fragmentation if VM sizes vary wildly. Calculation example: 8-node cluster, 400 GB total memory, 25% reserve = 100 GB reserved → can safely guarantee restart of largest 100 GB VM on any remaining host.
- Slot-Based Model (Advanced)
How it works:
Define a "slot" as the size of the largest VM (or typical large VM). Ensure cluster can accommodate N slots on N-1 hosts.
Pros:
Most accurate; accounts for actual workload sizes.
Cons:
Recalculation needed if new large VM added; assumes slot sizes don't change.
Example:
Largest VM = 32 GB RAM, 16 vCPU. Cluster 8 nodes, each 512 GB RAM, 56 vCPU. Slots = min(512/32, 56/16) = 16 slots per host → 7 hosts × 16 = 112 slots total. If 80 VMs total, we have 32 slots headroom for HA admission control.
Pitfall:
Percentage-based HA with DRS "consolidate" aggressive mode can violate admission control. Set DRS "Cluster imbalance threshold" to avoid over-consolidation. Recommendation: Threshold >40% for clusters with HA.
DRS Rules for HA Resilience
VM-VM Affinity (Should):
Group related VMs (e.g., app tier + cache tier) to minimize vMotion during failure:
Rule: "tier-1-app-vms" affinity "tier-1-cache-vms"
Effect: DRS prefers to run these VMs on same host (shared cache locality reduces latency).
Strength: Should (DRS can violate if capacity constrained).
VM-Host Affinity (Must):
Ensure tier-1 VMs never run on compute-constrained hosts (e.g., dedicated vSAN storage nodes if vSAN Max):
Rule: VM group "prod-vms" must run on host group "compute-hosts"
Effect: DRS rejects vMotion or initial placement on non-compute hosts.
Strength: Must (hard constraint; placement fails if violated).
Anti-Affinity (Should) for HA Resilience:
Spread redundant tier-1 VMs across hosts to tolerate node failure:
Rule: VM "erp-app-01" should NOT run on same host as "erp-app-02", "erp-app-03"
Effect: DRS spreads these VMs; if 1 host fails, 2 copies survive.
Risk: If anti-affinity breached, HA may restart both on same host → potential restart failure if host overloaded.
Mitigation: Combine with percentage HA (25%+ reserved).
Fault Domains: Rack & PDU Separation
vSAN stretched cluster requires fault domain isolation for witness placement:
Rack-level:
Assign 3 nodes to "rack-A", 3 to "rack-B", witness to "rack-C" (or separate site).
PDU-level:
If single rack, ensure nodes split across PDU A + PDU B for power independence.
vSAN implication:
Each fault domain must hold ≥1 copy of data (FTT1). Failure of 1 rack leaves data on other racks.
Fault Domain Configuration (vSphere):
vcenter.hostname> configure> set> cluster> fault domain
Node "esxi-1a" → Fault Domain: RackA Node "esxi-1b" → Fault Domain: RackA Node "esxi-2a" → Fault Domain: RackB Node "esxi-2b" → Fault Domain: RackB Node "esxi-3a" → Fault Domain: RackC
vSAN stretched witness → Place in RackC (3rd location)
Proactive HA: Hardware Support Manager Integration
VCF 9.0 integrates hardware health monitoring (IPMI, SNMP) with HA:
Feature:
Monitor hard drive predictive failure, power supply degradation, fan failures.
Action:
Trigger proactive VM migration off host BEFORE failure occurs (vMotion to healthy host).
Configuration:
Hardware Support Manager (HSM) + vSAN Health Service subscribe to hardware events.
Advantage:
Reduces unplanned HA restarts; maintains performance during proactive migration.
Key Takeaways
- 99.9% availability (SLA) allows ~43 minutes downtime/month: requires 1 N+1 host failure resilience minimum.
- 99.99% availability (four nines) demands N+2 resilience, multi-site failover, and <5 minute detection+recovery time.
- Admission control calculation: Reserve capacity for (Worst-case host failure + peak workload) to maintain SLA during incidents.
Security Architecture Design#
Defense in Depth: Seven Layers
- Layer
- Control Category
- VCF Technology
- Example Rule/Policy
- Physical
- Access control, intrusion detection
- Data center badge, camera, environmental monitoring
- Biometric access to server room; 24/7 camera
- Host/Hypervisor
- Secure boot, attestation, access control
- ESXi secure boot, TPM 2.0 attestation via HardwareSupport
- ESXi lockdown mode (strict); no SSH by default; audit logging all SSH access
- VM Runtime
- VM-level isolation, vTPM, secure boot
- vTPM 2.0, vSGX, Virtualization-Based Security (VBS) for Windows
- All Windows VMs: require Secure Boot + vTPM + VBS enabled
- Workload/App
- Authentication, authorization, encryption
- vDefend microsegmentation, RBAC in app, encrypted credentials
- Kubernetes RBAC per namespace; HashiCorp Vault for secret rotation
- Network
- Segmentation, DFW, IDS/IPS
- NSX Distributed Firewall, vDefend IDS, AVI WAF
- DFW: allow only port 443 (HTTPS) from Tier-1 to Tier-2; IDS: block known CVE-2024 exploits
- Storage
- Encryption at rest, access control
- vSAN native encryption, vSAN KMS, SPBM policy
- vSAN KMS policy: all tier-1 VMs encrypted; daily key rotation
- Data in Transit
- Encryption, DLP
- NSX IPSec, vMotion encryption, vSAN stretched replication TLS
- vMotion encrypted (default 9.0); vSAN stretched witness uses mTLS with certificate pinning
Mapping Controls to Frameworks
NIST Cybersecurity Framework (CSF) Alignment:
Identify:
Asset inventory (CMDB), data classification (PII, confidential, public). VCF: Tags in vCenter, vDefend security groups.
Protect:
Access control (SSO, RBAC), encryption, DFW rules. VCF: vDefend DFW policies tied to security tags.
Detect:
Logging, alerting, IDS. VCF: vDefend IDS inline on all tier-1 traffic; VCF Ops alert on policy violations.
Respond:
Incident response runbook, quarantine VMs. VCF: Automation to isolate compromised VM (move to quarantine VLAN, snapshot).
Recover:
Backup, DR, forensics. VCF: vSAN snapshots, SRM, audit logs → SIEM.
PCI-DSS Zone Design:
DMZ (Untrusted):
Web servers, load balancers exposed to internet. No direct DB access.
App Tier (Trusted):
Application servers, internal APIs. Only HTTP/443 ingress from DMZ.
Database Tier (Secure):
Databases, vault. Only app-specific ports (5432/PostgreSQL, 3306/MySQL, etc.) from App tier.
Admin Tier (Isolated):
Management VMs, backups. Only specific admin users from admin network.
DFW Rule Set for 3-Tier App:
Security Group "sg-dmz" (web servers):
Tag: tier=web, env=prod
Security Group "sg-app" (app servers):
Tag: tier=app, env=prod
Security Group "sg-db" (databases):
Tag: tier=db, env=prod
DFW Rules:
1. Allow: sg-dmz → sg-app, Port 443, Protocol TCP (HTTPS) 2. Allow: sg-app → sg-db, Port 5432, Protocol TCP (PostgreSQL) 3. Deny: sg-dmz → sg-db (zero-trust: no direct web-to-DB) 4. Allow: sg-db → sg-app, Port 5432, Protocol TCP (DB response)
- Deny: Any other traffic (implicit deny)
vDefend Micro-Segmentation for Zero-Trust
VCF 9.0 vDefend suite includes (previously NSX Security):
DFW (Distributed Firewall):
Inline at hypervisor layer; apply per-vNIC; context-aware rules (user, application, tag).
IDS/IPS (Intrusion Detection/Prevention):
Signature + behavior-based; deployed inline on TEPs (Tunnel Endpoints) or edges; inspect encrypted traffic if keys available.
Network Segmentation:
Logical segments (vxlan-backed) per tenant/app; microsegmentation via tags.
Zero-Trust Segmentation Model:
Default: DENY ALL traffic
Whitelist: Only explicitly allowed flows (source tag → dest tag, port, protocol)
Example for Kubernetes namespace "production":
Pod Security Group: tag=kube-namespace:production
Egress: Allow to:
- Internal kube-dns (port 53)
- External HTTPS endpoints (port 443, allowlist IPs)
- RDS database (port 5432)
Ingress: Allow from:
- Ingress controller (port 8080)
Result: Compromised pod cannot exfiltrate data laterally to other namespaces
Encryption Design: Key Hierarchy
Native vSAN KMS (Simplest, Recommended for Primary):
Master Key Storage:
TPM-based on vSAN host, or encrypted in vCenter.
Scope:
Per vSAN datastore. All VMs on datastore use same master key.
Advantage:
Operational simplicity; no external KMS dependency.
Disadvantage:
Key tied to vSAN cluster; harder for multi-site key rotation.
External KMIP 1.1 KMS (Recommended for DR/compliance):
Key Management:
Centralized KMIP server (Thales, Gemalto, AWS KMS, Azure Key Vault).
Scope:
Can scope keys per VM or per datastore; multi-cluster support.
Advantage:
Centralized audit trail; key rotation policy enforcement; compliance auditability.
Disadvantage:
Network latency risk (KMS unreachable → cluster degradation); additional operational overhead.
Key Hierarchy Design (Hybrid):
Primary Site:
KMS Server: Thales KMIP (on-prem)
vSAN Encryption: Master key cached on hosts, key rotation daily
VM-Level: Optional AES-256 per VM for HIPAA workloads
DR Site (Pilot-Light):
KMS Server: External KMIP read-only replica (Thales secondary)
vSAN Encryption: Same master key ID (key serv
Key Takeaways
- DFW rule consolidation: use groups and tags to reduce rule count to <100 rules; consolidation improves CPU performance.
- Zero-trust segmentation: design VCF with default-deny policies, explicitly allow required flows (East-West traffic).
- Encryption policy: data-at-rest (vSAN TMK) and data-in-transit (IPsec between sites) required for PCI-DSS compliance.
Multi-Site Design Patterns#
Active-Active (Stretched Cluster)
Topology:
Single vSAN cluster spanning 2 sites; 6 hosts (3 per site, 1 witness on 3rd site or primary).
Single HA domain; vMotion <10 ms RTT between sites.
NSX Tier-0 shared; Tier-1s local or stretched per site.
Characteristics:
RPO:
0 (synchronous replication).
RTO:
<1 minute (HA failover + DNS propagation).
Data locality:
Both sites hold copies; read from local, write synchronously across.
Latency requirement:
<5 ms RTT; lower = better (reduced replication latency).
vSAN Stretched Specifics:
Cluster Configuration:
Fault Domain A (Site-A): esxi-1a, esxi-1b, esxi-1c (3 nodes)
Fault Domain B (Site-B): esxi-2a, esxi-2b, esxi-2c (3 nodes)
Fault Domain C (Witness): witness-node (1 witness appliance, ~1 vCPU, 8 GB RAM)
vSAN Policy (FTT2, capacity):
- Minimum components: 3 (1 per FD → data on both sites + witness)
- Stripe width: 2
- Failures to tolerate: 2 (any 2 nodes fail, data survives)
Witness role:
- Does NOT hold data
- Breaks split-brain (partition → FD with witness wins, other FD fenced) - If witness fails → cluster continues but loses partition tolerance
Recovery procedure (1 node down):
- Automatic resync; rebuild rate limits to 500 MB/s (adjustable)
- Completion time: (10 TB / 500 MB/s) ÷ 3 (parallel) = ~6-7 hours
- During resync: cluster remains operational, may be slightly slower
Active-Passive (SRM-Based)
Topology:
Production site: active cluster (all VMs running).
DR site: passive cluster (smaller or standby), VMs powered off.
Replication: vSphere Replication (app-level snapshot) or native vSAN stretched (if <5 ms possible).
Characteristics:
RPO:
15 min – 4 h (depends on snapshot frequency).
RTO:
15 min – 1 h (manual failover via SRM; ~10–30 min to power on VMs and validate).
Cost:
Lower capex (DR site can be smaller; only prod cluster is full-scale).
Complexity:
SRM recovery plan testing required quarterly.
SRM Workflow:
- Create Recovery Plan (in SRM GUI):
- Select VMs to protect
- Define failover priority (tier-1 apps first, batch jobs last)
- Configure failover IP mapping (prod subnet → DR subnet)
- Enable Protection:
- vSphere Replication or array-based replication begins
- RPO monitoring: alert if replication lag exceeds threshold
- Test Failover (monthly):
- SRM creates isolated snapshot copies on DR
- Power on test VMs; validate application functionality
- Cleanup (no impact to production)
- Disaster Event:
- Initiate failover (SRM)
- SRM: powers off prod VMs, powers on DR VMs, updates DNS/load balancer
- Time: ~10–30 min depending on VM count and startup scripts
- Failback (recovery phase):
- Once prod restored, resync replication direction
- Power off DR VMs, power on prod VMs (controlled sequence)
Pilot-Light (Cost-Optimized)
Topology:
Minimal DR site (2–4 nodes); runs non-critical workloads or test VMs.
In disaster, DR site scales up (add cloud resources or pre-staged hardware) and receives production failover.
Characteristics:
RPO:
1–4 h.
RTO:
2–4 h (minimal if pre-staged hardware; longer if cloud scale-up).
Cost:
Lowest capex; minimal licensing on DR site.
Trade-off:
RTO/RPO not as tight as active-active or active-passive.
Automation Example (HCX RAV + Auto-Scale):
Normal state: Prod site (8 nodes), DR site (2 nodes)
Disaster trigger (prod site down):
- Auto-scale: AWS or Azure provisions 6 VMs (cloud-based compute)
- HCX RAV (replication acceleration): resume replication from off-site backup storage
- HCX network extension: stretch DR site subnet to cloud
- Failover: VMs powered on in cloud; restore to pilot-light + cloud hybrid
Cost during disaster: Higher (cloud resources active)
Timeline: RTO 2–4 h (cloud resource provisioning) vs. 30 min if physical standby
Backup-Only (Cost-Minimum)
Topology:
Production site: single cluster or multi-site within site.
DR: nightly backup to off-site storage (tape, cloud object store, or dedupe appliance).
Recovery: restore from backup (no standby cluster).
Characteristics:
RPO:
24 h (1 daily backup).
RTO:
4–8 h (restore + power on).
Cost:
Lowest total cost of ownership.
Risk:
Highest; suitable only for non-critical workloads or dev/test.
Hot / Warm / Cold Tier Decision Matrix
- Tier
- Definition
- RTO/RPO
- Tech
- Capex
- Use Case
- Hot
- DR site fully mirrored, VMs running or standby
- 0 / <1 min (active-active)
- Stretched vSAN, HA, GSLB
- Highest
- Tier-0 SLA (finance, healthcare)
- Warm
- DR site ready but VMs powered off; replication active
- 1 h / 15 min (active-passive)
- SRM, vSphere Replication
- Medium
- Tier-1 SLA (ERP, CRM)
- Cold
- DR site minimal; backup-only recovery
- 4 h / 24 h (backup-only)
- Commvault, Backint, tape
- Lowest
- Tier-2, dev/test, archive
Latency Budget Analysis for Stretched Architectures
vMotion Constraint: <10 ms RTT
Network Stack: TCP round-trip time for vMotion pre-copy phase
- Fiber latency (distance): ~5 µs per km
- Switch latency: ~5 µs per hop (typically 3–5 hops between sites)
- NIC latency: ~5
Key Takeaways
- Stretched cluster witness placement: third location (cloud provider, co-location) is mandatory; avoid witness on either primary site.
- Active-passive replication: site affinity rules pin critical VMs to primary site; failover to secondary for DR only.
- Bandwidth planning: vSAN resync uses 10-50% of WAN link; stretched cluster tolerates only 1 component loss, plan link capacity accordingly.
VCF 9.0 Automation & Operations Design Integration#
Catalog Item & Project Hierarchy Design (Upfront Planning)
VCF 9.0 self-service portal requires careful design of project structure, resource quotas, and approval workflows.
Example: SaaS Provider Multi-Tenant Structure
- Org Level
- Project/Folder
- Resource Quota
- Catalog Items Exposed
- Root Org
- N/A
- N/A (system level)
- N/A
- Org: CustomerA
- Project: Production
- vCPU: 100, RAM: 200 GB, Storage: 5 TB
- Web Server, DB Server, Load Balancer (tier-1 items)
- Project: Development
- vCPU: 50, RAM: 100 GB, Storage: 2 TB
- Test VM, Sandbox (tier-2 items)
- Org: CustomerB
- Project: Production
- vCPU: 80, RAM: 160 GB, Storage: 4 TB
- Web Server, App Server (CustomerB-specific)
Governance Policies & Customer Org Alignment
RBAC Design (aligned to customer org structure):
AppOwner role:
Can provision VMs from catalog, manage VM lifecycle, view utilization; cannot modify network/storage policies.
Security team role:
Can create DFW rules, vDefend security tags, audit policies; cannot provision VMs.
Ops team role:
Can modify catalog, adjust quotas, manage templates; full platform access.
Policy Enforcement Example (Kubernetes namespace isolation):
Policy: "All tier-1 VMs must have vTPM + Secure Boot + vSAN encryption"
Enforcement: Catalog item "Tier-1 Production VM" template includes:
- vTPM 2.0 device
- UEFI + Secure Boot
- SPBM policy "tier-1-encrypted" (vSAN KMS)
Consequence: User cannot override; if they try manual edit, policy blocks save.
Policy: "Development VMs max 4 vCPU, 16 GB RAM"
Enforcement: Catalog item "Dev VM" defaults capped; quota prevents oversizing.
VCF Ops Dashboard Design for Multiple Personas
Executive Dashboard (C-suite view):
Overall cluster utilization (CPU, Memory, Storage %).
Cost breakdown (capex amortized, opex power/cooling, licenses).
SLA compliance (% uptime, HA events count, planned maintenance windows).
Alerts: only critical (SLA breach, DR site failure, certification expiry).
NOC/Operations Dashboard:
Per-cluster health (vSAN, vCenter, NSX, Edge).
Workload distribution (VMs per cluster, top resource consumers).
Alerts: warnings + critical (DRS recommendations, space trending, failed backups).
Performance metrics (CPU %RDY, Memory contention, IOPS/throughput).
Quick actions: DRS recommendation approval, snapshot trigger, scaling buttons.
Security Dashboard:
DFW rule violations (attempted flows blocked).
vDefend IDS signature hits (top exploits detected).
Compliance status (PCI-DSS controls pass/fail, CIS benchmark scoring).
Encryption audit (VMs without encryption, KMS key age).
Alert-to-Action Chain Design
Automation reduces MTTR (Mean Time To Recovery). Example:
Alert: "vSAN object missing component (FTT1 violated)"
→ VCF Ops detects; triggers automation:
- Increase vSAN rebuild priority (adjust rate limit to 1000 MB/s)
- Migrate non-critical VMs off degraded host (if resync conflicts with workload)
- Notify NOC: "Component missing, auto-rebuild started, ETA 4 hours"
- Update ServiceNow ticket (ITSM integration via webhook)
- If rebuild fails after 2h: escalate to L2 engineer, create incident
Benefit: MTTR reduced from 8h (manual) to 1h (auto-remediation + escalation)
ITSM Integration (ServiceNow) via Webhook
VCF Ops → ServiceNow workflow:
Incident creation:
Critical vSAN alert → auto-create SNOW incident with: Alert timestamp, severity, component (cluster/host/datastore). Auto-assign to on-call group (e.g., "vSAN_Team"). SLA: P1 → 1 hour response. Change approval: "Cluster upgrade" request in VCF → SNOW change ticket; requires CAB approval before execution. Feedback loop: SNOW ticket resolution → VCF Ops cleared incident; closes feedback loop.
Metric Ingestion Design for Custom Metrics
VCF Ops collects native metrics (CPU, memory, storage) but custom app-level metrics enhance visibility.
Integration architecture:
Custom Metric Sources:
1. App container (Prometheus) → VCF Ops via vLCM metric adapter 2. Database (PostgreSQL pg_stat_statements) → Custom plug-in → VCF Ops 3. Load Balancer (Avi) → REST API → VCF Ops
Example: Track "transaction latency" per application
Prometheus exports: http_request_duration_seconds (histogram)
VCF Ops ingests: latency_p95, latency_p99 per app-tier
Dashboard: shows app latency trend alongside cluster CPU/memory
Alert rule: "If latency_p95 > 100ms AND cluster_cpu > 80% → recommend scaling"
Key Takeaways
- VCF Operations (formerly vROps) is now mandatory for policy compliance: design must include monitoring and alerting for every SLA metric.
- Ansible integration: design assumes day-2 operations via IaC; document automation dependencies (credentials, API endpoints) in architecture.
- Skyline Advisor: proactive risk identification; plan quarterly reviews to address recommendations before they become incidents.
Documentation Standards — VCDX-Grade Design Docs#
Complete Design Document Outline
A VCDX-grade design document follows this structure:
Executive Summary (1–2 pages)
High-level problem statement, proposed solution, key benefits, cost estimate, timeline.
Audience: CIO, CFO, business stakeholders (non-technical).
Scope (1 page)
What is IN scope: VMs, clusters, NSX, storage, security, HA/DR.
What is OUT of scope: existing legacy systems not migrated, application tuning, networking outside data center.
Example: "This design covers the new production VCF cluster and Tier-1 migration; legacy ESXi 5.5 systems remain in place until 2025."
Assumptions & Constraints (1 page)
Assumptions:
Network latency <5 ms between sites, existing SAN decommissioned by Q3 2025, 2 annual maintenance windows, customer IT staff trained.
Constraints:
Budget cap $2M, must use existing fiber (no new runs), certifications (PCI-DSS tier-1, HIPAA dev tier-2).
Current State Assessment (2–3 pages)
Existing infrastructure: 4 clusters (vSAN 6.x, vCenter 6.7, NSX 3.x).
Utilization snapshot: Avg CPU 35%, Memory 45%, Storage 60%.
Identified pain points: 3 HA failures/month, 2 missed RTO windows/quarter, manual capacity planning.
Conceptual Architecture (2 pages, 1 diagram)
Block diagram: Business domains (Production, Development, DR) → VCF constructs (clusters, NSX zones). Identify which business requirements map to which domains. Logical Architecture (5–8 pages, 3+ diagrams) Cluster layout (Management 4-node, VI 12-node, compute-dedicated, storage-dedicated if Max). NSX topology (Global Manager, Tier-0, Tier-1s per domain, security zones). Storage design (vSAN FTT/stripe, vSAN Max disaggregated, datastore capacity). HA/DR strategy (stretched cluster, SRM, pilot-light detail per domain). Security architecture (DFW zones, vDefend rules, encryption, compliance mapping). Physical Architecture (5–8 pages, 3+ diagrams) Rack layout: ESXi node placement, PDU groups, switch ports, VLAN/vxlan mappings. IP addressing scheme: Management network, vMotion, vSAN, NSX overlay, guest VLANs. Cable plant: uplink bonding (4× 25 Gbps), redundancy (dual switches). DR site physical layout (if applicable). Sizing Calculations (5–8 pages) CPU sizing math (workload class → subscription ratio → host count). Memory sizing with overhead breakdown. vSAN capacity (raw → usable), IOPS budget. NSX edge throughput, edge node count. HA admission control (slots, percentage reserved, impact). Growth headroom (next 3-year forecast). Transition / Migration Plan (3–5 pages) Phased approach: Phase 1 (Management cluster), Phase 2 (VI cluster), Phase 3 (NSX), Phase 4 (App migration), Phase 5 (Cutover). Downtime windows (if any), rollback triggers, success criteria per phase. Parallel run period (if applicable): legacy + new running concurrently for validation. Operational Plan (3–5 pages) Runbooks: power-on sequence, failover procedure, recovery from node failure, upgrade process. Monitoring: alerts, thresholds (e.g., CPU RDY >20% → warning), escalation policy.
Backup/recovery: snapshot frequency, retention, DR test schedule (monthly SRM test).
Capacity management: monthly reporting, quarterly forecasting, expansion triggers.
Exit Criteria / Success Metrics (1 page)
All HA/DR tests pass, SLA compliance >99.9%, performance baselines met (CPU RDY <10%, latency <10ms).
All staff trained, documentation updated, runbooks tested in dry-run.
Appendices
A. Bill of Materials (hardware, licenses, software).
B. Network diagram (detailed).
C. IP addressing spreadsheet.
D. vSAN sizing tool output.
E. DFW rule export (in table format).
F. HA admission control calculation sheet.
G. Design decision register.
H. Risk register with mitigations.
I. Glossary and acronyms.
Section Word Count Guidance
- Section
- Target Word Count
- Notes
- Executive Summary
- 500–1000
Concise; avoid jargon; focus on business impact.
Scope & Assumptions
800–1200
Be specific; use bullet lists.
Current State
1000–1500
Data-driven; include metrics, screenshots, charts.
Conceptual Architecture
800–1200
1 diagram; high-level business domains.
Logical Architecture
3000–4000
3–4 detailed diagrams (clusters, NSX, security).
Physical Architecture
3000–4000
Rack diagrams, IP tables, cable specifications.
Sizing Calculations
2000–3000
Show all math; include examples with numbers.
Transition Plan
2000–2500
Phased timeline; rollback procedures.
Operational Plan
2000–2500
Runbooks, alerts, capacity planning.
TOTAL
16,000–20,000
Professional design doc length.
Design Decision Register Template
Track every major architecture decision:
| Decision ID | Date | Area | Decision | Rationale | Alternatives | Trade-off | Owner |
|---|---|---|---|---|---|---|---|
| D-001 | 2024-01 | Cluster Design | 4-node mgmt + 12-node VI (split) |
Key Takeaways
- VCDX Insight:Panelists spend 20–30 minutes on your design doc before defense. Make it scannable: use headings, bullet lists, tables, diagrams. Include a 1-page summary at the front (separate from executive summary) so they can quickly grasp your architecture. If they ask "why did you choose FTT2 for
Lab Exercises & Reference Materials#
Lab Exercise Strategy for VCDX Architect Preparation
Lab exercises bridge theory and real-world architecture. Each lab in this section maps to specific 2V0-13.25 blueprint objectives and uses RCAR (Requirement → Constraint → Assumption → Risk) methodology to build defensible design decisions.
Lab Progression Path:
The labs follow a deliberate progression. Start with the RCAR Design Decision Framework (Lab 01) to establish methodology. Move to Compute Cluster Design (Lab 02) and vSAN Storage Design (Lab 03) for infrastructure foundations. Progress to Multi-Site Stretched Cluster Design (Lab 04) and NSX Micro-Segmentation (Lab 05) for availability and security. Advance to Multi-Workload Domain Design (Lab 06) for enterprise architecture. Then tackle Multi-Site DR Framework (Lab 10) and Self-Service Portal Design (Lab 11) for operational design. Culminate with the Design Document lab (Lab 12) and Healthcare Capstone (Lab 13) that integrate all prior knowledge into VCDX-grade deliverables.
Holodeck Toolkit Lab Environment:
All labs are designed for nested environments using VMware Holodeck Toolkit. Holodeck deploys a complete VCF stack — SDDC Manager, vCenter, ESXi hosts, NSX, vSAN — inside a single physical server. Resource requirements: minimum 128 GB RAM, 1 TB NVMe, 10-core CPU for a 4-host management domain. For stretched cluster labs, deploy two Holodeck pods with a simulated WAN link using Traffic Control (tc) to introduce 5ms RTT latency. The witness appliance runs as a separate nested VM on a third network segment.
Design Decision Documentation Standard:
Every lab produces design decisions in a consistent format: Decision ID (D-001 through D-NNN), Requirement Addressed (traced to blueprint objective), Options Evaluated (minimum 3 alternatives with pros/cons), Selected Option with justification, Constraints & Assumptions documented, Risk Impact (probability × impact scoring), and VCDX Defense preparation (anticipated panelist questions with pivot responses).
Reference Architecture Documents:
VMware Validated Design (VVD) — historical reference for VCF architecture patterns through VCF 4.x; VMware Cloud Foundation 9.0 Planning & Preparation Guide — current canonical sizing and network requirements; vSAN 9.0 Design & Sizing Guide — ESA storage pool calculations, FTT overhead multipliers, capacity planning worksheets; NSX 9.0 Design Guide — Tier-0/Tier-1 topology, micro-segmentation patterns, BGP/BFD design; VMware Cloud Foundation 9.0 Administration Guide — lifecycle management, upgrade sequencing, SDDC Manager operations.
Exam-Relevant Tools & Utilities:
vSAN Observer — real-time performance monitoring for storage design validation; NSX Traceflow — packet-level verification of DFW rules and routing; SDDC Manager API — programmatic domain deployment and lifecycle operations; Aria Operations — capacity modeling and what-if analysis for sizing decisions; VMware Compatibility Guide — hardware validation for HCL compliance; vSAN ReadyNode Sizer — capacity planning with workload profiles.
Version Compatibility Notes:
The 2V0-13.25 exam targets VCF 9.0 (released 2025). Key version dependencies: vSphere 9.0, vSAN 9.0 ESA, NSX 9.0, SDDC Manager 9.0, Aria Suite 8.x. When referencing older documentation (VVD, VCF 4.x/5.x guides), note architectural evolution: VCF 5.0 introduced VCF Import for brownfield adoption; VCF 5.1 added Async Patch Tool; VCF 9.0 moved to per-core licensing, mandated ESA, and unified the Broadcom product portfolio. Labs reference VCF 9.0 constructs but include version context showing evolution from earlier releases.
VCDX Defense Preparation:
Labs 12 and 13 specifically prepare for VCDX design defense. The design document lab (12) produces a complete design with executive summary, sizing calculations, architecture diagrams, 12-decision register, requirement traceability matrix, and risk register. The healthcare capstone (13) simulates a real customer engagement with regulatory constraints (HIPAA/HITRUST), clinical system requirements (Epic EHR, PACS), and multi-site HA/DR — the exact complexity level expected in VCDX submissions. Practice articulating each design decision in under 60 seconds with clear requirement traceability.
Key Takeaways
- Lab progression follows infrastructure-up approach: methodology → compute → storage → network → domain architecture → DR → operations → capstone integration.
- Every lab produces RCAR-formatted design decisions with requirement traceability — the exact format expected in VCDX design document submissions.
- Holodeck Toolkit enables full VCF stack validation in nested environments; stretched cluster labs require two pods with simulated WAN latency.
- Reference materials should be version-pinned to VCF 9.0 exam scope while understanding evolution from VVD/VCF 4.x/5.x for architectural context.
Business Requirements Analysis & Architecture Options (Objectives 1.1, 2.1, 3.1)#
Gathering & Analyzing Business Objectives (Obj 3.1)
Business requirements analysis is the starting point of every VCF design. The architect must translate business language into technical specifications.
Discovery Process:
├── Stakeholder interviews: C-level (budget, timeline, business drivers), IT Director (operational requirements, team skills), Application owners (workload characteristics, SLAs) ├── Current state assessment: inventory existing infrastructure, identify technical debt, document pain points ├── Business drivers documentation: cost reduction targets, compliance mandates, digital transformation goals, cloud strategy ├── SLA decomposition: translate "99.99% uptime" into concrete requirements (RPO, RTO, MTTR, maintenance windows) └── Growth projections: 1-year, 3-year, 5-year capacity planning with seasonal variation
Business vs Technical Requirements (Obj 1.1)
Business Requirements (non-technical stakeholders):
├── "Applications must be available 24x7" → translates to HA design, stretched cluster, SRM ├── "Must comply with data sovereignty regulations" → translates to data residency constraints, specific datacenter locations ├── "Reduce infrastructure costs by 30%" → translates to consolidation ratios, licensing optimization, shared infrastructure ├── "Support 500 new VMs within 6 months" → translates to capacity planning, procurement lead times └── "Enable developer self-service" → translates to VCF Automation, catalog items, approval policies
Technical Requirements (IT stakeholders):
├── CPU/memory/storage sizing for known workloads ├── Network bandwidth, latency, and throughput requirements ├── Integration with existing systems (AD, DNS, NTP, monitoring) ├── Backup and recovery infrastructure compatibility ├── Security posture (encryption, microsegmentation, compliance scanning) └── Automation and orchestration tool integration
Mapping Business to Technical:
Every business requirement must trace to one or more technical design decisions. Use a requirements traceability matrix:
| Business Req ID | Business Requirement | Technical Design Decision | Quality Attribute |
|---|---|---|---|
| BR-001 | 99.99% uptime | vSAN stretched cluster + vSphere HA + SRM | Availability |
| BR-002 | Data in-country only | Single-region deployment, no cross-border replication | Security |
| BR-003 | Developer self-service | VCF Automation with project-level tenancy | Manageability |
VCF Architecture Options (Obj 2.1)
VCF supports multiple deployment architectures depending on scale and requirements:
Standard Architecture:
├── Separate management domain (4 hosts) + one or more workload domains ├── Management domain runs: vCenter, SDDC Manager, NSX Manager, VCF Operations ├── Workload domains run customer VMs ├── Minimum: 4 management + 3 workload = 7 hosts ├── Best for: medium-large environments, multi-tenant, strong isolation └── Trade-off: higher host count, but clean separation of concerns
Consolidated Architecture:
├── Management and workload VMs share the same cluster ├── Single domain with management + workload collocated ├── Minimum: 4 hosts (all running management + workloads) ├── Best for: small environments, edge deployments, lab/dev ├── Trade-off: lower cost, but management contention risk under heavy workload └── Not recommended for production environments with >50 VMs
Multi-Site Architecture:
├── Multiple VCF instances across geographic locations ├── Options: stretched cluster (synchronous), SRM (asynchronous), federation ├── Stretched cluster: RPO=0, RTO<15min, requires <5ms RTT between sites ├── SRM replication: RPO=15min-1hr, RTO=1-4hr, works across WAN ├── Best for: disaster recovery, geographic redundancy, data sovereignty └── Trade-off: complexity + WAN costs vs. availability guarantees
Multi-AZ (Availability Zone) Architecture:
├── VCF 9.0 supports availability zone concept within single management domain ├── Each AZ maps to a fault domain (rack, room, building, site) ├── vSAN fault domains aligned with AZs ├── NSX Edge clusters can span AZs for north-south traffic resilience └── Best for: large single-site deployments with rack-level fault tolerance
Key Takeaways
- Every business requirement must trace to a technical design decision via requirements traceability matrix.
- Discovery process: stakeholder interviews → current state assessment → SLA decomposition → growth projections.
- Standard architecture (separate management/workload) is preferred for production; consolidated only for small/edge deployments.
- Architecture selection driven by: scale, isolation requirements, budget, availability targets, and compliance mandates.
- VCDX expects architects to defend architecture choice with quantified trade-offs, not just preferences.
VCF Consumption Strategy & Automation Tenant Design (Objectives 3.10, 3.11)#
Workload Migration & Onboarding Strategy (Obj 3.10)
Migration strategy defines how workloads move from source to VCF. The architect must plan migration waves, tooling, and validation.
Migration Assessment:
├── Workload inventory: categorize by type (VM, container, physical), criticality (P1-P4), complexity ├── Dependency mapping: application-to-application network flows, shared storage, database dependencies ├── Migration compatibility: HCX-compatible VMs, VMs requiring re-platform, applications needing refactoring ├── Risk categorization: low-risk (stateless web), medium (stateful app), high (database clusters) └── Success criteria per workload: performance baseline, functionality validation, rollback trigger
Migration Approaches:
├── Lift-and-shift (rehost): HCX bulk/RAV migration — fastest, minimal changes ├── Re-platform: move to VCF + modernize (e.g., move DB to vSAN, enable microsegmentation) ├── Refactor: redesign for cloud-native (Kubernetes on VKS/Supervisor) ├── Retire: decommission workloads no longer needed └── Retain: workloads that stay on legacy infrastructure (not VCF-compatible)
Wave Planning:
├── Wave 0: infrastructure (DNS, AD, monitoring) — migrated first or deployed fresh on VCF ├── Wave 1: non-production (dev/test) — validates migration process, trains team ├── Wave 2: low-criticality production — first real workloads with rollback plan ├── Wave 3+: production tiers by application group — migrate dependent apps together └── Final wave: decommission source infrastructure
Onboarding New Workloads:
├── New VMs deployed directly on VCF via vCenter or VCF Automation catalog ├── Container workloads deployed via Supervisor/VKS ├── Network onboarding: NSX segment assignment, DFW rules, load balancer config ├── Storage onboarding: vSAN storage policy assignment (FTT, encryption, compression) └── Monitoring onboarding: VCF Operations adapter auto-discovers new objects
VCF Automation Tenant Design (Obj 3.11)
VCF Automation (formerly vRealize Automation / Aria Automation) enables self-service infrastructure consumption.
Architecture Components:
├── Cloud Assembly: infrastructure provisioning engine (cloud templates, deployments) ├── Service Broker: catalog interface for consumers (catalog items, custom forms) ├── Orchestrator: workflow engine for custom automation ├── Code Stream: CI/CD pipeline integration (optional) └── Salt: configuration management (optional)
Multi-Tenant Design:
├── Organization: top-level container (typically one per VCF instance)
├── Project: tenant boundary — isolates resources, users, and policies
│ ├── Users/groups: RBAC — administrator, member, viewer roles
│ ├── Cloud zones: resource pools/clusters allocated to the project
│ ├── Network profiles: available network segments for the project
│ ├── Storage profiles: available vSAN policies for the project
│ └── Constraints: tags that control placement (e.g., "zone:prod", "tier:gold")
├── Cloud Zone: maps to vCenter cluster + resource pool + network config
│ ├── Compute: which clusters/resource pools are available
│ ├── Placement policy: default (spread), binpack (consolidate), spread (distribute)
│ └── Capability tags: match deployment constraints (e.g., "gpu:true", "region:us-east")
└── Blueprints/Cloud Templates: YAML-based infrastructure definitions
├── Resources: machines, networks, load balancers, disks
├── Inputs: parameterized values (VM size, OS, network)
├── Constraints: placement tags matching cloud zone capabilities
└── Version control: templates stored in Git, versioned, reviewedSelf-Service & Governance (Obj 3.11):
├── Catalog items: published cloud templates available to project members ├── Custom forms: UI customization for consumer-friendly deployment requests ├── Approval policies: require manager approval for large deployments (>X CPUs, >Y GB) ├── Lease policies: auto-expire deployments after X days (prevents resource sprawl) ├── Day-2 actions: resize, snapshot, power operations available to consumers ├── Quota management: limit CPU/memory/storage per project └── Cost tracking: tag-based chargeback integration with VCF Operations
Automating VCF Infrastructure (Obj 3.11):
├── Cloud templates (YAML): infrastructure-as-code for repeatable deployments ├── ABX actions (Action Based Extensibility): custom code (Python, Node.js) triggered by lifecycle events ├── Orchestrator workflows: complex multi-step automation (e.g., post-deploy AD join, monitoring setup) ├── Event subscriptions: trigger actions on deployment create/update/delete events ├── Terraform provider: manage VCF resources via HashiCorp Terraform └── REST API: full programmatic access to all VCF Automation capabilities
Modern Applications in VCF (Obj 3.11):
├── Supervisor: vSphere-integrated Kubernetes control plane
│ ├── vSphere Namespaces: resource-isolated K8s namespaces backed by vSphere
│ ├── VM Service: deploy VMs as K8s objects (VM Operator)
│ └── Storage: vSAN-backed persistent volumes via CSI driver
├── VKS (VMware Kubernetes Service): managed Kubernetes clusters
│ ├── Tanzu Kubernetes Grid (TKG) clusters provisioned via Supervisor
│ ├── Lifecycle management: cluster create, scale, upgrade via kubectl
│ └── Registry: Harbor for container image management
├── vSphere Pods: native pod execution on ESXi (CRX micro-VM per pod)
│ ├── No worker VM needed — pods run directly on hypervisor
│ ├── Best for: trusted workloads needing bare-metal-like performance
│ └── Limitation: requires NSX networking, ESA recommended
└── Design considerations:
├── Traditional VMs vs containers: choose based on workload characteristics
├── Stateful workloads: VMs or VMs-as-K8s-objects via VM Service
├── Stateless microservices: VKS clusters or vSphere Pods
└── Hybrid: mix VM and container workloads on same VCF infrastructureKey Takeaways
- Migration wave planning: infrastructure first (Wave 0), dev/test validation (Wave 1), low-criticality production (Wave 2), then critical workloads in dependency-ordered waves.
- VCF Automation tenant model: Organization → Project (tenant boundary) → Cloud Zone (resource mapping) → Cloud Templates (IaC).
- Governance pillars: approval policies (gate large deployments), lease policies (prevent sprawl), quotas (resource limits), chargeback (cost visibility).
- Modern app options on VCF: Supervisor (vSphere-native K8s), VKS (managed TKG clusters), vSphere Pods (native pod execution on ESXi), VM Service (VMs as K8s objects).
- VCDX expects consumption strategy to connect business self-service requirements to specific VCF Automation design decisions.
VCF Manageability Design — Objective 3.6 (All Sub-Objectives)#
Manageability determines how efficiently the platform can be operated, maintained, upgraded, and scaled over its lifecycle. The 2V0-13.25 blueprint Objective 3.6 tests manageability across lifecycle management, scalability, and capacity planning.
Obj 3.6.1 — Lifecycle Management Design
SDDC Manager is the lifecycle orchestrator for all VCF components. Design decisions:
Upgrade Strategy: Sequential (management domain first, then workload domains) vs parallel (multiple workload domains simultaneously). Sequential is safer — if management upgrade fails, workload domains are unaffected. Parallel reduces maintenance window duration but increases blast radius.
Bundle Management: Online (SDDC Manager pulls bundles from Broadcom depot) vs offline (bundles downloaded to local repository). Air-gapped environments require offline mode. Decision: maintain a local depot server with pre-validated bundles for controlled rollout.
Pre-check Validation: SDDC Manager runs automated pre-checks before every upgrade — compatibility matrix, disk space, health status, configuration drift. Design: build pre-check into change management workflow. If any pre-check fails, abort and remediate before proceeding.
Upgrade Sequencing for VCF 9.0: The mandatory upgrade order is: (1) SDDC Manager, (2) vCenter Server, (3) NSX Manager, (4) ESXi hosts (rolling), (5) vSAN on-disk format (if required), (6) VCF Operations and Automation. Never skip steps — SDDC Manager validates version compatibility matrix at each stage.
Rollback Strategy: vCenter supports snapshot-based rollback. NSX Manager supports cluster rollback (revert individual nodes). ESXi rolling upgrade allows stopping mid-cluster if issues arise. Design: take vCenter snapshot before upgrade, validate for 24 hours, then remove snapshot. Never run production with active snapshots beyond 72 hours (performance degradation from snapshot delta growth).
Configuration Drift Detection: SDDC Manager tracks desired state vs actual state. If an admin manually changes a setting (e.g., NTP server on ESXi), SDDC Manager flags drift. Design: enforce "no manual changes" policy — all configuration through SDDC Manager API or Automation workflows. Integrate drift detection alerts into VCF Operations dashboard.
Obj 3.6.2 — Scalability Design
Horizontal Scaling: Adding hosts to existing clusters or adding new clusters to workload domains. SDDC Manager automates host commissioning — the architect must design the process: hardware procurement lead time, firmware validation against HCL, network pre-provisioning (ToR switch ports, VLAN trunking), IP address allocation from IPAM.
Vertical Scaling: Not applicable to physical hosts. For VCF management appliances (vCenter, NSX Manager, VCF Operations), scale up by increasing vCPU/RAM allocation during maintenance window. Design: document maximum supported sizes per appliance and trigger points (e.g., "scale vCenter to Large when managed VMs exceed 5000").
Workload Domain Expansion: Adding new workload domains is a major scalability event. Design decisions: new vCenter instance (VCF 5.0+ model), new vSAN cluster, new NSX transport zone (isolated or shared), new IP subnets. Pre-plan 3-5 workload domain expansion slots in network and IP addressing design.
Cluster Scaling Limits: vSAN cluster maximum is 64 hosts. When approaching this limit, design splits into multiple clusters within the same workload domain. Decision: when to split — at 48 hosts (75% capacity) to allow growth headroom without disruption.
Multi-Instance Fleet Scaling: For large enterprises, scale by deploying additional VCF instances managed through Fleet Management. Design: federation topology (hub-spoke vs mesh), shared services (common NSX federation, shared monitoring), operational team assignment per instance.
Obj 3.6.3 — Capacity Planning Design
Capacity Monitoring Integration: VCF Operations provides capacity analytics — time-remaining projections, what-if modeling, and rightsizing recommendations. Design: configure capacity dashboards with 30/60/90-day projection windows. Set alerts at 70% utilization (plan procurement), 80% (expedite procurement), 90% (emergency action).
Growth Modeling: Capacity planning must account for organic growth (existing workloads growing) and project-driven growth (new application deployments). Design: maintain a capacity planning spreadsheet or VCF Operations custom dashboard that combines historical trend data with known project demand from the PMO.
Rightsizing Strategy: VCF Operations identifies oversized VMs (allocated resources significantly exceed actual usage). Design: implement monthly rightsizing review cycle. Create Automation workflow that generates rightsizing report and creates change tickets for approval. Target: reduce wasted capacity by 15-25% annually through rightsizing.
Storage Capacity Triggers: vSAN requires minimum 25% free space (slack space) for rebalancing and rebuild operations. Design: set procurement trigger at 65% vSAN utilization. Emergency trigger at 75%. Never exceed 80% — vSAN performance degrades significantly and rebuild operations may fail.
Capacity Reservation for HA/DR: Capacity planning must account for HA admission control overhead (typically 25% for N+1) and DR site capacity if active-passive. Design: usable capacity = total capacity × (1 - HA_reserve) × (1 - growth_buffer). Example: 100 hosts × 0.75 (HA) × 0.80 (20% growth buffer) = 60 hosts worth of workload capacity.
Automation-Driven Capacity Actions: Design closed-loop capacity management: VCF Operations detects threshold → triggers Automation workflow → provisions additional resources (add host, expand cluster) → validates health → notifies operations team. This reduces capacity response time from weeks (manual procurement) to hours (pre-staged hardware) or minutes (elastic cloud burst).
Key Takeaways
- Obj 3.6.1: Lifecycle management via SDDC Manager — sequential upgrade strategy (management first), offline bundle depot for air-gapped, mandatory pre-check validation, snapshot-based rollback for vCenter.
- Obj 3.6.2: Scalability design — horizontal (add hosts/clusters), workload domain expansion pre-planned in network design, cluster split at 48 hosts (75% of 64 max), multi-instance fleet for enterprise scale.
- Obj 3.6.3: Capacity planning — VCF Operations 30/60/90-day projections, procurement trigger at 70% utilization, vSAN never exceed 80%, rightsizing reviews reduce 15-25% waste, closed-loop automation for capacity actions.
- VCDX Defense: Panelists will ask 'how do you handle Day-2 operations at scale?' — answer with lifecycle automation, capacity triggers, and drift detection. Manageability is the quality that separates a design from a deployment.
VCF Recoverability Design — Objective 3.8 (All Sub-Objectives)#
Recoverability encompasses both business continuity (BC) and disaster recovery (DR). The 2V0-13.25 blueprint Objective 3.8 tests the architect's ability to design recovery strategies that align with business requirements.
Obj 3.8.1 — Business Continuity (BC) Design
Business Continuity vs Disaster Recovery: BC focuses on maintaining operations DURING a disruption (minimizing impact). DR focuses on restoring operations AFTER a disruption (recovery from failure). Both are required — BC reduces RTO exposure, DR provides the recovery mechanism.
BC Requirements Gathering: Interview business stakeholders to determine: Maximum Tolerable Downtime (MTD) per application — how long can the business survive without this application? Recovery Time Objective (RTO) — target time to restore service. Recovery Point Objective (RPO) — maximum acceptable data loss. Minimum Business Continuity Objective (MBCO) — minimum service level during disruption (e.g., "50% capacity is acceptable for 4 hours").
BC Design Patterns for VCF:
Pattern 1 — In-Site Resilience (BC within single datacenter): vSphere HA restarts VMs on surviving hosts after host failure. vSAN FTT ensures data survives disk/host failure. NSX DFW and Edge HA maintain network connectivity. Design: N+1 host capacity, vSAN FTT=1 minimum, dual Edge VMs. RTO: 1-5 minutes (HA restart time). RPO: 0 (no data loss for infrastructure failure).
Pattern 2 — Cross-Site Continuity (BC across datacenters): vSAN stretched cluster provides synchronous replication (RPO=0). vSphere HA spans both sites — VM restarts on surviving site automatically. NSX Tier-0 active-active provides continuous north-south connectivity. Design: identical infrastructure at both sites, < 5ms RTT, witness at third location. RTO: < 1 minute. RPO: 0. Cost: 2x infrastructure + WAN.
Pattern 3 — Degraded-Mode Operations: Design for graceful degradation when full capacity is unavailable. Example: if Site-A fails, Site-B handles only P1 workloads (50% capacity) while P2-P4 remain offline until Site-A recovers. Design: DRS resource pools with priority levels — P1 VMs get guaranteed resources, P2+ VMs are suspended if capacity is insufficient. Document MBCO per application tier.
Obj 3.8.2 — Disaster Recovery (DR) Design
DR Tier Classification: Align recovery technology to application criticality:
Tier 0 — Zero Data Loss (RPO=0, RTO<15min): Technology: vSAN stretched cluster (synchronous replication). Scope: mission-critical databases, financial systems. Cost: highest (2x infrastructure + low-latency WAN). Design: both sites sized for full workload capacity.
Tier 1 — Near-Zero Loss (RPO<15min, RTO<1hr): Technology: vSphere Replication or vSAN async replication + SRM. Scope: production applications, ERP, CRM. Cost: moderate (DR site can be smaller). Design: SRM recovery plans with automated failover sequencing.
Tier 2 — Standard Recovery (RPO<4hr, RTO<4hr): Technology: backup-based recovery (Veeam, Cohesity) with periodic replication. Scope: supporting applications, reporting systems. Cost: low (backup infrastructure only). Design: daily replication to DR site, SRM or manual failover.
Tier 3 — Best-Effort Recovery (RPO<24hr, RTO<24hr): Technology: nightly backup to offsite storage. Scope: development, test, archive workloads. Cost: minimal. Design: backup-only, no dedicated DR infrastructure.
DR Failover Orchestration: SRM recovery plans define: VM startup priority order (infrastructure first, then databases, then applications, then web tier), IP address remapping (production subnet to DR subnet), custom scripts (DNS update, load balancer reconfiguration, application validation), notification workflow (NOC alert, stakeholder communication).
DR Testing Strategy: Non-disruptive DR test monthly (SRM test failover — creates isolated copies, validates boot/application, cleans up). Full DR test quarterly (actual failover with planned downtime window). Failback test semi-annually (validate reverse replication and failback procedure). Design: document test procedures, success criteria, and results in DR test register.
Obj 3.8.3 — Management Domain Recovery
Management domain loss is the highest-impact failure scenario — all VCF operations depend on it. Design specific recovery procedures:
vCenter Recovery: Option A — vCenter HA (active/passive/witness). Automatic failover within 5 minutes. Design: anti-affinity rules for HA nodes, separate datastores. Option B — vCenter backup restore. HTTPS backup to external location, restore takes 30-60 minutes. Design: daily automated backup, quarterly restore test.
NSX Manager Recovery: 3-node cluster provides built-in HA. If all 3 nodes lost: restore from backup (configuration + policies). Design: daily NSX backup to external storage, document restore sequence (restore node-1 first, then join node-2 and node-3).
SDDC Manager Recovery: Restore from backup restores VCF configuration, domain inventory, and bundle cache. Design: daily automated backup, separate from vCenter/NSX backup schedule to avoid backup window conflicts.
Full Management Domain DR: If entire management domain is lost (catastrophic site failure): Step 1 — restore SDDC Manager from backup on DR site. Step 2 — restore vCenter from backup. Step 3 — restore NSX Manager cluster. Step 4 — reconnect to surviving workload domains. Step 5 — validate VCF health. Total estimated RTO: 2-4 hours with pre-staged DR hardware. Design: pre-deploy management domain DR hosts, maintain current backups, document and test recovery runbook quarterly.
Key Takeaways
- Obj 3.8.1: Business Continuity — maintain operations during disruption. Key metrics: MTD, RTO, RPO, MBCO. In-site resilience (HA + vSAN FTT), cross-site continuity (stretched cluster), degraded-mode operations (DRS priority pools).
- Obj 3.8.2: DR tier classification — Tier 0 (RPO=0, stretched cluster), Tier 1 (RPO<15min, SRM + async replication), Tier 2 (RPO<4hr, backup replication), Tier 3 (RPO<24hr, backup only). Each tier maps to application criticality and cost.
- Obj 3.8.3: Management domain recovery is highest priority — vCenter HA or backup restore, NSX 3-node cluster HA, SDDC Manager backup. Full management DR RTO: 2-4 hours with pre-staged hardware.
- VCDX Defense: Panelists will ask 'what happens when your management domain is lost?' — you must have a documented recovery sequence with realistic RTO estimates. Also expect 'how do you test DR without impacting production?' — answer with SRM non-disruptive test capability.
Modern Applications Architecture on VCF — Objective 3.11 (Design Deep-Dive)#
VCF 9.0 provides a unified platform for both traditional VMs and modern containerized workloads. The 2V0-13.25 blueprint Objective 3.11 tests the architect's ability to design Kubernetes and container infrastructure on VCF.
Supervisor Architecture Design
The Supervisor is the vSphere-integrated Kubernetes control plane. It runs as a set of VMs on ESXi hosts within a workload domain cluster.
Supervisor Deployment Model: Supervisor is enabled per vSphere cluster. Each Supervisor gets 3 control plane VMs (master nodes) distributed across hosts via anti-affinity. Design decisions: which clusters get Supervisor enabled (not all clusters need Kubernetes), sizing of control plane VMs (Small: 2 vCPU/8GB, Medium: 4 vCPU/16GB, Large: 8 vCPU/32GB — choose based on expected namespace count and API request volume).
Networking for Supervisor: Two options — NSX networking (full-featured, recommended) or vSphere Distributed Switch networking (VDS, simpler). NSX networking provides: per-namespace network isolation, NSX load balancer for Kubernetes Services, DFW micro-segmentation for pods, ingress/egress firewall policies. VDS networking provides: basic networking via HAProxy or Avi load balancer, simpler setup but less isolation. Design decision: NSX for production multi-tenant environments, VDS for single-team development clusters.
Storage for Supervisor: vSAN provides persistent storage via the vSphere CSI (Container Storage Interface) driver. Design decisions: StorageClass definitions map to vSAN storage policies. Example: "gold" StorageClass = vSAN FTT=2, RAID-6, encryption enabled; "silver" StorageClass = vSAN FTT=1, RAID-5; "bronze" StorageClass = vSAN FTT=1, RAID-1, no encryption.
vSphere Namespace Design: Namespaces are the tenant boundary on Supervisor. Each namespace gets: resource limits (CPU/memory/storage quotas), network isolation (NSX segment per namespace), storage quotas (per-StorageClass limits), RBAC (mapped to AD/LDAP groups).
Design Pattern — Namespace-per-Team: Create one namespace per development team. Each team gets: resource quota (e.g., 32 vCPU, 128GB RAM, 500GB storage), permission to deploy TKG clusters and VMs within their namespace, isolated network segment, access to approved StorageClasses. This mirrors the VCF Automation project model for VM consumers.
VKS (VMware Kubernetes Service) Design
VKS replaces Tanzu Kubernetes Grid (TKG) as the managed Kubernetes offering on VCF 9.0.
TKG Cluster Architecture: TKG clusters are deployed within vSphere Namespaces. Each TKG cluster has: control plane nodes (1 or 3, depending on HA requirement), worker nodes (1-N, auto-scalable), all running as VMs on ESXi via the Supervisor.
Cluster Sizing Design: Small development cluster: 1 control plane + 3 workers (8 vCPU, 16GB each). Production cluster: 3 control plane + 6-12 workers (16 vCPU, 32GB each). Design decision: worker node VM class selection — map to workload requirements (CPU-intensive, memory-intensive, GPU-enabled).
Multi-Cluster Strategy: Design decision: single large cluster vs multiple smaller clusters. Single large cluster: simpler operations, namespace isolation within cluster, risk of blast radius. Multiple clusters: environment isolation (dev/staging/prod), workload isolation (batch vs real-time), upgrade independence. Recommendation: separate clusters per environment, shared registry and CI/CD pipeline.
VM Service — VMs as Kubernetes Objects
VM Service allows deploying traditional VMs through the Kubernetes API using VM Operator. This bridges the gap between VM and container operations.
Use Cases: Legacy applications that cannot be containerized but need to be managed alongside containers. Database VMs that provide backend services to containerized frontends. Windows workloads in a predominantly Linux/container environment.
Design: Define VM Classes that map to hardware profiles (CPU/memory combinations). Create ContentLibrary sources for VM images (OVA/OVF templates). Deploy VMs using kubectl with YAML manifests — same GitOps workflow as containers.
Container Registry Design (Harbor)
Harbor provides container image management for VCF Kubernetes environments.
Deployment Model: Embedded Harbor (deployed automatically with Supervisor) for small environments. External Harbor (standalone deployment) for enterprise environments needing HA, multi-site replication, and advanced scanning.
Image Security: Harbor integrates vulnerability scanning (Trivy). Design: enforce "no deploy without scan" policy — Kubernetes admission controller rejects images with critical CVEs. Image signing with Notary/Cosign for supply chain security.
Design Pattern — Registry Hierarchy: Global registry (Harbor) hosts golden images (base OS, middleware). Team registries (project-level) host application images built from golden bases. Production promotion: images promoted from dev → staging → prod registries with approval gate.
Network Design for Modern Apps
Kubernetes networking on VCF requires careful design:
Pod Networking: With NSX, each pod gets an IP from the namespace's NSX segment. Pod-to-pod communication within namespace: direct L2. Cross-namespace: routed through Tier-1 gateway with DFW policy enforcement.
Service Networking: Kubernetes Services (ClusterIP, NodePort, LoadBalancer) mapped to NSX or Avi load balancer. Design: use NSX load balancer for L4 services, Avi for L7 (HTTP path-based routing, WAF, SSL termination).
Ingress Design: NSX or Avi Ingress controller. Design: single ingress controller per cluster, wildcard TLS certificate, path-based routing to backend services. For multi-cluster: use Avi GSLB for cross-cluster load balancing.
Network Policy: Kubernetes NetworkPolicy objects enforced by NSX DFW. Design: default-deny ingress per namespace, explicit allow rules per application. This extends the zero-trust model from VMs to containers.
Key Takeaways
- Obj 3.11: Supervisor = vSphere-integrated K8s control plane. 3 control plane VMs per cluster. NSX networking (production, multi-tenant) vs VDS (simple, single-team). vSphere CSI for persistent storage.
- Namespace-per-team design pattern mirrors VCF Automation project model. Each namespace gets: resource quota, network isolation (NSX segment), storage quota, RBAC mapped to AD/LDAP.
- VKS cluster strategy: separate clusters per environment (dev/staging/prod) for upgrade independence and blast radius containment. VM Service bridges VM and container operations via kubectl.
- Container security: Harbor registry with Trivy scanning, admission controller blocks unscanned images, image signing for supply chain security. Registry hierarchy: global golden images → team registries → production promotion.
- VCDX Defense: Expect 'how do you handle mixed VM and container workloads?' — answer with Supervisor + VM Service + VKS on shared VCF infrastructure, unified monitoring via VCF Operations, NSX providing consistent network policy across both.
VCF Monitoring & Observability Design — Objective 3.12 (All Sub-Objectives)#
Monitoring design ensures visibility into platform health, workload performance, and operational compliance. The 2V0-13.25 blueprint Objective 3.12 tests the architect's ability to design monitoring for both management components and workloads.
Obj 3.12.1 — Management Component Monitoring
VCF management components are the foundation — if monitoring fails to detect management plane issues, all downstream workloads are at risk.
SDDC Manager Monitoring: Monitor SDDC Manager health via API (/v1/system/health). Key metrics: task queue depth (stuck tasks indicate lifecycle issues), bundle download status, certificate expiry warnings, configuration drift count. Design: VCF Operations adapter for SDDC Manager, alerts on task failures and certificate expiry (90-day warning, 30-day critical).
vCenter Monitoring: Key metrics: VPXD service health, SSO authentication latency, database size and performance, managed object count (VMs, hosts, clusters). Critical alerts: VPXD service restart, SSO token failures, database backup failures. Design: VCF Operations vCenter adapter + custom dashboards for admin team. Monitor vCenter HA state (active/passive/witness) — alert if failover occurs.
NSX Manager Monitoring: Key metrics: cluster health (all 3 nodes UP), API response latency, control plane connectivity to transport nodes, certificate validity. Critical alerts: NSX Manager node failure (cluster degrades from 3→2), transport node disconnection (host loses overlay networking), DFW rule publish failures. Design: NSX adapter in VCF Operations, correlated alerts (NSX node down → check host connectivity → check network health).
ESXi Host Monitoring: Key metrics: CPU ready time (%RDY — performance indicator), memory ballooning/swapping (contention indicator), vSAN disk health and latency, network packet loss/errors, hardware health (sensors via IPMI/iLO/iDRAC). Critical alerts: host disconnected, hardware predictive failure (Proactive HA trigger), vSAN disk failure, network link down. Design: per-host health scorecard in VCF Operations, aggregated cluster health view.
vSAN Monitoring: Key metrics: cluster health (vSAN Health Service), resync/rebalance status, object compliance (all objects meeting storage policy), disk latency (read/write), cache hit ratio (OSA), capacity utilization. Critical alerts: object non-compliance (data at risk), resync stalled, capacity > 80%, disk failure. Design: vSAN adapter in VCF Operations + vSAN Health Service native alerts. Include vSAN performance trending for capacity planning.
Obj 3.12.2 — Workload Monitoring
Application-Level Monitoring: VCF Operations monitors infrastructure metrics (CPU, memory, storage, network) for each VM. For deeper application insight, integrate with application performance monitoring (APM): Telegraf agent in VMs exporting custom metrics, VCF Operations custom metric ingestion via REST API, correlation between infrastructure metrics and application KPIs (e.g., high CPU → slow response time).
Kubernetes Workload Monitoring: For Supervisor and VKS clusters: pod/container resource utilization (CPU/memory requests vs limits vs actual), persistent volume capacity and IOPS, Kubernetes API server health, etcd cluster health and latency, node readiness and conditions. Design: VCF Operations Kubernetes adapter + Prometheus integration for custom application metrics. Dashboard: cluster overview → namespace drill-down → pod detail.
Monitoring Topology Design:
Centralized Model: Single VCF Operations instance monitors all VCF instances, workload domains, and applications. Advantages: single pane of glass, unified alerting, simplified operations. Disadvantages: single point of failure for monitoring, potential performance bottleneck for large environments (>10,000 VMs). Design: deploy VCF Operations in HA mode (2-node minimum), place on management domain for maximum reliability.
Distributed Model: VCF Operations instance per VCF instance or per geographic site. Central analytics node aggregates data from distributed collectors. Advantages: monitoring survives site failure, lower WAN bandwidth for metrics, local team autonomy. Disadvantages: more complex operations, potential inconsistency in alert definitions. Design: remote collector nodes at each site, central analytics in primary datacenter.
Hybrid Model (Recommended for Enterprise): Central VCF Operations cluster in primary datacenter. Remote collectors at each remote site/VCF instance. Cloud proxy for Broadcom SaaS analytics (if internet-connected). Design: central cluster sized for total VM count, remote collectors sized for local VM count.
Obj 3.12.3 — Alerting & Escalation Design
Alert Classification: Informational (no action, logging only) → Warning (investigation needed within 4 hours) → Critical (immediate action required, SLA at risk) → Emergency (service down, escalation to management).
Alert Correlation: VCF Operations supports alert correlation — multiple related alerts grouped into a single incident. Design: define correlation rules to reduce alert noise. Example: if "ESXi host disconnected" + "vSAN object non-compliant" + "VM powered off" all fire within 5 minutes → correlate to single incident "Host failure with data impact."
Notification Channels: Email (all severities), Slack/Teams webhook (warning and above), PagerDuty/OpsGenie (critical and emergency), ServiceNow incident creation (critical and emergency). Design: notification routing by team — infrastructure alerts to platform team, application alerts to application owners, security alerts to SOC.
Escalation Matrix Design:
Level 1 (L1) — NOC On-Call: Responds to warnings within 30 minutes. Runs documented remediation runbooks. Escalates to L2 if not resolved within 1 hour.
Level 2 (L2) — Platform Engineer: Responds to critical alerts within 15 minutes. Deep troubleshooting using VCF Operations, esxtop, vSAN Observer, NSX Traceflow. Escalates to L3 if not resolved within 2 hours.
Level 3 (L3) — Architect/Vendor Support: Responds to emergency alerts immediately. Engages Broadcom GSS (Global Support Services) for product defects. Coordinates with application teams for cross-functional issues. Post-incident review and design improvement.
Obj 3.12.4 — Compliance & Audit Monitoring
CIS Benchmark Monitoring: VCF Operations includes CIS compliance dashboards for ESXi 8.x hardening benchmark. Design: enable CIS benchmark content pack, schedule weekly compliance scans, alert on any non-compliant findings. Track compliance score trending (target: >95% compliance).
Audit Log Design: All VCF components generate audit logs — vCenter events, NSX audit log, SDDC Manager task log, ESXi shell access log. Design: centralize all logs to VCF Operations for Logs (or external SIEM — Splunk, QRadar, Elastic). Retention: minimum 1 year for compliance (PCI-DSS, HIPAA), 90 days online searchable, remainder in cold storage.
Change Tracking: Monitor for unauthorized changes — ESXi configuration changes outside SDDC Manager, NSX DFW rule modifications, vCenter permission changes. Design: VCF Operations drift detection alerts + integration with change management system (ServiceNow). Any change without an approved change ticket triggers security alert.
Dashboard Design for Compliance Officers: Executive compliance dashboard showing: overall compliance posture (% compliant), top non-compliant findings with remediation status, audit readiness score, certificate expiry timeline, encryption coverage (% of VMs/datastores encrypted). Design: read-only access for compliance team, scheduled PDF reports (weekly), drill-down capability to specific control failures.
Key Takeaways
- Obj 3.12.1: Management component monitoring — SDDC Manager (task queue, cert expiry), vCenter (VPXD health, HA state), NSX (cluster health, transport node connectivity), ESXi (CPU %RDY, hardware health), vSAN (object compliance, capacity, latency).
- Obj 3.12.2: Workload monitoring — infrastructure metrics via VCF Operations adapters, application metrics via Telegraf/custom ingestion, Kubernetes monitoring via K8s adapter + Prometheus integration.
- Obj 3.12.3: Alerting design — 4-level classification (info/warning/critical/emergency), alert correlation reduces noise, notification routing by team, L1/L2/L3 escalation matrix with defined response times.
- Obj 3.12.4: Compliance monitoring — CIS benchmark scanning, audit log centralization (1-year retention minimum), configuration drift detection, executive compliance dashboards with remediation tracking.
- VCDX Defense: Panelists will ask 'how do you know when something is wrong?' — answer with layered monitoring (management → infrastructure → workload → application), correlated alerting, and closed-loop remediation via Automation integration.
Exam Mapping: 2V0-13.25 — VCP-VCF 9.0 Architect
- See VCP-VCF 9.0 Architect exam blueprint for detailed objectives
Labs in This Section
Design Decision Documentation — RCAR Framework & Decision Register
VCF 9.0IntermediateCompute Cluster Design with HA Admission Control
VCF 9.0IntermediatevSAN Storage Design with Erasure Coding
VCF 9.0IntermediateMulti-Site vSAN Stretched Cluster Design
VCF 9.0AdvancedNSX Micro-Segmentation Design
VCF 9.0AdvancedMulti-Workload Domain Design
VCF 9.0AdvancedDisaster Recovery Architecture Design
VCF 9.0AdvancedHA Admission Control Validation
VCF 9.0IntermediateZero-Trust DFW Design for 3-Tier App
VCF 9.0AdvancedMulti-Site Design Decision Framework
VCF 9.0AdvancedSelf-Service Portal Design (Catalog + Projects)
VCF 9.0IntermediateWrite a VCDX-Grade Design Document (Condensed)
VCF 9.0AdvancedEnd-to-End Architecture Design Scenario
VCF 9.0Advanced📝 Quiz — VCF 9.0 Architect
Section 1 — Architecture Design Methodology
- A bill of materials listing specific server SKUs, firmware versions, and switch part numbers
- A high-level diagram aligned to business goals showing only capability boxes without technology
- A technology-specific diagram showing vCenter, NSX Manager, and vSAN relationships but no specific vendor hardware
- A spreadsheet of business requirements translated into user stories
- Requirement
- Assumption
- Constraint
- Risk
- Availability
- Manageability
- Performance
- Recoverability
- Decision, cost, vendor
- Justification, implications, and alternatives considered
- Owner, date, and approver only
- A diagram and a configuration file
- A functional vs. non-functional requirement conflict
- A risk register entry
- An AMPRS trade-off between Availability and a cost constraint
- A validation checklist failure