Academy/VCP-VCF 9.0 Architect (2V0-13.25)/Write a VCDX-Grade Design Document (Condensed)
This lab targets VCF 9.0

Write a VCDX-Grade Design Document (Condensed)

VCF 9.0Advancedvcp-foundation⏱ 120 min

VCDX design document structure — executive summary, sizing, decisions, diagrams, and risk management

Objectives

  • Write a VCDX-quality executive summary that communicates design intent to non-technical stakeholders
  • Create detailed sizing calculations for compute, memory, storage, and network
  • Build a comprehensive design decision register with alternatives and traceability
  • Develop logical and physical architecture diagrams
  • Construct a risk register with quantified probability, impact, and mitigation strategies

Prerequisites

Completed labs vcp-architect-01 through -06 (design exercises)

Prior labs: vcp-architect-01, vcp-architect-02, vcp-architect-03, vcp-architect-04, vcp-architect-05, vcp-architect-06

Required skills:

  • All VCF architecture design skills from prior labs
  • Technical writing for mixed audiences
  • Visio/Draw.io diagramming (or text-based equivalent)
  • Risk assessment methodology

Lab Environment

Documentation exercise — produces a 15-20 page condensed design document using all prior lab outputs

Tasks

Task 1 Executive Summary & Scope Definition

The executive summary is the first thing VCDX panelists read. It must convey architectural maturity in 500 words — the problem being solved, the approach, and the key trade-offs. If the summary is generic, panelists assume the design is generic.

Write the opening sections of a VCDX design document that communicate the design intent, scope, and approach to both technical and business audiences.

Step 1

Write the Executive Summary (500-600 words):

Structure:

  1. Business context (2-3 sentences): What is the business challenge?

Example: 'GlobalFinance Corp operates 600+ virtual workloads across aging infrastructure with increasing compliance requirements (SOC 2 Type II). The current environment lacks standardized disaster recovery, has inconsistent security policies, and faces capacity constraints that limit business growth.'

  1. Solution overview (3-4 sentences): What are we building?

Example: 'This design deploys VMware Cloud Foundation 9.0 across two data center sites with a tiered architecture that aligns infrastructure investment to business criticality. Mission-critical Tier-1 workloads are protected by vSAN stretched cluster (RPO=0), while Tier-2 workloads use VMware Site Recovery with 15-minute RPO. NSX micro-segmentation enforces zero-trust security with tag-based Distributed Firewall policies.'

  1. Key design decisions (3-4 sentences): What are the most important trade-offs?

Example: 'The design uses two workload domains (Production and Development) to limit SOC 2 audit scope while minimizing vCenter licensing overhead. vSAN ESA with erasure coding provides 50% more usable storage than mirroring at equivalent protection. The self-service portal via Aria Automation reduces VM provisioning from 5 days to 10 minutes while enforcing governance through catalog-driven deployment.'

  1. Outcome (2-3 sentences): What does success look like?

Example: 'This architecture provides a 3-year platform supporting 20% annual growth, with full compliance attestation, automated DR capabilities, and a consumption-based operating model. Total infrastructure investment: $X for primary + $Y for DR, representing a Z% reduction in per-VM cost versus the current environment.'

Step 2

Write the Scope & Assumptions section (300-400 words):

  1. In Scope:
  • VCF 9.0 deployment: Management domain + 2 workload domains
  • vSAN ESA storage with tiered protection policies
  • NSX micro-segmentation with zero-trust DFW
  • Multi-site DR: Stretched cluster (Tier-1), SRM (Tier-2), backup (Tier-3)
  • Self-service portal with Aria Automation
  • Monitoring and operations with Aria Operations suite
  1. Out of Scope:
  • Physical data center construction (facilities assumed ready)
  • Application migration planning (separate workstream)
  • End-user device management
  • Third-party security tools (SIEM, EDR) — integration points documented only
  • Network outside the VCF overlay (physical switches, WAN, internet)
  1. Assumptions (reference RCAR from Lab 01):
  • List top 10 assumptions with cross-reference to full RCAR document
  • Flag the 3 highest-risk assumptions (ones most likely to be wrong)
  • Example high-risk assumption: 'Inter-site latency measured at 0.8ms RTT will remain stable under production load. If latency exceeds 5ms, stretched cluster design must be revised to SRM.'
  1. Dependencies:
  • Physical network: Spine-leaf fabric with BGP, MTU 9000 end-to-end
  • Identity: Active Directory with LDAPS available at both sites
  • Certificates: Enterprise PKI for VMCA subordinate CA
  • Procurement: Hardware on VCF 9.0 HCL (ordered and delivery confirmed)
Step 3

Write the Design Principles section (200-300 words):

Design principles guide every subsequent decision. List 6-8 principles:

  1. Security by default: Every VM is encrypted, tagged for DFW, and deployed from hardened golden images. Security is embedded in the platform, not bolted on.
  1. Failure domain isolation: Management, production, and development workloads are isolated by workload domain, transport zone, and storage. A failure in one domain does not propagate.
  1. Cost-aligned protection: DR investment matches business criticality. Tier-1 (RPO=0) costs 2× more than Tier-2 (RPO=15min) and 5× more than Tier-3 (backup).
  1. Automation-first operations: VM provisioning, patching, and monitoring are automated. Manual intervention is reserved for exceptions and incident response.
  1. Lifecycle-aware sizing: All capacity is sized for 3-year growth at 20% p.a. with N+1 HA. Quarterly capacity reviews trigger procurement before exhaustion.
  1. Compliance-native architecture: SOC 2 controls are implemented at the platform level (encryption, logging, access control, network segmentation). Compliance evidence is generated automatically.
  1. Operational simplicity: When two designs offer similar protection, the simpler one is preferred. Complexity adds operational risk.
  1. Defense in depth: No single control is relied upon. Network security = DFW + physical firewall. Data protection = encryption + access control + backup. Availability = HA + DR + monitoring.

Principle traceability: Each design decision in the register should reference which principles it upholds.

Step 4

Create a document outline (table of contents):

  1. Executive Summary
  2. Scope, Assumptions & Dependencies
  3. Design Principles
  4. Requirements (RCAR Framework)

4.1 Business Requirements
4.2 Technical Requirements
4.3 Constraints
4.4 Assumptions
4.5 Risks

  1. Architecture Overview

5.1 Conceptual Architecture
5.2 Logical Architecture
5.3 Physical Architecture

  1. Compute Design

6.1 Cluster Topology
6.2 Sizing Calculations
6.3 HA & DRS Configuration
6.4 EVC Mode

  1. Storage Design

7.1 vSAN ESA Architecture
7.2 Capacity Planning
7.3 Storage Policies
7.4 Encryption & Key Management

  1. Network Design

8.1 NSX Topology (Tier-0/Tier-1)
8.2 Segment Design & CIDR Plan
8.3 DFW Micro-Segmentation
8.4 BGP & Edge Design

  1. Multi-Site Design

9.1 Stretched Cluster
9.2 SRM Configuration
9.3 Failover/Failback Procedures
9.4 DR Testing Calendar

  1. Operations & Automation

10.1 Self-Service Portal
10.2 Monitoring & Alerting
10.3 Upgrade Lifecycle
10.4 Backup & Recovery

  1. Design Decision Register
  2. Risk Register
  3. Appendices

A. Detailed Sizing Worksheets
B. DFW Rule Set
C. VLAN & IP Address Plan
D. Bill of Materials

Target: 15-20 pages (excluding appendices). Each section 1-2 pages with diagrams.

Validation Gate

Check: Executive summary, scope, and principles written with design document outline

Expected: 500-600 word executive summary, scope with in/out/assumptions/dependencies, 6-8 design principles, complete table of contents

Common Errors

Writing a generic executive summary that could apply to any VCF deployment — personalize with specific business context, workload counts, and trade-offs
Omitting out-of-scope items — panelists will challenge boundaries; explicit out-of-scope prevents scope creep
Not having design principles — without them, design decisions appear arbitrary rather than principled
Making the document too long — VCDX panelists review hundreds of pages; concise documents that communicate key decisions are valued over verbose descriptions

Task 2 Sizing Calculations & Architecture Diagrams

Sizing sections must show the math, not just the results. Architecture diagrams must be layered: conceptual (business view), logical (component relationships), and physical (hardware placement).

Create the detailed sizing calculations and architecture diagrams that form the technical backbone of the VCDX design document.

Step 1

Create compute sizing summary (consolidate from Lab 02):

Compute Sizing Worksheet:

ItemTier-1Tier-2Tier-3ManagementTotal
VM Count (current)8030022015615
VM Count (3-year)138518380201,056
vCPU per VM (avg)4224-
Total vCPU5521,036760802,428
CPU Overcommit3:15:18:12:1-
Physical Cores Needed1842089540527
RAM per VM (avg)16 GB8 GB4 GB16 GB-
Total RAM2,208 GB4,144 GB1,520 GB320 GB8,192 GB
RAM Overcommit1.25:11.5:12:11:1-
Physical RAM Needed1,766 GB2,763 GB760 GB320 GB5,609 GB
Hosts (64 cores, 512 GB)352111
+ HA (N+1)463215
+ VCF Minimum (4 mgmt)463417
+ Stretch (2× Tier-1)863421

Governing constraint per tier: Document whether CPU or RAM is the limiting factor.
Licensing impact: 21 hosts × 64 cores × 2 sockets = 2,688 licensed cores total.

Step 2

Create storage sizing summary (consolidate from Lab 03):

Storage Sizing Worksheet:

ItemTier-1Tier-2/3ManagementTotal
VM Storage (current)16 TB67 TB3.7 TB86.7 TB
VM Storage (3-year)27.6 TB115.8 TB5 TB148.4 TB
FTT SchemeFTT=2 RAID-6FTT=1 RAID-5FTT=1 RAID-5-
FTT Multiplier1.5×1.33×1.33×-
Raw for Protection41.4 TB154 TB6.65 TB202 TB
+ Slack (30%)59.1 TB220 TB9.5 TB288.6 TB
+ Metadata (2%)60.3 TB224.4 TB9.7 TB294.4 TB
Dedup/Comp Ratio2:12.5:11.5:1-
Raw Needed (post dedup)30.2 TB89.8 TB6.5 TB126.5 TB
Disks/Host (3.84 TB NVMe)442-
Raw per Host15.36 TB15.36 TB7.68 TB-
Hosts (storage sizing)2619

Note: Compute sizing drives more hosts than storage (17 vs 9), so compute is the governing constraint.
Storage headroom with compute-driven host count: Significant (~2× surplus).

Step 3

Create the conceptual architecture diagram (text-based):

╔═══════════════════════════════════════════════════════════════╗
║                 CONCEPTUAL ARCHITECTURE                       ║
╠═══════════════════════════════════════════════════════════════╣
║                                                               ║
║  [Business Users]                                             ║
║       │                                                       ║
║       ▼                                                       ║
║  ┌──────────────────────────────────────┐                     ║
║  │      Self-Service Portal             │                     ║
║  │   (Aria Automation Catalog)          │                     ║
║  └──────────────┬───────────────────────┘                     ║
║                 │                                             ║
║  ┌──────────────▼───────────────────────┐                     ║
║  │      VCF 9.0 Platform                │                     ║
║  │  ┌─────────┐ ┌─────────┐ ┌────────┐  │                     ║
║  │  │Compute  │ │Storage  │ │Network │  │                     ║
║  │  │vSphere  │ │vSAN ESA │ │NSX     │  │                     ║
║  │  └─────────┘ └─────────┘ └────────┘  │                     ║
║  └──────────────┬───────────────────────┘                     ║
║                 │                                             ║
║  ┌──────────────▼───────────────────────┐                     ║
║  │      Workload Tiers                   │                     ║
║  │  [Tier-1: RPO=0]  [Tier-2: RPO=15m]  │                     ║
║  │  [Tier-3: Backup]  [Dev/Test]        │                     ║
║  └──────────────┬───────────────────────┘                     ║
║                 │                                             ║
║  ┌──────────────▼───────────────────────┐                     ║
║  │      Multi-Site DR                    │                     ║
║  │  [Site A]◄──Sync──►[Site B]          │                     ║
║  │              [Witness C]              │                     ║
║  └──────────────────────────────────────┘                     ║
╚═══════════════════════════════════════════════════════════════╝

Conceptual diagram purpose: Communicate the design intent to non-technical stakeholders (CIO, CFO). No IP addresses, no specific products — just capabilities and relationships.

Step 4

Create the logical architecture diagram (text-based):

┌─────────────────── Site A (Primary) ───────────────────┐
│                                                         │
│  ┌── Management Domain ──┐  ┌── WD-1 (Production) ────┐│
│  │ vCenter-mgmt           │  │ vCenter-prod             ││
│  │ SDDC Manager            │  │ ┌─Tier-1 Cluster (5h)──┐││
│  │ NSX Mgr (3-node)       │  │ │ 80 VMs, FTT=2        │││
│  │ Aria Suite              │  │ │ HA 25%, DRS Full     │││
│  │ 4 hosts, FTT=1 RAID-5  │  │ └───────────────────────┘││
│  └─────────────────────────┘  │ ┌─Tier-2 Cluster (8h)──┐││
│                               │ │ 300 VMs, FTT=1       │││
│  ┌── WD-2 (Dev/Test) ─────┐  │ │ HA 20%, DRS Full     │││
│  │ vCenter-devtest          │  │ └───────────────────────┘││
│  │ 4 hosts, FTT=1 RAID-5  │  └──────────────────────────┘│
│  │ 220 VMs, Force Prov OK  │                              │
│  └──────────────────────────┘                              │
│                                                         │
│  [Tier-0 Edge]──BGP──[Physical Network]                 │
│  [DFW: Zero-Trust, Tag-Based]                           │
└───────────┬─────────────────────────────────────────────┘
            │ 100Gbps Dark Fiber (0.8ms RTT)
            │ vSAN Sync Replication (Tier-1)
            │ vSphere Replication (Tier-2, 15min RPO)
┌───────────▼─────────────────────────────────────────────┐
│               Site B (DR)                                │
│  ┌─Tier-1 Mirror (8h)──┐  ┌─Tier-2 DR (4h)────────────┐│
│  │ Stretched cluster    │  │ SRM recovery targets       ││
│  │ (shared with Site A) │  │ Placeholder VMs            ││
│  └──────────────────────┘  └────────────────────────────┘│
└───────────┬─────────────────────────────────────────────┘
            │ 1Gbps MPLS (2.5ms RTT)
┌───────────▼─────┐
│  Site C (Witness) │
│  Medium OVA       │
│  ≤21,833 comp     │
└───────────────────┘

Logical diagram purpose: Show component relationships, domain boundaries, and data flow paths. Used by architects and senior engineers.

Validation Gate

Check: Sizing worksheets and architecture diagrams created

Expected: Compute and storage sizing tables with full math, conceptual diagram for business audience, logical diagram with domain boundaries and data flows

Common Errors

Showing only final host counts without the derivation math — VCDX panelists will ask 'how did you get to 21 hosts?'
Mixing diagram levels — putting IP addresses on a conceptual diagram or omitting components from a logical diagram
Not showing 3-year growth projections — sizing only for current workload ignores lifecycle planning
Forgetting to identify the governing constraint — is the host count driven by CPU, RAM, or storage?

Task 3 Design Decision Register (Complete)

The design decision register is arguably the most important artifact in a VCDX submission. Each decision must show: what was decided, what alternatives were considered, why they were rejected, and which requirements drove the choice.

Compile the comprehensive design decision register that consolidates all decisions from prior labs into a single, traceable artifact.

Step 1

Compile all design decisions into a register:

IDCategoryDecisionAlternatives RejectedPrimary RequirementRisk
D-001Compute4-node mgmt + split workload clustersConsolidated, 3-node mgmt, 5-node mgmtSOC 2 isolation, N+1 HAOver-provisioned mgmt
D-002StoragevSAN ESA FTT=1 RAID-5 for workloadOSA, FTT=2, external FCCost optimization, capacity efficiencyLower IOPS than mirroring
D-003StorageFTT=2 RAID-6 for Tier-1FTT=1 RAID-5, FTT=2 MirrorTier-1 SLA, dual-failure toleranceHigher capacity overhead
D-004NetworkNSX DFW zero-trust, tag-basedPerimeter only, host firewallSOC 2 CC6.6, east-west securityDFW complexity, rule sprawl
D-005Compute3 clusters (Mgmt + T1 + T2/3)1 cluster, 4 clustersSOC 2 isolation, cost balanceMgmt overhead
D-006SecurityvSAN encryption + external KMS (prod), NKP (dev)VM encryption, SED, no encryptionSOC 2 CC6.1, complianceKMS as SPOF
D-007OperationsEnsure accessibility as default maint. modeFull migration, no migrationBalance speed vs safetyReduced FTT during maint.
D-008Multi-SiteStretched cluster for Tier-1SRM, HCX DR, Pilot LightRPO=0 requirement2× host cost premium
D-009NetworkDFW + physical FW (defense-in-depth)DFW only, perimeter onlyZero-trust, complianceDual management
D-010Domain2 workload domains (Prod + Dev)1 domain, 3 domainsSOC 2 scope reductionAdditional vCenter license
D-011Multi-SiteTiered DR (Stretched/SRM/Backup)All stretched, all SRMCost optimizationComplexity of mixed DR
D-012OperationsAria Automation self-serviceDirect vCenter, ServiceNow onlyDeveloper velocityAria license cost

Each row links to the detailed design decision write-up in the relevant section of the document.

Step 2

Create a requirement traceability matrix:

Purpose: Every design decision must trace to at least one requirement. If a decision exists without traceability, it is either unjustified (remove it) or has an unrecorded requirement (add it).

Requirement IDRequirementDesign DecisionsSection
R-001SOC 2 complianceD-004, D-006, D-009, D-0104.1
R-002RPO=0 for Tier-1D-003, D-0084.1
R-0033-year lifecycle growthD-001, D-002, D-0054.2
R-004VM provisioning <10 minD-0124.2
R-005Management domain RTO <4hD-001, D-0074.2
C-001VCF per-core licensingD-001, D-002, D-005, D-0104.3
C-002vSAN ESA (NVMe only)D-002, D-0034.3
C-003Inter-site RTT ≤5msD-0084.3
C-004Mgmt domain min 4 hostsD-0014.3

Coverage check:

- Every requirement has at least 1 design decision → ✅
- Every design decision traces to at least 1 requirement → ✅
- No orphaned decisions or requirements → ✅
Step 3

Write detailed justification for the 3 most impactful decisions:

Decision D-008 (Stretched Cluster) — Deep justification:

Context: Customer requires RPO=0 for 80 Tier-1 financial trading VMs. Any data loss during a site failure could result in regulatory violations (SOC 2) and financial loss ($X million per hour of trading data loss).

Decision: vSAN stretched cluster with synchronous replication.

Alternatives evaluation:

  1. SRM with RPO=5min: Minimum data loss = 5 minutes of transactions. For trading workloads processing $Y million/hour, this represents unacceptable financial risk. Rejected.
  2. Array-based replication: Near-zero RPO possible but violates single-vendor constraint (C-001). Adds array management complexity. Rejected.
  3. HCX DR: RPO=5min minimum, same limitation as SRM. Rejected.

Risk acceptance: Stretched cluster adds 30-40% infrastructure cost for Tier-1 workloads. This cost is justified by: (a) regulatory requirement for RPO=0, (b) financial loss avoidance that exceeds infrastructure premium within 1 trading day, (c) automatic failover reducing RTO to seconds vs 30-60 minutes for SRM.

Dependencies: Inter-site RTT ≤5ms (measured 0.8ms), dark fiber connectivity, witness at third site. If any dependency fails, fallback to SRM with accepted RPO=5min.

[Repeat similar depth for D-004 (DFW) and D-010 (Workload Domains)]

Step 4

Validate the design decision register:

Quality checks:

  1. Completeness: Every major architecture choice has a DDR entry? Review checklist:
   - [ ] Cluster topology → D-001, D-005
   - [ ] Storage architecture → D-002, D-003
   - [ ] Network security → D-004, D-009
   - [ ] Encryption → D-006
   - [ ] Multi-site DR → D-008, D-011
   - [ ] Domain architecture → D-010
   - [ ] Operations model → D-007, D-012
  1. Consistency: No two decisions contradict each other?
  • D-002 (RAID-5 for cost) does not contradict D-003 (RAID-6 for Tier-1) — different tiers
  • D-010 (2 domains) aligns with D-009 (separate TZ per domain)
  • D-008 (stretched) aligns with D-011 (tiered DR)
  1. Traceability: Every decision links to a requirement?
   - Cross-reference matrix shows coverage → ✅
  1. Defensibility: Can each rejected alternative be explained in 30 seconds?
  • Practice: Say each rejection out loud in one sentence
  • If it takes more than 30 seconds, the reasoning is unclear — simplify
  1. Gap analysis: Are there obvious decisions NOT in the register?
  • Missing: Monitoring tool selection (Aria vs third-party)
  • Missing: Backup tool selection (Veeam, Commvault, etc.)
  • Missing: Certificate authority architecture (VMCA vs enterprise PKI)
  • Action: Add D-013 through D-015 for completeness

Validation Gate

Check: Design decision register compiled with traceability matrix and validation

Expected: 12+ design decisions in tabular format, requirement traceability matrix with full coverage, 3 deep-dive justifications, validation checklist with gap analysis

Common Errors

Design decisions without alternatives — stating what you chose without showing what you rejected suggests you did not evaluate options
No requirement traceability — decisions that cannot trace to a requirement appear arbitrary
Inconsistent decisions — e.g., choosing separate domains for isolation but sharing the same transport zone
Missing gap analysis — panelists will find missing decisions; proactively identifying gaps shows design maturity

Task 4 Risk Register & Defense Preparation

The risk register demonstrates that the architect has thought about what can go wrong. VCDX panelists respect designs that acknowledge risks and present specific mitigations — not designs that pretend everything will work perfectly.

Create the risk register and prepare for VCDX defense by identifying potential panel challenges and preparing responses.

Step 1

Create the risk register:

Risk IDRisk DescriptionProbabilityImpactRisk ScoreMitigationResidual Risk
R-001vSAN disk failure during upgrade windowMedium (3)High (4)12Schedule upgrades during low-IO; N+1 host capacityLow
R-002Inter-site link failure during DR eventLow (2)Critical (5)10Dual WAN links with auto-failover; monthly failover testMedium
R-003VCF license compliance audit findingMedium (3)Medium (3)9Quarterly license reconciliation; automated core countingLow
R-004ESXi PSOD from driver incompatibilityLow (2)High (4)8Test firmware/driver on Holodeck before production; HCL adherenceLow
R-005NSX Manager cluster quorum lossLow (2)Critical (5)103-node across failure domains; daily config backupMedium
R-006KMS failure renders encrypted data inaccessibleLow (2)Critical (5)10KMS HA cluster; quarterly recovery testMedium
R-007Growth exceeds 20% p.a. — capacity exhaustionMedium (3)High (4)12Quarterly capacity review; procurement trigger at 75%Low
R-008DFW misconfiguration blocks critical trafficMedium (3)High (4)12Monitor-only phase before enforcement; Traceflow validationLow
R-009SDDC Manager backup corruptionLow (2)High (4)8Weekly restore test; multiple backup copiesLow
R-010Stretched cluster witness site power failureLow (2)Medium (3)6UPS with 4hr runtime; documented replacement procedureLow

Scoring: Probability (1-5) × Impact (1-5) = Risk Score
Thresholds: Score ≥12 = High priority, 8-11 = Medium, <8 = Low priority

Step 2

Design risk monitoring integration:

  1. Map risks to monitoring alerts:
Risk IDMonitoring SourceAlert ConditionAction
R-001vSAN HealthDisk health degradedCreate change request for replacement
R-002Network monitoringInter-site link utilization >80%Capacity planning review
R-005NSX HealthNSX Manager node unreachableImmediate investigation
R-006KMS health checkKMS connection test failureSeverity 1 — restore KMS
R-007Aria OperationsCluster utilization >75% for 7 daysTrigger procurement workflow
R-008DFW logging>50 drops/min from production VMsInvestigate rule misconfiguration
  1. Risk review cadence:
  • Monthly: Automated risk dashboard review (capacity, health, replication)
  • Quarterly: Risk register review with infrastructure and security teams
  • Annual: Full risk assessment refresh with business stakeholders
  • Trigger-based: After any major incident, update risk register
Step 3

Prepare VCDX defense strategy:

  1. Know your top 3 risks and their mitigations cold:
  • R-001, R-007, R-008 are highest scoring (12 each)
   - Practice explaining each in 60 seconds: risk → impact → mitigation → residual
  1. Anticipate panel challenge patterns:
   - 'What if...?' → Failure scenarios (have the failure matrix ready)
   - 'Why not...?' → Alternative evaluation (have DDR with rejected alternatives)
   - 'How much...?' → Cost analysis (have the cost comparison tables)
   - 'Show me...' → Diagrams and traffic flows (have conceptual/logical/physical ready)
   - 'What changes if...?' → Requirement modification (know which decisions are affected)
  1. Prepare 'what if' pivot responses:
   - What if RPO=0 is relaxed to RPO=5min? → Remove stretched cluster, use SRM, save 30% cost
   - What if budget is cut by 30%? → Consolidate to 1 workload domain, reduce DR scope
   - What if the customer adds Kubernetes? → New workload domain with TKG, or vSphere Pods on existing
   - What if compliance changes from SOC 2 to PCI-DSS? → Isolate payment VMs in dedicated WD
  1. Practice timing:
  • Design presentation: 75 minutes (stick to schedule — going over is a red flag)
  • Focus on: Why, not what — panelists can read the document; they want to hear your reasoning
  • Structure: Follow the document order but prepare to jump to any section on demand
Step 4

Create the final document quality checklist:

  1. Content quality:
  • [ ] Executive summary is specific (not generic boilerplate)
   - [ ] All sizing shows full math (raw → usable chain)
  • [ ] Every design decision has ≥2 rejected alternatives
  • [ ] Risk register has ≥10 risks with specific mitigations
  • [ ] Diagrams exist at 3 levels: conceptual, logical, physical
  1. Traceability:
  • [ ] Every decision traces to ≥1 requirement
  • [ ] Every requirement has ≥1 design decision
  • [ ] Risk mitigations reference specific design elements
  • [ ] Design principles are referenced in decision justifications
  1. Technical accuracy:
  • [ ] VCF 9.0 licensing: per-core, minimum 16 cores per CPU socket
  • [ ] vSAN ESA: minimum 1 NVMe TLC device per storage pool (not 6 drives)
  • [ ] Witness: Official OVA tiers (Tiny/Medium/Large/XL), not generic specs
  • [ ] Geneve overhead: ~54 bytes (MTU ≥1700 recommended)
  • [ ] BGP failover: Sub-second with BFD (500ms×3 = 1.5s detection)
  • [ ] HCX concurrent migrations: 600 per HCX Manager (default config)
  • [ ] VCF upgrade: 19-step sequence per KB 390634
  1. Presentation readiness:
  • [ ] Document is 15-20 pages (excluding appendices)
  • [ ] Every page has a purpose — no padding
  • [ ] Key diagrams can be presented in 2 minutes each
  • [ ] Defense responses practiced for top 10 anticipated challenges

Validation Gate

Check: Risk register, monitoring integration, defense preparation, and quality checklist complete

Expected: 10+ risks with scoring and mitigation, risk monitoring mapped to alerts, defense strategy with pivot responses, comprehensive quality checklist

Common Errors

Risk register with generic mitigations ('monitor closely') — every mitigation must be a specific, actionable control
Not knowing the top 3 risks by heart — panelists will ask; hesitation signals weak risk understanding
Practicing the 'what' instead of the 'why' — panelists can read the document; they want to hear reasoning and judgment
Going over time in the design presentation — 75 minutes means 75 minutes; practice with a timer

Final Validation

Complete VCDX-grade design document with executive summary, sizing, decisions, diagrams, risks, and defense preparation

✓ Executive summary communicates design intent → 500-600 words covering business context, solution, key decisions, and outcomes

✓ Sizing calculations show full math → Compute and storage worksheets with governing constraint identified

✓ Design decision register is comprehensive → 12+ decisions with traceability matrix and gap analysis

✓ Risk register with monitoring integration → 10+ risks with scoring, mitigation, and alert mapping

Cleanup / Restore

• Save the complete design document draft

• Archive all supporting worksheets and diagrams

Design Reflection (VCDX)

This lab IS the VCDX preparation. The design document you produce here is the condensed version of what you would submit. The key insight: a VCDX design is not about showing everything you know — it is about showing that you can make decisions, justify them, and defend them under pressure. Every word in the document should serve one of three purposes: (1) Establish a requirement, (2) Present a decision with alternatives, or (3) Acknowledge a risk with mitigation. Everything else is padding.

Requirements

  • Produce a 15-20 page design document that covers all VCF architecture domains
  • Every design decision must trace to a documented requirement
  • Risk register must include monitoring integration, not just list risks
  • Document must be defensible in a 75-minute VCDX panel session

Constraints

  • Document length: 15-20 pages excluding appendices
  • All technical claims must be accurate to VCF 9.0 documentation
  • Design must be internally consistent (no contradictory decisions)
  • Presentation time: 75 minutes maximum

Assumptions

  • Prior labs (01-11) have produced the underlying design artifacts
  • The fictional scenario (600 VMs, 2 sites, SOC 2) is the basis
  • Panelists have read the document before the defense session
  • Supporting documentation (appendices, detailed worksheets) are available on demand

Risks

  • Document is too generic — does not demonstrate understanding of the specific scenario
  • Sizing math has errors — a single arithmetic mistake undermines credibility
  • Design decisions are inconsistent — one section contradicts another
  • Defense preparation is insufficient — panelist challenges catch the candidate off guard

Self-Assessment Discussion Prompts

  1. If you had to cut this document from 20 pages to 10, which sections would you keep and why?
  2. What is the single most important diagram in a VCDX design document?
  3. How would you adapt this design document for a different scenario — say, a hospital with HIPAA requirements instead of SOC 2?
  4. If a panelist points out an error in your sizing math during the defense, what is the best way to respond?

Extensions

Physical Architecture Diagram with Bill of Materials

Create the physical architecture layer showing rack layouts, server models, switch connections, and power/cooling requirements. Include a detailed Bill of Materials with part numbers, quantities, and unit costs for the complete VCF deployment.

VCDX Defense Mock Session

Conduct a mock VCDX defense session using the design document. Have a colleague or mentor act as panelist with prepared challenge questions. Record the session, review for improvement areas: timing, clarity of responses, ability to pivot when challenged, and confidence in design decisions.

Design Document Review Against VCDX Scoring Rubric

Review the design document against the publicly available VCDX scoring criteria. Score each section (architecture, implementation, operations) and identify gaps. Prioritize improvements based on scoring weight and current score.

⚠ Known Pitfalls (from Community KB)

Writing a design document that describes technology without making decisions — VCDX evaluators want to see YOUR design choices, not a product manual summary
Presenting the design as perfect with no risks — VCDX panelists respect designs that acknowledge weaknesses and present mitigations; pretending everything is perfect signals lack of real-world experience
Not practicing the defense presentation with a timer — going over 75 minutes or rushing through key sections both signal poor preparation
Including outdated technical specifications — every technical claim must be accurate to VCF 9.0; citing VCF 4.x or vSAN 7.x defaults when 9.0 has changed them undermines credibility

References

  • VCDX Application and Defense Guide (VMware/Broadcom)
  • VMware VCF 9.0 Architecture and Design Guide
  • VCDX Handbook (community resource) — Design Document Templates
  • VMware Validated Designs (VVD) Documentation — Reference Architecture
  • TOGAF Architecture Framework — ADM for Enterprise Architecture Documentation
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.