Write a VCDX-Grade Design Document (Condensed)
Objectives
- Write a VCDX-quality executive summary that communicates design intent to non-technical stakeholders
- Create detailed sizing calculations for compute, memory, storage, and network
- Build a comprehensive design decision register with alternatives and traceability
- Develop logical and physical architecture diagrams
- Construct a risk register with quantified probability, impact, and mitigation strategies
Prerequisites
Completed labs vcp-architect-01 through -06 (design exercises)
Prior labs: vcp-architect-01, vcp-architect-02, vcp-architect-03, vcp-architect-04, vcp-architect-05, vcp-architect-06
Required skills:
- All VCF architecture design skills from prior labs
- Technical writing for mixed audiences
- Visio/Draw.io diagramming (or text-based equivalent)
- Risk assessment methodology
Lab Environment
Documentation exercise — produces a 15-20 page condensed design document using all prior lab outputs
Tasks
Task 1 Executive Summary & Scope Definition
Write the opening sections of a VCDX design document that communicate the design intent, scope, and approach to both technical and business audiences.
Write the Executive Summary (500-600 words):
Structure:
- Business context (2-3 sentences): What is the business challenge?
Example: 'GlobalFinance Corp operates 600+ virtual workloads across aging infrastructure with increasing compliance requirements (SOC 2 Type II). The current environment lacks standardized disaster recovery, has inconsistent security policies, and faces capacity constraints that limit business growth.'
- Solution overview (3-4 sentences): What are we building?
Example: 'This design deploys VMware Cloud Foundation 9.0 across two data center sites with a tiered architecture that aligns infrastructure investment to business criticality. Mission-critical Tier-1 workloads are protected by vSAN stretched cluster (RPO=0), while Tier-2 workloads use VMware Site Recovery with 15-minute RPO. NSX micro-segmentation enforces zero-trust security with tag-based Distributed Firewall policies.'
- Key design decisions (3-4 sentences): What are the most important trade-offs?
Example: 'The design uses two workload domains (Production and Development) to limit SOC 2 audit scope while minimizing vCenter licensing overhead. vSAN ESA with erasure coding provides 50% more usable storage than mirroring at equivalent protection. The self-service portal via Aria Automation reduces VM provisioning from 5 days to 10 minutes while enforcing governance through catalog-driven deployment.'
- Outcome (2-3 sentences): What does success look like?
Example: 'This architecture provides a 3-year platform supporting 20% annual growth, with full compliance attestation, automated DR capabilities, and a consumption-based operating model. Total infrastructure investment: $X for primary + $Y for DR, representing a Z% reduction in per-VM cost versus the current environment.'
Write the Scope & Assumptions section (300-400 words):
- In Scope:
- VCF 9.0 deployment: Management domain + 2 workload domains
- vSAN ESA storage with tiered protection policies
- NSX micro-segmentation with zero-trust DFW
- Multi-site DR: Stretched cluster (Tier-1), SRM (Tier-2), backup (Tier-3)
- Self-service portal with Aria Automation
- Monitoring and operations with Aria Operations suite
- Out of Scope:
- Physical data center construction (facilities assumed ready)
- Application migration planning (separate workstream)
- End-user device management
- Third-party security tools (SIEM, EDR) — integration points documented only
- Network outside the VCF overlay (physical switches, WAN, internet)
- Assumptions (reference RCAR from Lab 01):
- List top 10 assumptions with cross-reference to full RCAR document
- Flag the 3 highest-risk assumptions (ones most likely to be wrong)
- Example high-risk assumption: 'Inter-site latency measured at 0.8ms RTT will remain stable under production load. If latency exceeds 5ms, stretched cluster design must be revised to SRM.'
- Dependencies:
- Physical network: Spine-leaf fabric with BGP, MTU 9000 end-to-end
- Identity: Active Directory with LDAPS available at both sites
- Certificates: Enterprise PKI for VMCA subordinate CA
- Procurement: Hardware on VCF 9.0 HCL (ordered and delivery confirmed)
Write the Design Principles section (200-300 words):
Design principles guide every subsequent decision. List 6-8 principles:
- Security by default: Every VM is encrypted, tagged for DFW, and deployed from hardened golden images. Security is embedded in the platform, not bolted on.
- Failure domain isolation: Management, production, and development workloads are isolated by workload domain, transport zone, and storage. A failure in one domain does not propagate.
- Cost-aligned protection: DR investment matches business criticality. Tier-1 (RPO=0) costs 2× more than Tier-2 (RPO=15min) and 5× more than Tier-3 (backup).
- Automation-first operations: VM provisioning, patching, and monitoring are automated. Manual intervention is reserved for exceptions and incident response.
- Lifecycle-aware sizing: All capacity is sized for 3-year growth at 20% p.a. with N+1 HA. Quarterly capacity reviews trigger procurement before exhaustion.
- Compliance-native architecture: SOC 2 controls are implemented at the platform level (encryption, logging, access control, network segmentation). Compliance evidence is generated automatically.
- Operational simplicity: When two designs offer similar protection, the simpler one is preferred. Complexity adds operational risk.
- Defense in depth: No single control is relied upon. Network security = DFW + physical firewall. Data protection = encryption + access control + backup. Availability = HA + DR + monitoring.
Principle traceability: Each design decision in the register should reference which principles it upholds.
Create a document outline (table of contents):
- Executive Summary
- Scope, Assumptions & Dependencies
- Design Principles
- Requirements (RCAR Framework)
4.1 Business Requirements
4.2 Technical Requirements
4.3 Constraints
4.4 Assumptions
4.5 Risks
- Architecture Overview
5.1 Conceptual Architecture
5.2 Logical Architecture
5.3 Physical Architecture
- Compute Design
6.1 Cluster Topology
6.2 Sizing Calculations
6.3 HA & DRS Configuration
6.4 EVC Mode
- Storage Design
7.1 vSAN ESA Architecture
7.2 Capacity Planning
7.3 Storage Policies
7.4 Encryption & Key Management
- Network Design
8.1 NSX Topology (Tier-0/Tier-1)
8.2 Segment Design & CIDR Plan
8.3 DFW Micro-Segmentation
8.4 BGP & Edge Design
- Multi-Site Design
9.1 Stretched Cluster
9.2 SRM Configuration
9.3 Failover/Failback Procedures
9.4 DR Testing Calendar
- Operations & Automation
10.1 Self-Service Portal
10.2 Monitoring & Alerting
10.3 Upgrade Lifecycle
10.4 Backup & Recovery
- Design Decision Register
- Risk Register
- Appendices
A. Detailed Sizing Worksheets
B. DFW Rule Set
C. VLAN & IP Address Plan
D. Bill of Materials
Target: 15-20 pages (excluding appendices). Each section 1-2 pages with diagrams.
Validation Gate
Check: Executive summary, scope, and principles written with design document outline
Expected: 500-600 word executive summary, scope with in/out/assumptions/dependencies, 6-8 design principles, complete table of contents
Common Errors
Task 2 Sizing Calculations & Architecture Diagrams
Create the detailed sizing calculations and architecture diagrams that form the technical backbone of the VCDX design document.
Create compute sizing summary (consolidate from Lab 02):
Compute Sizing Worksheet:
| Item | Tier-1 | Tier-2 | Tier-3 | Management | Total |
|---|---|---|---|---|---|
| VM Count (current) | 80 | 300 | 220 | 15 | 615 |
| VM Count (3-year) | 138 | 518 | 380 | 20 | 1,056 |
| vCPU per VM (avg) | 4 | 2 | 2 | 4 | - |
| Total vCPU | 552 | 1,036 | 760 | 80 | 2,428 |
| CPU Overcommit | 3:1 | 5:1 | 8:1 | 2:1 | - |
| Physical Cores Needed | 184 | 208 | 95 | 40 | 527 |
| RAM per VM (avg) | 16 GB | 8 GB | 4 GB | 16 GB | - |
| Total RAM | 2,208 GB | 4,144 GB | 1,520 GB | 320 GB | 8,192 GB |
| RAM Overcommit | 1.25:1 | 1.5:1 | 2:1 | 1:1 | - |
| Physical RAM Needed | 1,766 GB | 2,763 GB | 760 GB | 320 GB | 5,609 GB |
| Hosts (64 cores, 512 GB) | 3 | 5 | 2 | 1 | 11 |
| + HA (N+1) | 4 | 6 | 3 | 2 | 15 |
| + VCF Minimum (4 mgmt) | 4 | 6 | 3 | 4 | 17 |
| + Stretch (2× Tier-1) | 8 | 6 | 3 | 4 | 21 |
Governing constraint per tier: Document whether CPU or RAM is the limiting factor.
Licensing impact: 21 hosts × 64 cores × 2 sockets = 2,688 licensed cores total.
Create storage sizing summary (consolidate from Lab 03):
Storage Sizing Worksheet:
| Item | Tier-1 | Tier-2/3 | Management | Total |
|---|---|---|---|---|
| VM Storage (current) | 16 TB | 67 TB | 3.7 TB | 86.7 TB |
| VM Storage (3-year) | 27.6 TB | 115.8 TB | 5 TB | 148.4 TB |
| FTT Scheme | FTT=2 RAID-6 | FTT=1 RAID-5 | FTT=1 RAID-5 | - |
| FTT Multiplier | 1.5× | 1.33× | 1.33× | - |
| Raw for Protection | 41.4 TB | 154 TB | 6.65 TB | 202 TB |
| + Slack (30%) | 59.1 TB | 220 TB | 9.5 TB | 288.6 TB |
| + Metadata (2%) | 60.3 TB | 224.4 TB | 9.7 TB | 294.4 TB |
| Dedup/Comp Ratio | 2:1 | 2.5:1 | 1.5:1 | - |
| Raw Needed (post dedup) | 30.2 TB | 89.8 TB | 6.5 TB | 126.5 TB |
| Disks/Host (3.84 TB NVMe) | 4 | 4 | 2 | - |
| Raw per Host | 15.36 TB | 15.36 TB | 7.68 TB | - |
| Hosts (storage sizing) | 2 | 6 | 1 | 9 |
Note: Compute sizing drives more hosts than storage (17 vs 9), so compute is the governing constraint.
Storage headroom with compute-driven host count: Significant (~2× surplus).
Create the conceptual architecture diagram (text-based):
╔═══════════════════════════════════════════════════════════════╗
║ CONCEPTUAL ARCHITECTURE ║
╠═══════════════════════════════════════════════════════════════╣
║ ║
║ [Business Users] ║
║ │ ║
║ ▼ ║
║ ┌──────────────────────────────────────┐ ║
║ │ Self-Service Portal │ ║
║ │ (Aria Automation Catalog) │ ║
║ └──────────────┬───────────────────────┘ ║
║ │ ║
║ ┌──────────────▼───────────────────────┐ ║
║ │ VCF 9.0 Platform │ ║
║ │ ┌─────────┐ ┌─────────┐ ┌────────┐ │ ║
║ │ │Compute │ │Storage │ │Network │ │ ║
║ │ │vSphere │ │vSAN ESA │ │NSX │ │ ║
║ │ └─────────┘ └─────────┘ └────────┘ │ ║
║ └──────────────┬───────────────────────┘ ║
║ │ ║
║ ┌──────────────▼───────────────────────┐ ║
║ │ Workload Tiers │ ║
║ │ [Tier-1: RPO=0] [Tier-2: RPO=15m] │ ║
║ │ [Tier-3: Backup] [Dev/Test] │ ║
║ └──────────────┬───────────────────────┘ ║
║ │ ║
║ ┌──────────────▼───────────────────────┐ ║
║ │ Multi-Site DR │ ║
║ │ [Site A]◄──Sync──►[Site B] │ ║
║ │ [Witness C] │ ║
║ └──────────────────────────────────────┘ ║
╚═══════════════════════════════════════════════════════════════╝Conceptual diagram purpose: Communicate the design intent to non-technical stakeholders (CIO, CFO). No IP addresses, no specific products — just capabilities and relationships.
Create the logical architecture diagram (text-based):
┌─────────────────── Site A (Primary) ───────────────────┐
│ │
│ ┌── Management Domain ──┐ ┌── WD-1 (Production) ────┐│
│ │ vCenter-mgmt │ │ vCenter-prod ││
│ │ SDDC Manager │ │ ┌─Tier-1 Cluster (5h)──┐││
│ │ NSX Mgr (3-node) │ │ │ 80 VMs, FTT=2 │││
│ │ Aria Suite │ │ │ HA 25%, DRS Full │││
│ │ 4 hosts, FTT=1 RAID-5 │ │ └───────────────────────┘││
│ └─────────────────────────┘ │ ┌─Tier-2 Cluster (8h)──┐││
│ │ │ 300 VMs, FTT=1 │││
│ ┌── WD-2 (Dev/Test) ─────┐ │ │ HA 20%, DRS Full │││
│ │ vCenter-devtest │ │ └───────────────────────┘││
│ │ 4 hosts, FTT=1 RAID-5 │ └──────────────────────────┘│
│ │ 220 VMs, Force Prov OK │ │
│ └──────────────────────────┘ │
│ │
│ [Tier-0 Edge]──BGP──[Physical Network] │
│ [DFW: Zero-Trust, Tag-Based] │
└───────────┬─────────────────────────────────────────────┘
│ 100Gbps Dark Fiber (0.8ms RTT)
│ vSAN Sync Replication (Tier-1)
│ vSphere Replication (Tier-2, 15min RPO)
┌───────────▼─────────────────────────────────────────────┐
│ Site B (DR) │
│ ┌─Tier-1 Mirror (8h)──┐ ┌─Tier-2 DR (4h)────────────┐│
│ │ Stretched cluster │ │ SRM recovery targets ││
│ │ (shared with Site A) │ │ Placeholder VMs ││
│ └──────────────────────┘ └────────────────────────────┘│
└───────────┬─────────────────────────────────────────────┘
│ 1Gbps MPLS (2.5ms RTT)
┌───────────▼─────┐
│ Site C (Witness) │
│ Medium OVA │
│ ≤21,833 comp │
└───────────────────┘Logical diagram purpose: Show component relationships, domain boundaries, and data flow paths. Used by architects and senior engineers.
Validation Gate
Check: Sizing worksheets and architecture diagrams created
Expected: Compute and storage sizing tables with full math, conceptual diagram for business audience, logical diagram with domain boundaries and data flows
Common Errors
Task 3 Design Decision Register (Complete)
Compile the comprehensive design decision register that consolidates all decisions from prior labs into a single, traceable artifact.
Compile all design decisions into a register:
| ID | Category | Decision | Alternatives Rejected | Primary Requirement | Risk |
|---|---|---|---|---|---|
| D-001 | Compute | 4-node mgmt + split workload clusters | Consolidated, 3-node mgmt, 5-node mgmt | SOC 2 isolation, N+1 HA | Over-provisioned mgmt |
| D-002 | Storage | vSAN ESA FTT=1 RAID-5 for workload | OSA, FTT=2, external FC | Cost optimization, capacity efficiency | Lower IOPS than mirroring |
| D-003 | Storage | FTT=2 RAID-6 for Tier-1 | FTT=1 RAID-5, FTT=2 Mirror | Tier-1 SLA, dual-failure tolerance | Higher capacity overhead |
| D-004 | Network | NSX DFW zero-trust, tag-based | Perimeter only, host firewall | SOC 2 CC6.6, east-west security | DFW complexity, rule sprawl |
| D-005 | Compute | 3 clusters (Mgmt + T1 + T2/3) | 1 cluster, 4 clusters | SOC 2 isolation, cost balance | Mgmt overhead |
| D-006 | Security | vSAN encryption + external KMS (prod), NKP (dev) | VM encryption, SED, no encryption | SOC 2 CC6.1, compliance | KMS as SPOF |
| D-007 | Operations | Ensure accessibility as default maint. mode | Full migration, no migration | Balance speed vs safety | Reduced FTT during maint. |
| D-008 | Multi-Site | Stretched cluster for Tier-1 | SRM, HCX DR, Pilot Light | RPO=0 requirement | 2× host cost premium |
| D-009 | Network | DFW + physical FW (defense-in-depth) | DFW only, perimeter only | Zero-trust, compliance | Dual management |
| D-010 | Domain | 2 workload domains (Prod + Dev) | 1 domain, 3 domains | SOC 2 scope reduction | Additional vCenter license |
| D-011 | Multi-Site | Tiered DR (Stretched/SRM/Backup) | All stretched, all SRM | Cost optimization | Complexity of mixed DR |
| D-012 | Operations | Aria Automation self-service | Direct vCenter, ServiceNow only | Developer velocity | Aria license cost |
Each row links to the detailed design decision write-up in the relevant section of the document.
Create a requirement traceability matrix:
Purpose: Every design decision must trace to at least one requirement. If a decision exists without traceability, it is either unjustified (remove it) or has an unrecorded requirement (add it).
| Requirement ID | Requirement | Design Decisions | Section |
|---|---|---|---|
| R-001 | SOC 2 compliance | D-004, D-006, D-009, D-010 | 4.1 |
| R-002 | RPO=0 for Tier-1 | D-003, D-008 | 4.1 |
| R-003 | 3-year lifecycle growth | D-001, D-002, D-005 | 4.2 |
| R-004 | VM provisioning <10 min | D-012 | 4.2 |
| R-005 | Management domain RTO <4h | D-001, D-007 | 4.2 |
| C-001 | VCF per-core licensing | D-001, D-002, D-005, D-010 | 4.3 |
| C-002 | vSAN ESA (NVMe only) | D-002, D-003 | 4.3 |
| C-003 | Inter-site RTT ≤5ms | D-008 | 4.3 |
| C-004 | Mgmt domain min 4 hosts | D-001 | 4.3 |
Coverage check:
- Every requirement has at least 1 design decision → ✅ - Every design decision traces to at least 1 requirement → ✅ - No orphaned decisions or requirements → ✅
Write detailed justification for the 3 most impactful decisions:
Decision D-008 (Stretched Cluster) — Deep justification:
Context: Customer requires RPO=0 for 80 Tier-1 financial trading VMs. Any data loss during a site failure could result in regulatory violations (SOC 2) and financial loss ($X million per hour of trading data loss).
Decision: vSAN stretched cluster with synchronous replication.
Alternatives evaluation:
- SRM with RPO=5min: Minimum data loss = 5 minutes of transactions. For trading workloads processing $Y million/hour, this represents unacceptable financial risk. Rejected.
- Array-based replication: Near-zero RPO possible but violates single-vendor constraint (C-001). Adds array management complexity. Rejected.
- HCX DR: RPO=5min minimum, same limitation as SRM. Rejected.
Risk acceptance: Stretched cluster adds 30-40% infrastructure cost for Tier-1 workloads. This cost is justified by: (a) regulatory requirement for RPO=0, (b) financial loss avoidance that exceeds infrastructure premium within 1 trading day, (c) automatic failover reducing RTO to seconds vs 30-60 minutes for SRM.
Dependencies: Inter-site RTT ≤5ms (measured 0.8ms), dark fiber connectivity, witness at third site. If any dependency fails, fallback to SRM with accepted RPO=5min.
[Repeat similar depth for D-004 (DFW) and D-010 (Workload Domains)]
Validate the design decision register:
Quality checks:
- Completeness: Every major architecture choice has a DDR entry? Review checklist:
- [ ] Cluster topology → D-001, D-005 - [ ] Storage architecture → D-002, D-003 - [ ] Network security → D-004, D-009 - [ ] Encryption → D-006 - [ ] Multi-site DR → D-008, D-011 - [ ] Domain architecture → D-010 - [ ] Operations model → D-007, D-012
- Consistency: No two decisions contradict each other?
- D-002 (RAID-5 for cost) does not contradict D-003 (RAID-6 for Tier-1) — different tiers
- D-010 (2 domains) aligns with D-009 (separate TZ per domain)
- D-008 (stretched) aligns with D-011 (tiered DR)
- Traceability: Every decision links to a requirement?
- Cross-reference matrix shows coverage → ✅
- Defensibility: Can each rejected alternative be explained in 30 seconds?
- Practice: Say each rejection out loud in one sentence
- If it takes more than 30 seconds, the reasoning is unclear — simplify
- Gap analysis: Are there obvious decisions NOT in the register?
- Missing: Monitoring tool selection (Aria vs third-party)
- Missing: Backup tool selection (Veeam, Commvault, etc.)
- Missing: Certificate authority architecture (VMCA vs enterprise PKI)
- Action: Add D-013 through D-015 for completeness
Validation Gate
Check: Design decision register compiled with traceability matrix and validation
Expected: 12+ design decisions in tabular format, requirement traceability matrix with full coverage, 3 deep-dive justifications, validation checklist with gap analysis
Common Errors
Task 4 Risk Register & Defense Preparation
Create the risk register and prepare for VCDX defense by identifying potential panel challenges and preparing responses.
Create the risk register:
| Risk ID | Risk Description | Probability | Impact | Risk Score | Mitigation | Residual Risk |
|---|---|---|---|---|---|---|
| R-001 | vSAN disk failure during upgrade window | Medium (3) | High (4) | 12 | Schedule upgrades during low-IO; N+1 host capacity | Low |
| R-002 | Inter-site link failure during DR event | Low (2) | Critical (5) | 10 | Dual WAN links with auto-failover; monthly failover test | Medium |
| R-003 | VCF license compliance audit finding | Medium (3) | Medium (3) | 9 | Quarterly license reconciliation; automated core counting | Low |
| R-004 | ESXi PSOD from driver incompatibility | Low (2) | High (4) | 8 | Test firmware/driver on Holodeck before production; HCL adherence | Low |
| R-005 | NSX Manager cluster quorum loss | Low (2) | Critical (5) | 10 | 3-node across failure domains; daily config backup | Medium |
| R-006 | KMS failure renders encrypted data inaccessible | Low (2) | Critical (5) | 10 | KMS HA cluster; quarterly recovery test | Medium |
| R-007 | Growth exceeds 20% p.a. — capacity exhaustion | Medium (3) | High (4) | 12 | Quarterly capacity review; procurement trigger at 75% | Low |
| R-008 | DFW misconfiguration blocks critical traffic | Medium (3) | High (4) | 12 | Monitor-only phase before enforcement; Traceflow validation | Low |
| R-009 | SDDC Manager backup corruption | Low (2) | High (4) | 8 | Weekly restore test; multiple backup copies | Low |
| R-010 | Stretched cluster witness site power failure | Low (2) | Medium (3) | 6 | UPS with 4hr runtime; documented replacement procedure | Low |
Scoring: Probability (1-5) × Impact (1-5) = Risk Score
Thresholds: Score ≥12 = High priority, 8-11 = Medium, <8 = Low priority
Design risk monitoring integration:
- Map risks to monitoring alerts:
| Risk ID | Monitoring Source | Alert Condition | Action |
|---|---|---|---|
| R-001 | vSAN Health | Disk health degraded | Create change request for replacement |
| R-002 | Network monitoring | Inter-site link utilization >80% | Capacity planning review |
| R-005 | NSX Health | NSX Manager node unreachable | Immediate investigation |
| R-006 | KMS health check | KMS connection test failure | Severity 1 — restore KMS |
| R-007 | Aria Operations | Cluster utilization >75% for 7 days | Trigger procurement workflow |
| R-008 | DFW logging | >50 drops/min from production VMs | Investigate rule misconfiguration |
- Risk review cadence:
- Monthly: Automated risk dashboard review (capacity, health, replication)
- Quarterly: Risk register review with infrastructure and security teams
- Annual: Full risk assessment refresh with business stakeholders
- Trigger-based: After any major incident, update risk register
Prepare VCDX defense strategy:
- Know your top 3 risks and their mitigations cold:
- R-001, R-007, R-008 are highest scoring (12 each)
- Practice explaining each in 60 seconds: risk → impact → mitigation → residual
- Anticipate panel challenge patterns:
- 'What if...?' → Failure scenarios (have the failure matrix ready) - 'Why not...?' → Alternative evaluation (have DDR with rejected alternatives) - 'How much...?' → Cost analysis (have the cost comparison tables) - 'Show me...' → Diagrams and traffic flows (have conceptual/logical/physical ready) - 'What changes if...?' → Requirement modification (know which decisions are affected)
- Prepare 'what if' pivot responses:
- What if RPO=0 is relaxed to RPO=5min? → Remove stretched cluster, use SRM, save 30% cost - What if budget is cut by 30%? → Consolidate to 1 workload domain, reduce DR scope - What if the customer adds Kubernetes? → New workload domain with TKG, or vSphere Pods on existing - What if compliance changes from SOC 2 to PCI-DSS? → Isolate payment VMs in dedicated WD
- Practice timing:
- Design presentation: 75 minutes (stick to schedule — going over is a red flag)
- Focus on: Why, not what — panelists can read the document; they want to hear your reasoning
- Structure: Follow the document order but prepare to jump to any section on demand
Create the final document quality checklist:
- Content quality:
- [ ] Executive summary is specific (not generic boilerplate)
- [ ] All sizing shows full math (raw → usable chain)
- [ ] Every design decision has ≥2 rejected alternatives
- [ ] Risk register has ≥10 risks with specific mitigations
- [ ] Diagrams exist at 3 levels: conceptual, logical, physical
- Traceability:
- [ ] Every decision traces to ≥1 requirement
- [ ] Every requirement has ≥1 design decision
- [ ] Risk mitigations reference specific design elements
- [ ] Design principles are referenced in decision justifications
- Technical accuracy:
- [ ] VCF 9.0 licensing: per-core, minimum 16 cores per CPU socket
- [ ] vSAN ESA: minimum 1 NVMe TLC device per storage pool (not 6 drives)
- [ ] Witness: Official OVA tiers (Tiny/Medium/Large/XL), not generic specs
- [ ] Geneve overhead: ~54 bytes (MTU ≥1700 recommended)
- [ ] BGP failover: Sub-second with BFD (500ms×3 = 1.5s detection)
- [ ] HCX concurrent migrations: 600 per HCX Manager (default config)
- [ ] VCF upgrade: 19-step sequence per KB 390634
- Presentation readiness:
- [ ] Document is 15-20 pages (excluding appendices)
- [ ] Every page has a purpose — no padding
- [ ] Key diagrams can be presented in 2 minutes each
- [ ] Defense responses practiced for top 10 anticipated challenges
Validation Gate
Check: Risk register, monitoring integration, defense preparation, and quality checklist complete
Expected: 10+ risks with scoring and mitigation, risk monitoring mapped to alerts, defense strategy with pivot responses, comprehensive quality checklist
Common Errors
Final Validation
Complete VCDX-grade design document with executive summary, sizing, decisions, diagrams, risks, and defense preparation
✓ Executive summary communicates design intent → 500-600 words covering business context, solution, key decisions, and outcomes
✓ Sizing calculations show full math → Compute and storage worksheets with governing constraint identified
✓ Design decision register is comprehensive → 12+ decisions with traceability matrix and gap analysis
✓ Risk register with monitoring integration → 10+ risks with scoring, mitigation, and alert mapping
Cleanup / Restore
• Save the complete design document draft
• Archive all supporting worksheets and diagrams
Design Reflection (VCDX)
This lab IS the VCDX preparation. The design document you produce here is the condensed version of what you would submit. The key insight: a VCDX design is not about showing everything you know — it is about showing that you can make decisions, justify them, and defend them under pressure. Every word in the document should serve one of three purposes: (1) Establish a requirement, (2) Present a decision with alternatives, or (3) Acknowledge a risk with mitigation. Everything else is padding.
Requirements
- Produce a 15-20 page design document that covers all VCF architecture domains
- Every design decision must trace to a documented requirement
- Risk register must include monitoring integration, not just list risks
- Document must be defensible in a 75-minute VCDX panel session
Constraints
- Document length: 15-20 pages excluding appendices
- All technical claims must be accurate to VCF 9.0 documentation
- Design must be internally consistent (no contradictory decisions)
- Presentation time: 75 minutes maximum
Assumptions
- Prior labs (01-11) have produced the underlying design artifacts
- The fictional scenario (600 VMs, 2 sites, SOC 2) is the basis
- Panelists have read the document before the defense session
- Supporting documentation (appendices, detailed worksheets) are available on demand
Risks
- Document is too generic — does not demonstrate understanding of the specific scenario
- Sizing math has errors — a single arithmetic mistake undermines credibility
- Design decisions are inconsistent — one section contradicts another
- Defense preparation is insufficient — panelist challenges catch the candidate off guard
Self-Assessment Discussion Prompts
- If you had to cut this document from 20 pages to 10, which sections would you keep and why?
- What is the single most important diagram in a VCDX design document?
- How would you adapt this design document for a different scenario — say, a hospital with HIPAA requirements instead of SOC 2?
- If a panelist points out an error in your sizing math during the defense, what is the best way to respond?
Extensions
Physical Architecture Diagram with Bill of Materials
Create the physical architecture layer showing rack layouts, server models, switch connections, and power/cooling requirements. Include a detailed Bill of Materials with part numbers, quantities, and unit costs for the complete VCF deployment.
VCDX Defense Mock Session
Conduct a mock VCDX defense session using the design document. Have a colleague or mentor act as panelist with prepared challenge questions. Record the session, review for improvement areas: timing, clarity of responses, ability to pivot when challenged, and confidence in design decisions.
Design Document Review Against VCDX Scoring Rubric
Review the design document against the publicly available VCDX scoring criteria. Score each section (architecture, implementation, operations) and identify gaps. Prioritize improvements based on scoring weight and current score.
⚠ Known Pitfalls (from Community KB)
References
- VCDX Application and Defense Guide (VMware/Broadcom)
- VMware VCF 9.0 Architecture and Design Guide
- VCDX Handbook (community resource) — Design Document Templates
- VMware Validated Designs (VVD) Documentation — Reference Architecture
- TOGAF Architecture Framework — ADM for Enterprise Architecture Documentation