Design Decision Documentation — RCAR Framework & Decision Register
Objectives
- Document architecture design decisions using the RCAR framework
- Create a design decision register with alternatives and justifications
- Develop risk register with mitigation strategies
Prerequisites
Access to VCF 9.0 documentation and architecture templates
Required skills:
- Understanding of VCF 9.0 components
- Basic architecture concepts
Lab Environment
Documentation exercise — requires VCF 9.0 reference architecture knowledge, no live environment needed
Tasks
Task 1 Requirements Gathering & RCAR Framework
Establish the foundation of any VCDX-grade design: structured requirements, constraints, assumptions, and risks.
Define business requirements for a fictional 500-VM VCF deployment:
- RTO: 4 hours, RPO: 1 hour for Tier-1 applications
- Budget constraint: $2M capex for 3-year lifecycle
- Compliance: SOC 2 Type II, data encryption at rest and in transit
- Growth projection: 20% annual VM growth
- Multi-site requirement: Primary DC + DR site, 50km apart
Translate business requirements into technical requirements:
- Compute: Size for 500 VMs × 1.2 growth × 3 years = 1,036 VMs at end of lifecycle
- Storage: Average 200GB per VM × 1,036 = ~207 TB usable (plan for FTT=1 mirroring = 414 TB raw)
- Network: 25 Gbps minimum per host for vSAN + workload traffic
- DR: SRM or vSphere Replication with 1-hour RPO
- Security: NSX DFW micro-segmentation, vTPM for all VMs, vSAN encryption
Document constraints (minimum 7 items):
- Existing network infrastructure is spine-leaf with BGP (cannot change to OSPF)
- Single vendor requirement for storage (no third-party arrays)
- Maximum 8 hosts per cluster (rack space limitation)
- ESXi version must support vSAN ESA (NVMe-only)
- All management components must be HA (no single points of failure)
- Air-gapped environment — no internet access for SDDC Manager depot
- VMware licensing: VCF per-core, minimum 16 cores per CPU socket
Document assumptions (minimum 5 items):
- Network latency between DC sites < 5ms RTT (measured and verified)
- Host hardware is on VCF 9.0 HCL (Dell VxRail or HPE SimpliVity)
- NTP and DNS services available and reliable at both sites
- Active Directory available for vCenter SSO integration
- Sufficient cooling and power capacity at both data centers for planned host count
Create risk register with 5 risks:
| Risk ID | Risk Description | Probability | Impact | Mitigation |
|---|---|---|---|---|
| R-001 | vSAN disk failure during upgrade window | Medium | High | Schedule upgrades during low-IO periods; ensure N+1 host capacity |
| R-002 | WAN link failure between sites during DR test | Low | Critical | Dual WAN links with automatic failover; test monthly |
| R-003 | VCF license compliance audit finding | Medium | Medium | Quarterly license reconciliation; automated core counting |
| R-004 | ESXi host PSOD due to driver incompatibility | Low | High | Test firmware/driver matrix on Holodeck before production; keep HCL current |
| R-005 | NSX Manager cluster quorum loss | Low | Critical | 3-node NSX cluster across failure domains; backup NSX config daily |
Validation Gate
Check: RCAR document contains all four sections with minimum item counts
Expected: Requirements (business + technical), Constraints (≥7), Assumptions (≥5), Risks (≥5 with probability/impact/mitigation)
Common Errors
Task 2 Design Decision Register — Core Architecture Decisions
Create a formal design decision register documenting each major architecture choice with alternatives considered and rejection reasoning.
Create Design Decision D-001 — Cluster Architecture:
Decision: 4-node management cluster + 8-node VI workload cluster (split architecture)
Alternatives considered:
A) Consolidated architecture (management + workload on same cluster) — Rejected: violates isolation requirement for SOC 2; upgrade risk affects management and workload simultaneously
B) 3-node management + 12-node workload — Rejected: 3-node management cluster cannot tolerate simultaneous host failure + maintenance; vSAN quorum at risk during rolling upgrades
C) 5-node management + 8-node workload — Acceptable but over-provisioned for management workload; 4-node provides N+1 with 25% HA admission control
Justification: 4-node management provides N+1 availability with 25% HA reservation while minimizing idle resources. 8-node workload supports 500+ VMs with growth headroom.
Create Design Decision D-002 — Storage Architecture:
Decision: vSAN ESA with FTT=1 RAID-5 (erasure coding) for workload domain
Alternatives:
A) vSAN OSA with disk groups — Rejected: ESA eliminates cache/capacity tier complexity; better dedup/compression efficiency for mixed workloads
B) vSAN ESA with FTT=2 RAID-6 — Rejected: requires minimum 6 hosts for optimal stripe width; FTT=1 RAID-5 sufficient for single-site (DR provides cross-site protection)
C) External FC storage — Rejected: violates single-vendor constraint; adds fabric complexity; higher operational cost
Justification: ESA provides 67% usable capacity (vs 50% for mirroring), native encryption, and simplified operations. Combined with cross-site DR, single-site FTT=1 meets availability requirements.
Create Design Decision D-003 — Network Architecture:
Decision: NSX with dedicated Tier-0 per workload domain, ECMP with 2 edge nodes
Alternatives:
A) Shared Tier-0 across domains — Rejected: blast radius too large; single T0 failure affects all domains
B) Physical load balancer instead of NSX — Rejected: cannot enforce micro-segmentation at VM level; no integration with VCF automation
C) NSX Federation across sites — Considered for future phase; adds complexity not justified for initial 2-site deployment
Justification: Per-domain T0 provides isolation aligned with SOC 2 requirements. ECMP with 2 edges gives active-active north-south traffic with automatic failover.
Create Design Decision D-004 — DR Strategy:
Decision: vSphere Replication with SRM for automated failover (RPO 1 hour)
Alternatives:
A) vSAN stretched cluster — Rejected: requires <5ms latency between sites; 50km distance may exceed this depending on fiber path; witness site complexity
B) Array-based replication — Rejected: using vSAN (no external arrays); not applicable
C) Zerto or third-party — Rejected: additional licensing cost; VCF-native solution preferred for operational simplicity
Justification: vSphere Replication provides RPO ≥15 minutes (meets 1-hour requirement with margin). SRM automates failover runbooks, reducing RTO. Native VCF integration, no additional licensing.
Create Design Decision D-005 — Security Architecture:
Decision: NSX DFW with zero-trust model (deny-by-default, explicit allow per application tier)
Alternatives:
A) Traditional VLAN-based segmentation — Rejected: cannot enforce intra-VLAN security; VMs on same VLAN can communicate freely
B) DFW with allow-by-default — Rejected: violates SOC 2 requirement for least-privilege access
C) Third-party firewall (Palo Alto VM-Series) — Considered as addition for north-south; DFW handles east-west
Justification: DFW provides kernel-level enforcement at every vNIC. Zero-trust model ensures all traffic is explicitly authorized. Audit logging satisfies SOC 2 requirement 10 (monitoring).
Validation Gate
Check: Decision register contains ≥5 decisions, each with ≥2 alternatives and rejection reasoning
Expected: All decisions trace back to requirements/constraints documented in Task 1
Common Errors
Task 3 Architecture Diagram Suite — Conceptual, Logical, Physical
Create the three-tier architecture diagram set that forms the visual backbone of any VCDX design document.
Conceptual Architecture Diagram:
Create a high-level block diagram showing:
- Business domains: 'Production', 'Development', 'DR Site'
- Interconnections: WAN link between Primary and DR
- User access: Corporate LAN → Load Balancer → Application Tier
- Management plane: Centralized management at Primary site
- Key callouts: 'RTO 4h / RPO 1h', 'SOC 2 Compliant', '500+ VMs'
Audience: CIO, business stakeholders
No technical details (no IP addresses, VLANs, or host counts)
Logical Architecture Diagram:
Map conceptual domains to VCF constructs:
- Management Domain: vCenter, NSX Manager (3-node), VCF Operations, VCF Automation
- VI Workload Domain: 8-node cluster, vSAN ESA datastore, NSX segments
- DR Domain: Replica vCenter, SRM, vSphere Replication appliance
- Security zones: DMZ tier (NSX T1), App tier (NSX T1), Database tier (NSX T1), Admin tier (isolated)
- NSX topology: Tier-0 (ECMP, 2 edges) → Tier-1 per zone → Segments per application - Data flows: North-South (client→app), East-West (app→db), Management (admin→vCenter)
Audience: Architects, security team
Physical Architecture Diagram:
Specify real hardware and network details:
- Rack layout: Rack A (hosts 1-4, mgmt cluster), Rack B (hosts 5-8, workload), Rack C (DR site)
- Network: Spine-leaf topology, 25G ToR (Cisco 93180YC-FX3), 100G spine
- VLANs: 10 (mgmt), 20 (vMotion), 30 (vSAN), 40 (TEP), 50-59 (workload segments)
- IP ranges: 10.10.10.0/24 (mgmt), 10.10.20.0/24 (vMotion), 10.10.30.0/24 (vSAN)
- Storage: 2× 1.6TB NVMe (cache) + 8× 3.84TB NVMe (capacity) per host
- Edge nodes: 2× dedicated hosts or VMs, 4 vCPU/32GB each
Audience: Implementation team, NetOps
Validate diagram consistency:
- Every component in Physical must map to a Logical construct
- Every Logical construct must trace to a Conceptual domain
- IP addressing must not overlap between VLANs
- vSAN network (VLAN 30) must have MTU 9000 on all switches
- TEP network (VLAN 40) must have MTU ≥1700 for Geneve encapsulation
- Management network must be routable from all sites for vCenter/NSX access
Cross-check: Count hosts in Physical (12) matches sizing calculation from requirements (4 mgmt + 8 workload)
Validation Gate
Check: Three diagrams created at increasing detail levels with full traceability
Expected: Conceptual → Logical → Physical diagrams are internally consistent; every component is traceable
Common Errors
Task 4 Design Validation & VCDX Defense Preparation
Validate the complete design against requirements and prepare for defense-style questioning.
Requirements Traceability Matrix:
Create a matrix mapping every requirement to a design decision:
| Requirement | Decision ID | How Satisfied | Verification Method |
|---|---|---|---|
| RTO 4h | D-004 (SRM) | Automated failover runbook | Quarterly DR test |
| RPO 1h | D-004 (vSphere Replication) | 15-min replication interval | Replication lag monitoring |
| SOC 2 | D-005 (DFW zero-trust) | Micro-segmentation audit | Annual penetration test |
| 500 VMs | D-001 (8-node workload) | 62 VMs/host capacity | VCF Operations capacity dashboard |
| Encryption | D-002 (vSAN encryption) | AES-256 at rest | Compliance scan via VCF Ops |
Identify single points of failure (SPOF) analysis:
- vCenter: Single instance → Mitigation: vCenter HA (active/passive) + file-level backup - NSX Manager: 3-node cluster → No SPOF (quorum-based) - vSAN: FTT=1 → Can tolerate 1 host failure → SPOF if 2 hosts fail simultaneously (mitigated by cross-site DR) - WAN link: Single fiber path? → Mitigation: Dual diverse-path WAN circuits - DNS/NTP: External dependency → Mitigation: Redundant DNS/NTP servers at each site - Power: Single UPS? → Mitigation: Dual PDU per rack, generator backup
Prepare defense responses for common VCDX panel questions:
Q: 'Why didn't you use a stretched cluster instead of SRM?'
A: 'The 50km distance between sites introduces latency risk exceeding the 5ms vSAN stretched cluster requirement. Measured RTT was 7ms on our fiber path. SRM with vSphere Replication provides the required RPO of 1 hour without the latency constraint. Additionally, stretched clusters require a witness site — adding a third location was outside our constraint of two data centers.'
Q: 'What happens if you lose the management cluster entirely?'
A: 'Workload VMs continue running on the VI workload domain — they are independent of management plane availability. We restore management from backup within 4-hour RTO using documented runbook. VCF Operations data is preserved on its own datastore.'
Q: 'How do you handle a scenario where vSAN is at 85% capacity and growing?'
A: 'VCF Operations capacity policies alert at 70% utilization with 90-day projection. Our growth plan allows adding 2 hosts to the workload cluster within the 8-host rack limit. Beyond that, we create a new workload domain. The per-core licensing model means additional hosts are licensed by adding cores to the subscription.'
Final design review checklist:
□ All requirements have at least one design decision addressing them
□ All design decisions have alternatives documented with rejection reasoning
□ All diagrams are internally consistent (no IP conflicts, correct host counts)
□ Risk register has mitigation for every High/Critical risk
□ SPOF analysis covers compute, storage, network, management, and facilities
□ Capacity planning accounts for 3-year growth projection
□ Compliance requirements mapped to specific technical controls
□ Backup and DR strategy documented with RTO/RPO commitments
□ Licensing model calculated and within budget constraint
□ Upgrade/lifecycle strategy documented (VCF LCM rolling upgrade process)
Validation Gate
Check: Complete traceability matrix + SPOF analysis + defense responses prepared
Expected: Every requirement maps to a decision; no unmitigated high risks; defense responses are specific and data-driven
Common Errors
Final Validation
Complete VCDX-grade design documentation package created
✓ RCAR framework document with all 4 sections populated → Requirements, Constraints (≥7), Assumptions (≥5), Risks (≥5)
✓ Design decision register with ≥5 decisions → Each decision has alternatives and rejection reasoning
✓ Three-tier architecture diagram set → Conceptual, Logical, Physical — internally consistent
✓ Requirements traceability matrix complete → Every requirement maps to at least one design decision
Cleanup / Restore
• Save all documentation artifacts for future reference
• Archive design decision register as template for future designs
Design Reflection (VCDX)
This lab covers the documentation cornerstone of VCDX certification. The design defense is won or lost on the quality of documentation — panelists evaluate your ability to justify decisions, not just make them.
Requirements
- Document business and technical requirements
- Create traceable design decisions
Constraints
- Budget, rack space, vendor, compliance constraints shape all decisions
Assumptions
- Network latency, hardware availability, service dependencies
Risks
- Single points of failure, capacity exhaustion, upgrade failures
Self-Assessment Discussion Prompts
- How would your design change if the budget was halved?
- What if the RPO requirement changed from 1 hour to zero (RPO=0)?
- How would you handle a requirement for 99.999% availability?
Extensions
Add VKS workload domain to the design
Extend the architecture to include a Kubernetes workload domain with Supervisor and TKG clusters, adding container networking and storage class design decisions.
Multi-tenant design variant
Redesign for 3 business units sharing the same VCF instance with tenant isolation via NSX Projects and VCF Automation multi-tenancy.
Cost optimization exercise
Calculate TCO for 3-year lifecycle including licensing, hardware refresh, operational costs. Compare VCF vs traditional 3-tier architecture.
⚠ Known Pitfalls (from Community KB)
References
- VMware Validated Design Guide (VVD) for VCF 9.0
- VCDX Application and Defense Guide
- Broadcom TechDocs — VCF 9.0 Planning and Preparation