Glossary & Quick Reference for VCDX Certification
Comprehensive reference for VMware certifications, technical acronyms, CLI commands, design decision templates, and study resources. Essential companion for all VCP, VCAP, and VCDX exam preparation.
Version Evolution
Glossary expanded with 3 new theory topics (VCDX Defense Prep, VCF Component Interaction Map, Anti-Patterns), key_takeaways added to all 5 core topics, learning_outcomes made specific, and references curated for certification rigor.
Learning Outcomes
- Understand VCDX prerequisites: VCP → 4 VCAP exams → design document submission → 75-minute panel defense with consensus approval
- Apply the Design Decision Template to structure every major decision in your VCDX document with requirement, constraint, alternatives, and risk analysis
- Demonstrate operational competence by knowing esxcli, dcli, nsxcli, and kubectl commands appropriate to each VCF component context
- Identify common design anti-patterns (over-engineering, under-specifying, ignoring day-2 ops) and articulate mitigation strategies in your defense
- Prepare for VCDX defense by whiteboarding design rationale, addressing panelist challenges, and managing 75 minutes across executive summary and technical deep-dive
Certification Path Quick Reference#
All VMware Certifications (2025 Roadmap)
- Certification
- Exam Code
- Questions
- Time (min)
- Pass Score
- Level
- Focus Area
- VCP-Cloud
- 3V0-11.24
- 85
- 150
- 300/500 (60%)
- Associate
- Cloud Concepts
- VCP-Data Center
- 3V0-12.24
- 85
- 150
- 300/500 (60%)
- Associate
- vSphere Fundamentals
- VCP-VCF
- 3V0-13.24
- 85
- 150
- 300/500 (60%)
- Associate
- VCF Fundamentals
- VCAP-DCV (Deploy)
- 3V0-21.24
- 25 scenario
- 180
- 300/500 (60%)
- Professional
- vSphere Deployment
- VCAP-DCP (Design)
- 3V0-20.24
- 25 scenario
- 180
- 300/500 (60%)
- Professional
- vSphere Design
- VCAP-VCF (Deploy)
- 3V0-22.24
- 25 scenario
- 180
- 300/500 (60%)
- Professional
- VCF Deployment
- VCAP-VKS (Deploy)
- 3V0-24.25
- 25 scenario
- 180
- 300/500 (60%)
- Professional
- Kubernetes (VKS)
- VCAP-NV (Deploy)
- 3V0-25.25
- 25 scenario
- 180
- 300/500 (60%)
- Professional
- Advanced Networking
- VCDX (Architect)
- Remote/In-Person
- Design Doc
- 75 min
- Panel Consensus
- Expert
- Enterprise Design
Prerequisite Chain for VCDX (Architect Path)
Level 1 - Associate
: Any VCP exam (VCP-Cloud, VCP-DCV, or VCP-VCF). Validates foundational knowledge.
Level 2 - Professional
: 4 VCAP exams minimum (recommended: 3V0-21, 3V0-22, 3V0-24, 3V0-25). Each must be passed within 3 years of VCDX application.
VCDX Application
: Submit design document + resume. Requires all VCAP exams passed; VCP not checked at application stage but assumed pre-requisite.
Panel Review
: 2-4 week document evaluation. ~40% of submissions require revision before defense scheduling.
Defense
: 75-minute panel Q&A. Panel consensus required (all 4 panelists vote; typically 2-3 positive votes sufficient for pass).
Most competitive VCDX candidates take 3-5 years from VCP to VCDX due to design complexity and VCAP depth. Fast-track candidates (high exam scores, strong architecture background) achieve VCDX in 18-24 months.
Key Takeaways
- VCDX prerequisite chain is strict: VCP → all 4 VCAP exams (passed within 3 years) → design doc submission → panel consensus approval. No shortcuts; panelists expect mastery across deployment (DCV, VCF, VKS, NV) contexts.
- Space VCAP exams 2-3 weeks apart and target 75%+ scores; lower scores invite deeper panelist scrutiny during defense. High VCAP performance signals design depth and speeds defense acceptance.
- Panel consensus required: all 4 panelists vote; typically 3+ affirmative votes = pass, 2 affirmative may trigger revision request. Panelists assess design quality, technical depth, communication clarity, and scalability reasoning.
- Document review takes 2-4 weeks post-submission; ~40% of candidates are asked to revise before defense scheduling. Address feedback systematically and resubmit rather than defending weak designs.
Technical Acronyms & Glossary (50+ Terms)#
A
ABX
: Aria Business eXtensions. Automation language for vSphere with Tanzu deployments; bridge Kubernetes and vSphere abstractions.
AFI/SAFI
: Address Family Identifier / Subsequent AFI. BGP routing family (IPv4 unicast = AFI 1, SAFI 1; IPv6 unicast = AFI 2, SAFI 1).
ALB
: Application Load Balancer. Avi load balancer platform (acquired by VMware 2020); provides virtual services, WAF, GSLB.
ARP
: Address Resolution Protocol. Layer-2 protocol for IP-to-MAC mapping. Proxy ARP used in stretched segments (Global Manager responds for remote site IPs).
ASN
: Autonomous System Number. BGP identifier; unique per routing domain (65000-65535 = private range).
B
BGP
: Border Gateway Protocol. Exterior routing protocol; used for multi-site and upstream DC connectivity (Tier-0 router as BGP speaker).
BUM
: Broadcast, Unknown Unicast, Multicast. Network traffic requiring replication (flooded to all destinations).
BYOA
: Bring Your Own Appliance. NSX feature allowing external vSphere VMs as edge nodes (cost-effective for greenfield deployments).
C
CLOM
: Cloud Lifecycle Operations Manager. VCF component orchestrating cluster deployment, patching, upgrades.
CNI
: Container Network Interface. Kubernetes plugin (NSX NCP, Antrea, Flannel) responsible for pod IP assignment and networking.
CCP
: Cloud Center Platform. Prelude to Aria (pre-Broadcom product).
CSI
: Container Storage Interface. Kubernetes plugin (vSphere CSI driver) bridging Kubernetes PVC requests to vSAN/external storage.
CRUD
: Create, Read, Update, Delete. Standard database operations; used in API design.
D
DFW
: Distributed Firewall. NSX stateful packet filter (microsegmentation enforcement).
DHCP
: Dynamic Host Configuration Protocol. Auto-assignment of IP addresses; vSAN uses DHCP for replica rebalancing IPs.
DOM
: Distributed Object Manager. vSAN metadata layer; tracks objects, replicas, fault domains.
DR
: Disaster Recovery. Recovery from catastrophic failure; backup/restore, replication, or failover to alternate site.
DRS
: Distributed Resource Scheduler. vSphere VM load balancing across cluster; respects affinity rules.
E
ECMP
: Equal-Cost Multi-Path. Routing load balancing across multiple equal-cost paths (NSX supports 8-way ECMP).
ESA
: Enterprise Storage Architecture. vSAN all-flash design (v6.6+); better performance/efficiency than hybrid.
ESXi
: Enterprise vSphere Hypervisor. VMware bare-metal hypervisor (kernel + management services).
etcd
: Distributed key-value store. Kubernetes control plane state storage (API objects, cluster config).
F
FTT
: Failures To Tolerate. vSAN redundancy parameter (FTT=1 = tolerate 1 disk/node failure; FTT=2 = tolerate 2 failures).
FQDN
: Fully Qualified Domain Name. Complete hostname with domain (e.g., vcenter.example.com).
G
GSLB
: Global Service Load Balancing. Avi/ALB feature for multi-site service redundancy; DNS-based failover or BGP anycast.
GUID
: Globally Unique Identifier. UUID format for vCenter objects (clusters, VMs, etc.).
H
HA
: High Availability. vSphere cluster feature; automatic VM restart on host failure (4-host per-VM priority, restarts up to 4 VMs concurrently).
HCI
: Hyper-Converged Infrastructure. Compute + storage + networking on same nodes (vSAN architecture).
HCX
: Hybrid Cloud eXtension. Workload migration tool; enables cross-version vMotion, bulk VM replication.
HTTP/2
: Hypertext Transfer Protocol v2. Multiplexing, server push; used by Avi/ALB for faster client connections.
I
IDFW
: Identity Firewall. NSX feature using Active Directory context (user/group) for microsegmentation rules.
IDS/IPS
: Intrusion Detection/Prevention System. NSX network threat prevention (passive IDS or active IPS).
IP CIDR
: IP address range notation (e.g., 10.100.0.0/16 = 10.100.0.0 to 10.100.255.255).
L
LAG
: Link Aggregation. Bundling multiple NIC links for higher throughput (LACP protocol for dynamic negotiation).
LCM
: Lifecycle Management. VCF component orchestrating cluster patching, upgrades (vSAN, NSX, vSphere concurrent coordination).
LDAP
: Lightweight Directory Access Protocol. vCenter uses LDAP for external identity provider integration (Active Directory, OpenLDAP).
LSOM
: Local Storage Object Manager. vSAN component on ESXi managing local disk pool, object replicas.
LSA
: Link-State Advertisement. OSPF routing packet flooding Link-State Database across area (Type 1-5).
LSP
: Link-State Packet. OSPF equivalent of BGP UPDATE message (carries routing changes).
M
MCTLSZ
: Multi-Cloud To Local Storage Zone. vSAN feature; manages data across multiple storage silos (capacity/performance tiers).
MTU
: Maximum Transmission Unit. Frame size (1500 bytes standard, 1550 for NSX overlays); jumbo frames = 9000 bytes.
mTLS
: Mutual TLS. Service mesh encryption; client and server both present certificates (e.g., Istio mTLS).
N
NDR
: Network Detection and Response. NSX threat detection (behavioral analysis, ML-based anomaly detection).
NCP
: NSX Container Plugin. Kubernetes CNI binding Kubernetes PVCs to NSX logical networks.
Key Takeaways
- Use correct terminology in your VCDX defense: panelists expect you to distinguish FTT (vSAN fault tolerance) from HA (vSphere cluster restart), DFW (NSX micro-segmentation) from IDFP (identity-based rules), and understand ASN/AFI/SAFI context for BGP design decisions.
- Acronyms panelists expect you to know cold: VCAP mastery requires fluent understanding of HCI/DOM/LSOM (vSAN internals), DRS/HA (vSphere availability), DFW/NSX (networking), and CSI/CNI (Kubernetes integration). Fumbling on definitions signals shallow design knowledge.
- Acronym misuse reveals design gaps: saying 'I'll use HA for DR' (conflating restart with disaster recovery) or 'DFW for north-south security' (confusing distributed firewall role) undermines panelist confidence. Demonstrate clarity by using precise terminology in every design section.
Key CLI Commands Reference#
ESXi (esxcli & dcli)
vSAN cluster info
esxcli vsan cluster get
esxcli vsan storage list
Host networking
esxcli network ip interface ipv4 get
esxcli network nic list
esxcli network vswitch standard listService management
esxcli system service list
/etc/init.d/hostd restart
PowerCLI (vSphere automation)
Connect-VIServer vcenter.example.com -User admin@example.com
Get-Cluster | Get-VM | Where-Object {$_.PowerState -eq "PoweredOn"}
vCenter (dcli - Data Center CLI)
vCenter service health
dcli com vmware vcenter services management list
dcli com vmware vcenter services management get --service wcp
Licensing
dcli com vmware cis license get
DNS/NTP config
dcli com vmware cis networkmanager hostname set --name vcenter.example.com
NSX (nsxcli)
NSX Manager cluster status
get cluster status
get managers
Routing info
- get logical-routers
- get logical-routers
- routing bgp
DFW rules
get firewall rules | grep production
Segment (logical switch) info
get logical-switches
show logical-switch
Kubernetes (kubectl)
Cluster info
kubectl cluster-info
kubectl get nodes -o wide
Pod operations
kubectl get pods -n kube-system
kubectl describe pod
-n
kubectl logs
-n
--tail=50Resource management
kubectl get pvc -n production
kubectl describe pvc
-n productionNetwork policies
kubectl get networkpolicies -A
kubectl describe networkpolicy
-nvSphere login
kubectl vsphere login --server=supervisor-ip --vsphere-username=admin@example.com
kubectl get clusters # List available TKG clusters
vSAN Performance & Health
vSAN Health check
/usr/lib/vmware/vsan/bin/vsan-cluster-debug.py
esxcli vsan cluster get
vSAN Performance Simulator
esxcli vsan perf collect
Object inspection
esxcli vsan object list
esxcli vsan object get --uuid=
Avi/ALB (ALB CLI)
SSH into SE (Service Engine)
ssh admin@
Virtual service status
show virtualservice
show pool
Analytics
show opsdb analytics
BGP neighbor status
show bgp neighbor
Key Takeaways
- Demonstrate operational competence by knowing the right CLI context for each layer: esxcli for ESXi host troubleshooting (vSAN cluster health, networking), dcli for vCenter service management (licensing, DNS/NTP), nsxcli for NSX routing and DFW rules, kubectl for Kubernetes pod and PVC inspection.
- In your VCDX defense, reference CLI examples to show you've operationalized your design: 'I validated vSAN FTT=1 with `esxcli vsan cluster get`, monitored DFW rule hits with `get firewall rules`, verified pod networking with `kubectl get networkpolicies`.' Panelists respect hands-on depth.
- Know the difference between tools: esxcli runs on ESXi, dcli/vCenter CLI runs on vCenter appliance, nsxcli/NSX CLI runs on NSX Manager, kubectl assumes a configured kubeconfig. In your design, specify which tool operators use for each operational task (day-2 monitoring, troubleshooting, patching) to show operational thinking.
Design Decision Template (Reusable)#
Copy this template for every major design decision in your VCDX document:
- ---
- Design Decision: [Title]
- Decision ID: DD-XXX
- Date: YYYY-MM-DD
Requirement
-----------
Business/technical requirement driving this decision.
Reference: REQ-YYY (Priority: P0/P1/P2)
Constraint
----------
Budget, timeline, regulatory, skill, or technology constraints limiting options.
Assumption
----------
What are we taking for granted about the environment, customer commitment, or vendor support?
Risk
----
What could go wrong? Include: Likelihood (High/Medium/Low), Impact (High/Medium/Low), Mitigation, Residual Risk.
Alternatives Considered
------------------------
- Option A: Description, pros, cons.
- Option B: Description, pros, cons.
- Chosen option: Justification.
Decision Impact
---------------
Cost impact: $ amount or percentage change.
Complexity impact: learning curve, operational overhead.
Performance impact: latency, throughput, resource utilization.
Scalability impact: future growth, multi-site expansion.
- Approval
- --------
- Approver: Name, Title
- Date: YYYY-MM-DD
- Signature: [if formal]
- Review & Update
- ---------------
- Last reviewed: YYYY-MM-DD
- Revision history: [track changes over time]
- ---
Key Takeaways
- In your VCDX design document, include a Design Decision Register (table of all DD-XXX decisions) as front-matter. Each decision should be 0.5-1 page in main document + full details in appendix. Panelists want to see systematic decision-making, not ad-hoc choices.
- Go deep on alternative analysis: don't just state 'we chose vSAN over external SAN.' Show quantitative comparison (cost, latency, IOPS, complexity), articulate which constraints ruled out each alternative (e.g., 'external SAN requires 10GbE infra we don't have'), and link back to requirements (e.g., 'HA requirement + budget <$X favored HCI').
- Quantify risk explicitly: assign Likelihood and Impact (H/M/L), state residual risk after mitigation, and map risks to design decisions. Panelists probe unquantified risk ('What if N+1+buffer fails?'); show you've thought through failure scenarios and have a mitigation story.
Study Resources & References#
Official VMware Documentation
- 📄 VCF 9.0 Official Documentation (authoritative source)
- 📄 vSphere 8.0 Documentation
- 📄 vSAN 8.0 Documentation
- 📄 NSX 4.x Documentation
- 📄 vSphere with Tanzu (VKS) Documentation
Exam Preparation Resources
VMware Learning Network (VLN)
: Official exam registration, study guides, practice tests. https://education.vmware.com/
Official Study Guides
: VMware publishes exam-specific prep guides (PDF downloads on VLN). 50-100 pages each; outlines exam objectives.
Hands-On Labs (HOL)
: Free 24-hour lab environments; practice scenarios on real VCF/NSX/vSAN. https://www.vmware.com/resources/hands-on-labs
Exam Objectives
: Each exam (VCAP, VCDX) publishes 50-80 objectives. Study materials mapped to objectives.
Practice Exams
: VCE (VCE Software) offers practice tests; ~60% coverage of real exam difficulty. Available on Pearson Vue.
Community & Blogs
VMware Communities
: https://communities.vmware.com/ (vSphere, VCF, Kubernetes forums)
Broadcom VCF Community — Holodeck Toolkit
: Authoritative source for the Holodeck release bits, JSON reference, and known-issue list (Broadcom-owned).
William Lam's Blog (williamlam.com)
: Community deep-dives, automation snippets and Holodeck walk-throughs — valuable external contributor content.
NSX & Security Blog
: NSX federation, micro-segmentation best practices.
Kubernetes on vSphere Blog
: TKG/VKS operational guides, Velero backup procedures.
Reddit /r/vmware
: Community Q&A avoid relying on for certification prep (variable accuracy).
Books & Whitepapers
"vSphere 8.0 Best Practices" (VMware Press)
: Design and operational best practices; suitable for VCDX study.
"NSX-T for Cloud-Native Infrastructure" (O'Reilly)
: Covers micro-segmentation, federation, automation.
"vSAN Design & Deployment" (VMware Whitepaper)
: Detailed vSAN sizing, performance tuning, fault domain design.
"Cloud Foundation Adoption Best Practices" (VMware Whitepaper)
: Multi-site, hybrid-cloud, licensing strategy.
Training Courses
VMware Official Training
: "VCF Foundations", "VCF Advanced", "NSX Advanced", "Kubernetes on vSphere". https://education.vmware.com/
Linux Academy / A Cloud Guru
: Self-paced video courses (cheaper than official, but less comprehensive).
vBrownBag Community
: Free weekly tech talks (YouTube); guest experts discuss VCF, NSX, Kubernetes.
VMware Hands-On Labs (HOL)
: Free 24-hour interactive labs; practice real deployments and troubleshooting.
Study Schedule Recommendation (6-Month VCDX Prep Path)
Months 1-2: VCAP Mastery
Complete 4 VCAP exams (3V0-21, 3V0-22, 3V0-24, 3V0-25) if not already passed.
Schedule exams 2 weeks apart to space out study intensity.
Use official study guides + practice exams; target 75%+ score (shows depth).
Months 2-4: Design Document Draft
Month 2: Outline structure (executive summary, current state, requirements, constraints).
Month 3: Deep-dive on logical and physical design (60% of document).
Month 4: Complete implementation, validation, appendices. Peer review (colleague review, not VMware).
Months 4-5: Document Refinement & Panelist Simulation
Revise document based on peer feedback.
Simulate panelist Q&A: Role-play with colleagues posing scenarios, challenges, trade-off questions.
Record yourself presenting; review for clarity, pace, confidence.
Month 5-6: Application & Defense Prep
Month 5: Submit application (document review takes 2-4 weeks).
Month 6: Prepare for defense (whiteboarding practice, final Q&A prep, logistics).
Panelist Q&A simulated (5+ mock defense sessions)
Key Takeaways
- Follow a structured 6-month VCDX prep timeline: 1-2 months finishing any remaining VCAP exams (target 75%+ scores), 2-4 months drafting and refining your design document with peer review, 4-5 months simulating panelist Q&A and recording your presentation for self-critique, 5-6 months submitting the application, undergoing document review (2-4 weeks), and preparing final defense logistics.
- Mock defense sessions are critical: conduct 5+ simulated defenses with colleagues acting as panelists, asking trade-off questions, challenging alternatives, and testing your whiteboarding ability. Record and review for clarity, pace, and confidence gaps. Panelists spend 75 minutes probing your design; practice preparation directly correlates with pass rates.
- Leverage free resources strategically: VMware Hands-On Labs (HOL) provides 24-hour free sandbox access to real VCF/NSX/vSAN deployments; use these to validate operational assumptions in your design (CLI commands, scaling scenarios, failure modes). Whitepaper-guided study (vSAN sizing, VCF adoption, NSX federation) ensures your design aligns with VMware best practices panelists expect.
VCDX Defense Preparation Guide#
Document Structure (50-60 Pages, 75-Minute Defense)
Executive Summary (1-2 pages)
: High-level overview of business context, customer profile, key design drivers, and expected outcomes. No technical jargon; this is for executive stakeholders and panel context-setting.
Current State Analysis (2-3 pages)
: Existing infrastructure, gaps, pain points. Include capacity, performance baseline, security posture, operational maturity. Link each gap to a design requirement (e.g., "Limited monitoring → RTM requirement for 99.99% availability SLA").
Requirements & Constraints (3-4 pages)
: Functional (workload performance, scalability), non-functional (availability, RTO/RPO, security), operational (day-2 patching, monitoring, cost), regulatory (compliance, licensing). Prioritize (P0=blocking, P1=high, P2=nice-to-have).
Logical Design (15-20 pages)
: Multi-diagram section showing layers (compute, storage, networking, security, management, monitoring). Address each requirement explicitly (e.g., "Req-022: HA Across Zones → vSphere cluster design with DRS + affinity rules + N+1 sizing"). Use decision IDs (DD-XXX) for major choices.
Physical Design (10-15 pages)
: Hardware specs, network topology, cable diagrams, IP addressing, zoning. Address capacity, power, cooling. Show how logical maps to physical (e.g., "Tier-0 GW logical → 2x NSX Edge nodes across availability zones physical").
Implementation Plan (5-8 pages)
: Phased approach (Phase 1: SDDC Manager + ESXi, Phase 2: vSAN cluster, Phase 3: NSX). Timeline, resource requirements, rollback points. Risk-driven (highest risk deployed first for validation).
Validation & Testing (2-3 pages)
: Post-implementation tests (HA failover, DFW rule validation, vSAN rebuild, vMotion scenarios). Success criteria for each (e.g., "VM restarts <5 min, zero data loss").
Appendices
: Design Decision details (full DD-XXX analysis), network diagrams, CLI command examples, cost breakdown, licensing summary, monitoring/alerting plan, runbook snippets.
Common Panelist Questions & Responses
Why This & Not That?
Q: Why vSAN instead of external SAN?
A: (Link to DD-XXX) HCI simplifies operationalization, reduces CAPEX by ~30% vs. SAN+fabric, and co-locates storage with compute for lower latency. Customer has no existing SAN investment, so external SAN adds OpEx complexity without benefit.
Q: Why NSX-T instead of VDS?
A: NSX provides east-west microsegmentation (identity-based DFW rules) required by Req-018 (identity-driven security). VDS cannot enforce policy based on Active Directory context; NSX's IDFW feature is mandatory for our compliance requirements.
What's Your Rollback Plan?
Q: If the new design fails, how do you roll back?
A: Phase 1 (vSphere upgrade) uses parallel environments; we migrate pilot workloads to new cluster, validate for 2 weeks, then migrate remaining. Old cluster remains operational during Phase 2-3 (NSX, vSAN).
Q: What if NSX deployment fails mid-phase?
A: NSX implementation doesn't impact compute until we migrate Tier-0 gateway. VDS remains active; we simply retain old gateways until NSX is production-ready. Rollback window: 4 hours to revert Tier-0 failover and resume old gateway traffic.
How Does This Scale?
Q: What's your growth capacity?
A: Design sized for N+1+20% buffer: 36 nodes current + 4 hot-spare equivalent = 40-node headroom. vSAN stretches to 64 nodes per VSAN cluster; we stop scaling at 32 nodes (half capacity) and deploy second cluster if growth continues.
Q: How do you scale across geographies?
A: Design includes stretched NSX segment and vSAN over IP for multi-site HA. Requires <10ms RTT; our sites meet this (8ms measured). If future sites exceed RTT, we deploy local SDDC and use HCX for workload mobility.
What Happens When Y Fails?
Q: What if a vSAN node fails mid-rebuild?
A: FTT=1 protects against 1 simultaneous failure. If a second node fails during rebuild (unlikely; rebuild is ~2-4 hours), that VM is at risk. Mitigation: 1. Prioritize critical VMs during rebuild (vSAN observer API). 2. Monitor rebuild progress; trigger manual remediation if >50% complete. 3. Document SLA: brief customer during rebuilds that HA is temporarily degraded.
Q: What if NSX Manager cluster quorum is lost?
A: 3-node Manager cluster; quorum survives 1 failure. If 2+ Managers fail: (1) NSX data plane (routing, DFW) continues operating independently. (2) Control plane operations (new segment creation, rule changes) blocked until quorum restored. Mitigation: geographic distribution across 3 zones; nightly backup to external storage; restore RTO <2 hours.
Presentation Tips for 75-Minute Defense
Whiteboarding (40 min)
Panelist will ask you to whiteboard key decisions (e.g., "Walk us through your compute cluster design"). Come prepared with 3-5 diagrams you know by heart:
- Compute cluster (DRS, HA, vSAN node distribution, storage tiers).
- Networking (Tier-0/1 gateways, segments, DFW enforcement points).
- Multi-site replication (NSX stretched segment, vSAN replication, RTO/RPO).
4. Failure scenario (node failure → HA restart → vSAN rebuild; explain at 90 seconds/diagram).
Time Management (75 min total)
0-5 min: Introduction (3 min presentation, 2 min panelists intro).
5-40 min: Your presentation (executive summary + design overview = 15 min, then whiteboarding = 20 min). Leave pauses for panelist clarifications.
40-75 min: Deep-dive Q&A. Panelists interrupt; expect 3-4 questions. Answer concisely (2-3 min/answer), then ask "Shall I elaborate?" to gauge depth desired.
Tip: Speak at 70% normal pace. Pause after each key point. If asked "Why?", answer in 30 seconds; if asked "How?", whiteboard it.
Scoring Criteria (What Panelists Grade)
Design Quality (40% of score)
Does the design meet all stated requirements? Are constraints addressed? Is it operationally sound?
Red flag: Missing RTO/RPO, no HA story, over-engineered for use case.
Green flag: Every requirement mapped to a design decision; constraints explicitly documented; scalability shown.
Justification Depth (35% of score)
Can you articulate WHY you chose this design? Did you consider alternatives? Can you defend against panelist challenges?
Red flag: "I chose vSAN because it's popular" or vague alternatives ("I considered SAN but it seemed worse").
Green flag: "Per DD-008, we evaluated SAN vs. vSAN using TCO analysis (cost $X, OpEx $Y, complexity score Z). vSAN won because HCI reduces OpEx 30% and simplifies patching (single LCM pipeline vs. separate storage upgrade)."
Communication Clarity (25% of score)
Do you explain concepts in plain English? Are diagrams legible? Can you whiteboard under pressure?
Red flag: Jargon soup; disorganized whiteboarding; long-winded answers.
Green flag: Concise explanations (1-2 min), clean diagrams, direct answers to questions, willingness to pivot topics.
Architectural Thinking (bonus, 10% buffer)
Do you think holistically? Do you connect design decisions (e.g., "NSX federation affects vSAN replication timing, so we must synchronize LCM patches across sites")?
Red flag: Siloed thinking ("Here's my compute, here's my storage" with no connection).
Green flag: Cross-layer insights ("vSAN rebuild impacts network I/O, so I designed NSX with QoS rules to protect vSAN traffic" or "High-frequency monitoring (Prometheus scrape every 10s) + stretched cluster = network overhead, so we use edge observability instead.").
Key Takeaways
- Your VCDX document is a 50-60 page structured case study: Executive Summary (business context) → Current State (gaps) → Requirements → Logical Design (layers, decisions) → Physical Design (hardware, topology) → Implementation Plan (phases, rollback) → Validation (tests, success criteria) → Appendices (deep-dive on each DD-XXX decision). Panelists expect systematic decision-making, not scattered ideas.
- Anticipate panelist challenges by pre-emptively addressing trade-offs in your document: 'Why vSAN over external SAN? Because HCI reduces OpEx and simplifies LCM.' 'What if a node fails during rebuild? FTT=1 tolerate 1 failure; simultaneous failures are 0.001% probability given our health monitoring and runbooks.' This proactive defense saves whiteboarding time during the 75-minute panel.
- Score highest by connecting design decisions across layers: show how NSX federation affects vSAN replication timing, how monitoring strategy (centralized vs. edge observability) impacts network load, how HA/DRS cluster design shapes vSAN fault domain strategy. Panelists award bonus points for architectural thinking that shows you understand system interactions, not just individual components.
VCF Component Interaction Map#
Core VCF Architecture
VCF 9.0 Orchestration Stack
VCF SDDC Manager (Orchestration Layer)
: Serves as the control plane for the entire SDDC. Manages lifecycle (LCM) for vCenter, ESXi, NSX, vSAN. Provides REST APIs for automation, health dashboards, and workload provisioning.
Depends on: PostgreSQL database (embedded), identity provider (vIDM or external AD), vCenter connectivity.
Version coupling: VCF 9.0 ships with vCenter 8.0 U1, ESXi 8.0 U1, NSX 4.1.0, vSAN 8.0 U1 (individual component versions may be patched independently).
vCenter (Compute Orchestration)
: Manages vSphere cluster, ESXi host licensing, resource pools. Enforces DRS (load balancing) and HA (automatic VM restart on failure).
Managed by: SDDC Manager via LCM (updates, patches, ha config).
Scaling: vCenter can manage up to 1000 hosts per instance; 4-host HA election quorum per cluster.
HA integration: Each cluster elects an HA master (typically highest host ID). Master monitors ESXi hosts; if a host fails, master triggers VM restart on healthy host per HA restart priority.
NSX (Networking & Security Orchestration)
: Tier-0 router handles north-south (external) routing (BGP, static routes, NAT). Tier-1 router handles east-west (inter-workload) routing and application connectivity. DFW enforces microsegmentation (identity-based rules via IDFW). Segments replace VLAN concept (overlay-based, encrypted).
Orchestation: SDDC Manager via LCM manages NSX cluster health, licensing, and patching. NSX Manager is deployed as 3-node appliance cluster (quorum required for control plane; data plane operates independently if Manager is down).
Version coupling: NSX 4.1.0 (VCF 9.0) supports 128 Edge clusters, 100 Tier-0 gateways, 1000 segments per Tier-1.
Integration with vCenter: NSX registers with vCenter as a plugin; vCenter web UI can provision NSX segments and DFW rules.
vSAN (Distributed Storage)
: Pools local disk capacity (SSD cache tier + HDD capacity tier in hybrid, or all-SSD in ESA mode) into a shared cluster. Each cluster has a witness node (optional for stretched cluster over IP). Distributed Object Manager (DOM) tracks objects and replicas. LSOM (Local Storage Object Manager) manages local disk pool on each ESXi host.
Ochestration: SDDC Manager via LCM coordinates vSAN patches with vCenter and NSX (serialized to prevent cascading failures).
Fault domains: You assign hosts to fault domains (zone, rack, host level). FTT=1 means tolerate 1 fault domain failure. Objects replicated across domains (e.g., FTT=1 with 3 hosts → each object has 2 replicas in 2 different domains).
Scaling: Single cluster up to 64 nodes. Stretched cluster (2 sites) supported with <10ms RTT between sites.
VCF Operations (Monitoring & Alerting)
: Dashboards for SDDC health, LCM updates, capacity trending. Built on Prometheus (metrics) and Wavefront (multi-tenant SaaS). Provides alerts for vCenter alarms, NSX alerts (DFW rule violations, BGP neighbor down), vSAN health (rebuild in progress, capacity threshold), ESXi warnings.
Integration: Ingests metrics from vCenter, NSX, vSAN, Avi (if deployed). Can feed to external monitoring (Splunk, ELK) via REST APIs.
Cloud Builder (Initial Deployment Automation)
: Orchestrates first-time VCF SDDC deployment using an OVA template. Inputs: vCenter IP, NSX Manager IPs, ESXi node list, network config (management, vSAN, vMotion, overlay VLANs). Outputs: Deployed SDDC with vCenter, NSX, vSAN, SDDC Manager ready for operations.
Runs once per SDDC; decommissioned after deployment. Used for VMware Cloud on AWS and on-premises VCF.
Component Interactions
LCM Update Flow (Concurrent Patching)
SDDC Manager initiates LCM update for vSAN → vCenter → NSX → ESXi (serialized order).
Each component update in SDDC Manager triggers that component's update manager:
- vCenter update: vCenter Update Manager orchestrates ESXi host patches, waits for vMotion to complete.
- NSX update: NSX Manager updates Managers (quorum preserved), then Edge nodes (rolling update).
- vSAN update: Rolling rebalance triggers as each host is patched; rebuild can be bandwidth-constrained (monitor vSAN resync).
- ESXi update: vMotion drains VMs, applies patch, reboots, brings VMs back.
Risk: If multiple patches run out of order, HA/DRS cluster state and vSAN replica placement can be inconsistent. SDDC Manager enforces serialized order to prevent cascades.
vCenter Cluster + vSAN Integration
vCenter creates a resource pool hierarchy; ESXi hosts are assigned to clusters.
Each cluster gets a vSAN datastore (automatically created when vSAN is enabled on the cluster).
DRS scheduler respects vSAN replica placement (tries not to put both replicas on same host).
HA restart priority queue: vCenter ranks VMs by priority (4 levels) and restarts up to 4 concurrently; remaining VMs queue. If vSAN is degraded (rebuild in progress), HA may trigger slower restarts to avoid network storms.
NSX Tier-0/Tier-1 + vSAN Overlay
NSX Tier-1 router provides default gateway for workload VMs on a segment.
Each segment is a logical network; VM traffic is encapsulated in Geneve tunnels (UDP 6081) to other ESXi hosts.
If VM on Host A communicates with VM on Host B, NSX kernel module on Host A encapsulates traffic → sends to Host B over vSAN network (or dedicated overlay network). NSX MTU consideration: geneve header adds ~50 bytes; typical MTU 1550 (support jumbo 9000 for higher throughput clusters).
vSAN and NSX can share the same physical network (vSAN VMkernel + NSX overlay on same vNIC) if designed for QoS separation; or dedicated vNICs for vSAN and overlay (reduces contention, higher cost).
Kubernetes Integration (Tanzu/VKS)
vSphere with Tanzu enables WCP (Workload Control Plane) on a vSphere cluster.
Kubernetes Supervisor cluster is provisioned on the vSphere cluster; it controls vSAN storage (CSI driver) and NSX segments (CNI plugin).
vSAN → Kubernetes: vSphere CSI driver creates Persistent Volumes (PVs) backed by vSAN storage policies (e.g., "high-perf SSD tier" → vSAN high-priority pool). NSX → Kubernetes: NSX Container Plugin (NCP) creates Kubernetes segments for pod networks. Each pod gets an IP from NSX segment, and DFW rules enforce network policies (e.g., "deny traffic between production and dev namespaces").
Scaling: A single Supervisor cluster can host 100s of TKG workload clusters. Each TKG cluster consumes vSAN and NSX resources; sizing must account for total workload cluster memory + storage footprint.
Monitoring Integration
vCenter Alarms → VCF Operations: vCenter publishes alarms (vSAN rebuild health, cluster DRS score, licensing) to Operations dashboard. NSX Alerts → Operations: DFW rule violations, BGP neighbor state changes, logical router traffic anomalies. vSAN Health → Operations: Object compliance status (replicas placed correctly), disk rebuilds in progress, capacity utilization.
Avi/ALB integration: If Avi is deployed (virtual service load balancing), its metrics (virtual service response time, pool health) also feed Operations.
Version Coupling Examples
VCF 9.0 Stack (Feb 2025)
SDDC Manager 9.0 → vCenter 8.0 U1, ESXi 8.0 U1, NSX 4.1.0, vSAN 8.0 U1 (shipped in Broadcom release package).
You can patch each component independently (vCenter to 8.0 U2, etc.), but SDDC Manager tracks supported combinations (versions listed in LCM support matrix).
NSX 4.1.0 requires NSX Managers on vCenter 7.0+. If you're on vCenter 6.5, NSX 3.x is required.
vSAN 8.0 U1 requires ESXi 8.0 U1. If you want to test vSAN on ESXi 7.0 U3, compatibility matrix shows it's supported (backward-compatible 1 major ESXi version).
VCF 5.2 Stack (Legacy Support)
vCenter 6.7 U3, ESXi 6.7 U3, NSX 6.4.0, vSAN 6.7 U3 (shipped together).
Components are tightly coupled; upgrading NSX to 7.0 requires vCenter 7.0 and ESXi 7.0 simultaneously (no piecemeal upgrades).
LCM in older SDDC Managers enforces version alignment; newer LCM (VCF 9.0) allows more flexibility (e.g., NSX 4.0 on vCenter 8.0, vSAN 8.0 on ESXi 7.0 U3).
Key Takeaways
- SDDC Manager orchestrates the entire stack via LCM, serializing vSAN → vCenter → NSX → ESXi patches in a strict order to prevent cascading failures. In your VCDX design, specify patch windows (e.g., 'monthly LCM updates on 2nd Sunday, 4-hour maintenance window') and document rollback points if a patch fails mid-sequence (e.g., 'if NSX patch fails, revert SDDC Manager LCM to previous snapshot and restart vCenter').
- Version coupling is critical: VCF 9.0 ships with vCenter 8.0 U1, NSX 4.1.0, vSAN 8.0 U1, ESXi 8.0 U1. You can patch components independently within SDDC Manager's support matrix, but NSX upgrades often require vCenter/ESXi alignment (e.g., NSX 4.1 requires vCenter 7.0+). In your design, state your target component versions and clarify any constraints (e.g., 'We're upgrading NSX 4.0 → 4.1, which requires vCenter 8.0; we must upgrade vCenter first and validate for 1 week before NSX upgrade').
- Component interactions span layers: vSAN replica placement respects vCenter DRS cluster hierarchy; NSX overlay (Geneve tunnels) shares bandwidth with vSAN replication traffic; Kubernetes CSI driver consumes vSAN storage policies; Kubernetes CNI enforces NSX DFW rules. In your design, address how these interactions are managed during operations (QoS priorities, monitoring thresholds, scaling limits) and document design decisions for each interaction point (e.g., 'DD-015: Dedicated vNICs for vSAN and NSX overlay to prevent contention during high-frequency pod churn + VM provisioning peaks').
Common Design Anti-Patterns#
Anti-Pattern 1: Over-Engineering for Non-Critical Workloads
Definition: Designing 99.99% availability (4x9s, stretched cluster, active-active replication, Avi GSLB) for dev/test or non-critical workloads that have 99.0% (1x9s) SLA.
Why Panelists Hate It
Adds 30-40% cost (extra cluster, Avi licensing, dual NSX edge nodes, stretched vSAN).
Complexity (multi-site failover, stretched vSAN requires <10ms RTT, operational overhead).
No business justification; architecture doesn't match requirement.
Red Flag Statements
- "We're deploying a stretched cluster for all our workloads because that's best practice."
- "Every application tier gets dual Avi load balancers for redundancy."
- "We're building for 99.99% availability across the board."
Corrective Guidance
Tier workloads by SLA: mission-critical (4x9s) vs. standard (3x9s) vs. development (2x9s).
Design at tier level: 4x9s gets stretched cluster + active-active, 3x9s gets local cluster + backup, 2x9s gets local cluster + manual recovery.
Example DD: "DD-004: Tiered Availability Design. Req-012 (Business SLA by workload) → 20% of workloads are mission-critical (4x9s), 60% standard (3x9s), 20% dev (2x9s). Mitigation: Mission-critical gets stretched vSAN + active-active NSX, standard gets N+1 local cluster + daily backup, dev gets N local cluster + snapshot recovery. Cost savings vs. universal stretch: 40% CAPEX reduction."
Anti-Pattern 2: Under-Specifying RTO/RPO or Business Requirements
Definition: No explicit RTO (Recovery Time Objective) or RPO (Recovery Point Objective) for workloads; assumes one-size-fits-all disaster recovery.
Why Panelists Hate It
Panelists will ask "What's your RTO for this workload?" If you answer "Uh... 4 hours?", they'll probe: "Why 4 hours? Who set that? Did the business agree?" Vague answers suggest you don't understand the customer's pain.
Design decisions (stretched cluster, backup frequency, replication lag tolerance) are justified by RTO/RPO. Without them, your design is arbitrary.
Red Flag Statements
"We'll back up everything daily." (Not tying backup to RTO/RPO.)
"Disaster recovery is handled by vSAN replication." (Not specifying failover time, data loss tolerance.)
"We'll use HCX for migration to the backup site." (Not clarifying time-to-failover, cost of active-active vs. backup-only.)
Corrective Guidance
Build a requirement table:
Workload Tier | RTO (min) | RPO (min) | Cost Justification
Mission-Critical | 15 | 5 | Stretched cluster (async replication = 5-min lag) + active-active NSX + GSLB = 15-min failover
Standard | 240 | 60 | Local cluster + hourly backup (1-hour RPO) + HCX migration on failure (4-hour RTO)
Dev | 1440 | 1440 | Snapshot-based recovery; data loss acceptable if restoring from daily snapshot
Map each design decision to RTO/RPO: "DD-008: vSAN Sync Replication. Req-005 (RTO 15 min for tier-1) + Req-011 (5-min data loss acceptable) → Sync replication ensures zero data loss but adds 10% write latency. Async replication would be 1% latency but risks 5-min data loss; sync is acceptable for this tier."
Anti-Pattern 3: Ignoring Day-2 Operations (Beautiful Design, No Operational Plan)
Definition: Design is logically and physically sound, but no monitoring, alerting, runbook, or operational maturity documented. Panelist asks "What's your monitoring strategy for vSAN health?" and you have no answer.
Why Panelists Hate It
A design is only as good as its operations. You've built a vSAN cluster, but if you don't monitor disk health, rebuild progress, or capacity trends, the cluster will fail silently (one day you're 99% full and unaware).
Day-2 is where 80% of operational cost lives. A beautiful design that's expensive or error-prone to operate is a failed design.
Red Flag Statements
"We'll monitor it with vCenter alarms." (Vague; which alarms? What's the SLA for alert response?)
"Operations will handle the details." (Deflecting; architecture should include operational thinking.)
"We don't have a runbook yet." (Suggests you haven't thought through failure recovery.)
Corrective Guidance
Include an Operational Maturity section:
Monitoring: Specify metrics and dashboards for each layer.
- vSAN: Disk health, rebuild progress, capacity utilization (by datastore/tier), object compliance (replica placement), latency (read/write I/O).
- NSX: BGP neighbor state, DFW rule hits (count violations/min), segment utilization, east-west latency p50/p99.
- vCenter: HA cluster health (master election, VM restart queue), DRS score, licensing expiration.
- Kubernetes (if deployed): Pod scheduling latency, PVC provisioning time (CSI driver), network policy enforcement.
Alerting: Define thresholds and escalation.
- vSAN rebuild time >4 hours → PAGE oncall (potential cascading failure). - DFW rule drop rate >1000 violations/min → Alert (potential security incident or misconfigured policy). - NSX BGP neighbor down → Alert within 5 min (routing instability). - vCenter HA master election in progress → Alert (cluster instability; typically resolves in <2 min, but investigate if >5 min).
Runbooks: Step-by-step procedures for common tasks.
- "vSAN Disk Failure Recovery": Detect failed disk → Check rebuild status → If >70% complete, monitor; if <30%, consider replacing disk to accelerate rebuild → Verify object compliance after rebuild. - "NSX BGP Neighbor Recovery": Check link status (ping, MTU, VLAN config) → Verify BGP config (ASN match, timers) → Restart BGP neighbor → Monitor convergence (5-10 min typical). - "vCenter HA Master Election": Trigger if current master is unresponsive → Verify HA agent status on each host → Force re-election (vCenter CLI command) → Confirm new master within 2 min.
Cost of Operations: Estimate hourly operational cost (alerts to resolve, runbook steps, mean time to resolution [MTTR]).
- Example: "DFW rule violation alerts 5x/week → 15 min each to investigate (70 min/week opex). Mitigation: Automated runbook (DFW rule violation detector + auto-remediate via NSX API) reduces MTTR to 1 min, freeing 60 min/week opex."
Anti-Pattern 4: Single Vendor Lock-In Without Justification
Definition: Design relies entirely on VMware stack (vSAN, NSX, Aria) with no escape route if licensing costs escalate or vendor strategy changes.
Why Panelists Hate It
Panelists want to see you've evaluated alternatives: What if customer budget drops 20%? What if VMware licensing becomes unaffordable? A single-vendor design with no exit strategy is risky.
You've trapped the customer into a costly renewal cycle.
Red Flag Statements
"We're using vSAN because it's VMware." (No alternative analysis.)
"NSX is the only option for security; no alternatives." (Dismissive; Kubernetes NetworkPolicy + Istio exist.)
"We can't use open-source storage because VMware doesn't support it." (Incorrect; you can use external storage with vSphere.)
Corrective Guidance
Evaluate alternatives transparently, even if VMware wins:
DD-009: Storage Architecture. Req-003 (HCI performance, simplicity). Alternatives: (A) vSAN (Broadcom, tight vSphere integration, proven at scale). (B) Ceph (open-source, lower cost, steeper ops learning curve). (C) External SAN (Dell EMC, NetApp, higher cost, better multi-tenancy). (D) Hyperconverged competitors (Nutanix, Simplivity, require separate support contract).
Analysis:
| Criteria | vSAN | Ceph | External SAN | Nutanix |
|---|---|---|---|---|
| Cost (CAPEX) | $80k | $60k | $150k | $120k |
| OpEx (annual support) | $15k | $5k | $25k | $20k |
| Ops complexity (1-5 scale) | 2 | 4 | 3 | 3 |
| Performance (IOPS/TB) | 50k | 40k | 80k | 55k |
| Integration with vSphere | 5 (native) | 2 (plugin) | 3 (iSCSI/NFS) | 1 (separate) |
| Scaling (max nodes) | 64 | 300+ | Unlimited | 128 |
| Exit cost (if switching) | High (vSAN-specific configs) | Low (portable data) | Low (portable SAN) | High (proprietary) |
Conclusion: vSAN chosen because (1) OpEx+CAPEX over 5 years ($155k vSAN vs. $180k Ceph vs. $275k SAN) favors vSAN, (2) ops complexity is lowest (tight vSphere integration = fewer separate monitoring tools), (3) performance sufficient for this workload (50k IOPS vs. 40k Ceph), (4) customer already has vSphere expertise (ops team trained). Trade-off: Exit cost is high if licensing becomes prohibitive; mitigation = plan 3-5 year refresh cycle and evaluate alternatives at renewal (budget time/cost to evaluate Ceph or external SAN options).
Justify VMware decisions with crisp reasoning, not vendor loyalty.
Anti-Pattern 5: Not Sizing for N+1+Buffer
Definition: Design spec is exactly what's needed for current workload (e.g., 36-node cluster for 100% utilization); no headroom for failure recovery or growth.
Why Panelists Hate It
If one node fails, cluster is at 98% utilization (no room for HA restart, vSAN rebuild, or DRS rebalancing).
Growth expectations (even 10-20% over 2 years) will exceed capacity within 18 months.
Operational incidents (planned maintenance, unplanned failures) have no slack.
Red Flag Statements
"We've sized the cluster to exactly match current demand." (No buffer.)
"If a node fails, we'll manually migrate workloads to the backup site." (Defeating HA automation.)
"We'll add nodes later if needed." (Reactive; design should be proactive.)
Corrective Guidance
Size for N+1+buffer:
N = nodes needed for current workload (100% utilization assumption).
N+1 = tolerate 1 simultaneous failure (e.g., ESXi host crash or vSAN disk rebuild): cluster capacity drops to N-1 nodes; all VMs + vSAN objects must still fit.
Buffer = 10-20% headroom for growth and operational slack.
Example:
Current workload: 50 TB, 200 VMs, 30 vCPU/TB (150 TB vCPU equivalent).
Current sizing: 36 nodes × 4 TB per node = 144 TB capacity. At 50 TB used, we're 35% utilized.
N+1+buffer analysis: If 1 node fails, 35 nodes remain. Can we fit 50 TB? Yes (35×4 TB = 140 TB, 50 TB is 36% utilized). BUT HA will trigger VM restarts, and DRS will need to place VMs on 35 nodes (no room for N+1 constraint). Add 4 more nodes (40 total).
Buffer analysis: Current 35% util on 40 nodes = 28% util (14 TB slack). That's 28% growth headroom before hitting 64% util (safety threshold for DRS/HA to operate comfortably). With 20% annual growth assumption, we hit 64% util in ~2 years → plan upgrade by year 2.
Result design decision: "DD-012: Cluster Sizing for N+1+20% buffer. Req-008 (HA resilience, growth accommodation). Current 50 TB → N+1 requires 40 nodes (35 can accommodate workload on failure, plus 5 hot-spare equivalent). 20% buffer = 64% target utilization threshold for DRS/HA; at that point, performance degrades if another node fails. Timeline: Monitor growth quarterly; upgrade by year 2 if trending >60% util."
Anti-Pattern 6: Mixing Design & Implementation Concerns
Definition: Design document conflates "what the system does" (design) with "how we build it" (implementation). Example: "We'll deploy vSAN using Cloud Builder OVA, which requires 100 GB disk on the appliance and 30 minutes to initialize." (Implementation detail; design should abstract this.)
Why Panelists Hate It
Design should be tool-agnostic. If your design says "Cloud Builder", you're locked into that deployment method. What if Cloud Builder fails? Do you have a manual deployment procedure?
Implementation details (CLI commands, file paths, deployment tooling) belong in runbooks, not the design document.
Panelists want to see architecture thinking, not step-by-step procedures.
Red Flag Statements
"We'll use Terraform to deploy NSX segments because it's faster." (Tool-specific; what's the design principle?)
"Cloud Builder will initialize the cluster in 30 minutes." (Implementation; design just needs to state that Cloud Builder or equivalent is used for initial setup.)
"We'll SSH into ESXi hosts and run esxcli commands to configure vSAN." (Procedure; design should abstract to "vSAN is configured via vCenter or equivalent CLI tools.")
Corrective Guidance
Separate concerns:
Design Layer: "SDDC Manager + Cloud Builder orchestrate initial deployment of vCenter, NSX, vSAN, and ESXi cluster. Design is tool-agnostic (could use Terraform, Ansible, or manual vCenter GUI); the key requirement is that initial deployment is automated to reduce human error and deployment time."
Implementation Layer (Appendix Runbook): "Initial Deployment Procedure: 1. Download Cloud Builder OVA. 2. Deploy OVA on management ESXi host. 3. Power on OVA; it starts initialization. 4. SSH into Cloud Builder: ssh root@cloud-builder.example.com. 5. Run cd /var/lib/vcf/...; ./initialize_sddc.py (30-minute process). 6. Verify SDDC status via SDDC Manager dashboard. If failure, consult troubleshooting runbook."
Anti-Pattern 7: Not Addressing Security at Every Layer
Definition: Design focuses on compute/storage/networking performance but relegates security to a single NSX DFW rule set. Missing: encryption (in-transit, at-rest), identity/access control, audit logging, compliance (FIPS 140-2, PCI-DSS).
Why Panelists Hate It
A performant system that's insecure is a liability. Panelists assume you've thought about security in every layer: compute (vSphere privilege separation, VM isolation), storage (vSAN encryption, backup encryption), networking (NSX DFW, encryption, RBAC), and management (vCenter SSO, audit logging).
Missing security layer design suggests shallow thinking.
Red Flag Statements
"NSX DFW will handle all security." (Implies compute/storage security is out-of-scope.)
"We don't need encryption because traffic stays in the data center." (Ignoring insider threats, regulatory requirements.)
"Identity management is just vCenter AD integration." (Missing RBAC per-workload, audit trail requirements.)
Corrective Guidance
Include a Security Design section (5-8 pages):
Compute Security:
- vSphere VM isolation (hypervisor security patches, vTPM for attestation).
- VM privilege separation (run non-privileged VMs in restrictive resource pools).
- Trusted execution (TPM attestation if required by PCI/HIPAA).
Storage Security:
- vSAN encryption-at-rest (AES-256 per FIPS 140-2 if required).
- Snapshot encryption (encrypted copies for backup/DR).
- Secure erase (data destruction if hardware decommissioned).
Networking Security:
- NSX DFW (microsegmentation; identity-based rules if using IDFW).
- Encryption in-transit (Geneve tunnel encryption for overlay traffic; TLS for management traffic).
- North-south (Tier-0 security appliance, WAF on Avi if exposed to internet).
Identity & Access:
- vCenter SSO (single sign-on via AD/LDAP).
- RBAC (role-based access control; example: "cluster admin" role vs. "storage viewer" role).
- Audit logging (vCenter audit logs, NSX API audit, vSAN object access logs).
- Service account management (API tokens, API key rotation, least-privilege tokens).
Compliance:
- Map design to compliance requirements (e.g., "PCI-DSS Req-2.3 requires encrypted management access" → NSX DFW limits SSH to bastion host; NSX encryption protects management traffic). - Encryption standards (AES-256 for data at-rest, TLS 1.2+ for data in-transit). - Audit retention (logs retained for 1 year per SOC 2 requirement; archived to external storage weekly). - Incident response (procedures for security event containment; example: "DFW rule violation detected → automated runbook disables segment, escalates to security team, collects NSX flow logs for forensics").
Anti-Pattern 8: Ignoring Licensing & Cost in Design Decisions
Definition: Design assumes unlimited licensing; no cost justification for dual vSAN clusters, Avi load balancers, or stretched NSX federations.
Why Panelists Hate It
VMware licensing is complex (per-socket, per-core, per-VM, subscription vs. perpetual). A design that ignores licensing can blow budget 3x over.
Panelist question: "You've proposed 2 stretched vSAN clusters + 3 NSX Managers + 2 Avi load balancers. What's the annual licensing cost?" If you have no answer, panelists lose confidence.
Red Flag Statements
"We'll deploy VMware products without considering licensing." (Naive.)
"Licensing is a finance question, not architecture." (Abdicating responsibility; architecture drives licensing cost.)
"Enterprise licenses include everything." (Wrong; different editions and feature packs have different costs.)
Corrective Guidance
Include a Licensing & Cost section (2-3 pages):
Component Licensing Model:
- vSAN: Per-socket (2 sockets per ESXi host) or per-core (more expensive, but scales better). Example: 40 nodes × 2 sockets × $10k/socket = $800k perpetual license.
- NSX: Per-managed-object (Tier-1 GW, segments). Example: 3 Tier-1 GWs + 100 segments = 103 objects × $2k = $206k.
- vCenter: Per-instance (1 vCenter can manage 1000 hosts). Example: 2 vCenters (primary + HA) × $20k = $40k.
- Avi ALB: Per-virtual-service (each load balancer virtual IP). Example: 10 virtual services × $5k = $50k.
- Subscription support (SnS): ~15-20% of license cost annually. Example: ($800k + $206k + $40k + $50k) × 18% = $266k/year SnS.
Total 5-Year Cost:
- Year 1: $1.096M licenses + $266k SnS = $1.362M.
- Years 2-5: $266k/year SnS = $1.064M.
- Total: $2.426M over 5 years.
Cost Justification (via DD):
"DD-018: Licensing & Cost Model. Req-021 (budget <$500k annual OpEx including support). Analysis: vSAN licensing is $800k upfront (perpetual; spreads to $160k/year over 5 years). NSX licensing is $206k upfront ($41k/year over 5 years). SnS is $266k/year. Total year 1: $1.362M (initial), years 2-5: $266k/year. Cost per VM: $1.362M / 500 VMs = $2.7k per VM year 1; $266k / 500 VMs = $532 per VM years 2-5. Benchmark: industry average is $2k-5k per VM annually; we're within range. Mitigation: If budget drops, reduce scope (single cluster instead of stretched, remove Avi, defer NSX advanced features)."
Trade-Offs & Mitigation:
If customer budget is $300k/year, you can't afford the full stack. Offer tiered approach:
- Year 1: vCenter + vSAN (compute/storage only) = $400k licenses + $80k SnS = $480k.
- Year 2-3: Upgrade vSAN licensing, add NSX (add $206k + $40k SnS).
- Year 4+: Add Avi if security requirements change.
This phased approach respects budget constraints and shows you've thought about cost realities.
Deferring licensing decisions to "finance will figure it out" is a sign of shallow design thinking. Panelists expect architects to balance cost, performance, and feature requirements.
Key Takeaways
- Tier workloads by RTO/RPO and SLA before designing availability: mission-critical gets stretched cluster + 4x9s availability, standard tier gets N+1 local cluster + 3x9s, dev gets cost-optimized local cluster + 2x9s. Over-engineering non-critical workloads wastes 30-40% budget; panelists spot this immediately.
- Every design decision must map to a business requirement or operational constraint, articulated in a DD-XXX decision document. Panelists ask 'Why did you choose vSAN over external SAN?' and expect a crisp answer citing RTO/RPO, cost trade-offs, and compliance drivers. Weak alternatives analysis ('vSAN seemed better') signals shallow thinking.
- Day-2 operations is 80% of cost; include monitoring strategy (metrics and thresholds), alerting (SLA for response), runbooks (MTTR targets), and OPEX estimate in your design. Panelists will ask 'What happens when a vSAN disk fails?' and expect a detailed runbook response, not 'We'll call support.'
References
- VCDX Program Overview & RequirementsTier 1 — Official
- VCF 9.0 Official Documentation HubTier 1 — Official
- VMware Cloud Foundation Design Methodology & Best PracticesTier 1 — Official