VNA Cluster Networking on Holodeck 9.1 — Distributed Networking for Management and Workload Domains, Replacing the Single-Router Model
Objectives
- Grade every available claim about VNA clusters into documented / community-verified / verify_live before deploying, and plan a VNA topology and addressing scheme on that graded evidence
- Deploy a Holodeck 9.1 management domain with both a VNA cluster and an NSX Edge cluster (-VnaClusterMgmtDomain -NsxEdgeClusterMgmtDomain), the community-verified coexistence configuration
- Empirically characterize what the VNA cluster is on the deployed estate: its appliances or components, placement, addressing (IP Spaces, Distributed VLAN Connections, gateway CIDRs), and its relationship to NSX and the HoloRouter
- Build the dependency map: which networking functions are now distributed across the VNA cluster versus which still traverse the HoloRouter or the NSX Edge cluster
- Characterize availability: blast radius of a VNA component failure versus the single-router model's total-loss failure mode, with a controlled failure probe
- Map the single-router-versus-distributed decision to production architecture — NSX edge cluster design, SPOF elimination, scale-out versus scale-up — in a full RCAR design exercise ready for panel defense
Prerequisites
Holodeck 9.1 control plane operational (holodeck91-02 completed: HoloRouter 9.1 with service stack up, VCF 9.1.0.0 content staged and checksummed, snapshot 'holodeck91-02-complete' exists). This lab deploys its OWN nested instance with VNA enabled, because the VNA switches are Day-0-only: if a non-VNA instance from holodeck91-03 occupies the host, either remove it first (Remove-HoloDeckInstance, coordinating on shared hosts) or ensure genuinely non-overlapping CIDR/VLAN allocations and sufficient capacity for parallel instances — the second option needs a large host and is the riskier path. Physical sizing: the official 9.1 single-site VCF figure is 48 CPU / 650 GB / 2.2 TB for the standard architecture; a ManagementOnly deployment (this lab's shape) runs smaller, and 9.1's soft validation will warn rather than block below recommendations — treat warnings as data, not noise. Budget extra beyond your 9.0-era ManagementOnly experience: the VNA cluster and the heavier 9.1 router are both additions whose exact cost you will measure.
Prior labs: holodeck91-02
Required skills:
- Holodeck 9.1 deployment drivers: cmdlet path and/or Click-to-Deploy UI (holodeck91-02/03)
- The 9.0-era single-router model in depth — HoloRouter as L3 gateway, DNS, DHCP, BGP speaker (holodeck-04); this lab's entire contrast rests on knowing what is being replaced
- NSX constructs: Tier-0/Tier-1 gateways, edge clusters, transport zones, segments, BGP peering
- IP addressing design: CIDR planning, IP spaces, gateway addressing
- Failure domain and blast-radius analysis (holodeck-06 Task 1 methodology)
- PowerCLI and the Holodeck query cmdlets (Get-HoloDeckSubnet, Get-HoloDeckBGPConfig, Get-HoloDeckAppNetwork)
Lab Environment
Single physical host running: (1) the HoloRouter 9.1 control plane from holodeck91-02 (external gateway, Technitium DNS, service stack — the router does NOT disappear in the VNA model; what changes is the networking INSIDE the nested estate), and (2) a fresh ManagementOnly VCF 9.1.0.0 instance deployed in this lab with -VnaClusterMgmtDomain and -NsxEdgeClusterMgmtDomain, giving the management domain BOTH networking models side by side for comparison. Community-observed distributed-networking objects on such a deployment: management-domain IP Space 'ipspace-mgmt-a' (observed value 10.1.8.0/25 at default CIDR), Distributed VLAN Connection 'distributed-vlan-connection-mgmt-a' with gateway CIDR 10.1.7.129/25, and a distributed-transit-gateway object model in which Tier-0 retrieval is explicitly skipped for distributed-mode consumers. The concrete VNA runtime footprint (appliance VMs and/or host-level components) is discovered in Task 3, not assumed here.
graph TB
WS[Operator Workstation] --> HR[HoloRouter 9.1<br/>external gateway + Technitium DNS + BGP<br/>UNCHANGED role at the estate boundary]
HR --> MGMT[Management Domain VCF 9.1.0.0<br/>vc-mgmt-a / sddcmanager-a / nsx-mgmt-a]
subgraph Domain networking - two models side by side
EDGE[NSX Edge Cluster<br/>centralized: T0/T1, edge-hosted LB<br/>the classic path]
VNA[VNA Cluster<br/>distributed: IP Spaces, Distributed VLAN Connections,<br/>distributed transit gateway - T0 skipped]
end
MGMT --> EDGE
MGMT --> VNA
EDGE -->|BGP peering| HR
VNA -->|boundary relationship = Task 3 discovery| HRIP Addressing
| Network | Purpose | VLAN |
|---|---|---|
10.1.0.0/20 | Site A management supernet (9.1 documented default MasterCIDR; /20 mandatory). Choose a non-overlapping /20 if any other instance shares the host. | Site A VLANs 0, 10-25 (9.1 documented defaults; >=16 consecutive IDs required) |
10.1.8.0/25 | Management-domain IP Space (ipspace-mgmt-a) as observed on a community-verified default-CIDR VNA deployment — verify the generated value on your build | Distributed VLAN Connection (distributed-vlan-connection-mgmt-a) |
10.1.7.129/25 | Gateway CIDR of the Distributed VLAN Connection as observed on the same deployment. NOTE: the observed mismatch between this and the IP Space range is the root of the open All Apps Org defect — see pitfalls. | Distributed VLAN Connection gateway |
10.2.0.0/20 | Site B supernet — auto-generated by New-HoloDeckConfig in 9.1 even for single-site work; parked, not deployed, in this lab | Site B VLANs 40-58 (documented defaults) |
Credentials
| System | Username | Password |
|---|---|---|
| SDDC Manager | administrator@vsphere.local | Documented master password (doubled form); doubled convention for API calls |
| Management vCenter | administrator@vsphere.local | Same SSO |
| NSX Manager | admin | Documented master password (doubled form) |
| HoloRouter (SSH / pwsh) | root | Set at OVA deployment (holodeck91-02) |
| VNA appliances (if the VNA materializes as appliances with a login surface) | verify against live 9.1 estate | Undocumented — expect the master-password convention but verify; record the actual mechanism in the findings log |
Tasks
Task 1 Grade the evidence and plan the VNA topology
manageabilityThis lab covers a capability with no 9.0 ancestor, thin documentation, and exactly one public deployment record. The discipline that keeps such a lab honest is evidence grading BEFORE action: sort every claim about VNA clusters into documented (official 9.1 docs), community-verified (the one public deployment), and verify_live (everything else), then plan the deployment so that each unknown has a task step designed to convert it into evidence. This is also the VCDX skill in miniature — architects are routinely asked to design against partially documented features, and the difference between diligence and recklessness is exactly this ledger.
Build the evidence ledger with three columns. DOCUMENTED (official 9.1 docs): the -VnaClusterMgmtDomain and -VnaClusterWkldDomain switches exist in the Management-Only and Full Stack parameter sets of New-HoloDeckInstance (not the VVF set); VNA clusters are supported for both management and workload domains; SupervisorDeploymentMode 'Distributed' rides the VNA and requires VCF 9.1.0.0+. COMMUNITY-VERIFIED (single public deployment): Edge + VNA coexist in one management domain; Day-0 succeeds with -ManagementOnly -DeployVcfAutomation -DeploySupervisorMgmtDomain Distributed -NsxEdgeClusterMgmtDomain -VnaClusterMgmtDomain; distributed-mode consumers skip Tier-0 retrieval; the object model includes IP Spaces, Distributed VLAN Connections, and a distributed transit gateway; observed values ipspace-mgmt-a 10.1.8.0/25 and gateway CIDR 10.1.7.129/25 at default CIDR. VERIFY_LIVE: VNA sizing and host minimums (explicitly open in the release intake); what the VNA physically is (VMs? per-host?); its failure behavior; its boundary relationship to the HoloRouter; its resource cost.
Document the baseline being replaced, from your own 9.0-era knowledge (holodeck-04): in the single-router model, EVERY domain function converges on one appliance — L3 gateway for all nested subnets, DNS, DHCP, BGP peering with NSX Tier-0. Write the failure-mode table for that model: router loss = total loss of external access, inter-VLAN routing, name resolution, and route advertisement simultaneously.
Plan the deployment shape: ManagementOnly, OSA vSAN mode (smaller footprint), with BOTH -NsxEdgeClusterMgmtDomain and -VnaClusterMgmtDomain so the two models coexist for direct comparison — the community-verified combination. Decide your CIDR (default 10.1.0.0/20 if the host is otherwise clean; a non-overlapping /20 if not) and VLAN range (>=16 consecutive IDs for Site A).
Run the capacity gate: record physical host CPU/memory/datastore free, subtract the running HoloRouter's actual consumption, and compare the remainder against your 9.0-era ManagementOnly experience PLUS explicit unknown-cost margin for the VNA cluster and edge cluster. The official full-stack 9.1 figure (48 CPU / 650 GB / 2.2 TB) is the ceiling reference; your ManagementOnly shape sits below it by an amount you will measure, not assume.
Decide the deployment driver and pre-stage it: cmdlet path (the community-verified shape — SSH to router, pwsh, the New-HoloDeckInstance command from step 3) or the Click-to-Deploy UI (holodeck91-03 method — locate where the UI exposes the VNA option and record the control's exact wording as a findings entry). Verify content staging for VCF 9.1.0.0 (/holodeck-runtime/bin/9.1.0.0/) is intact from holodeck91-01/02.
Snapshot the control plane before the deployment run (if anything changed since 'holodeck91-02-complete'), and on a shared host, announce the deployment window and your CIDR/VLAN claims to co-tenants.
Validation Gate
Check: Evidence ledger complete with all claims graded; single-router failure-mode baseline written; deployment plan (parameters, CIDR, VLANs, driver) documented; capacity gate passed with explicit margins; content staging verified; snapshot and shared-host coordination done.
Expected: A deployment plan in which every unknown is named and assigned a step that will convert it to evidence.
Common Errors
Task 2 Deploy the VNA-enabled management domain
manageabilityThe deployment itself is procedurally familiar from holodeck91-02/03 — what makes this run different is what you watch for: every VNA-attributable stage, object, and resource claim is a first recording for your environment. Treat the deployment log as a primary source to be captured, not a progress bar to be waited out.
Start the deployment via your chosen driver. Cmdlet form (community-verified shape, adjusted to your plan): New-HoloDeckInstance -Version 9.1.0.0 -InstanceID <id> -vSANMode OSA -Site a -ManagementOnly -NsxEdgeClusterMgmtDomain -VnaClusterMgmtDomain [-DepotType <Online|Offline>] — capture the full transcript. UI form: trigger the pipeline with the equivalent configuration committed to Git (holodeck91-03 method) and follow the run in GitLab.
While the run progresses, watch three planes (the holodeck91-03 observation discipline): the driver transcript/pipeline log, vCenter on the physical host (VM materialization order), and the HoloRouter (Technitium record creation, BGP state via vtysh or Get-HoloDeckBGPConfig once available). Specifically note WHEN VNA-related objects first appear and what stage names the log uses for them.
On completion, run the standard bring-up validation gate before touching anything VNA-specific: Get-HoloDeckInstance shows Running; SDDC Manager, vCenter, and NSX Manager UIs reachable and green; nested hosts commissioned; vSAN healthy.
Capture the post-deployment resource claim: physical host CPU/memory/datastore consumed, per-VM inventory with sizes (vCenter > physical host view), and the delta against your Task 1 capacity-gate figures.
Snapshot the deployed estate: all instance VMs plus the router, labeled 'pre-vna-probe'. Task 4 deliberately degrades things; this is its safety line.
Validation Gate
Check: Deployment completed with full transcript captured; three-plane timeline with VNA-attributable stages highlighted; standard health gate green; per-VM resource ledger recorded; 'pre-vna-probe' snapshot taken.
Expected: Healthy VNA+Edge management domain with the deployment itself already yielding first-recording evidence (timings, stage names, soft-validation behavior).
Common Errors
Task 3 Discover and map the VNA data plane
manageabilityThis is the lab's investigative core: answering 'what IS a VNA cluster on this build?' with inventory evidence. The method is a structured sweep across every surface that could hold VNA state — vCenter inventory, NSX object model, the toolkit's query cmdlets, the router's BGP view, and the distributed-networking objects the community deployment exposed (IP Spaces, Distributed VLAN Connections, distributed transit gateway). The output is the topology diagram and dependency map that everything in Tasks 4-5 argues from.
vCenter sweep: diff the deployed VM inventory against a known 9.0-era ManagementOnly layout (hosts, VCF Installer, SDDC Manager, vCenter, NSX, edge nodes). Identify every VM you cannot attribute to the classic layout — candidate VNA appliances. Record names, sizes, host placement, and network attachments. Also sweep host-level: any new per-host agents/components visible in the host clients.
NSX sweep: in NSX Manager, inventory what the VNA model added versus the classic edge model — examine Tier-0/Tier-1 gateways (which exist, which are attached to what), segments, transport zones and transport node types, and any distributed-gateway or VPC-style constructs your NSX 9.1 build surfaces. Explicitly locate where the edge cluster's objects end and the VNA-attributable objects begin.
Toolkit sweep: from the router pwsh session (Import-HoloDeckConfig -ConfigId <id> -Site a first), run the query cmdlets and capture output: Get-HoloDeckSubnet, Get-HoloDeckAppNetwork, Get-HoloDeckAppIpPools, Get-HolodeckServiceIPPools, Get-HoloDeckOverlaySubnet, Get-HoloDeckBGPConfig. Identify which entries are VNA-attributable (new pools, new service IPs, changed BGP neighbor set).
Distributed-networking object sweep: locate the IP Space and Distributed VLAN Connection objects for the management domain (surfaces to check: NSX, VCF Operations/Automation networking views, and the SDDC Manager API). Record their names, ranges, and gateway CIDRs on your build, and compare against the community-observed values (ipspace-mgmt-a = 10.1.8.0/25, gateway CIDR 10.1.7.129/25 at default CIDR).
Boundary sweep: establish the VNA-to-HoloRouter relationship. From the router: routing table (ip route / vtysh), BGP neighbors, and which nested networks the router still originates or forwards for. From the estate: trace the north-south path of a management-network packet and (conceptually, from the object model) of a distributed-path packet.
Synthesize: draw the full topology diagram (workstation -> HoloRouter -> [edge path | VNA path] -> management components), attribute the Task 2 resource ledger rows to VNA vs edge vs classic components, and write the one-paragraph 'what a VNA cluster is on this build' statement with every clause traceable to a sweep finding.
Validation Gate
Check: All five sweeps completed with captured output (vCenter diff, NSX two-column inventory, toolkit query outputs, distributed-object values, boundary map); topology diagram drawn; resource ledger attributed; definition paragraph written; evidence ledger updated.
Expected: The VNA cluster characterized empirically end to end — components, addressing, object model, boundary behavior, and measured cost.
Common Errors
Task 4 Characterize availability: blast radius, dependency map, and a controlled failure probe
availabilityThe whole point of distributed networking is what happens when something breaks. This task converts the architecture claim ('replaces the single-router SPOF model') into a measured availability characterization: a dependency table saying which functions die with which component, and one controlled failure probe to test the sharpest prediction. This is rn-9-1-015's defense-relevance-HIGH content: SPOF elimination and blast radius are exactly what a panel probes on any networking design.
Build the dependency table from Task 3's boundary map. Rows: external access to management UIs, inter-VLAN routing inside the estate, DNS resolution (infrastructure + estate zones), DHCP, route advertisement, distributed-path consumer networking (e.g. a Distributed-mode Supervisor if present), edge-path consumer networking. Columns: HoloRouter fails / one VNA member fails / whole VNA cluster fails / edge cluster fails. Fill each cell with the PREDICTED outcome and mark confidence.
Choose the probe: the sharpest testable prediction is the single-VNA-member failure (IF Task 3 found a multi-member VNA cluster). Verify the 'pre-vna-probe' snapshot line exists, confirm you are NOT on a shared host mid-someone-else's-work, and define the probe protocol: what you will power off, what you will measure (the dependency table rows), your observation window, and your restore procedure.
Execute the probe: power off the chosen component (VNA member or edge node per your protocol). Work through the dependency table rows measuring actual outcomes: management UI reachability, name resolution, distributed-path consumer behavior, edge-path consumer behavior, NSX alarm state. Record detection latency — how long before each affected surface visibly degraded — and which surfaces stayed clean.
Restore: power the component back on, verify convergence (NSX health, dependency rows back to baseline), and record recovery time and whether any function needed manual intervention to recover.
Write the blast-radius comparison artifact: single-router model (total loss, one failure domain) versus the measured 9.1 estate (per-component columns with tested and predicted cells marked distinctly). Add the honest limits: what remains a SPOF even in the distributed model (the HoloRouter at the boundary; the physical host under everything), and which cells remain predictions.
Validation Gate
Check: Predicted dependency table complete; probe protocol written with rollback; probe executed with measured outcomes and detection/recovery timings; estate restored to green; blast-radius comparison artifact written with tested/predicted cells distinguished and residual SPOFs named.
Expected: A measured availability characterization of distributed VNA networking versus the single-router model — the core defense evidence of this lab.
Common Errors
Task 5 Design exercise — single-router versus distributed VNA as a production architecture decision (RCAR)
availabilityThis is the VCDX defense moment and the reason rn-9-1-015 carries HIGH defense relevance. The lab's measurements become a production design argument: when does a networking design concentrate services on a shared appliance tier, and when does it distribute them? The Holodeck contrast (single HoloRouter vs VNA cluster) is the same decision production architects make about NSX edge cluster design — dedicated versus shared edges, scale-up versus scale-out, service placement, failure domains. The task follows the holodeck-06 pattern: scenario, RCAR, quality-matrix trade-offs, and the panel questions asked out loud.
Define the production scenario your lab contrast maps to. Example frame: a VCF platform serving 40 workload teams, currently with all north-south, load balancing, and VPN services concentrated on one shared 2-node NSX edge cluster (the 'single-router model' at production scale); the proposal is distributing networking services closer to consumers (per-domain edge/VNA-style placement). State the business drivers: blast-radius reduction after a company-wide outage caused by an edge failure, and scale headroom for 3x team growth.
REQUIREMENTS (5+): e.g. (a) failure of any single networking component must not affect more than one domain's consumers; (b) networking capacity must scale horizontally with domain count without forklift upgrades of a central tier; (c) per-domain networking changes must be independently deployable and reversible; (d) the platform team must retain central observability across distributed components; (e) route propagation between distributed components and the physical network must be dynamic and convergent within stated SLOs.
CONSTRAINTS (4+), drawn honestly from the lab: (a) distributed placement is a deploy-time decision on this platform generation — retrofit means redeploy-and-migrate (your Task 1 Day-0-only finding, generalized); (b) the distributed model is newer with a thinner operational corpus and known early defects (the IP Space/gateway CIDR mismatch class); (c) distributed components multiply the fleet to patch, monitor, and certify — operational cost scales with distribution; (d) some functions remain necessarily centralized (boundary routing, external peering — your residual-SPOF finding); (e) sizing guidance for the distributed tier is immature and must be established by measurement.
ASSUMPTIONS (4+) and RISKS (5+, each with impact and mitigation). Assumptions: e.g. distributed components fail independently (partially TESTED in Task 4 — cite the cell); central observability tooling covers the distributed tier; teams consuming distributed networking do not require cross-domain L2. Risks must include: distributed-tier version/config drift across many components (mitigation: the same automation-owns-everything discipline from the Day-2 labs); operational-maturity gap on the new model (mitigation: staged adoption — one domain first, measured, then expand); hidden coupling between distributed and central tiers (mitigation: the Task 4 probe methodology promoted to a production failure-testing program); monitoring gap turning distributed into invisible (mitigation: per-component health SLOs before go-live).
Write the quality-matrix trade-off row for all five dimensions. Availability: distributed wins on blast radius (measured), but multiplies components that can individually fail — state the net honestly. Manageability: central is one thing to run; distributed is many things to run identically — automation is the deciding factor. Performance: distribution removes a choke point and shortens paths; it also forfeits the central tier's easy traffic-engineering single point. Recoverability: smaller failure domains recover faster, but fleet-wide recovery (bad config pushed everywhere) is now a coordinated operation. Security: distribution multiplies enforcement points (more places to get policy wrong) while shrinking the value of any single compromised component.
Answer the panel questions in writing, out loud, or to a study partner — these are the questions this design walks into: (1) 'Your distributed model still has a central boundary router. Have you eliminated the SPOF or relocated it?' (2) 'Show me the failure test that proves domain isolation. Now show me the one that proves the central tier failing does not take the domains with it.' (3) 'You cited blast radius. What is your MTTR story when the failure is a bad configuration deployed to ALL distributed components simultaneously?' (4) 'The distributed feature is one release old. Defend choosing it over the proven centralized design for a customer whose primary requirement is operational stability.' (5) 'What telemetry tells you a distributed component is degraded before its consumers do?' (6) 'Scale this to 200 domains: what breaks first in your design — addressing, route table size, monitoring, or the operations team?'
Close the loop on the workload-domain dimension: rn-9-1-015 covers VNA for BOTH domain types, and -VnaClusterWkldDomain remains untested in this lab (ManagementOnly shape). Write the extension plan: what a workload-domain VNA deployment would add to the evidence (per-domain isolation between mgmt and WLD distributed tiers — the closest lab analog to the production per-domain argument), and its cost. File it as the follow-on.
Validation Gate
Check: Scenario defined with quantified drivers; RCAR complete with lab-evidence citations or explicit untested flags; five-row quality trade-off matrix with net positions; all six panel questions answered in writing; workload-domain follow-on plan filed.
Expected: A defensible, evidence-graded production design argument for the single-router-versus-distributed decision — the artifact rn-9-1-015's HIGH defense relevance demanded.
Common Errors
Final Validation
A VCF 9.1.0.0 management domain was deployed with both a VNA cluster and an NSX Edge cluster (the community-verified coexistence configuration), and the new distributed networking model was characterized end to end: components and placement discovered empirically, addressing and object model (IP Spaces, Distributed VLAN Connections, distributed transit gateway path with Tier-0 skipped) recorded from the live build, the VNA-to-HoloRouter boundary mapped, resource cost measured against the open sizing question, and availability characterized through a predicted dependency table plus a controlled failure probe. The blast-radius comparison against the single-router model and the full RCAR design exercise convert rn-9-1-015 — the flagged content gap with HIGH defense relevance — into evidence-graded VCDX defense material.
✓ Task 1: Evidence ledger (documented / community-verified / verify_live), single-router baseline, deployment plan, capacity gate → Every unknown named and assigned a converting step
✓ Task 2: VNA+Edge ManagementOnly instance deployed with transcript, three-plane timeline, resource ledger, 'pre-vna-probe' snapshot → Healthy estate with first-recording deployment evidence
✓ Task 3: Five-sweep discovery (vCenter, NSX, toolkit, distributed objects, boundary), topology diagram, attributed resource ledger, definition paragraph → VNA cluster empirically characterized; sizing open question closed for this environment
✓ Task 4: Dependency table (predicted + measured), controlled failure probe with recovery, blast-radius comparison artifact with residual SPOFs named → Measured availability characterization, honestly graded
✓ Task 5: Scenario, RCAR with evidence citations, quality trade-off matrix, six panel answers, workload-domain follow-on plan → Panel-ready design argument for centralized versus distributed networking
Cleanup / Restore
Snapshot: holodeck91-10-complete
• Snapshot all instance VMs plus the HoloRouter as 'holodeck91-10-complete' — this instance is the VNA-enabled estate holodeck91-09 Task 4 (Path B) uses for Distributed-mode Supervisor work; keep it unless capacity forces teardown
• Delete the 'pre-vna-probe' snapshot line only after 'holodeck91-10-complete' is boot-verified
• Archive the artifact set to the VCDX evidence folder: evidence ledger, topology diagram, attributed resource ledger, dependency table with probe measurements, blast-radius comparison, RCAR worksheet, panel answers, deployment transcript
• Update the release-intake status for rn-9-1-015: record which open questions your environment closed (VNA sizing measured, component form identified, boundary behavior mapped) and which remain (workload-domain VNA, member-level resilience if untested on your build) — as 9.1-scoped annotations, never edits to 9.0-era docs
• If the estate must come down instead (capacity or shared-host reasons): Remove-HoloDeckInstance with shared-host coordination, then verify the router's global state (Technitium records, BGP, port groups) survived the removal cleanly — and record that observation, since instance-removal hygiene under the 9.1 router model is itself unverified
• Brief co-tenants and colleagues: the VNA findings log is the only field documentation of this surface in your environment — file it where the team's 9.1 references live
Design Reflection (VCDX)
This lab exists because rn-9-1-015 is the rare toolkit change that maps one-to-one onto a first-order production design decision: WHERE DO NETWORK SERVICES LIVE — concentrated on a shared appliance tier, or distributed toward their consumers? The single-router Holodeck model is the concentrated design in miniature: one appliance carrying L3, DNS, DHCP, and BGP for everything, with the total-loss failure mode your Task 1 table documents.
The VNA cluster is the distributed counter-design: per-domain networking, measured smaller blast radius, horizontal scale — at the price of more components, a Day-0-only adoption path, immature sizing guidance, and a first-release defect record.
The production parallel a panel will expect you to draw is NSX edge cluster design: the shared-edge-cluster-for-everything pattern versus dedicated per-domain edge clusters (and, in current VCF, VPC-style distributed constructs) is the same decision at scale, with the same trade-off skeleton — blast radius and scale-out versus operational simplicity and maturity. Three positions to hold under pressure.
(1) SPOF ELIMINATION IS REALLY SPOF RELOCATION-AND-SHRINKING: your estate still has the HoloRouter at the boundary and the physical host under everything; the honest claim is that distribution removes the SHARED-FATE coupling for the most common failures while leaving a smaller, simpler boundary SPOF whose own HA is the next design iteration. Candidates who claim total SPOF elimination get dismantled by the first follow-up.
(2) BLAST RADIUS VERSUS FLEET RISK: distribution converts one big failure domain into many small ones — but it also creates a new fleet-wide failure mode (a bad change deployed to every distributed component at once) that the centralized design literally could not have. The mitigation is change-management maturity (staged rollout, canary domains), which means the distributed design's availability is contingent on operational discipline in a way the centralized design's is not. That contingency belongs IN the design document.
(3) ADOPTION TIMING: the feature is one release old, its first public deployment surfaced a defect within weeks, and its sizing is unpublished — your evidence-grading ledger is the professional answer to 'should we adopt now?': adopt where the blast-radius requirement is binding and the operational team can absorb a measurement-driven rollout; hold where stability is the binding requirement.
Bring the ledger itself to the defense — the method of grading claims into documented / community-verified / measured / assumed IS the demonstration of architectural maturity, independent of which way the decision goes.
Requirements
- The nested lab must model distributed domain networking (VNA cluster) alongside the classic edge path in one management domain, so both models are directly comparable on identical infrastructure
- Every VNA claim entering design documentation must be graded: documented, community-verified, measured-in-lab, or assumed
- The availability characterization must include at least one controlled, snapshot-protected failure probe with measured detection and recovery
- Resource cost of the VNA cluster must be measured and attributed, closing the release-intake sizing question for this environment
- The estate must remain usable as the VNA-enabled platform for holodeck91-09's Distributed-mode Supervisor work
- All findings must be filed as 9.1-scoped documentation without modifying any 9.0-era reference (version coexistence policy)
Constraints
- VNA cluster deployment is Day-0 only: -VnaClusterMgmtDomain / -VnaClusterWkldDomain exist solely on New-HoloDeckInstance parameter sets — no documented Day-2 retrofit, so the decision is irreversible short of redeploy
- VNA sizing and host-count minimums are unpublished (release-intake open question) — capacity planning is by margin and measurement, not by table
- The VNA's concrete runtime form is undocumented; all characterization is empirical against one build and graded accordingly
- Exactly one community-verified VNA deployment existed at authoring time — there is no failure-mode corpus; novel breakage has no external answers
- The distributed path has a known open defect (IP Space range vs Distributed VLAN Connection gateway CIDR mismatch affecting the VCFA All Apps Org flow) even at default CIDR
- Physical host capacity bounds the comparison: VNA cluster + edge cluster + management domain + 9.1 router must coexist under the ~90% ceiling; 9.1 soft validation warns but does not block undersizing
- The VVF parameter set carries no VNA switches — this lab's content requires a VCF (not VVF) deployment
Assumptions
- The community-verified parameter combination (ManagementOnly + Edge + VNA) deploys successfully on this build as it did in the public record — re-verified by this lab's own Task 2 run
- The HoloRouter retains its estate-boundary role (external gateway, Technitium DNS, BGP) under the VNA model — tested by the Task 3 boundary sweep rather than assumed silently
- Distributed-networking objects observed publicly (IP Spaces, Distributed VLAN Connections, transit gateway path) exist in comparable form on this build — matched or divergence-noted in Task 3
- A single controlled component failure is recoverable via power-on or the snapshot line — probe protocol assumes no destructive persistence
- Single-operator host, or explicit shared-host coordination for the deployment window, CIDR/VLAN claims, and the failure probe
Risks
- Undocumented-surface misinterpretation: naming a VNA component or path wrongly in the findings — IMPACT: wrong mental model propagates into design documents and defense answers; MITIGATION: evidence grading, per-claim citations, and describing the build's own vocabulary rather than imported terminology
- Day-0 irreversibility: deploying without VNA (or with wrong CIDR) and discovering the need later — IMPACT: full redeploy of a multi-hour estate; MITIGATION: Task 1's parameter plan review, and the coexistence configuration that keeps both models available from the start
- Failure probe cascade: the probe degrades more than predicted on an undocumented component — IMPACT: estate instability, potential rebuild; MITIGATION: snapshot-first protocol, single-action probes with defined observables, restore-then-analyze discipline
- Capacity overrun: VNA + edge + domain exceeds the host under soft validation's permissive gate — IMPACT: estate-wide degradation misattributed to the new feature; MITIGATION: explicit headroom margins for unknown-cost components, warnings treated as decisions not noise
- Known-defect adjacency: the IP Space / gateway CIDR mismatch pre-staged in the generated config — IMPACT: later VCFA workflows fail mysteriously; MITIGATION: values recorded in Task 3 with the mismatch flagged, workaround referenced but applied only when the affected workflow is exercised
- Evidence staleness: this lab's findings harden into 'facts' as the product evolves — IMPACT: future designs cite superseded observations; MITIGATION: findings filed version-scoped with build identifiers and dates, per the coexistence policy, and re-graded at the next toolkit release
Self-Assessment Discussion Prompts
- State precisely what 'the VNA cluster replaces the single-router model' means on your build — what moved off the router, what stayed, and what evidence backs each clause. Now defend the claim that this is a model replacement rather than just an additional component.
- Your blast-radius table shows the distributed model containing per-component failures. A panelist responds: 'You built a new fleet-wide failure mode — one bad config pushed to every VNA. Compare the expected frequency and impact of that against the central-router failure it replaced.' Answer with the change-management contingency stated explicitly.
- The VNA decision is Day-0-only on this platform generation. Generalize: how do you handle architecturally significant capabilities that cannot be retrofitted? What does that constraint do to your Day-0 decision checklist, and where else in VCF design does the same pattern appear?
- Map your lab contrast onto production NSX edge design: shared edge cluster for all domains versus dedicated per-domain edges versus distributed constructs. Which lab measurement transfers to that decision directly, which transfers as hypothesis, and which does not transfer at all?
- You proved (or predicted) that losing one VNA member spares other consumers. Design the production monitoring that would detect that member's failure before any consumer notices — and state what you would measure to prove the monitoring works.
- The feature is one release old with one public deployment and a known defect. Write the two-paragraph adoption recommendation for a risk-averse customer whose outage history argues FOR distribution — hold the tension honestly rather than resolving it by fiat.
- Scale the distributed model to 200 domains: walk through addressing (IP Space proliferation), route propagation, monitoring fan-out, and operational load, and name what breaks first with your reasoning.
- Your estate retains the HoloRouter as boundary SPOF. Propose its HA design as the next iteration, and explain why fixing the boundary AFTER distributing the domains is (or is not) the right order.
Extensions
Workload-Domain VNA — the Second Half of rn-9-1-015
Execute the Task 5 step 7 plan: deploy a full-stack instance with -VnaClusterWkldDomain (alongside or instead of the management-domain VNA) and repeat the discovery and dependency methodology on the workload domain. The mgmt-vs-WLD isolation measurement is the closest lab analog to the production per-domain blast-radius argument, and the switch is entirely community-unverified — you would produce its first public-quality record.
harderDistributed Supervisor on Your VNA (Bridge to holodeck91-09)
Run holodeck91-09's Task 4 against this instance: deploy a Supervisor in Distributed mode and extend your topology diagram with the consumer layer — where the control plane VIP lands in the VNA model, and how the 'Skipping T0' path terminates. The two labs' artifacts merge into one full-stack distributed networking story.
sameWalk Into the Known Defect Deliberately
Exercise the VCFA All Apps Org flow against your distributed-mode domain with the IP Space / gateway CIDR values you recorded in Task 3, reproducing (or failing to reproduce) the open mismatch defect, then apply the published workaround and document the full cycle. Converts a pitfall entry into first-hand defect literacy.
harderFull Failure Matrix
Extend the single probe into the complete dependency-table campaign: fail each component class (VNA member, whole VNA, edge node, HoloRouter) under active load, measuring every table cell. Produces the measured availability matrix the design exercise currently grades as part-predicted — the strongest possible artifact for the defense.
much harderBGP and Route-Propagation Deep Dive
Characterize the routing control plane end to end under the VNA model: every BGP speaker, every peering, route origination for the distributed networks, and convergence timing when a path fails. Compares directly against the 9.0-era router-to-T0 peering model and answers the 'how do routes actually move' question your boundary sweep opened.
harder⚠ Known Pitfalls (from Community KB)
References
- VCF 9.1 Release Intake — rn-9-1-015 (Holodeck VNA cluster networking)
academy-v2/content/releases/vcf-9.1.jsonlocal fileTier 1 — Official
The canonical changelog entry this lab exists to cover: VNA clusters for management and workload domains, distributed networking replacing the single-router model, defense relevance HIGH, sizing/host minimums flagged open. - Holodeck 9.1 Documentation — New-HoloDeckInstance parameter setsTier 1 — Official
Authoritative for the -VnaClusterMgmtDomain / -VnaClusterWkldDomain switches (Management-Only and Full Stack sets; absent from the VVF set), the string-typed Supervisor mode parameters, /20 CIDR and VLAN-range rules, and 9.1 resource tables with the new soft validation. Use the versioned /9.1/ URLs. - HoloDeck 9.1 Release Announcement (verbatim capture)
Holodeck-9.1/01-release-announcement.mdlocal fileTier 1 — Official
Primary capture of the headline claim: 'New Virtual Network Appliance (VNA) cluster support for both management and workload domains.' - HoloDeck Toolkit 9.1 — Structured Reference & Changelog
Holodeck-9.1/02-structured-changelog.mdlocal fileTier 1 — Official
Section 2.2 (VNA as the headline architectural change) and section 7 (VNA sizing/host minimums as an open question this lab measures). - Holodeck GitHub Issue #143 — the community-verified VNA deployment recordTier 1 — Official
The single public VNA deployment at authoring time: verified Day-0 command (Edge + VNA + Distributed Supervisor + VCFA, ManagementOnly), observed object values (ipspace-mgmt-a 10.1.8.0/25, gateway CIDR 10.1.7.129/25, distributed-vlan-connection-mgmt-a), the 'Skipping T0 retrieval as distributed mode selected' log line, and the open IP Space defect with workaround. - holodeck-04 — External Access, DNS, and Routing Configuration (9.0 single-router baseline)
holodeck-04.jsonlocal fileTier 2 — VMware Press
The definitive record of the single-router model this lab contrasts against: HoloRouter as L3 gateway, DNS, DHCP, BGP speaker, and total-loss SPOF. Remains authoritative for 9.0.x. - holodeck-06 — Dual-Site Holodeck Deployment (structural model for the design exercise)
holodeck-06.jsonlocal fileTier 2 — VMware Press
Source of this lab's Task 1/Task 5 methodology: topology planning, failure-domain matrices, and the RCAR design exercise pattern with its anti-generic-RCAR guidance. - holodeck91-09 — Day-2 Part 2: Supervisors in Edge and VNA modes
holodeck91-09.jsonlocal fileTier 2 — VMware Press
The consumer-layer companion: Distributed-mode Supervisors ride this lab's VNA cluster; its Task 4 Path B uses this instance. - VMware Cloud Foundation 9.x Networking Documentation (Broadcom Techdocs)Tier 2 — VMware Press
Production-side reference for NSX edge cluster design, VPC-style constructs, and the distributed networking concepts the design exercise maps onto — read AFTER forming your lab-evidence view, to avoid importing vocabulary the build does not use. - VCF Holodeck Toolkit — Broadcom Community ForumTier 1 — Official
Community intake. No VNA thread existed at authoring time beyond the GitHub record — this lab's findings log is candidate seed content for the first one.