Architect-level delta between VCF 9.0 and VCF 9.1 (GA May 2026): the unified VCF management services runtime, the API-first platform story, patching and LCM improvements (vCenter quick patch, single-API resize, ESX live patch on TPM hosts), NVMe memory tiering, fleet password policies, the VCF Operations pillar reorg, VKS 3.6 and multi-vNIC cluster networking, self-service CaaS namespaces, native S3 object storage (Tech Preview), LACP in the deployment UI, and Holodeck 9.1 changes. Each topic covers what changed, why it matters for design and operations, and how it shows up in exams and VCDX-style defenses.
Version Evolution
This section is the 9.0→9.1 delta companion to section 18 (VCF Evolution 2.x→9.0). VCF 9.1 GA'd May 2026. Headline deltas: VCF management services common runtime; API-first with consistent OpenAPI across Python/Java SDKs, PowerCLI, Terraform; vCenter quick patch (<1 min, sometimes zero downtime) and single-API resize; ESX Live Patch on TPM-enabled hosts (~80% patch coverage); NVMe memory tiering with software mirroring; fleet/instance password policies; VCF Operations Build/Manage/Operate/Protect pillars; VKS 3.6 (500 clusters/Supervisor, ~70% faster linked-clone provisioning); multi-vNIC cluster node networking; self-service CaaS namespaces; native S3 object storage (Tech Preview); LACP in the deployment UI. Holodeck 9.1: VNA cluster networking, Click-to-Deploy GitOps, HoloRouter Vault/Authentik/Technitium, dual-site, min nested ESX 8.0 U3.
Learning Outcomes
Explain the VCF management services common runtime introduced in 9.1 and its impact on management-plane availability, sizing, and lifecycle design
Contrast 9.0 and 9.1 patching models (in-place vs quick patch, live patch on TPM hosts, single-API vCenter resize) and quantify the maintenance-window impact
Design a memory-tiering-enabled cluster (NVMe tier sizing, mirroring, VCP-driven configuration) and defend the VM-density/TCO trade-off
Design fleet- or instance-level password policies in VCF Operations and map them to compliance requirements
Position VKS 3.6 (500-cluster scale, linked-clone fast deploy), multi-vNIC node networking, and self-service CaaS namespaces in a developer-platform design
Identify which 9.1 features are GA versus Tech Preview (native S3 object storage) and state the defense-safe way to reference each
VCF 9.1 introduces 'VCF management services' — a common runtime and shared set of components that unifies the architecture of lifecycle and operational capabilities across the fleet. In 9.0, the management plane was already consolidating (VCF Installer deploying VCF Operations, SDDC Manager functions absorbed), but individual services still ran as loosely coupled appliances with their own runtimes, update cadences, and HA stories. In 9.1 those capabilities converge onto one runtime layer with shared components.
9.0 → 9.1 Contrast
VCF 9.0: VCF Operations + Fleet Manager + per-function appliances; separate lifecycle for each management component; manual failover for some fleet components. VCF 9.1: VCF management services = common runtime hosting lifecycle + operational services; one architectural unit to size, patch, protect, and monitor.
Why It Matters (Design/Ops Implications)
Availability design: the management-services runtime becomes a single logical dependency for Day-2 operations. Your AMPRS availability analysis must treat it like you treated SDDC Manager + Aria stack in 5.2 — what breaks when it is down (LCM, monitoring, automation intake) and what keeps running (data plane, vCenter, NSX, workloads).
Sizing and placement: one runtime means one capacity model instead of per-appliance sizing sprawl; place it in the management domain and protect with HA + backup like any other management component.
Upgrade sequencing simplifies: fewer independently versioned management appliances shrinks the compatibility matrix — the historical #1 source of VCF upgrade failures.
Exam/Defense Relevance
Defense panels probe single points of failure: be ready to state explicitly that workload data plane survives management-services outage, and quantify RTO for restoring it from backup. Delta-style exam items test that 'VCF management services' is a 9.1 term — it does not exist in 9.0 documentation.
Key Takeaways
VCF management services (new in 9.1) = common runtime unifying lifecycle and operational capabilities that were separate components in 9.0
Treat the runtime as one availability/sizing unit in design docs — analyze what fails when it is offline (LCM, ops) vs what survives (workloads, vCenter, NSX data plane)
Defense tip: never claim the management plane is 'not a SPOF' without stating the workload-continuity and restore-time story
API-First Platform: Consistent OpenAPI Across SDKs#
What Changed
VCF 9.1 formalizes the API-first platform story: a consistent OpenAPI surface across the stack, with aligned SDKs for Python and Java, PowerCLI coverage, and a Terraform provider — all tracking the same API definitions. In 9.0 the APIs existed but tooling parity was uneven (some operations UI-only or PowerCLI-only, SDKs trailing the API).
9.0 → 9.1 Contrast
VCF 9.0: APIs available but inconsistent coverage per tool; automation designs had to document per-tool gaps. VCF 9.1: one OpenAPI contract; Python/Java SDKs, PowerCLI, and Terraform generated/aligned from it — pick the tool that fits the team, not the gap.
New API capabilities land alongside: the vCenter deployment/size resize API, a vCenter maintenance-notification API (components can query whether vCenter maintenance is planned or underway; Envoy returns a 503 header with estimated completion during maintenance), and a vCenter utilization-monitoring API with new alarms (High Session Count — default limit 3000 sessions, reports top-5 users over 100 sessions; Increased Request Load — endpoint limit 1024 active requests for most endpoints).
Why It Matters (Design/Ops Implications)
Automation strategy becomes a real design decision with alternatives: Terraform for declarative infra, Python SDK for integration glue, PowerCLI for ops teams — all first-class. Document the choice and its skills implication.
The maintenance-notification API lets you design automation that degrades gracefully during vCenter patching instead of failing mid-run — pair it with quick patch for near-zero-impact maintenance.
Utilization APIs + session alarms give you enforceable guardrails against integration sprawl (backup tools, CMDB sync, monitoring all hammering vCenter).
Exam/Defense Relevance
Expect scenario items on choosing the right SDK/tool for a requirement and on the new alarms' thresholds. In defense, an 'automation-first operating model' claim is only credible if you can name the API contract and the failure-mode handling (503 during maintenance).
Key Takeaways
9.1 delivers consistent OpenAPI + aligned Python/Java SDKs, PowerCLI, and Terraform — tool choice is now preference/skills-driven, not gap-driven
New vCenter maintenance-notification API: Envoy returns 503 with estimated completion so automation can pause instead of fail
High Session Count alarm fires near the 3000-session default limit and names top-5 heavy users (>100 sessions); Increased Request Load alarm fires near 1024 active requests per endpoint
Design decision to document: which automation tool per persona, with rationale and skills implication
Patching & LCM: Quick Patch, Single-API Resize, Live Patch on TPM#
What Changed
Three lifecycle improvements headline 9.1:
vCenter quick patch — patches only the RPMs/binaries with actual code changes instead of updating every RPM in-place. Downtime drops to under 1 minute, and to zero for some service payloads. VM/K8s deployments and API workflows continue during the patch. Targeted at rapid security-fix rollout.
Single-API vCenter resize — the deployment/size API (PATCH method, callable from the vCenter Developer Center API Explorer) scales vCenter compute and disk with one call plus a reboot. In 9.0 resizing was a multi-step manual procedure.
ESX Live Patch on TPM-enabled hosts — 9.0 live patch excluded TPM-enabled hosts, forcing security-hardened fleets to choose between TPM attestation and reboot-free patching. 9.1 removes that conflict. Live patch also now covers more of the vmkernel plus additional user-space daemons (vSAN daemons, core storage daemons/libraries) and covers up to ~80% of patches. It is enabled by default with automatic fallback to maintenance-mode remediation; an 'enforce' setting (cluster- or vCenter-level) allows live-patch-only remediation.
Supporting changes: reduced-downtime upgrade (RDU) can now use the online depot (not just mounted ISO); vLCM images carry a SHA256 checksum of the image definition for cross-vCenter export/import integrity; RDU-based 8.x→9.1 or 9.0.x→9.1 upgrades automatically raise vCenter VM hardware to version 17 (in-place updates require a manual, powered-off HW upgrade afterwards); VMCA-managed certs auto-renew (vCenter TLS 5 days before expiry, ESX 30 days, threshold configurable via vpxd.certmgmt.certs.autoRenewThreshold — external CA certs are NOT auto-renewed).
Why It Matters (Design/Ops Implications)
Maintenance-window math changes: a monthly vCenter security patch used to cost a scheduled outage; now it is a sub-minute blip. That directly lowers the operational-cost argument against aggressive patch cadence — a compliance win you can quantify. Live patch on TPM ends the 'hardening vs patchability' trade-off in security-sensitive designs. The 'enforce' live-patch setting is itself a design decision: it guarantees no surprise reboots but blocks non-live-patchable remediations, so you must define the exception process.
Exam/Defense Relevance
Classic distractor set: quick patch (vCenter, changed-binaries-only) vs RDU (vCenter, migration-based upgrade) vs live patch (ESX kernel/daemons, no reboot). Know which applies where, the TPM change, and the ~80% patch coverage figure. Defense: bake quick patch + live patch into your patch-cadence SLA and be ready to state the fallback path for the other ~20%.
Key Takeaways
vCenter quick patch updates only changed RPMs/binaries: <1 min downtime, sometimes zero; workflows and deployments keep running
deployment/size PATCH API resizes vCenter compute+disk in one call + reboot (was multi-step manual in 9.0)
ESX Live Patch now works on TPM-enabled hosts, covers ~80% of patches, is on by default with maintenance-mode fallback, and supports vSAN/core-storage daemons
VMCA certs auto-renew (vCenter TLS 5 days, ESX 30 days pre-expiry); external CA certs remain the admin's responsibility
RDU upgrades to 9.1 auto-upgrade vCenter VM hardware to v17; in-place updates need a manual powered-off HW upgrade
VCF 9.1 enhances memory tiering: the hypervisor keeps active (hot) pages in DRAM and tiers cold pages to local NVMe SSDs, presenting an expanded effective memory capacity per host. 9.1 adds software-based mirroring of the NVMe tier (claim a second NVMe device as mirror) and cost-savings analysis in VCF Operations. Configuration is done through vSphere Configuration Profiles: NVMe devices are claimed for memory tiering (plus an optional mirror device) as part of the cluster's desired state.
Why It Matters (Design/Ops Implications)
VM density and TCO: memory is the dominant constraint (and cost) in most consolidation designs. Tiering raises usable memory per host without DRAM purchases — that changes host sizing, consolidation ratios, and the cost model you present to stakeholders. The built-in cost-savings analysis gives you defensible numbers instead of hand-waving.
Performance risk: cold-page tiering assumes a working-set smaller than DRAM. Latency-sensitive workloads with large, hot working sets (in-memory DBs) are poor candidates. A defensible design states the workload profile, the DRAM:NVMe ratio chosen, and the monitoring that validates the assumption (active-memory vs tiered-page metrics in VCF Operations).
Failure domain: an NVMe tier device failure without mirroring impacts the VMs whose pages live there — the 9.1 software mirroring option exists precisely to close that risk; claiming it doubles the NVMe capacity requirement.
Operational consistency: because configuration flows through vSphere Configuration Profiles, tier config is desired-state and survives host additions — new hosts inherit it automatically.
Exam/Defense Relevance
Expect items on what tiering does (DRAM=hot, NVMe=cold), what 9.1 adds (mirroring, cost analysis, VCP-driven config), and candidate-workload selection. In a defense, memory tiering is a strong cost-requirement answer — but the panel will attack with 'what about your Tier-1 database?' Have the exclusion/placement answer ready (host groups or dedicated clusters without tiering).
Key Takeaways
Memory tiering: DRAM holds hot pages, local NVMe holds cold pages — more effective memory per host without buying DRAM
9.1 adds software mirroring of the NVMe tier (optional second device) and cost-savings analysis in VCF Operations
Configured via vSphere Configuration Profiles (desired state) — new cluster hosts inherit tiering config automatically
Design defensibly: state workload profile, DRAM:NVMe ratio, mirroring decision, and the monitoring that validates the working-set assumption
Poor fit for latency-critical large-working-set workloads — plan exclusions explicitly
VCF 9.1 adds centralized password-policy management: policies for length, complexity, account lockout, change interval (expiration), and password history can be defined once and applied at fleet level or per VCF instance, from VCF Operations. In 9.0 (and painfully in 5.2), these policies were configured per component — ESX advanced settings per host, vCenter SSO policies per instance, NSX via CLI/API, appliance-local PAM settings — exactly the sprawl the VVS IAM solution documented as dozens of individual design decisions.
9.0 → 9.1 Contrast
VCF 9.0: password ROTATION was centralized (Password Manager), but password POLICY remained per-component. VCF 9.1: policy definition (length, complexity, lockout, interval, history) is centralized with fleet/instance scoping; rotation remains centralized.
Why It Matters (Design/Ops Implications)
Compliance mapping collapses from N decisions to one: 'fleet password policy aligned to <framework>' with per-instance exceptions where a regulator demands stricter settings. Audit evidence comes from one control point.
Drift elimination: per-component policies drift (a rebuilt host missing its lockout setting is a classic finding). Central policy pushed from VCF Operations removes the drift class, and pairs with 9.1 continuous compliance enforcement (Advanced Cyber Compliance) which remediates posture drift automatically.
Scoping is a real decision: fleet-level = consistency; instance-level = sovereignty/tenancy boundaries (e.g., a regulated instance with 14-char minimum and 30-day interval inside a fleet defaulting to 12/90). Document which and why.
Exam/Defense Relevance
Delta items will contrast rotation (9.0) vs policy (9.1) centralization — a subtle, testable distinction. For VCDX-style security sections, replace the old per-component policy decision cluster with one fleet-policy decision plus documented exceptions; panels reward the simplification but will ask how legacy/unmanaged components outside VCF Operations scope are handled.
Key Takeaways
9.1 centralizes password POLICY (length, complexity, lockout, change interval, history) at fleet or instance scope via VCF Operations
Contrast with 9.0: rotation was already central (Password Manager); policy definition was still per-component
In 9.0 the navigation was organized more around underlying products/functions inherited from the Aria consolidation; operators had to know which tool-area owned a task. 9.1 aligns navigation to the operator's intent.
Why It Matters (Design/Ops Implications)
Operating-model mapping: the pillars map cleanly onto team responsibilities (platform build team, LCM owner, NOC/SRE, security). Your operations design and RACI can now reference product navigation directly — 'Protect pillar tasks are owned by the security operations role' — which shortens runbooks and training.
Documentation churn: any runbook, SOP, or training material written against the 9.0 navigation needs a refresh pass; budget it in the upgrade project (a real migration-cost line item architects often forget).
It is UX reorganization, not RBAC: pillar visibility does not itself constitute a security boundary — roles/scopes still control authorization. Do not present the pillar split as a control in a defense.
Exam/Defense Relevance
Expect 'which pillar contains X?' items (password policies → Manage; continuous compliance → Protect; domain deployment with LACP → Build; capacity/diagnostics → Operate). In defenses, use the pillars as your operational-handoff narrative, but keep authorization claims anchored in roles, not UI layout.
Pillars align navigation to operator intent instead of product heritage — map them to your RACI and runbooks
UX reorg is not RBAC: roles and scopes still define authorization boundaries
Upgrade projects must budget a runbook/training refresh for the new navigation
VKS 3.6: Scale, Linked-Clone Fast Deploy, and Multi-vNIC Node Networking#
What Changed
vSphere Kubernetes Service 3.6 (the VKS release paired with VCF 9.1) brings:
Scale: up to 500 clusters per Supervisor — consolidating what previously required multiple Supervisors or careful cluster rationing.
Fast deploy/upgrade via linked clones: cluster node VMs deploy from linked clones instead of full clones, cutting provisioning time roughly 70% and speeding rolling upgrades (each replaced node comes up faster). Faster node replacement = shorter upgrade windows and quicker scale-out under load.
Multi-vNIC cluster node networking: cluster nodes can attach multiple vNICs to separate application, storage, and management traffic onto distinct networks. Before 9.1, node networking was effectively single-homed — storage traffic (e.g., to external arrays or high-throughput CSI backends) shared the node's one network path with app and control traffic.
Why It Matters (Design/Ops Implications)
The 500-cluster ceiling changes multi-tenancy topology: 'one Supervisor, many clusters' becomes viable for larger estates, moving isolation decisions from 'separate Supervisors' toward namespaces + network policy + vSphere Zones. Document the blast-radius trade-off explicitly: one Supervisor is now a bigger failure/upgrade domain.
Linked-clone deploy trades disk-time for a shared-base dependency: know the storage implication (clones reference a base image; monitor delta growth) — it is the same architectural trade-off Holodeck users know from nested labs.
Multi-vNIC unlocks conformance with traffic-separation NFRs (regulated environments demanding storage/mgmt/app isolation) inside Kubernetes clusters, and lets you honor existing physical network segmentation designs rather than collapsing them at the container layer.
Exam/Defense Relevance
Numbers are testable: 500 clusters, ~70% faster provisioning, VKS 3.6. Design items will probe when to use multiple Supervisors anyway (independent failure domains, different vCenters/zones, hard tenancy). For defense, multi-vNIC is your answer when a panelist asks how the containerized tier honors the same traffic-separation rules as the VM tier.
Key Takeaways
VKS 3.6: up to 500 clusters per Supervisor and ~70% faster provisioning via linked-clone fast deploy
Linked clones speed deploy/upgrade but introduce a shared base-image dependency — monitor delta growth
Multi-vNIC node networking separates app/storage/mgmt traffic per cluster node — extends physical traffic-separation NFRs into Kubernetes
Bigger Supervisors = bigger failure/upgrade domains: keep separate Supervisors for hard tenancy or independent failure domains
VCF 9.1 adds a simplified CaaS consumption path: self-service namespace provisioning where a developer gets a namespace with registry access, ingress, resource quotas, and identity — all inherited from existing VCF constructs (projects, identity providers, network/storage policy) — without anyone standing up or managing a full Kubernetes cluster for them.
9.0 → 9.1 Contrast
VCF 9.0: 'a few containers' still meant requesting a VKS cluster (or sharing one via manually carved namespaces with hand-built registry/ingress/quota plumbing). VCF 9.1: namespace-as-the-unit-of-consumption, provisioned self-service with guardrails pre-wired from the platform.
Why It Matters (Design/Ops Implications)
Consumption-model design gains a middle tier. Your service catalog now has: full clusters (teams owning K8s versions/operators), self-service namespaces (teams that just need to run containers), and VMs. Matching workload personas to tiers is a design decision with cost and operational-load consequences — namespaces avoid per-team cluster sprawl and its patching burden.
Governance inheritance is the killer feature: quotas, registry policy, ingress, and identity come from VCF-level constructs, so security review happens once at the platform level, not per namespace request. This is the CaaS analog of the fleet password-policy story — centralize the control, inherit everywhere.
Capacity management shifts: many small namespaces on shared Supervisor capacity need quota discipline and utilization monitoring (Operate pillar) to avoid noisy-neighbor effects.
Exam/Defense Relevance
Scenario items will describe a team needing 'a few containers, fast, with corporate registry and SSO' — the 9.1-correct answer is self-service namespaces, not a dedicated VKS cluster. In defense, this is your answer to 'how do you stop cluster sprawl?' — but expect the follow-up on isolation limits: namespace isolation is softer than cluster isolation; state where you draw the line (compliance workloads → dedicated clusters).
Key Takeaways
9.1 CaaS: self-service namespace provisioning with registry, ingress, quotas, and identity inherited from VCF constructs
New middle tier in the consumption model: full VKS clusters vs namespaces vs VMs — match personas to tiers deliberately
Governance is inherited from platform level — one security review covers all namespaces
Namespace isolation < cluster isolation: keep hard-compliance workloads on dedicated clusters and say so in the defense
VCF 9.1 introduces native S3-compatible object storage as a Tech Preview: developers self-serve buckets while IT keeps the guardrails, and object storage is deployed, scaled, and managed with the same VCF workflows used for block and file storage.
Why It Matters (Design/Ops Implications)
Tech Preview status is the headline design fact: no production support, no SLA, subject to change. In any real design it may appear only as a labeled future consideration or an isolated pilot — never on the dependency path of a production requirement.
Strategically it signals platform direction: modern app patterns (backups, AI/ML datasets, logs, static assets) assume an S3 endpoint; today that forces an external MinIO/array/vendor dependency that sits outside VCF lifecycle and identity. A native service would collapse that integration seam — worth a roadmap note in designs with heavy S3 consumption.
Evaluation criteria to track toward GA: durability model, multi-tenancy/quota controls, encryption and key management, protocol conformance level, and how capacity is carved from vSAN.
Exam/Defense Relevance
The testable fact is its status: 9.1 items love a distractor that treats object storage as GA. In a VCDX-style defense, citing a Tech Preview feature as the solution to a requirement is an instant credibility hit — the defense-safe framing is 'requirement met today by <supported alternative>; native object storage tracked as a roadmap simplification.' Know the difference between GA (memory tiering, live patch on TPM, quick patch) and Tech Preview (native S3) in the 9.1 feature set cold.
Key Takeaways
Native S3 object storage in 9.1 is TECH PREVIEW — self-service for developers, IT guardrails, managed like block/file
Never place Tech Preview features on a production requirement's dependency path; frame as roadmap consideration
Exam trap: distractors presenting 9.1 object storage as GA/supported
Deployment & Networking Delta: LACP in the UI, vSAN Data Services, Elastic Provisioning#
What Changed
A cluster of smaller but design-relevant 9.1 deltas:
LACP in the workload-domain deployment UI: link aggregation (LACP) can now be configured directly in the workload-domain deployment workflow. In 9.0, LACP-based host networking required post-deployment vDS surgery or API-driven workarounds that fought the automated bring-up. Customers whose network standards mandate LACP to the leaf switches can now comply at Day-0 without breaking supportability of the automated path.
vSphere Elastic (Zero Touch) Provisioning: ESX bootstraps onto bare metal via UEFI HTTP/S network boot — no TFTP server, supports Secure Boot and TPM; the target cluster's image and vSphere Configuration Profile determine what the host becomes; DHCP required unless UEFI has a static IP. Host-onboarding at scale stops being an imaging project.
vSAN data services expansion: deduplication and compression broadened across cluster types and workload profiles, dedupe now supported WITH data-at-rest encryption, improved compression — reopens data-reduction assumptions in vSAN capacity models.
DRS/vMotion refinements: non-disruptive vMotion evacuation option for maintenance mode (only evacuate when destination capacity avoids contention), rolling vMotion slot release (a new migration starts as each of the 8 concurrent slots frees, instead of batch-of-8 barriers), and encrypted vMotion offload to Intel QAT.
Why It Matters (Design/Ops Implications)
LACP-at-deploy removes a classic constraint-vs-automation conflict — previously you either bent the network standard or accepted manual deviation from the reference deployment. ZTP + VCP + vDS bootstrap makes 'server arrives → cluster member' a pipeline, which changes your host-replacement RTO and scale-out lead-time numbers. Dedupe+encryption coexistence removes an old either/or from storage design (compliance no longer costs you data reduction).
Exam/Defense Relevance
These are precision-distractor magnets: LACP is now IN the deployment UI (9.1) vs post-deploy workaround (9.0); dedupe WITH encryption is 9.1-new; maintenance-mode evacuation has TWO options (standard vs non-disruptive) and 'non-disruptive' refers to avoiding resource contention, not implying the standard method disrupts.
Key Takeaways
9.1 deployment UI supports LACP for workload-domain host networking — Day-0 compliance with LACP-mandating network standards
Zero Touch Provisioning: UEFI HTTP/S boot (no TFTP), Secure Boot/TPM compatible, cluster VCP defines the host's config
The Holodeck toolkit tracking VCF 9.1 is a major release (9.0.1 and 9.0.2 were maintenance), and it resets several baseline assumptions. Docs are dated May 2026; the GA blog is dated 2026-07-01 — an unresolved conflict. Note that vmware/Holodeck publishes no GitHub releases, so vmware.github.io/Holodeck/9.1/ is the sole source of truth.
1. The resource envelope doubles. This is the change that actually constrains lab planning. VCF 9.1 single site needs 48 CPU / 650 GB / 2.2 TB against VCF 9.0's 32 / 325 GB / 1.1 TB — +50% CPU and exactly 2.0× on both RAM and disk. Dual site is 96 / 1536 GB / 5 TB. VVF 9.1 is gentler at 24 / 512 GB / 2 TB single, 48 / 1024 GB / 4 TB dual. The HoloRouter appliance itself grows to 6 vCPU / 12 GB / 75 GB. New in 9.1, soft validation lets pre-checks warn rather than block below-spec deployments — convenient on constrained hosts, but the fail-fast guardrail is gone and the architect owns the degradation. On a 1 TB host: VCF 9.1 single site fits, VVF 9.1 dual site sits exactly on the RAM line, and VCF 9.1 dual site is simply not achievable.
2. VNA: distributed vs centralized networking. The Virtual Network Appliance cluster is the headline architectural addition, supported in both management and workload domains. The clean mental model is VNA = distributed, NSX Edge cluster = centralized — Edge remains fully supported and is still the default. Day 0 exposes -VnaClusterMgmtDomain and -VnaClusterWkldDomain; Supervisor placement gains a mode parameter taking Centralized or Distributed. Caveat: the published cmd_reference Day 0 table shows no VNA switch and its -Version list omits 9.1.0.0 — that page is stale, so verify with Get-Help New-HoloDeckInstance on a live router.
3. GitOps is an OVA toggle, not a cmdlet. Enabling GitOps at HoloRouter OVA deployment — on by default — auto-provisions GitLab, pre-loads a repo with the deployment pipeline, and registers the router as a runner, so a full deployment can be triggered from the GitLab UI with no PowerShell. There are zero GitOps cmdlets. The consequence matters: GitOps cannot be retrofitted onto an existing router, so it is a decision made once, at deploy time.
4. HoloRouter service stack. Vault (:8200), Authentik SSO with SCIM (:9443), Technitium DNS (:5380) and an authenticated Webtop (:30000) now sit behind an HTTPS reverse proxy with an internal CA. Technitium becomes the DNS management surface and the HoloDeckDNSConfig cmdlets are deprecated — but the docs do not establish that dnsmasq is gone underneath, so 9.0-era DNS and shared-host cautions stay in force until re-baselined live. See the Holodeck section for the full evidence.
5. Six breaking changes. Mandatory -Site a|b on Import-HoloDeckConfig (this fails every existing 9.0.x invocation), dual config files from New-HoloDeckConfig, authenticated Webtop, deprecated DNS cmdlets, VCF Installer 9.0.0.0 unsupported, and HTTP endpoints moved to HTTPS. Every 9.0-era runbook needs a pass.
**6. No documented upgrade path.** There is no 9.0→9.1 migration guidance anywhere — treat 9.1 as a clean redeploy alongside, not an in-place upgrade. Related: 9.1 publishes **no Known Issues section at all**, which is a documentation gap rather than evidence of a clean release; source issues from GitHub Issues and the Broadcom community.
Why it matters for VCDX. The lab is the evidence factory. Dual-site support means DR runbooks and failover timings can be validated rather than asserted; Vault and Authentik make the 9.1 password-policy and federation stories demonstrable end-to-end; GitOps makes lab states reproducible artifacts you can cite. Most valuable of all, VNA (distributed) versus Edge (centralized) is a genuine design-decision axis — a real trade-off with justification, alternatives considered, and implications, which is exactly the shape of reasoning a panel probes. Holodeck itself is not exam content, but 'I failed site A over to site B and measured RTO' beats any datasheet citation. Track the nested ESX 8.0 U3 floor as a lab-refresh action item.
Key Takeaways
VCF 9.1 doubles the lab resource envelope: 48 CPU / 650 GB / 2.2 TB single site (2.0× RAM and disk vs 9.0); soft validation now warns instead of blocking, so you own the degradation
VNA = distributed networking, NSX Edge = centralized — supported in both management and workload domains, and a genuine VCDX design-decision axis
GitOps is an OVA deploy-time toggle (default on) that cannot be retrofitted — decide before you deploy the router
Six breaking changes; mandatory -Site on Import-HoloDeckConfig fails every existing 9.0.x call
No documented 9.0→9.1 upgrade path and no published Known Issues — plan a clean redeploy and source issues from GitHub/community
Exam Mapping: 2V0-13.25 — VCP-VCF 9.0 Architect
Delta awareness — current exam targets 9.0, but design-methodology items (AMPRS trade-offs, GA vs Tech Preview discipline, maintenance-window quantification) are exercised here against 9.1 features
Watch for exam refresh to 9.1: patching/LCM (quick patch, live patch on TPM), memory tiering, and fleet password policies are the most examinable deltas
Operational deltas most likely to appear on a 9.1-refreshed exam: VCF Operations Build/Manage/Operate/Protect navigation, password policy scoping, quick patch vs RDU vs live patch, VKS 3.6 provisioning
VCF 9.1 introduces VCF management services as a common runtime plus shared components unifying how lifecycle and operational capabilities are architected — in 9.0 these ran as more loosely coupled components. Option A is wrong: SDDC Manager's functions were already absorbed into the 9.0 platform; this is a new runtime concept, not a rebrand. Option C is wrong: it is an architectural construct inside the product, not an operated/managed service offering. Option D describes the traditional management-domain appliance inventory — the management services runtime is the unifying layer for lifecycle/operational capabilities, not a name for the existing appliance set.
An architect claims the VCF management services runtime is 'not a single point of failure' for the platform. A review panel pushes back. Which statement is the most defensible correction?
The runtime is fully active-active across sites, so no failure analysis is required
Workloads, vCenter, and the NSX data plane continue operating during a runtime outage, but lifecycle and fleet operations stop — so its restore RTO must be designed and stated
Because it is a common runtime, a failure affects only the UI; all APIs remain available
The runtime cannot fail because it is protected by vSphere HA in the management domain
The defensible position distinguishes what survives (data plane: running VMs, vCenter, NSX forwarding, vSAN) from what stops (LCM, fleet monitoring/operations, automation intake) and quantifies restore time. Option A asserts a topology claim that isn't a given and dodges failure analysis — panels reject 'no analysis required' on principle. Option C is wrong: the runtime hosts services, not just UI; its APIs go down with it. Option D confuses a mitigation with immunity — vSphere HA restarts VMs after host failure but does not prevent runtime outages (patching, corruption, config error) and still implies downtime.
The 9.1 API-first story is one OpenAPI contract with aligned SDKs (Python/Java), PowerCLI, and Terraform — in 9.0, coverage differed per tool, so designs had to document gaps. Option A inverts the point: more tools are first-class, not fewer. Option C invents a protocol change that did not happen — the alignment is about consistent OpenAPI definitions. Option D is nonsense from a security standpoint: authentication remains mandatory; identity inheritance in 9.1 refers to constructs like CaaS namespaces, not unauthenticated APIs.
9.1 adds a maintenance-notification capability: components can query whether vCenter maintenance is planned or underway, and during maintenance Envoy returns a 503 with estimated completion — automation can pause and retry rather than fail. Option A (404) would falsely signal a missing resource and give no completion hint. Option C is wrong: this mechanism is not an HA redirect, and vCenter HA failover is a different capability entirely. Option D would be actively harmful — a 200 with empty payload looks like success and would corrupt automation logic; the design goal is an explicit, machine-readable 'unavailable, back at T' signal.
The new vCenter 'High Session Count' alarm in 9.1 fires as sessions approach the default limit. What are the default session limit and the reporting detail included?
Limit 1024; reports the top 10 API endpoints by call volume
Limit 3000; reports the IPs/usernames of the top 5 users each holding more than 100 sessions
Limit 500; reports only the total session count
Limit 3000; reports the top 5 endpoints nearing their request limits
High Session Count fires near the 3000-session default and names the top 5 users (IPs and usernames, service accounts eligible) that each hold over 100 sessions — enough to identify the offending integration. Option A mixes in 1024, which is the per-endpoint active-request limit associated with the separate 'Increased Request Load' alarm. Option C invents a limit and understates the diagnostic detail. Option D correctly states 3000 but attaches the endpoint-level reporting that belongs to Increased Request Load — the two alarms are a deliberate distractor pair: sessions/users vs endpoints/requests.
In the reorganized VCF Operations console (9.1), where would an operator go to configure fleet password policies, and where to review continuous compliance posture?
Both under Operate
Password policies under Manage; compliance posture under Protect
Password policies under Protect; compliance posture under Manage
Password policies under Build; compliance posture under Operate
The 9.1 pillars are task-oriented: Manage owns lifecycle concerns (updates, certificates, passwords/policies, licensing) and Protect owns compliance, security posture, and recovery. Option A is wrong — Operate is monitoring/capacity/diagnostics, not configuration of credentials or compliance. Option C reverses the two: password policy is administrative lifecycle (Manage), while compliance enforcement is protective posture (Protect). Option D misplaces both: Build is provisioning (domains, clusters, networking such as the new LACP option), and Operate is observability.
A security reviewer proposes using the VCF Operations pillar layout (Build/Manage/Operate/Protect) as the authorization boundary between teams. What is the correct architectural response?
Agree — pillar visibility is enforced per user and constitutes RBAC
Disagree — the pillars are a navigation/UX reorganization; authorization must still be enforced through roles and scopes
Agree, but only if each pillar is deployed as a separate appliance
Disagree — pillars should be disabled entirely in secure environments
The pillar reorg aligns navigation to operator intent; it is not itself a security control — roles and scopes remain the authorization mechanism, and a design must anchor access claims there. Option A treats UI layout as enforcement, a classic audit finding in the making. Option C is architecturally wrong: pillars are console navigation within VCF Operations, not deployable units. Option D is a non-sequitur — there is no 'disable pillars' hardening step, and removing navigation would not change authorization anyway.
Quick patch updates only the RPMs/binaries that actually changed, instead of reinstalling every RPM in place. techdocs classifies the result into three service groups: zero downtime (e.g. vlcm, vsphere-ui, eam), near-zero up to 180 seconds (e.g. vmware-envoy, vmware-cis-license), and 3-5 minutes for vpxd/vmca/vmon-class services. Teach zero-to-5-minutes: the '<1 minute, sometimes zero' figure comes from the launch blog and describes only the best-case groups.
9.1 introduces a single-API resize: one deployment/size PATCH call scales vCenter compute and disk, followed by a reboot — replacing 9.0's multi-step manual procedure. Option A is the heavyweight legacy workaround the API makes unnecessary. Option C is incomplete and unsupported as a full resize path: vCenter sizing involves internal service/disk layout, not just VM hardware values, which is exactly why a dedicated API exists. Option D misuses RDU — RDU is an upgrade mechanism (new appliance, data migration); running an upgrade to resize is neither its purpose nor the supported resize path.
The 9.1 deltas: live patch now works on TPM-enabled hosts (removing the 9.0-era hardening-vs-patchability conflict) and covers up to ~80% of patches, with expanded vmkernel and user-space daemon coverage (including vSAN and core storage daemons). Option A states the old 9.0 restriction that 9.1 removes. Option C overreaches — the ~20% of patches that are not live-patch-capable still fall back to maintenance-mode/reboot, which is why a fallback process must remain in the design. Option D is wrong: live patch is enabled by default at cluster level with automatic fallback; configuration is at cluster or vCenter level, not a per-host pre-remediation toggle.
What does the 'enforce' setting for ESX live patch do, and what operational consequence must the architect plan for?
Forces immediate patching of all hosts; consequence: unplanned reboots
Allows live-patch-only remediation and blocks remediations requiring maintenance mode; consequence: an exception process is needed for non-live-patchable payloads
'Enforce' permits only live-patch remediation and blocks any remediation that would need maintenance mode and a reboot — guaranteeing no surprise reboots, but stranding the ~20% of payloads that are not live-patch-capable unless the design defines an exception/maintenance process. Option A inverts the behavior: enforce prevents reboot-requiring remediation rather than forcing patching. Option C confuses remediation policy with compliance reporting — enforce governs how patches may be applied, not what is logged. Option D describes turning live patch off, the opposite of enforcing it; the actual default is enabled-with-fallback, with enforce as the stricter option.
RDU creates a NEW vCenter VM, so the VM hardware version moves automatically (version 10 → 17). An in-place update keeps the old VM, so the admin must upgrade VM hardware afterwards — a powered-off operation, i.e., a small extra downtime window to schedule. Option A is not a standard post-update step for either path. Option C is wrong: certificate handling is unchanged by the update path, and 9.1 actually auto-renews VMCA-managed certs near expiry (VMCA root renewal during upgrade only occurs if it is within 1 year of expiry — not a manual re-issue-all). Option D misunderstands the resize API — sizing is preserved through updates; no restoration call exists or is needed.
Which statement correctly describes certificate auto-renewal behavior introduced with 9.1?
All certificates, including those from external CAs, auto-renew 30 days before expiry
VMCA-managed certificates auto-renew — vCenter TLS 5 days and ESX 30 days before expiry (ESX threshold configurable); externally issued certificates remain the administrator's responsibility
Only ESX certificates auto-renew; vCenter TLS must be renewed manually
Auto-renewal applies only during upgrades, not during steady-state operation
9.1 auto-renews VMCA-managed certs: vCenter TLS at 5 days pre-expiry, ESX at 30 days (tunable via vpxd.certmgmt.certs.autoRenewThreshold). External-CA certificates are explicitly excluded — the design must still own that renewal process. Option A dangerously extends the behavior to external CAs, exactly the assumption that causes outages. Option C reverses reality — both vCenter TLS and ESX VMCA-managed certs auto-renew. Option D confuses the steady-state renewal feature with the separate upgrade-time behavior (VMCA root + solution-user leaf renewal during vCenter upgrade when the root is <1 year from expiry).
How does 9.1 Zero Touch Provisioning (ZTP) of ESX differ from classic vSphere Auto-Deploy?
ZTP requires a dedicated TFTP server per rack
ZTP uses UEFI HTTP/S network boot with no TFTP server, supports Secure Boot and TPM, and takes image plus configuration from the target cluster's deploy rule and configuration profile
ZTP images hosts from a USB key prepared by vLCM
ZTP only works on hosts that already have ESX installed
ZTP builds on Auto-Deploy foundations but modernizes the boot path: UEFI HTTP/S boot pointed at vCenter, no external TFTP dependency, compatible with Secure Boot and TPM; the deploy rule's cluster determines the ESX image and vSphere Configuration Profile (vDS config can be bootstrapped), with DHCP required unless UEFI has a static boot IP. Option A describes the classic PXE/TFTP dependency ZTP removes. Option C contradicts the point of network-based zero-touch imaging. Option D inverts the use case — ZTP exists precisely to bootstrap bare-metal servers with no OS; booting to a non-VCP cluster simply joins with a default configuration.
Memory tiering is page-granular placement: hot pages stay in DRAM, cold pages tier to local NVMe, so the host presents more effective memory without added DRAM — the basis of the VM-density/TCO case. Option A describes hypervisor swapping, a last-resort overcommit response with severe performance cliffs; tiering is a proactive, page-level mechanism, not VM-level swap. Option C confuses memory tiering with vSAN caching — the NVMe tier serves guest memory pages, not storage metadata. Option D invents a compression-offload mechanism; the 9.1 additions are software mirroring of the tier and cost-savings analysis, not compression engines.
The documented 9.1 enhancements are software-based mirroring (claim an additional NVMe device as a mirror, protecting tiered pages from device failure) and cost-savings analysis surfacing the TCO benefit. Option A substitutes hardware RAID for the actual software mirroring approach — the mirror is claimed via vSphere Configuration Profiles, not a RAID controller. Option C invents remote tiering; the tier uses LOCAL NVMe devices. Option D lists plausible-sounding features that are not part of the 9.1 tiering delta — treat 'predictive prefetching' style distractors with suspicion unless documented.
A panel asks: 'You enabled memory tiering to raise VM density. What happens to your Tier-1 in-memory database?' What is the architect's strongest answer?
Memory tiering is transparent, so no workload is affected
The database's working set exceeds typical DRAM-resident assumptions, so it is explicitly excluded — placed on hosts/clusters without tiering — and tiering telemetry validates the working-set assumption for everything else
The database will be migrated to a namespace with a higher quota
vSphere HA will restart the database if tiering degrades it
Tiering assumes the hot working set fits in DRAM; a large in-memory database violates that, so a defensible design excludes it (dedicated non-tiered cluster or host group) and monitors active vs tiered memory to validate assumptions for remaining workloads. Option A is the trap word 'transparent' — mechanically transparent, but performance-wise workload-dependent; asserting no impact is indefensible. Option C is a category error: namespaces/quotas are a CaaS construct irrelevant to VM memory placement. Option D misapplies HA — availability restart does nothing for latency degradation and signals the candidate cannot distinguish performance design from availability design.
9.1 configures tiering through vSphere Configuration Profiles: NVMe devices are claimed for the tier, optionally a second device as software mirror — and because VCP is desired-state, incoming hosts inherit the configuration during remediation (drift eliminated by design). Option A describes the legacy imperative pattern VCP exists to replace; its drift-prone property is exactly what this delta removes. Option C misplaces the setting at VM scope — tiering is host/cluster infrastructure config. Option D is word salad — the password vault manages credentials, not memory configuration.
The maximum stays at 8, but 9.1 removes the batch barrier: previously, in a batch operation no new vMotion started until all 8 in flight completed; now each freed slot immediately admits the next task, evening out network/storage load and reducing peak concurrent load per host. Option A is the classic 'bigger number' distractor — the limit did not change; the scheduling did. Option C misuses QAT, which offloads encrypted vMotion crypto to hardware but does not lift concurrency limits. Option D describes the opposite direction of the change.
Non-Disruptive evacuation is contention-aware placement: migrate only when landing capacity avoids resource contention, with DRS optionally rebalancing first. Option A misreads 'non-disruptive' as a network guarantee — vMotion's momentary switchover characteristics are unchanged. Option C falls for the naming trap the docs explicitly warn about: standard evacuation is NOT disruptive; the new term only means avoiding post-landing resource contention, and both options coexist. Option D invents suspend-based evacuation — the mechanism is still vMotion; only admission criteria changed.
9.1 can offload encrypted vMotion cryptography to Intel QAT hardware, returning CPU cores to workloads — the direct answer to 'encrypt everything without the CPU cost'. Option B undermines the security requirement rather than meeting it, and selective weakening of a mandated control is a compliance failure. Option C misapplies SR-IOV — it bypasses the virtual switch for guest NIC traffic; it neither performs nor avoids vMotion stream encryption. Option D is operational avoidance, not a solution: the CPU cost still occurs, just at night, and emergency evacuations don't wait for windows.
The 9.1 delta is centralized password POLICY management with fleet or per-instance scoping. Option A is the trap: automated rotation was already centralized in 9.0 (Password Manager) — the rotation-vs-policy distinction is the key delta fact. Option C describes credential storage, present well before 9.1. Option D is the legacy per-component approach (ESX advanced settings, vCenter SSO policy, NSX CLI, appliance PAM) that 9.1 specifically supersedes.
A financial-services customer requires 14-character minimum passwords and 30-day change intervals for one regulated VCF instance, while the rest of the fleet uses 12 characters and 90 days. How should this be designed in 9.1?
Apply the strictest policy fleet-wide to keep configuration uniform
Set the fleet-level policy to the common baseline and configure an instance-level policy override on the regulated instance, documenting the exception and its audit evidence
Configure the regulated settings per component on the affected instance using ESX advanced settings and vCenter SSO policies
Decline the requirement because VCF supports only one global policy
9.1's scoping model exists exactly for this: fleet baseline for consistency plus instance-level policy where regulation demands more — a clean, auditable design decision. Option A is a valid-sounding alternative but imposes the regulated burden (30-day rotations) on every user fleet-wide without a requirement driving it; defensible only if the customer accepts the operational cost, and the question implies they scoped it to one instance. Option C regresses to per-component sprawl, recreating the drift and audit burden the 9.1 feature eliminates. Option D is factually wrong — instance-level policies are supported.
The 9.1 delta is from detect-and-report to continuously ENFORCE: drift is remediated automatically per policy, with unified posture management across the stack. The new design decision is which controls auto-remediate vs alert-only, since auto-remediation intersects change control. Option A trivializes the change to report formatting. Option C confuses compliance posture with network security enforcement — vDefend/DFW still owns traffic control. Option D invents a SaaS shift; the capability operates within the customer's VCF Operations context, which matters for sovereignty-sensitive designs.
Why does ESX Live Patch supporting TPM-enabled hosts (9.1) matter to the SECURITY architecture, not just to lifecycle operations?
TPM chips accelerate the application of live patches
It removes the 9.0-era disincentive against enabling TPM: the baseline can now mandate TPM attestation everywhere without giving up reboot-free patching
It allows patches to be signed by the TPM instead of VMware
It makes host attestation optional since patches are now verified in memory
In 9.0, TPM-enabled hosts were excluded from live patch, so mandating TPM cost you patch agility — some designs left TPM off for operational reasons. 9.1 dissolves the trade-off: hardened baseline AND live patching. Option A invents a hardware acceleration role for TPM — it is an attestation/root-of-trust device, not a patch engine. Option C misstates code signing: patch payloads are vendor-signed; TPM does not sign patches. Option D gets attestation backwards — TPM attestation remains exactly as valuable; live patch simply now coexists with it.
Which guest OS customization capabilities were added in 9.1, primarily to reach parity for VMware Cloud Director → VCF Automation migrations?
GPU passthrough configuration during customization
Setting/resetting Linux root passwords, resetting Windows administrators-group passwords, and executing Windows customization scripts
Automatic application installation via SCCM integration
In-guest antivirus deployment
The GOSC VIM API gained vCD-parity capabilities: set root password in Linux, reset root passwords in Linux, reset Windows administrators-group account passwords, and execute customization scripts in Windows — plus (separately) IPv6-only guest networking with IPv4 explicitly disabled, and network-only re-customization of powered-on VMs. Options A, C, and D are all plausible-sounding platform features that are not GOSC functions: GPU assignment is VM hardware configuration, software deployment belongs to configuration-management tooling, and security agents are a guest-management concern — none are part of the 9.1 customization delta.
Q27
What is the correct characterization of VCF 9.1's on-premises ransomware recovery capability?
It requires replicating workloads to a cloud-hosted recovery service
It integrates cyber recovery using isolated clean rooms hosted on on-premises VCF
It is an NSX feature that quarantines infected VMs automatically
9.1 brings the clean-room recovery pattern on-premises: integrated cyber recovery into isolated clean-room environments on VCF itself — relevant to sovereignty-constrained customers who could not adopt cloud-based ransomware recovery. Option A describes the older cloud-centric model this delta specifically moves beyond. Option C confuses recovery with containment — vDefend-style quarantine is a different control in the kill chain; clean rooms are for validating and restoring workloads after compromise. Option D is wrong because immutability supports recovery but a clean room's essence is an isolated environment to inspect/restore into, not a snapshot format change.
VKS 3.6 supports up to 500 clusters per Supervisor with roughly 70% faster provisioning from linked-clone fast deploy (which also accelerates rolling upgrades). Option A understates both figures. Option C swaps the unit — the 500 figure is CLUSTERS per Supervisor, not nodes per cluster, and the speedup is cluster provisioning, not pod scheduling. Option D inflates the cluster count and invents an upgrade multiplier; upgrade improvement comes from faster node replacement, not a quoted '2x'.
What is the architectural trade-off introduced by VKS linked-clone fast deploy?
Higher DRAM usage per node in exchange for faster boot
Nodes share a base-image dependency and accumulate delta disks — provisioning speed in exchange for base-image lifecycle management and delta-growth monitoring
Linked clones reference a shared base image; each node carries a delta disk. You gain ~70% faster provisioning and quicker rolling upgrades but must manage base-image lifecycle and watch delta growth — the same trade nested-lab practitioners know well. Option A misattributes the cost to memory; the mechanism is storage-level. Option C is false — clone type does not remove HA protection. Option D is backwards: node REPLACEMENT (the rolling-upgrade pattern) is exactly what linked clones accelerate; VKS upgrades were already replacement-based.
A regulated customer requires storage traffic isolated from application traffic for containerized workloads, matching their VM estate's segmentation. Which 9.1 capability addresses this directly?
Self-service namespaces with resource quotas
Multi-vNIC cluster node networking — attaching multiple vNICs per VKS node to separate application, storage, and management traffic
vSAN deduplication with data-at-rest encryption
LACP configuration in the workload-domain deployment UI
Multi-vNIC node networking is the delta that lets Kubernetes cluster nodes home different traffic classes to different networks, extending the VM estate's traffic-separation NFR into the container tier — pre-9.1, node networking was effectively single-homed. Option A governs consumption (quotas/identity), not traffic paths. Option C is a storage-efficiency/encryption feature orthogonal to network segregation. Option D operates at the physical host uplink layer during domain deployment — it aggregates links; it does not separate traffic classes per Kubernetes node.
Scale headroom changes the default but not the isolation logic: one big Supervisor is one big failure and upgrade domain, so hard tenancy, blast-radius, or lifecycle-divergence requirements still argue for separate Supervisors — an explicit trade-off to document. Option A commits the classic 'new limit = new rule' fallacy. Option C invents a threshold with no basis — 500 is the ceiling, and count alone is not the decision driver. Option D is a red herring; CPU architecture concerns are addressed at cluster/host level, not by Supervisor multiplication.
What does VCF 9.1's simplified Container-as-a-Service provide that 9.0 did not?
A managed public-cloud Kubernetes control plane
Self-service namespace provisioning with registry, ingress, quotas, and identity inherited from VCF constructs — containers without owning a full cluster
Automatic conversion of VMs into containers
A replacement for VKS clusters, which are deprecated in 9.1
The 9.1 CaaS delta makes the namespace the self-service unit of consumption, pre-wired with registry, ingress, quota, and identity guardrails inherited from platform constructs — 'a few containers, fast' without a dedicated cluster. In 9.0 that need meant a cluster request or manual namespace plumbing. Option A relocates the answer off-platform; this is on-prem VCF capability. Option C describes workload transformation tooling that is not what CaaS provisioning does. Option D is flatly wrong — VKS clusters remain the tier for teams needing cluster-scoped control; namespaces are the new middle tier beside them, not a replacement.
A developer team requests 'a namespace for a small internal app — needs corporate registry, SSO, and modest quotas, running next week.' A separate compliance team runs PCI-scoped payment workloads. Applying 9.1 constructs, what is the defensible placement?
Both teams get self-service namespaces on the shared Supervisor
The internal app gets a self-service namespace; the PCI workloads get a dedicated VKS cluster because namespace isolation is weaker than cluster isolation
Both teams get dedicated VKS clusters to keep policy uniform
The internal app gets a VM; namespaces are not appropriate for production apps
This is the tiering rule in action: namespaces fit the low-friction internal case (guardrails inherited, no cluster ownership), while compliance-scoped workloads warrant cluster-level isolation — namespace boundaries share a control plane and kernel surface with neighbors, which most PCI scoping arguments won't accept. Option A ignores the isolation differential for regulated workloads. Option C is defensible only at a cost — it recreates cluster sprawl and patch burden for a team with no cluster-scoped needs; 'uniformity' is not a requirement here. Option D is dogma, not analysis: namespaces are production-appropriate within their isolation envelope.
The inheritance is the point: governance (quota, registry, ingress, identity) flows from VCF-level constructs, so security review happens once at platform level and every namespace arrives compliant — the same centralize-then-inherit pattern as fleet password policies. Option A defeats the governance model; developer-set guardrails are no guardrails. Option C confuses application packaging with platform policy. Option D overextends DFW — network microsegmentation complements namespace policy but does not define quotas, registry, or identity.
A design requirement calls for S3-compatible object storage for application artifacts in production. How should VCF 9.1's native object storage be treated in the design?
Adopt it as the primary solution since it ships with 9.1
Meet the requirement with a supported alternative and record native object storage (Tech Preview) as a tracked roadmap simplification
Use it in production but only for non-critical data
Reject S3 requirements entirely until the feature reaches GA
Native S3 object storage in 9.1 is a Tech Preview — no production support or SLA — so it cannot sit on a production requirement's dependency path; the defense-safe framing is 'supported alternative today, TP feature tracked toward GA'. Option A is the credibility-killer: citing a TP feature as the solution to a stated requirement. Option C still puts an unsupported service in production — 'non-critical' does not manufacture a support statement. Option D confuses feature status with requirement validity: the S3 requirement is real and must be met; only the native feature is premature.
The 9.1 delta: inline data reduction broadened across cluster types and workload profiles, dedupe support added for data-at-rest encryption, and compression improved — so compliance-driven encryption no longer forfeits deduplication savings, and capacity models should be recalculated. Option A is fabricated — RAID-6 has host-count minimums that two nodes cannot meet. Option C contradicts ESA's architectural basis on high-performance NVMe. Option D turns an expanded capability into a mandate; data services remain policy choices, not forced settings.
What deployment-networking friction does the 9.1 LACP enhancement remove?
It enables LACP between NSX Edge nodes and Tier-0 gateways
LACP can now be configured within the workload-domain deployment UI, whereas 9.0 required post-deployment vDS changes or API workarounds outside the automated bring-up
It removes the need for redundant uplinks entirely
It converts all uplinks to active-active without switch configuration
The delta is Day-0 compliance: environments whose network standards mandate LACP to the leaf can configure it during workload-domain deployment instead of deviating from (and fighting) the automated path afterwards. Option A relocates the feature to Edge/T0 routing — a different layer with its own LACP considerations that this UI change does not address. Option C is backwards: LACP is about USING redundant uplinks as an aggregate, not removing them. Option D omits the physical reality — LACP requires matching switch-side port-channel configuration; nothing is 'without switch configuration'.
Holodeck 9.1 delivers exactly the five items in option A: VNA cluster networking aligned to 9.1 multi-vNIC patterns, Click-to-Deploy GitOps lab definitions, HoloRouter gaining Vault (secrets), Authentik (identity), and Technitium (DNS), dual-site deployments, and a nested ESX floor of 8.0 U3. Option B is wrong on every element — and the HoloRouter is extended, not removed. Option C swaps in plausible-sounding but incorrect components (Keycloak/BIND instead of Authentik/Technitium) and invents triple-site. Option D gets two items right but misses the HoloRouter service additions and understates the ESX floor (U1 vs U3) — the precision trap.
Why is Holodeck 9.1's dual-site support strategically valuable for a VCDX candidate?
It doubles available lab compute capacity
DR and failover claims in the design can be validated with measured results (e.g., rehearsed failover with observed RTO) instead of asserted from documentation
It removes the need for a DR section in the design document
Defense panels weight validated evidence: with two nested sites in one deployment, site protection and failover scenarios become executable experiments — 'I failed over and measured RTO' beats datasheet citations. Option A misreads the feature: a second site consumes MORE host resources; it adds topology, not capacity. Option C is backwards — it strengthens the DR section rather than removing it. Option D confuses dual-site with multi-tenancy of the lab host; instance separation is a different mechanism (and shared-host HoloRouter changes still require coordination).
A colleague plans to adopt Holodeck 9.1 using existing nested ESX 8.0 U1 base templates and to reconfigure the HoloRouter on a shared lab host. What two corrections apply?
U1 templates are fine, but the HoloRouter must be redeployed per user
Templates must be rebuilt to at least ESX 8.0 U3 (the 9.1 minimum), and HoloRouter reconfiguration on a shared host must be coordinated because it overwrites global configuration
Templates must be ESX 9.0, and HoloRouter changes are per-instance so no coordination is needed
Both plans are acceptable as long as GitOps deploy is used
Two independent gotchas: Holodeck 9.1's nested ESX floor is 8.0 U3, so U1 templates are below the minimum and must be rebuilt; and Set-HoloRouter-style reconfiguration rewrites global configuration (dnsmasq/routing state) on the shared router, so on a multi-user host it is destructive without coordination. Option A accepts the invalid template and prescribes a redeploy model that is not how the shared HoloRouter works. Option C overshoots the floor (8.0 U3, not 9.0) and dangerously asserts per-instance isolation of router changes. Option D treats GitOps as a shield — declarative deploy does not change the template floor or the shared-router blast radius.
New in 9.1: a common runtime and shared set of components unifying the architecture of lifecycle and operational capabilities across the fleet. In 9.0 these were loosely coupled appliances with separate runtimes and update cadences. Design impact: one availability/sizing unit to protect, back up, and sequence in upgrades — and a smaller compatibility matrix.
What survives an outage of the VCF management services runtime?
The workload data plane: running VMs, vCenter, NSX data path, vSAN. What stops: lifecycle operations, fleet monitoring/operations, automation intake. Defense-critical distinction — quantify the restore RTO for the management runtime rather than claiming 'no SPOF'.
API-first in 9.1 — what specifically changed vs 9.0?
9.1 broadens OpenAPI coverage (VODAP) and expands Python/Java SDK and PowerCLI coverage per techdocs; Terraform is referenced only directionally in blogs — techdocs claims no provider parity. In 9.0, APIs existed but tool coverage was uneven (some tasks UI-only or PowerCLI-only). Tool choice becomes persona/skills-driven, not gap-driven.
vCenter maintenance-notification API (9.1)
New API other components can query to learn whether vCenter maintenance is planned or in progress. During maintenance, the Envoy reverse proxy returns a 503 header with estimated completion time. Lets automation pause gracefully instead of failing mid-run.
High Session Count alarm (vCenter 9.1)
New alarm firing as vCenter approaches its session limit (3000 by default). The message names the top 5 users with more than 100 sessions each (service accounts included) — pinpointing which integration is overloading vCenter.
Increased Request Load alarm (vCenter 9.1)
New alarm firing when a service endpoint nears its active-request limit (1024 for most endpoints). Message identifies impacted services and endpoints. Pairs with the new vCenter utilization-monitoring API for tracking request volumes against thresholds.
vCenter scale improvements in 9.1
Up to 20% more operations per minute in extra-large vCenter deployments (techdocs; the launch blog says 25%); concurrent VM backup operations scale to 500–1000 depending on vCenter size; file transfers use dedicated threads so backups no longer starve other vCenter operations.
Design decision: choosing an automation tool in 9.1
With OpenAPI-aligned Python/Java SDKs, PowerCLI, and Terraform all first-class, document tool choice per persona: Terraform for declarative infra teams, Python/Java SDKs for integration development, PowerCLI for vSphere-native ops teams. State the skills implication and the handling of vCenter-maintenance 503s.
VCF Operations pillar reorg (9.1)
Console navigation reorganized into four task-oriented pillars: Build (provision domains/clusters/networking incl. LACP), Manage (updates, certificates, passwords/policies, licensing, capacity planning & cost metering), Operate (monitoring, metrics, logs, diagnostics, sustainability), Protect (compliance, security posture, backup, cyber recovery). UX change only — not an RBAC boundary.
Which pillar owns: password policies? continuous compliance? domain deployment? capacity?
Password policies → Manage. Continuous compliance → Protect. Workload-domain deployment (incl. LACP config) → Build. Capacity planning & cost metering → Manage; monitoring/diagnostics/logs → Operate. Map these pillars to your RACI; budget a runbook refresh when upgrading from 9.0 navigation.
vCenter quick patch
9.1 patching mode that updates only the RPMs/binaries with actual code changes (traditional in-place patching updated every RPM). Downtime: zero to 5 minutes in three service groups per techdocs (zero / 0-3 min / 3-5 min); the launch blog says 'under 1 minute, sometimes zero'. VM/K8s deployments and API workflows keep running. Aimed at rapid security-fix rollout.
quick patch vs RDU vs live patch — disambiguate
Quick patch: vCenter, changed-binaries-only, zero-to-5-min downtime per techdocs. RDU (reduced downtime upgrade): vCenter, migration-based upgrade to a new appliance VM; in 9.1 usable with the online depot, not just ISO. Live patch: ESX hosts, patches running kernel memory with no maintenance window. Three different mechanisms, three different targets.
Single-API vCenter resize (9.1)
The deployment/size API (PATCH method, invokable from vCenter Developer Center API Explorer) scales vCenter compute and disk with one call plus a reboot. Replaces the multi-step manual resize procedure of 9.0. Makes vCenter right-sizing a low-risk, scriptable operation.
ESX Live Patch + TPM: what changed in 9.1?
9.1 supports live patch on TPM-enabled hosts. In 9.0, TPM-enabled hosts were excluded, forcing a choice between TPM attestation and reboot-free patching. Security-hardened fleets no longer sacrifice either.
ESX live patch coverage in 9.1
Covers up to ~80% of patches. Expanded to more of the vmkernel (with better performance) and additional user-space daemons: vSAN daemons, core storage daemons and libraries. The remaining ~20% falls back to maintenance-mode + reboot remediation.
ESX live patch default behavior and 'enforce' mode
Enabled by default on all clusters: uses live patch when the payload is capable, otherwise automatically falls back to maintenance-mode/reboot. 'Enforce' mode allows live-patch-only remediation and blocks remediations requiring reboot. Configurable at vCenter level (applies to all clusters) or overridden per cluster.
Certificate auto-renewal in 9.1
VMCA-managed certs auto-renew: vCenter TLS 5 days before expiry, ESX 30 days before (configurable via vpxd.certmgmt.certs.autoRenewThreshold). External-CA certificates are NOT auto-renewed — still the administrator's job. If the VMCA root is <1 year from expiry, it (plus solution-user leaf certs) renews during vCenter upgrade.
vCenter VM hardware version after upgrading to 9.1
RDU-based upgrades (8.x→9.1 or 9.0.x→9.1) create a new vCenter VM, automatically moving it from HW version 10 to version 17. In-place updates do NOT: the admin must upgrade VM hardware afterwards, which requires powering off the vCenter VM.
vLCM image checksum (9.1)
vLCM images now carry a SHA256 checksum of the image definition, so images exported/imported across vCenter instances can be integrity-verified by comparing source and destination checksums. Note: it checksums the image definition, not the ESX VIBs themselves.
vLCM HCL validation without an HSM (9.1)
For vSAN clusters, vLCM now reports current device driver/firmware and validates against the HCL even when no third-party Hardware Support Manager is present (9.0 required an HSM). Caveat: some devices cannot report firmware without an HSM — it is first-level validation.
Zero Touch (Elastic) Provisioning of ESX
9.1: bare-metal hosts network-boot via UEFI HTTP/S directly against vCenter — no TFTP server (unlike classic Auto-Deploy), compatible with Secure Boot and TPM. The deploy rule's target cluster determines image and configuration (vSphere Configuration Profile, including bootstrapped vDS config). DHCP needed unless UEFI has a static boot IP.
Design implication of quick patch + live patch together
The maintenance-window cost of monthly security patching drops sharply for both vCenter (quick patch, zero-to-5-min per techdocs) and ESX (live patch, ~80% coverage, now incl. TPM hosts). Architects can commit to aggressive patch-cadence SLAs in security designs — but must document the fallback path (maintenance-mode rotation) for the non-live-patchable remainder.
NVMe Memory Tiering — mechanism
The hypervisor keeps active (hot) pages in DRAM and tiers cold pages to local NVMe SSDs, expanding effective memory per host without additional DRAM. Primary business case: higher VM density and lower TCO, since memory is usually the binding constraint in consolidation designs.
What did 9.1 ADD to memory tiering (vs 9.0)?
Two enhancements: (1) software-based mirroring — an additional NVMe device can be claimed as a mirror of the tier, closing the device-failure risk; (2) cost-savings analysis in VCF Operations, giving defensible TCO numbers. Configuration flows through vSphere Configuration Profiles.
Memory tiering — which workloads are poor candidates?
Latency-sensitive workloads with large hot working sets (in-memory databases, low-latency trading). Tiering assumes the active working set fits in DRAM; if hot pages land on NVMe, latency spikes. A defensible design states workload profile, DRAM:NVMe ratio, and the monitoring that validates the assumption — plus explicit exclusions (dedicated non-tiered clusters or host groups).
How is memory tiering configured in 9.1?
Via vSphere Configuration Profiles (desired state): NVMe devices are claimed for tiering, optionally a second NVMe device as software mirror. Because it is desired-state, hosts added to the cluster inherit tiering configuration automatically.
Memory tiering mirroring — the trade-off
Mirroring protects tiered pages from a single NVMe device failure (availability up) but doubles the NVMe capacity you must install per host (cost up) and consumes another drive bay. Classic AMPRS trade: availability vs cost. State the decision either way — panels ask what happens when the tier device dies.
DRS non-disruptive vMotion evacuation (9.1)
New maintenance-mode option: evacuate VMs only when destination hosts can absorb their current compute demand without contention; DRS may pre-rebalance to create room. 'Non-disruptive' = avoids resource contention on landing — it does NOT imply standard evacuation disrupts VMs. Two options now: Standard and Non-Disruptive.
vMotion concurrency change in 9.1
Default max concurrent vMotions is still 8, with slots releasing on a rolling basis (as each migration completes, the next begins) — and 9.1 adds Dynamic Concurrency Control, which can raise concurrency beyond 8 when resources allow. In earlier releases a batch of 8 had to fully complete before the next batch started. Result: better utilization, more even vMotion load distribution across hosts.
Encrypted vMotion + Intel QAT (9.1)
Encrypted vMotion operations can offload encryption to Intel QuickAssist Technology hardware, returning CPU cores to workloads. Relevant when security policy mandates encrypted vMotion fleet-wide and CPU headroom is a constraint.
Topology Aware Scheduler (9.1)
NUMA scheduler evolution: event-driven inline updates for more consistent placements; considers cache and memory-bandwidth contention (not just CPU ready time); optimized for high-density CPUs; on asymmetric NUMA topologies it places sibling NUMA clients of one VM on nodes that are close to each other.
AI/accelerator platform changes in 9.1
Enhanced DirectPath I/O gains virtualization benefits (Storage vMotion, snapshots incl. memory, hot add/remove, ESX Live Patch compatibility) and direct GPU-to-GPU RDMA/RoCE; AMD vIOMMU enables PCI passthrough on AMD systems; NVIDIA vGPU can combine time-slicing and MIG modes; Flow Processing Offload moves network rule processing to hardware.
Fleet/instance password policies (9.1)
VCF Operations centrally defines password policies — length, complexity, account lockout, change interval (expiration), and password history — applied at fleet level or per VCF instance. Replaces the per-component sprawl (ESX advanced settings, vCenter SSO policies, NSX CLI, appliance PAM) documented as dozens of separate VVS design decisions in 5.2.
Password ROTATION vs password POLICY — which was centralized when?
Rotation: centralized in 9.0 (Password Manager, automated credential rotation). Policy (length/complexity/lockout/interval/history): centralized in 9.1 at fleet or instance scope. A favorite delta-exam distinction.
Fleet-level vs instance-level password policy — the design decision
Fleet scope maximizes consistency and single-point audit evidence. Instance scope handles sovereignty/tenancy/regulatory exceptions (e.g., one regulated instance with stricter length and shorter change interval). Defensible answer: fleet default + documented instance exceptions, plus a stated approach for components outside VCF Operations scope.
Continuous Compliance Enforcement (9.1)
Advanced Cyber Compliance now supports continuous compliance REMEDIATION (not just detection/reporting) and unified security posture management across the VCF stack. Design shift: drift gets corrected automatically per policy, so the design decision becomes which controls are auto-remediated vs alert-only (change-control implications).
On-prem ransomware recovery (9.1)
Integrated cyber recovery to on-premises VCF isolated clean rooms — the clean-room recovery pattern (previously cloud-centric in VMware's portfolio) lands on-prem. Design relevance: recovery isolation domain, clean-room capacity sizing, and periodic recovery testing become in-platform design items.
Guest OS customization changes (9.1)
GOSC reaches parity with VMware Cloud Director capabilities (supporting vCD→VCF Automation migration): set/reset Linux root password, reset Windows administrators-group passwords, run Windows customization scripts. Also: IPv6-only guest networking (explicitly disable IPv4 — no parallel IPv4 config needed) and network-only re-customization of powered-on VMs.
Why does live patch on TPM hosts matter for SECURITY design (not just LCM)?
Before 9.1, enabling TPM (host attestation, key sealing) disqualified hosts from live patch — so hardened fleets patched slower or rebooted more. 9.1 removes the disincentive to enable TPM everywhere, letting the security baseline mandate TPM without an operational-cost argument against it.
External vault integration in the credential story
VCF supports external secrets backends (e.g., HashiCorp Vault, OIDC-compatible vaults) for platform credentials — delivered in VCF 9.1. Holodeck 9.1's HoloRouter now ships Vault (plus Authentik for IdP), so the integration is exercisable in the nested lab.
VKS 3.6 headline numbers
Up to 500 clusters per Supervisor, and roughly 70% faster cluster provisioning through linked-clone fast deploy. Ships with VCF 9.1.
VKS linked-clone fast deploy — mechanism and trade-off
Cluster node VMs deploy as linked clones of a base image instead of full clones — ~70% faster provisioning and faster rolling upgrades (each replacement node boots sooner). Trade-off: nodes share a base-image dependency and accumulate delta disks; monitor delta growth and manage base-image lifecycle.
Multi-vNIC cluster node networking (9.1)
VKS cluster nodes can attach multiple vNICs (secondary NICs on workers, pod interfaces via Antrea, SR-IOV), separating application, storage, and management traffic. Shipped with VKS 3.6.1 (2026-04-13) alongside VCF 9.1. Lets Kubernetes clusters honor the same traffic-separation NFRs as the VM estate (regulated segmentation, dedicated storage paths).
Design impact of the 500-cluster Supervisor ceiling
'One Supervisor, many clusters' becomes viable at enterprise scale, shifting isolation from separate Supervisors toward namespaces + network policy + zones. Counterweight: a big Supervisor is a bigger failure and upgrade domain. Keep separate Supervisors for hard tenancy, independent failure domains, or divergent lifecycle needs.
Self-service namespace CaaS (9.1)
Developers self-provision a namespace with registry access, ingress, resource quotas, and identity — all inherited from existing VCF constructs — without owning a Kubernetes cluster. Positioned as the 'simpler, faster way to run a few containers' without full K8s management.
CaaS namespaces vs dedicated VKS clusters — when each?
Namespaces: teams needing containers with corporate guardrails, no K8s version/operator ownership — avoids cluster sprawl and its patch burden. Dedicated clusters: teams needing cluster-scoped control (CRDs/operators/versions), hard isolation, or compliance workloads where namespace isolation is insufficient. Document the tiering rule in the consumption model.
What does 'identity inherited from VCF constructs' buy the architect?
Security review happens once at platform level: namespace consumers get quotas, registry policy, ingress, and IdP-backed identity from VCF projects/policies rather than per-request manual plumbing. Same centralize-then-inherit pattern as 9.1 fleet password policies — a strong governance story in defenses.
VKS upgrade-window math in 9.1
Rolling upgrades replace nodes one by one; linked-clone fast deploy makes each replacement node available faster, compressing total upgrade duration per cluster. At 500 clusters per Supervisor, per-cluster minutes saved multiply into real maintenance-calendar relief — quantify it in ops designs.
Namespace noisy-neighbor control
Many namespaces on shared Supervisor capacity require quota discipline (CPU/memory/storage quotas per namespace) plus utilization monitoring in the Operate pillar. Without quotas, self-service becomes capacity roulette — call out the quota policy as a design decision.
Multi-vNIC + VNA cluster networking in Holodeck 9.1
Holodeck 9.1's Virtual Network Appliance supports cluster-style networking aligned with multi-vNIC patterns, so traffic-separation designs (app/storage/mgmt on distinct nested networks) can be built and demonstrated in the nested lab rather than asserted on paper.
Native S3 object storage (9.1) — status and scope
TECH PREVIEW (not GA): self-service S3-compatible object storage for developers with IT guardrails, deployed/scaled/managed via the same workflows as block and file storage. Never place it on a production requirement's dependency path; frame as roadmap. Exam trap: distractors calling it GA/supported.
vSAN data services changes in 9.1
Deduplication and compression extended across more cluster types and workload profiles; deduplication now supported together with data-at-rest encryption (previously an either/or); improved compression. Design impact: revisit data-reduction assumptions in vSAN capacity models — compliance encryption no longer costs you dedupe.
LACP in the workload-domain deployment UI (9.1)
LACP link aggregation is configurable directly in the workload-domain deployment workflow. In 9.0 LACP required post-deployment vDS changes or API workarounds outside the automated bring-up. Customers whose network standards mandate LACP to the leaf can now comply at Day-0 without deviating from the supported automated path.
GA vs Tech Preview in the 9.1 feature set — sort them
GA: memory tiering (+mirroring), vCenter quick patch, resize API, ESX live patch on TPM, fleet password policies, VKS 3.6 scale/fast-deploy, multi-vNIC, CaaS namespaces, LACP UI, ZTP, vSAN dedupe+encryption. Tech Preview: native S3 object storage. Citing TP features as requirement solutions is a defense credibility hit.
Holodeck 9.1 — the headline changes
(1) RESOURCE ENVELOPE DOUBLES — VCF 9.1 single site 48 CPU / 650 GB / 2.2 TB, vs 9.0's 32 / 325 GB / 1.1 TB (2.0x RAM and disk), plus new soft validation that warns instead of blocking; (2) VNA cluster networking = distributed, vs NSX Edge = centralized, in both management and workload domains; (3) Click-to-Deploy GitOps as an OVA deploy-time toggle (default on, cannot be retrofitted); (4) HoloRouter services stack — Vault :8200, Authentik SSO :9443, Technitium DNS :5380, authenticated Webtop :30000 behind an HTTPS reverse proxy; (5) dual-site config auto-generation; (6) minimum nested ESX 8.0 U3. Plus six breaking changes and NO documented 9.0->9.1 upgrade path.
Holodeck 9.1 minimum nested ESX version
8.0 U3 — unchanged since Holodeck 9.0 (not a new 9.1 requirement). Nested base templates on 8.0 GA/U1/U2 were already below the floor and need rebuilding regardless of 9.1 adoption.
Holodeck dual-site support — why it matters for VCDX prep
Two nested sites in one deployment make DR, site protection, and failover scenarios practicable in the lab. 'I failed over site A→B in Holodeck and measured RTO' is validated evidence — categorically stronger in a defense than datasheet citations.
Click-to-Deploy GitOps (Holodeck 9.1)
Lab environment definitions live in a Git repo and the toolkit reconciles the nested environment to them — lab-as-code. Rebuilds become deterministic, diffable, reviewable; lab states become citable artifacts for design validation.
HoloRouter service additions — what does each enable?
Vault (:8200): external-vault credential integration end-to-end, and it auto-provisions certificates for the other services. Authentik (:9443): external IdP/SSO federation with SCIM — Initialize-Authentik and Set-VCFSSOConfiguration wire a real OIDC chain into VCF Operations. Technitium (:5380): the new DNS management UI (the Get-/Set-/Remove-HoloDeckDNSConfig cmdlets are deprecated from 9.1). All behind an HTTPS reverse proxy with an internal CA. CAUTION: the docs do NOT establish that Technitium retires dnsmasq — Set-HoloRouter and Reset-HoloRouter still reference DNSMASQ/FRR — so treat 9.0-era DNS fixes and the shared-host Set-HoloRouter caution as still applicable until re-baselined on a live 9.1 router.
vSphere Configuration Profiles expansion (9.1)
VCP now: honors vSAN maintenance-mode and object-accessibility policies during remediation; applies advanced vSAN config cluster-wide; configures memory tiering (incl. mirror device claim); bootstraps host config and vDS during Zero Touch Provisioning; and can auto-remediate incoming hosts (host-specific attributes auto-extracted; automatic remediation disabled by default, enable per vCenter or cluster).
Holodeck 9.1 resource requirements — the numbers
VCF 9.1: 48 CPU / 650 GB / 2.2 TB single site; 96 / 1536 GB / 5 TB dual site. VVF 9.1: 24 / 512 GB / 2 TB single; 48 / 1024 GB / 4 TB dual. Versus VCF 9.0 single site (32 / 325 GB / 1.1 TB) that is +50% CPU and exactly 2.0x on RAM and disk. HoloRouter appliance itself: 6 vCPU / 12 GB / 75 GB (was 4/8). On a 1 TB host: VCF 9.1 single site fits, VVF 9.1 dual site sits exactly on the RAM line, VCF 9.1 dual site is unachievable.
Soft validation (Holodeck 9.1) — what it changes operationally
New in 9.1: pre-checks WARN rather than block when resources fall below recommendation. Verbatim: 'allows deployment with lower resources, but performance may be impacted.' The operational consequence is that an under-spec deployment starts and then DEGRADES rather than failing fast — the guardrail that used to catch sizing errors early is gone, so the architect owns the risk explicitly. Size before deploying; do not rely on pre-checks to stop you.
VNA vs NSX Edge cluster in Holodeck 9.1 — the design axis
VNA (Virtual Network Appliance) cluster = DISTRIBUTED networking. NSX Edge cluster = CENTRALIZED networking. Both supported in both management and workload domains; Edge remains the default and is fully supported. Day 0 switches -VnaClusterMgmtDomain / -VnaClusterWkldDomain; Supervisor placement takes Centralized or Distributed. This is a genuine VCDX design-decision axis — a real trade-off with justification, alternatives and implications, which is exactly what a panel probes. Caveat: the published cmd_reference is stale and shows no VNA switch — verify with Get-Help on a live router.
Holodeck 9.1 breaking changes — what breaks a 9.0.x runbook
Six: (1) Import-HoloDeckConfig now requires MANDATORY -Site a|b, failing every existing call; (2) New-HoloDeckConfig emits two config files (Site A + Site B), Site A auto-loaded; (3) Webtop requires authentication; (4) Get-/Set-/Remove-HoloDeckDNSConfig deprecated after 9.0.2 (use the Technitium UI); (5) VCF Installer 9.0.0.0 no longer supported; (6) services moved to HTTPS behind a reverse proxy, breaking old HTTP automation. Still carried from 9.0.2: -Interactive removed, -InstanceID mandatory, MasterCIDR must be /20.
Is there a Holodeck 9.0 -> 9.1 upgrade path?
No. There is no 9.0-to-9.1 upgrade or migration guidance anywhere in the documentation — the only migration section covers 9.0/9.0.1 to 9.0.2 and is a single image with no textual steps. Evidence for clean redeploy: Vault/Authentik/Technitium/CA/reverse-proxy are OVA-resident; GitOps is a deploy-time-only property; HoloRouter sizing changed; new /holodeck-runtime/bin/9.1.0.0/ staging path; DNS substrate and config model both changed. Plan a parallel 9.1 deployment and keep the 9.0.x instance intact.
Why does Holodeck 9.1 publish no Known Issues?
It does not — and that is a DOCUMENTATION GAP, not evidence of a clean release. Versions 9.0, 9.0.1 and 9.0.2 each carry 5-9 documented issues; 9.1 has 'What's New' only. Source 9.1 issues from GitHub Issues and the Broadcom community instead. Issues almost certainly still applicable (documented under 9.0.x, repeated in the 9.1 FAQ): memory tiering destabilizes nested workloads; nested vSAN ESA on non-vSAN backing over-consumes storage; VCFA All Apps Org hard-codes a /28 from 10.1.0.0/20. Related: vmware/Holodeck publishes no GitHub releases at all.
Holodeck 9.1 GitOps — the one thing you cannot undo
GitOps is enabled as an Extra Property at HoloRouter OVA deployment time and is ON BY DEFAULT. It auto-provisions GitLab, pre-loads a repo with the deployment pipeline, and registers the router as a runner, so deployments can be triggered from the GitLab UI with no PowerShell. There are ZERO GitOps cmdlets. Critically it CANNOT be retrofitted onto an existing HoloRouter — the decision is made once, at deploy time, which makes it a genuine pre-deployment design choice rather than a Day-2 option.