Academy/VCF 9.0 → 9.1: What Changed
VCF

VCF 9.0 → 9.1: What Changed

VCF 9.1vcdx-distinguished

Architect-level delta between VCF 9.0 and VCF 9.1 (GA May 2026): the unified VCF management services runtime, the API-first platform story, patching and LCM improvements (vCenter quick patch, single-API resize, ESX live patch on TPM hosts), NVMe memory tiering, fleet password policies, the VCF Operations pillar reorg, VKS 3.6 and multi-vNIC cluster networking, self-service CaaS namespaces, native S3 object storage (Tech Preview), LACP in the deployment UI, and Holodeck 9.1 changes. Each topic covers what changed, why it matters for design and operations, and how it shows up in exams and VCDX-style defenses.

Version Evolution

This section is the 9.0→9.1 delta companion to section 18 (VCF Evolution 2.x→9.0). VCF 9.1 GA'd May 2026. Headline deltas: VCF management services common runtime; API-first with consistent OpenAPI across Python/Java SDKs, PowerCLI, Terraform; vCenter quick patch (<1 min, sometimes zero downtime) and single-API resize; ESX Live Patch on TPM-enabled hosts (~80% patch coverage); NVMe memory tiering with software mirroring; fleet/instance password policies; VCF Operations Build/Manage/Operate/Protect pillars; VKS 3.6 (500 clusters/Supervisor, ~70% faster linked-clone provisioning); multi-vNIC cluster node networking; self-service CaaS namespaces; native S3 object storage (Tech Preview); LACP in the deployment UI. Holodeck 9.1: VNA cluster networking, Click-to-Deploy GitOps, HoloRouter Vault/Authentik/Technitium, dual-site, min nested ESX 8.0 U3.

Learning Outcomes

  • Explain the VCF management services common runtime introduced in 9.1 and its impact on management-plane availability, sizing, and lifecycle design
  • Contrast 9.0 and 9.1 patching models (in-place vs quick patch, live patch on TPM hosts, single-API vCenter resize) and quantify the maintenance-window impact
  • Design a memory-tiering-enabled cluster (NVMe tier sizing, mirroring, VCP-driven configuration) and defend the VM-density/TCO trade-off
  • Design fleet- or instance-level password policies in VCF Operations and map them to compliance requirements
  • Position VKS 3.6 (500-cluster scale, linked-clone fast deploy), multi-vNIC node networking, and self-service CaaS namespaces in a developer-platform design
  • Identify which 9.1 features are GA versus Tech Preview (native S3 object storage) and state the defense-safe way to reference each

VCF Management Services: The Unified Runtime#

What Changed

VCF 9.1 introduces 'VCF management services' — a common runtime and shared set of components that unifies the architecture of lifecycle and operational capabilities across the fleet. In 9.0, the management plane was already consolidating (VCF Installer deploying VCF Operations, SDDC Manager functions absorbed), but individual services still ran as loosely coupled appliances with their own runtimes, update cadences, and HA stories. In 9.1 those capabilities converge onto one runtime layer with shared components.

9.0 → 9.1 Contrast

VCF 9.0: VCF Operations + Fleet Manager + per-function appliances; separate lifecycle for each management component; manual failover for some fleet components.
VCF 9.1: VCF management services = common runtime hosting lifecycle + operational services; one architectural unit to size, patch, protect, and monitor.

Why It Matters (Design/Ops Implications)

  1. Availability design: the management-services runtime becomes a single logical dependency for Day-2 operations. Your AMPRS availability analysis must treat it like you treated SDDC Manager + Aria stack in 5.2 — what breaks when it is down (LCM, monitoring, automation intake) and what keeps running (data plane, vCenter, NSX, workloads).
  2. Sizing and placement: one runtime means one capacity model instead of per-appliance sizing sprawl; place it in the management domain and protect with HA + backup like any other management component.
  3. Upgrade sequencing simplifies: fewer independently versioned management appliances shrinks the compatibility matrix — the historical #1 source of VCF upgrade failures.

Exam/Defense Relevance

Defense panels probe single points of failure: be ready to state explicitly that workload data plane survives management-services outage, and quantify RTO for restoring it from backup. Delta-style exam items test that 'VCF management services' is a 9.1 term — it does not exist in 9.0 documentation.

Key Takeaways

  • VCF management services (new in 9.1) = common runtime unifying lifecycle and operational capabilities that were separate components in 9.0
  • Treat the runtime as one availability/sizing unit in design docs — analyze what fails when it is offline (LCM, ops) vs what survives (workloads, vCenter, NSX data plane)
  • Fewer independently versioned management appliances = smaller compatibility matrix and simpler upgrade sequencing
  • Defense tip: never claim the management plane is 'not a SPOF' without stating the workload-continuity and restore-time story

API-First Platform: Consistent OpenAPI Across SDKs#

What Changed

VCF 9.1 formalizes the API-first platform story: a consistent OpenAPI surface across the stack, with aligned SDKs for Python and Java, PowerCLI coverage, and a Terraform provider — all tracking the same API definitions. In 9.0 the APIs existed but tooling parity was uneven (some operations UI-only or PowerCLI-only, SDKs trailing the API).

9.0 → 9.1 Contrast

VCF 9.0: APIs available but inconsistent coverage per tool; automation designs had to document per-tool gaps.
VCF 9.1: one OpenAPI contract; Python/Java SDKs, PowerCLI, and Terraform generated/aligned from it — pick the tool that fits the team, not the gap.

New API capabilities land alongside: the vCenter deployment/size resize API, a vCenter maintenance-notification API (components can query whether vCenter maintenance is planned or underway; Envoy returns a 503 header with estimated completion during maintenance), and a vCenter utilization-monitoring API with new alarms (High Session Count — default limit 3000 sessions, reports top-5 users over 100 sessions; Increased Request Load — endpoint limit 1024 active requests for most endpoints).

Why It Matters (Design/Ops Implications)

  1. Automation strategy becomes a real design decision with alternatives: Terraform for declarative infra, Python SDK for integration glue, PowerCLI for ops teams — all first-class. Document the choice and its skills implication.
  2. The maintenance-notification API lets you design automation that degrades gracefully during vCenter patching instead of failing mid-run — pair it with quick patch for near-zero-impact maintenance.
  3. Utilization APIs + session alarms give you enforceable guardrails against integration sprawl (backup tools, CMDB sync, monitoring all hammering vCenter).

Exam/Defense Relevance

Expect scenario items on choosing the right SDK/tool for a requirement and on the new alarms' thresholds. In defense, an 'automation-first operating model' claim is only credible if you can name the API contract and the failure-mode handling (503 during maintenance).

Key Takeaways

  • 9.1 delivers consistent OpenAPI + aligned Python/Java SDKs, PowerCLI, and Terraform — tool choice is now preference/skills-driven, not gap-driven
  • New vCenter maintenance-notification API: Envoy returns 503 with estimated completion so automation can pause instead of fail
  • High Session Count alarm fires near the 3000-session default limit and names top-5 heavy users (>100 sessions); Increased Request Load alarm fires near 1024 active requests per endpoint
  • Design decision to document: which automation tool per persona, with rationale and skills implication

Patching & LCM: Quick Patch, Single-API Resize, Live Patch on TPM#

What Changed

Three lifecycle improvements headline 9.1:

  1. vCenter quick patch — patches only the RPMs/binaries with actual code changes instead of updating every RPM in-place. Downtime drops to under 1 minute, and to zero for some service payloads. VM/K8s deployments and API workflows continue during the patch. Targeted at rapid security-fix rollout.
  2. Single-API vCenter resize — the deployment/size API (PATCH method, callable from the vCenter Developer Center API Explorer) scales vCenter compute and disk with one call plus a reboot. In 9.0 resizing was a multi-step manual procedure.
  3. ESX Live Patch on TPM-enabled hosts — 9.0 live patch excluded TPM-enabled hosts, forcing security-hardened fleets to choose between TPM attestation and reboot-free patching. 9.1 removes that conflict. Live patch also now covers more of the vmkernel plus additional user-space daemons (vSAN daemons, core storage daemons/libraries) and covers up to ~80% of patches. It is enabled by default with automatic fallback to maintenance-mode remediation; an 'enforce' setting (cluster- or vCenter-level) allows live-patch-only remediation.
Supporting changes: reduced-downtime upgrade (RDU) can now use the online depot (not just mounted ISO); vLCM images carry a SHA256 checksum of the image definition for cross-vCenter export/import integrity; RDU-based 8.x→9.1 or 9.0.x→9.1 upgrades automatically raise vCenter VM hardware to version 17 (in-place updates require a manual, powered-off HW upgrade afterwards); VMCA-managed certs auto-renew (vCenter TLS 5 days before expiry, ESX 30 days, threshold configurable via vpxd.certmgmt.certs.autoRenewThreshold — external CA certs are NOT auto-renewed).

Why It Matters (Design/Ops Implications)

Maintenance-window math changes: a monthly vCenter security patch used to cost a scheduled outage; now it is a sub-minute blip. That directly lowers the operational-cost argument against aggressive patch cadence — a compliance win you can quantify. Live patch on TPM ends the 'hardening vs patchability' trade-off in security-sensitive designs. The 'enforce' live-patch setting is itself a design decision: it guarantees no surprise reboots but blocks non-live-patchable remediations, so you must define the exception process.

Exam/Defense Relevance

Classic distractor set: quick patch (vCenter, changed-binaries-only) vs RDU (vCenter, migration-based upgrade) vs live patch (ESX kernel/daemons, no reboot). Know which applies where, the TPM change, and the ~80% patch coverage figure. Defense: bake quick patch + live patch into your patch-cadence SLA and be ready to state the fallback path for the other ~20%.

Key Takeaways

  • vCenter quick patch updates only changed RPMs/binaries: <1 min downtime, sometimes zero; workflows and deployments keep running
  • deployment/size PATCH API resizes vCenter compute+disk in one call + reboot (was multi-step manual in 9.0)
  • ESX Live Patch now works on TPM-enabled hosts, covers ~80% of patches, is on by default with maintenance-mode fallback, and supports vSAN/core-storage daemons
  • VMCA certs auto-renew (vCenter TLS 5 days, ESX 30 days pre-expiry); external CA certs remain the admin's responsibility
  • RDU upgrades to 9.1 auto-upgrade vCenter VM hardware to v17; in-place updates need a manual powered-off HW upgrade

Advanced Memory Tiering with NVMe#

What Changed

VCF 9.1 enhances memory tiering: the hypervisor keeps active (hot) pages in DRAM and tiers cold pages to local NVMe SSDs, presenting an expanded effective memory capacity per host. 9.1 adds software-based mirroring of the NVMe tier (claim a second NVMe device as mirror) and cost-savings analysis in VCF Operations. Configuration is done through vSphere Configuration Profiles: NVMe devices are claimed for memory tiering (plus an optional mirror device) as part of the cluster's desired state.

Why It Matters (Design/Ops Implications)

  1. VM density and TCO: memory is the dominant constraint (and cost) in most consolidation designs. Tiering raises usable memory per host without DRAM purchases — that changes host sizing, consolidation ratios, and the cost model you present to stakeholders. The built-in cost-savings analysis gives you defensible numbers instead of hand-waving.
  2. Performance risk: cold-page tiering assumes a working-set smaller than DRAM. Latency-sensitive workloads with large, hot working sets (in-memory DBs) are poor candidates. A defensible design states the workload profile, the DRAM:NVMe ratio chosen, and the monitoring that validates the assumption (active-memory vs tiered-page metrics in VCF Operations).
  3. Failure domain: an NVMe tier device failure without mirroring impacts the VMs whose pages live there — the 9.1 software mirroring option exists precisely to close that risk; claiming it doubles the NVMe capacity requirement.
  4. Operational consistency: because configuration flows through vSphere Configuration Profiles, tier config is desired-state and survives host additions — new hosts inherit it automatically.

Exam/Defense Relevance

Expect items on what tiering does (DRAM=hot, NVMe=cold), what 9.1 adds (mirroring, cost analysis, VCP-driven config), and candidate-workload selection. In a defense, memory tiering is a strong cost-requirement answer — but the panel will attack with 'what about your Tier-1 database?' Have the exclusion/placement answer ready (host groups or dedicated clusters without tiering).

Key Takeaways

  • Memory tiering: DRAM holds hot pages, local NVMe holds cold pages — more effective memory per host without buying DRAM
  • 9.1 adds software mirroring of the NVMe tier (optional second device) and cost-savings analysis in VCF Operations
  • Configured via vSphere Configuration Profiles (desired state) — new cluster hosts inherit tiering config automatically
  • Design defensibly: state workload profile, DRAM:NVMe ratio, mirroring decision, and the monitoring that validates the working-set assumption
  • Poor fit for latency-critical large-working-set workloads — plan exclusions explicitly

Password Policy Management at Fleet Scale#

What Changed

VCF 9.1 adds centralized password-policy management: policies for length, complexity, account lockout, change interval (expiration), and password history can be defined once and applied at fleet level or per VCF instance, from VCF Operations. In 9.0 (and painfully in 5.2), these policies were configured per component — ESX advanced settings per host, vCenter SSO policies per instance, NSX via CLI/API, appliance-local PAM settings — exactly the sprawl the VVS IAM solution documented as dozens of individual design decisions.

9.0 → 9.1 Contrast

VCF 9.0: password ROTATION was centralized (Password Manager), but password POLICY remained per-component.
VCF 9.1: policy definition (length, complexity, lockout, interval, history) is centralized with fleet/instance scoping; rotation remains centralized.

Why It Matters (Design/Ops Implications)

  1. Compliance mapping collapses from N decisions to one: 'fleet password policy aligned to <framework>' with per-instance exceptions where a regulator demands stricter settings. Audit evidence comes from one control point.
  2. Drift elimination: per-component policies drift (a rebuilt host missing its lockout setting is a classic finding). Central policy pushed from VCF Operations removes the drift class, and pairs with 9.1 continuous compliance enforcement (Advanced Cyber Compliance) which remediates posture drift automatically.
  3. Scoping is a real decision: fleet-level = consistency; instance-level = sovereignty/tenancy boundaries (e.g., a regulated instance with 14-char minimum and 30-day interval inside a fleet defaulting to 12/90). Document which and why.

Exam/Defense Relevance

Delta items will contrast rotation (9.0) vs policy (9.1) centralization — a subtle, testable distinction. For VCDX-style security sections, replace the old per-component policy decision cluster with one fleet-policy decision plus documented exceptions; panels reward the simplification but will ask how legacy/unmanaged components outside VCF Operations scope are handled.

Key Takeaways

  • 9.1 centralizes password POLICY (length, complexity, lockout, change interval, history) at fleet or instance scope via VCF Operations
  • Contrast with 9.0: rotation was already central (Password Manager); policy definition was still per-component
  • Fleet scope = consistency; instance scope = regulatory/tenancy exceptions — an explicit, defendable design decision
  • Pairs with 9.1 continuous compliance enforcement to remediate posture drift, not just report it

VCF Operations UX: Build / Manage / Operate / Protect Pillars#

What Changed

VCF Operations in 9.1 reorganizes its console around four task-oriented pillars:

Build — provisioning infrastructure: workload domains, clusters, networking (including the new LACP option in the deployment UI)
Manage — lifecycle: inventory, updates/patching, certificates, passwords/policies, licensing

Operate — monitoring, capacity, performance, diagnostics, logs

Protect — compliance (continuous enforcement), security posture, backup/restore, ransomware/cyber recovery

In 9.0 the navigation was organized more around underlying products/functions inherited from the Aria consolidation; operators had to know which tool-area owned a task. 9.1 aligns navigation to the operator's intent.

Why It Matters (Design/Ops Implications)

  1. Operating-model mapping: the pillars map cleanly onto team responsibilities (platform build team, LCM owner, NOC/SRE, security). Your operations design and RACI can now reference product navigation directly — 'Protect pillar tasks are owned by the security operations role' — which shortens runbooks and training.
  2. Documentation churn: any runbook, SOP, or training material written against the 9.0 navigation needs a refresh pass; budget it in the upgrade project (a real migration-cost line item architects often forget).
  3. It is UX reorganization, not RBAC: pillar visibility does not itself constitute a security boundary — roles/scopes still control authorization. Do not present the pillar split as a control in a defense.

Exam/Defense Relevance

Expect 'which pillar contains X?' items (password policies → Manage; continuous compliance → Protect; domain deployment with LACP → Build; capacity/diagnostics → Operate). In defenses, use the pillars as your operational-handoff narrative, but keep authorization claims anchored in roles, not UI layout.

Key Takeaways

  • 9.1 VCF Operations navigation = Build (provision), Manage (lifecycle), Operate (monitor/capacity), Protect (compliance/recovery)
  • Pillars align navigation to operator intent instead of product heritage — map them to your RACI and runbooks
  • UX reorg is not RBAC: roles and scopes still define authorization boundaries
  • Upgrade projects must budget a runbook/training refresh for the new navigation

VKS 3.6: Scale, Linked-Clone Fast Deploy, and Multi-vNIC Node Networking#

What Changed

vSphere Kubernetes Service 3.6 (the VKS release paired with VCF 9.1) brings:

  1. Scale: up to 500 clusters per Supervisor — consolidating what previously required multiple Supervisors or careful cluster rationing.
  2. Fast deploy/upgrade via linked clones: cluster node VMs deploy from linked clones instead of full clones, cutting provisioning time roughly 70% and speeding rolling upgrades (each replaced node comes up faster). Faster node replacement = shorter upgrade windows and quicker scale-out under load.
  3. Multi-vNIC cluster node networking: cluster nodes can attach multiple vNICs to separate application, storage, and management traffic onto distinct networks. Before 9.1, node networking was effectively single-homed — storage traffic (e.g., to external arrays or high-throughput CSI backends) shared the node's one network path with app and control traffic.

Why It Matters (Design/Ops Implications)

  1. The 500-cluster ceiling changes multi-tenancy topology: 'one Supervisor, many clusters' becomes viable for larger estates, moving isolation decisions from 'separate Supervisors' toward namespaces + network policy + vSphere Zones. Document the blast-radius trade-off explicitly: one Supervisor is now a bigger failure/upgrade domain.
  2. Linked-clone deploy trades disk-time for a shared-base dependency: know the storage implication (clones reference a base image; monitor delta growth) — it is the same architectural trade-off Holodeck users know from nested labs.
  3. Multi-vNIC unlocks conformance with traffic-separation NFRs (regulated environments demanding storage/mgmt/app isolation) inside Kubernetes clusters, and lets you honor existing physical network segmentation designs rather than collapsing them at the container layer.

Exam/Defense Relevance

Numbers are testable: 500 clusters, ~70% faster provisioning, VKS 3.6. Design items will probe when to use multiple Supervisors anyway (independent failure domains, different vCenters/zones, hard tenancy). For defense, multi-vNIC is your answer when a panelist asks how the containerized tier honors the same traffic-separation rules as the VM tier.

Key Takeaways

  • VKS 3.6: up to 500 clusters per Supervisor and ~70% faster provisioning via linked-clone fast deploy
  • Linked clones speed deploy/upgrade but introduce a shared base-image dependency — monitor delta growth
  • Multi-vNIC node networking separates app/storage/mgmt traffic per cluster node — extends physical traffic-separation NFRs into Kubernetes
  • Bigger Supervisors = bigger failure/upgrade domains: keep separate Supervisors for hard tenancy or independent failure domains

Simplified Container-as-a-Service: Self-Service Namespaces#

What Changed

VCF 9.1 adds a simplified CaaS consumption path: self-service namespace provisioning where a developer gets a namespace with registry access, ingress, resource quotas, and identity — all inherited from existing VCF constructs (projects, identity providers, network/storage policy) — without anyone standing up or managing a full Kubernetes cluster for them.

9.0 → 9.1 Contrast

VCF 9.0: 'a few containers' still meant requesting a VKS cluster (or sharing one via manually carved namespaces with hand-built registry/ingress/quota plumbing).
VCF 9.1: namespace-as-the-unit-of-consumption, provisioned self-service with guardrails pre-wired from the platform.

Why It Matters (Design/Ops Implications)

  1. Consumption-model design gains a middle tier. Your service catalog now has: full clusters (teams owning K8s versions/operators), self-service namespaces (teams that just need to run containers), and VMs. Matching workload personas to tiers is a design decision with cost and operational-load consequences — namespaces avoid per-team cluster sprawl and its patching burden.
  2. Governance inheritance is the killer feature: quotas, registry policy, ingress, and identity come from VCF-level constructs, so security review happens once at the platform level, not per namespace request. This is the CaaS analog of the fleet password-policy story — centralize the control, inherit everywhere.
  3. Capacity management shifts: many small namespaces on shared Supervisor capacity need quota discipline and utilization monitoring (Operate pillar) to avoid noisy-neighbor effects.

Exam/Defense Relevance

Scenario items will describe a team needing 'a few containers, fast, with corporate registry and SSO' — the 9.1-correct answer is self-service namespaces, not a dedicated VKS cluster. In defense, this is your answer to 'how do you stop cluster sprawl?' — but expect the follow-up on isolation limits: namespace isolation is softer than cluster isolation; state where you draw the line (compliance workloads → dedicated clusters).

Key Takeaways

  • 9.1 CaaS: self-service namespace provisioning with registry, ingress, quotas, and identity inherited from VCF constructs
  • New middle tier in the consumption model: full VKS clusters vs namespaces vs VMs — match personas to tiers deliberately
  • Governance is inherited from platform level — one security review covers all namespaces
  • Namespace isolation < cluster isolation: keep hard-compliance workloads on dedicated clusters and say so in the defense

Native S3 Object Storage (Tech Preview)#

What Changed

VCF 9.1 introduces native S3-compatible object storage as a Tech Preview: developers self-serve buckets while IT keeps the guardrails, and object storage is deployed, scaled, and managed with the same VCF workflows used for block and file storage.

Why It Matters (Design/Ops Implications)

  1. Tech Preview status is the headline design fact: no production support, no SLA, subject to change. In any real design it may appear only as a labeled future consideration or an isolated pilot — never on the dependency path of a production requirement.
  2. Strategically it signals platform direction: modern app patterns (backups, AI/ML datasets, logs, static assets) assume an S3 endpoint; today that forces an external MinIO/array/vendor dependency that sits outside VCF lifecycle and identity. A native service would collapse that integration seam — worth a roadmap note in designs with heavy S3 consumption.
  3. Evaluation criteria to track toward GA: durability model, multi-tenancy/quota controls, encryption and key management, protocol conformance level, and how capacity is carved from vSAN.

Exam/Defense Relevance

The testable fact is its status: 9.1 items love a distractor that treats object storage as GA. In a VCDX-style defense, citing a Tech Preview feature as the solution to a requirement is an instant credibility hit — the defense-safe framing is 'requirement met today by <supported alternative>; native object storage tracked as a roadmap simplification.' Know the difference between GA (memory tiering, live patch on TPM, quick patch) and Tech Preview (native S3) in the 9.1 feature set cold.

Key Takeaways

  • Native S3 object storage in 9.1 is TECH PREVIEW — self-service for developers, IT guardrails, managed like block/file
  • Never place Tech Preview features on a production requirement's dependency path; frame as roadmap consideration
  • Track toward GA: durability, tenancy/quotas, encryption/KMS, S3 conformance, vSAN capacity model
  • Exam trap: distractors presenting 9.1 object storage as GA/supported

Deployment & Networking Delta: LACP in the UI, vSAN Data Services, Elastic Provisioning#

What Changed

A cluster of smaller but design-relevant 9.1 deltas:

  1. LACP in the workload-domain deployment UI: link aggregation (LACP) can now be configured directly in the workload-domain deployment workflow. In 9.0, LACP-based host networking required post-deployment vDS surgery or API-driven workarounds that fought the automated bring-up. Customers whose network standards mandate LACP to the leaf switches can now comply at Day-0 without breaking supportability of the automated path.
  2. vSphere Elastic (Zero Touch) Provisioning: ESX bootstraps onto bare metal via UEFI HTTP/S network boot — no TFTP server, supports Secure Boot and TPM; the target cluster's image and vSphere Configuration Profile determine what the host becomes; DHCP required unless UEFI has a static IP. Host-onboarding at scale stops being an imaging project.
  3. vSAN data services expansion: deduplication and compression broadened across cluster types and workload profiles, dedupe now supported WITH data-at-rest encryption, improved compression — reopens data-reduction assumptions in vSAN capacity models.
  4. DRS/vMotion refinements: non-disruptive vMotion evacuation option for maintenance mode (only evacuate when destination capacity avoids contention), rolling vMotion slot release (a new migration starts as each of the 8 concurrent slots frees, instead of batch-of-8 barriers), and encrypted vMotion offload to Intel QAT.

Why It Matters (Design/Ops Implications)

LACP-at-deploy removes a classic constraint-vs-automation conflict — previously you either bent the network standard or accepted manual deviation from the reference deployment. ZTP + VCP + vDS bootstrap makes 'server arrives → cluster member' a pipeline, which changes your host-replacement RTO and scale-out lead-time numbers. Dedupe+encryption coexistence removes an old either/or from storage design (compliance no longer costs you data reduction).

Exam/Defense Relevance

These are precision-distractor magnets: LACP is now IN the deployment UI (9.1) vs post-deploy workaround (9.0); dedupe WITH encryption is 9.1-new; maintenance-mode evacuation has TWO options (standard vs non-disruptive) and 'non-disruptive' refers to avoiding resource contention, not implying the standard method disrupts.

Key Takeaways

  • 9.1 deployment UI supports LACP for workload-domain host networking — Day-0 compliance with LACP-mandating network standards
  • Zero Touch Provisioning: UEFI HTTP/S boot (no TFTP), Secure Boot/TPM compatible, cluster VCP defines the host's config
  • vSAN 9.1: dedupe supported with data-at-rest encryption, broader dedupe/compression coverage — revisit capacity models
  • vMotion: rolling slot release (8 concurrent, no batch barrier), optional non-disruptive (contention-aware) evacuation, encrypted vMotion offload to Intel QAT

Holodeck 9.1: Lab Environment Delta#

The Holodeck toolkit tracking VCF 9.1 is a major release (9.0.1 and 9.0.2 were maintenance), and it resets several baseline assumptions. Docs are dated May 2026; the GA blog is dated 2026-07-01 — an unresolved conflict. Note that vmware/Holodeck publishes no GitHub releases, so vmware.github.io/Holodeck/9.1/ is the sole source of truth.

1. The resource envelope doubles. This is the change that actually constrains lab planning. VCF 9.1 single site needs 48 CPU / 650 GB / 2.2 TB against VCF 9.0's 32 / 325 GB / 1.1 TB — +50% CPU and exactly 2.0× on both RAM and disk. Dual site is 96 / 1536 GB / 5 TB. VVF 9.1 is gentler at 24 / 512 GB / 2 TB single, 48 / 1024 GB / 4 TB dual. The HoloRouter appliance itself grows to 6 vCPU / 12 GB / 75 GB. New in 9.1, soft validation lets pre-checks warn rather than block below-spec deployments — convenient on constrained hosts, but the fail-fast guardrail is gone and the architect owns the degradation. On a 1 TB host: VCF 9.1 single site fits, VVF 9.1 dual site sits exactly on the RAM line, and VCF 9.1 dual site is simply not achievable.

2. VNA: distributed vs centralized networking. The Virtual Network Appliance cluster is the headline architectural addition, supported in both management and workload domains. The clean mental model is VNA = distributed, NSX Edge cluster = centralized — Edge remains fully supported and is still the default. Day 0 exposes -VnaClusterMgmtDomain and -VnaClusterWkldDomain; Supervisor placement gains a mode parameter taking Centralized or Distributed. Caveat: the published cmd_reference Day 0 table shows no VNA switch and its -Version list omits 9.1.0.0 — that page is stale, so verify with Get-Help New-HoloDeckInstance on a live router.

3. GitOps is an OVA toggle, not a cmdlet. Enabling GitOps at HoloRouter OVA deployment — on by default — auto-provisions GitLab, pre-loads a repo with the deployment pipeline, and registers the router as a runner, so a full deployment can be triggered from the GitLab UI with no PowerShell. There are zero GitOps cmdlets. The consequence matters: GitOps cannot be retrofitted onto an existing router, so it is a decision made once, at deploy time.

4. HoloRouter service stack. Vault (:8200), Authentik SSO with SCIM (:9443), Technitium DNS (:5380) and an authenticated Webtop (:30000) now sit behind an HTTPS reverse proxy with an internal CA. Technitium becomes the DNS management surface and the HoloDeckDNSConfig cmdlets are deprecated — but the docs do not establish that dnsmasq is gone underneath, so 9.0-era DNS and shared-host cautions stay in force until re-baselined live. See the Holodeck section for the full evidence.

5. Six breaking changes. Mandatory -Site a|b on Import-HoloDeckConfig (this fails every existing 9.0.x invocation), dual config files from New-HoloDeckConfig, authenticated Webtop, deprecated DNS cmdlets, VCF Installer 9.0.0.0 unsupported, and HTTP endpoints moved to HTTPS. Every 9.0-era runbook needs a pass.

**6. No documented upgrade path.** There is no 9.0→9.1 migration guidance anywhere — treat 9.1 as a clean redeploy alongside, not an in-place upgrade. Related: 9.1 publishes **no Known Issues section at all**, which is a documentation gap rather than evidence of a clean release; source issues from GitHub Issues and the Broadcom community.

Why it matters for VCDX. The lab is the evidence factory. Dual-site support means DR runbooks and failover timings can be validated rather than asserted; Vault and Authentik make the 9.1 password-policy and federation stories demonstrable end-to-end; GitOps makes lab states reproducible artifacts you can cite. Most valuable of all, VNA (distributed) versus Edge (centralized) is a genuine design-decision axis — a real trade-off with justification, alternatives considered, and implications, which is exactly the shape of reasoning a panel probes. Holodeck itself is not exam content, but 'I failed site A over to site B and measured RTO' beats any datasheet citation. Track the nested ESX 8.0 U3 floor as a lab-refresh action item.

Key Takeaways

  • VCF 9.1 doubles the lab resource envelope: 48 CPU / 650 GB / 2.2 TB single site (2.0× RAM and disk vs 9.0); soft validation now warns instead of blocking, so you own the degradation
  • VNA = distributed networking, NSX Edge = centralized — supported in both management and workload domains, and a genuine VCDX design-decision axis
  • GitOps is an OVA deploy-time toggle (default on) that cannot be retrofitted — decide before you deploy the router
  • Six breaking changes; mandatory -Site on Import-HoloDeckConfig fails every existing 9.0.x call
  • No documented 9.0→9.1 upgrade path and no published Known Issues — plan a clean redeploy and source issues from GitHub/community

Exam Mapping: 2V0-13.25 — VCP-VCF 9.0 Architect

  • Delta awareness — current exam targets 9.0, but design-methodology items (AMPRS trade-offs, GA vs Tech Preview discipline, maintenance-window quantification) are exercised here against 9.1 features
  • Watch for exam refresh to 9.1: patching/LCM (quick patch, live patch on TPM), memory tiering, and fleet password policies are the most examinable deltas

Exam Mapping: 2V0-17.25 — VCP-VCF 9.0 Administrator

  • Operational deltas most likely to appear on a 9.1-refreshed exam: VCF Operations Build/Manage/Operate/Protect navigation, password policy scoping, quick patch vs RDU vs live patch, VKS 3.6 provisioning
📝 Quiz (40)
🃏 Flashcards (67)

📝 Quiz — VCF 9.0 → 9.1 Delta Assessment

0/40 correct

Section 1 — Platform, Management Services & API-First

Q1
What does the term 'VCF management services', introduced in VCF 9.1, refer to?
  • A rebranding of SDDC Manager for fleet-scale deployments
  • A common runtime and set of components that unifies the architecture of lifecycle and operational capabilities
  • A managed service offering where Broadcom operates the customer's management domain
  • The set of NSX Manager, vCenter, and VCF Operations appliances deployed in the management domain
VCF 9.1 introduces VCF management services as a common runtime plus shared components unifying how lifecycle and operational capabilities are architected — in 9.0 these ran as more loosely coupled components. Option A is wrong: SDDC Manager's functions were already absorbed into the 9.0 platform; this is a new runtime concept, not a rebrand. Option C is wrong: it is an architectural construct inside the product, not an operated/managed service offering. Option D describes the traditional management-domain appliance inventory — the management services runtime is the unifying layer for lifecycle/operational capabilities, not a name for the existing appliance set.
Q2
An architect claims the VCF management services runtime is 'not a single point of failure' for the platform. A review panel pushes back. Which statement is the most defensible correction?
  • The runtime is fully active-active across sites, so no failure analysis is required
  • Workloads, vCenter, and the NSX data plane continue operating during a runtime outage, but lifecycle and fleet operations stop — so its restore RTO must be designed and stated
  • Because it is a common runtime, a failure affects only the UI; all APIs remain available
  • The runtime cannot fail because it is protected by vSphere HA in the management domain
The defensible position distinguishes what survives (data plane: running VMs, vCenter, NSX forwarding, vSAN) from what stops (LCM, fleet monitoring/operations, automation intake) and quantifies restore time. Option A asserts a topology claim that isn't a given and dodges failure analysis — panels reject 'no analysis required' on principle. Option C is wrong: the runtime hosts services, not just UI; its APIs go down with it. Option D confuses a mitigation with immunity — vSphere HA restarts VMs after host failure but does not prevent runtime outages (patching, corruption, config error) and still implies downtime.
Q3
What is the practical significance of VCF 9.1's 'consistent OpenAPI' claim for automation design?
  • Terraform is now the only supported automation tool for VCF
  • Python, Java, PowerCLI, and Terraform tooling align to the same API contract, so tool selection follows team skills rather than per-tool feature gaps
  • The REST APIs were replaced by gRPC endpoints for performance
  • Automation no longer requires authentication because the API gateway handles identity
The 9.1 API-first story is one OpenAPI contract with aligned SDKs (Python/Java), PowerCLI, and Terraform — in 9.0, coverage differed per tool, so designs had to document gaps. Option A inverts the point: more tools are first-class, not fewer. Option C invents a protocol change that did not happen — the alignment is about consistent OpenAPI definitions. Option D is nonsense from a security standpoint: authentication remains mandatory; identity inheritance in 9.1 refers to constructs like CaaS namespaces, not unauthenticated APIs.
Q4
During vCenter maintenance in 9.1, what does the Envoy reverse proxy return to clients, enabling graceful automation behavior?
  • A 404 status indicating the endpoint does not exist
  • A 503 header indicating maintenance is in progress, with the estimated time of completion
  • A 301 redirect to the standby vCenter node
  • A 200 response with an empty payload until maintenance completes
9.1 adds a maintenance-notification capability: components can query whether vCenter maintenance is planned or underway, and during maintenance Envoy returns a 503 with estimated completion — automation can pause and retry rather than fail. Option A (404) would falsely signal a missing resource and give no completion hint. Option C is wrong: this mechanism is not an HA redirect, and vCenter HA failover is a different capability entirely. Option D would be actively harmful — a 200 with empty payload looks like success and would corrupt automation logic; the design goal is an explicit, machine-readable 'unavailable, back at T' signal.
Q5
The new vCenter 'High Session Count' alarm in 9.1 fires as sessions approach the default limit. What are the default session limit and the reporting detail included?
  • Limit 1024; reports the top 10 API endpoints by call volume
  • Limit 3000; reports the IPs/usernames of the top 5 users each holding more than 100 sessions
  • Limit 500; reports only the total session count
  • Limit 3000; reports the top 5 endpoints nearing their request limits
High Session Count fires near the 3000-session default and names the top 5 users (IPs and usernames, service accounts eligible) that each hold over 100 sessions — enough to identify the offending integration. Option A mixes in 1024, which is the per-endpoint active-request limit associated with the separate 'Increased Request Load' alarm. Option C invents a limit and understates the diagnostic detail. Option D correctly states 3000 but attaches the endpoint-level reporting that belongs to Increased Request Load — the two alarms are a deliberate distractor pair: sessions/users vs endpoints/requests.

Section 2 — Patching, Upgrades & vCenter Lifecycle

Section 3 — Memory Tiering & Platform Performance

Section 4 — Security, Password Policies & Compliance

Section 5 — VKS 3.6, Multi-vNIC & Container-as-a-Service

Section 6 — Storage, Deployment Networking & Holodeck 9.1

🃏 Flashcards — VCF 9.0 → 9.1 Delta

67 cards
Card 1 of 67
VCF Management Services
New in 9.1: a common runtime and shared set of components unifying the architecture of lifecycle and operational capabilities across the fleet. In 9.0 these were loosely coupled appliances with separate runtimes and update cadences. Design impact: one availability/sizing unit to protect, back up, and sequence in upgrades — and a smaller compatibility matrix.

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.