Academy/VCP-VCF 9.0 Architect (2V0-13.25)/Multi-Workload Domain Design
This lab targets VCF 9.0

Multi-Workload Domain Design

VCF 9.0Advancedvcp-foundation⏱ 105 min

VCF 9.0 workload domain architecture — domain isolation, vCenter allocation, lifecycle management, and cost optimization

Objectives

  • Design multi-workload domain architecture with proper isolation boundaries
  • Determine vCenter allocation strategy (federated SSO vs isolated)
  • Size and secure the management domain for VCF infrastructure services
  • Design storage, network, and security isolation between workload domains
  • Plan independent lifecycle management and upgrade sequencing across domains

Prerequisites

Access to VCF 9.0 architecture documentation and SDDC Manager guides

Prior labs: vcp-architect-01, vcp-architect-02, vcp-architect-03

Required skills:

  • VCF management and workload domain concepts
  • vCenter SSO architecture
  • VCF licensing model (per-core)
  • SDDC Manager workflows

Lab Environment

Design exercise — validate domain creation workflow on Holodeck VCF environment if available

Tasks

Task 1 Workload Domain Strategy & Isolation Design

VCDX panelists probe domain boundary decisions heavily. Each domain must be justified by a requirement — compliance isolation, lifecycle independence, or resource governance. Domains for the sake of domains is bad design.

Determine how many workload domains are needed and what isolation boundaries each domain provides.

Step 1

Document what a VCF workload domain provides:

A workload domain in VCF 9.0 is an isolated unit with:

  • Dedicated vCenter Server instance
  • Dedicated vSAN datastore (or external storage)
  • Dedicated NSX transport zone (or shared)
  • Independent cluster(s) within the domain
  • Independent lifecycle (upgrade independently of other domains)
  • SDDC Manager manages ALL domains from the management domain

Isolation boundaries:

LayerIsolatedShared
Compute (vCenter)Yes — separate vCenter per domainN/A
Storage (vSAN)Yes — separate vSAN datastoreOptional: HCI Mesh sharing
Network (NSX)Configurable — separate or shared transport zoneShared Tier-0 possible
Security (DFW)Configurable — independent or shared policiesEmergency category shared
Lifecycle (upgrades)Yes — independent upgrade scheduleManagement domain must upgrade first
Identity (SSO)Federated via Enhanced Linked ModeSeparate SSO domains possible
Step 2

Evaluate domain strategies for the 600-VM scenario:

Option A — Single Workload Domain (all 600 VMs):

  • Pros: Simplest operations, single vCenter, unified DRS, fewer licenses
  • Cons: No isolation between tiers, single upgrade window affects all VMs, compliance audit scope includes everything
  • Verdict: Acceptable for non-regulated environments

Option B — Two Workload Domains (Production + Dev/Test):

  • WD-1 (Production): Tier-1 + Tier-2 VMs (380 VMs, 13 hosts)
  • WD-2 (Dev/Test): Tier-3 VMs (220 VMs, 4 hosts minimum)
  • Pros: Production isolation from dev/test, independent upgrade schedules, separate vSAN datastores
  • Cons: Additional vCenter license, management overhead
  • Verdict: Recommended for SOC 2 environments

Option C — Three Workload Domains (Tier-1 + Tier-2 + Dev):

  • WD-1 (Tier-1): Mission-critical (80 VMs, 5 hosts)
  • WD-2 (Tier-2): Business apps (300 VMs, 8 hosts)
  • WD-3 (Dev/Test): Development (220 VMs, 4 hosts)
  • Pros: Maximum isolation, PCI scope limited to WD-1 only, independent overcommit policies
  • Cons: 3 vCenter instances, 3 vSAN datastores, highest licensing cost, operational complexity
  • Verdict: Required only for PCI-DSS or government classification environments

Design decision: Option B (Production + Dev/Test) — provides compliance isolation at reasonable cost.

Step 3

Design domain-to-cluster mapping:

Management Domain:

  • Cluster: mgmt-cluster-01 (4 hosts)
  • Contains: vCenter-mgmt, SDDC Manager, NSX Manager cluster (3 nodes), Aria Suite
  • Storage: vSAN ESA, FTT=1 RAID-5, ~15 TB usable
  • Network: Management NSX transport zone

Workload Domain 1 — Production:

  • Cluster A: prod-tier1-cluster (5 hosts) — Mission-critical VMs
  • Cluster B: prod-tier2-cluster (8 hosts) — Business application VMs
  • Contains: vCenter-prod (manages both clusters)
  • Storage: vSAN ESA, Tier-1 FTT=2 RAID-6, Tier-2 FTT=1 RAID-5
  • Network: Production NSX transport zone
  • Note: Multiple clusters in one workload domain share the same vCenter and vSAN datastore per cluster

Workload Domain 2 — Dev/Test:

  • Cluster: devtest-cluster (4 hosts) — All dev/test VMs
  • Contains: vCenter-devtest
  • Storage: vSAN ESA, FTT=1 RAID-5, Force Provisioning allowed
  • Network: Dev/Test NSX transport zone (isolated from production)

Total: 1 management domain + 2 workload domains = 3 vCenter instances, 21 hosts

Step 4

Design vCenter identity architecture:

VCF 9.0 SSO architecture:

  • Each vCenter has an embedded Platform Services Controller (PSC)
  • Enhanced Linked Mode (ELM): Connect vCenters for cross-domain visibility
    • All vCenters share the same SSO domain (e.g., vsphere.local)
    • Single-pane management: Log into any vCenter, see all inventories
    • RBAC: Cross-domain roles possible (e.g., read-only on production from dev vCenter)
  • Separate SSO domains: Complete identity isolation
    • Each vCenter has its own SSO domain
    • No cross-domain visibility
    • Use case: Multi-tenant environments, classified networks

Design decision: Enhanced Linked Mode for all three vCenters
Justification: Enables unified monitoring via Aria Operations, simplifies RBAC management, allows content library sharing across domains
Constraint: ELM requires network connectivity and time synchronization between all vCenters

RBAC design:

RoleManagement DomainProduction WDDev/Test WD
VCF AdminAdministratorRead-OnlyRead-Only
Prod AdminRead-OnlyAdministratorNo Access
Dev AdminNo AccessRead-OnlyAdministrator
SecurityRead-OnlyRead-Only (VM Console)Read-Only
AuditorRead-OnlyRead-OnlyRead-Only

Validation Gate

Check: Workload domain strategy selected with cluster mapping, isolation analysis, and identity design

Expected: Domain count justified with alternatives, cluster-to-domain mapping documented, vCenter SSO architecture designed, RBAC matrix defined

Common Errors

Creating workload domains without a requirements-driven justification — each additional domain adds vCenter licensing cost and operational overhead
Forgetting that management domain must upgrade BEFORE any workload domain — this creates an upgrade dependency chain
Not considering Enhanced Linked Mode limitations — ELM has a maximum of 15 vCenters in VCF 9.0
Mixing compliance-scoped and non-compliant VMs in the same workload domain — defeats the isolation purpose

Task 2 Management Domain Sizing & Protection

The management domain is the most critical domain in VCF. If it fails, no other domain can be managed. VCDX panelists challenge candidates on management domain sizing, HA, and recovery procedures.

Size the management domain correctly for all VCF infrastructure services and design its backup/recovery strategy.

Step 1

Size management domain components:

ComponentvCPURAM (GB)Storage (GB)HANotes
SDDC Manager416500No HA (single VM)File-level backup critical
vCenter-mgmt420290vCenter HA (active-passive)Embedded PSC
NSX Manager 16243003-node clusterAnti-affinity with other NSX nodes
NSX Manager 2624300""
NSX Manager 3624300""
NSX Edge 1832200Active-Standby pairLarge form factor
NSX Edge 2832200""
Aria Operations416500Cluster (optional)Analytics node
Aria Ops for Logs4161,000Optional HALog retention sizing
Aria Automation416100Cluster (optional)If private cloud portal needed
Total~54~220 GB~3.7 TB

Host requirement:

  • 4 hosts × 64 cores × 512 GB RAM = 256 cores, 2,048 GB RAM
  • After HA (25% reserve): 192 cores, 1,536 GB usable
- Management workload: 54 vCPU, 220 GB RAM → well within capacity
  • Headroom: Significant — allows for additional monitoring, backup agents, jump hosts
Step 2

Design management domain backup strategy:

  1. SDDC Manager backup:
  • Type: File-level backup via built-in backup mechanism
  • Schedule: Daily at 02:00
  • Target: External SFTP server (NOT on management domain vSAN)
  • Retention: 14 days (rolling)
  • Content: Configuration database, credentials vault, domain inventory
  • Recovery: Restore from backup to fresh SDDC Manager VM; re-register components
  1. vCenter backup:
   - Type: File-level backup (VAMI → Backup)
  • Schedule: Daily at 01:00
  • Target: Same external SFTP server
  • Content: vCenter configuration, inventory, statistics, events
  • Recovery time: 15-30 minutes from backup restore
  1. NSX Manager backup:
  • Type: NSX backup via NSX Manager UI or API
  • Schedule: Daily at 03:00
  • Target: External SFTP server
  • Content: NSX configuration, DFW rules, segments, gateway configs
  • Recovery: Restore to fresh NSX Manager cluster; verify DFW rules intact
  1. Aria Suite:
  • Lifecycle Manager (LCM): Manages Aria component deployment and updates
  • Backup via LCM snapshots or file-level backup
  • Lower priority: Aria can be redeployed from LCM if data loss is acceptable

Critical: ALL backups must be stored OUTSIDE the management domain — if management domain fails, backups must be accessible independently.

Step 3

Design management domain recovery procedure:

Recovery priority order (sequential — each depends on previous):

1. vCenter-mgmt → Provides ESXi management plane
2. NSX Manager → Restores networking and DFW
3. SDDC Manager → Restores VCF lifecycle management
4. Aria Suite → Restores monitoring (lower priority)

Recovery scenarios:

Scenario A — Single management VM failure:

  • HA restarts VM on another host
  • RTO: 2-5 minutes (HA restart)
  • Action: None required; automatic recovery

Scenario B — Management domain host failure:

  • HA restarts affected VMs; anti-affinity rules redistribute NSX Manager nodes
  • RTO: 5-10 minutes
  • Verify: NSX Manager cluster quorum maintained (2 of 3 nodes up)

Scenario C — Complete management domain failure (disaster):

  • Workload VMs continue running (data plane independent of management)
  • Recovery: Rebuild management domain from backups
  1. Deploy 4 ESXi hosts from ISO
  2. Restore vCenter-mgmt from file-level backup (15-30 min)
  3. Restore NSX Manager from backup (30-60 min)
  4. Restore SDDC Manager from backup (30-60 min)
  • Total RTO: 2-4 hours for management plane restoration
  • Workload impact: VMs run throughout; no management changes possible during recovery
Step 4

Design management domain security:

  1. Network isolation:
  • Management VLAN: Isolated from workload networks
  • Firewall: Only admin jump hosts can reach management interfaces
  • NSX Edge: Separate management Edge cluster (not shared with workload)
  1. Access control:
  • SDDC Manager: Local admin + AD integration; MFA required
  • vCenter: AD/LDAPS integration; SSO admin password in sealed envelope
  • NSX Manager: Local admin + AD integration; API access restricted to automation service accounts
  • Audit: All management actions logged to Aria Operations for Logs
  1. Certificate management:
  • VMCA as subordinate CA to enterprise PKI
  • Management VM certificates: Auto-renewed by VMCA
  • External-facing certificates: Signed by enterprise CA (vCenter UI, NSX Manager UI)
  • Rotation: Annual certificate renewal; automated via SDDC Manager
  1. Physical security:
  • Management hosts in locked racks
  • IPMI/iLO interfaces on dedicated management VLAN (not routable from workload)
  • Console access: KVM-over-IP with audit logging

Validation Gate

Check: Management domain sized, backup strategy defined, recovery procedures documented, security controls in place

Expected: Component sizing table, daily backup schedule to external target, recovery procedure with RTO per scenario, security controls covering network/access/certificate/physical layers

Common Errors

Storing management domain backups ON the management domain vSAN — if the domain fails, backups are lost
Forgetting that SDDC Manager has no native HA — it is a single VM; backup and recovery procedure is the only protection
Not documenting the recovery order — restoring NSX before vCenter wastes time because NSX needs vCenter for registration
Undersizing management domain — 3 hosts is the absolute minimum but leaves no room for maintenance + failure; 4 hosts recommended

Task 3 Cross-Domain Network & Storage Design

VCDX panelists challenge whether domains are truly isolated or have hidden dependencies. The design must show clear boundaries AND controlled cross-domain pathways.

Design the network and storage isolation between workload domains while enabling necessary cross-domain communication.

Step 1

Design cross-domain network architecture:

  1. NSX transport zone strategy:

Option A — Shared transport zone across workload domains:

  • Single overlay TZ shared by WD-1 and WD-2
  • Pro: Simpler segment management, cross-domain routing possible
  • Con: Blast radius — NSX issue affects both domains
  • Use when: Domains need frequent east-west communication

Option B — Separate transport zones per domain:

  • WD-1: prod-overlay-tz
  • WD-2: devtest-overlay-tz
  • Pro: Complete network isolation, independent NSX lifecycle
  • Con: Cross-domain traffic must route through physical network (north-south hairpin)
  • Use when: Compliance requires network isolation

Design decision: Separate transport zones
Justification: SOC 2 audit scope is limited to WD-1; separate TZ ensures dev/test traffic never touches production overlay

  1. Cross-domain communication (when needed):
   - Route through physical firewall: WD-1 Tier-0 → Physical FW → WD-2 Tier-0
   - Inspect and log at physical FW (compliance evidence)
   - Permitted flows: Monitoring (Aria → both domains), backup, DNS, NTP
   - Blocked flows: Any dev/test → production workload traffic
Step 2

Design cross-domain storage architecture:

  1. vSAN datastore per domain (VCF default):
  • WD-1: prod-vsan-datastore (managed by vCenter-prod)
  • WD-2: devtest-vsan-datastore (managed by vCenter-devtest)
  • Complete storage isolation — VMs cannot access cross-domain datastores
  1. Shared services storage:
  • ISO library: NFS datastore accessible from all domains (read-only)
  • Template library: vSphere Content Library with publisher-subscriber model
    • Publisher: Management domain Content Library
    • Subscribers: WD-1 and WD-2 sync templates from publisher
    • Benefit: Golden images managed centrally, deployed locally
  1. Backup target:
  • Shared NFS/SMB target for backups from both domains
  • Accessible via management network (not workload network)
  • Separate backup jobs per domain for compliance audit trail
  1. HCI Mesh (capacity sharing):
  • If WD-2 has excess storage: WD-1 can mount remote datastore via HCI Mesh
  • Use case: Temporary overflow during capacity procurement
  • Constraint: Remote datastore performance is lower (~1-2ms additional latency)
  • Production policy: Do NOT run Tier-1 VMs on remote HCI Mesh datastores
Step 3

Design content library and template sharing:

  1. Content Library architecture:
  • Master library: mgmt-content-library (on management domain vCenter)
  • Published items: OS templates (Windows 2022, RHEL 9, Ubuntu 22.04), hardened images
  • Subscriber libraries:
    • prod-content-library (WD-1 vCenter) — syncs from master
    • devtest-content-library (WD-2 vCenter) — syncs from master
  • Sync policy: On-demand (download template when first used in target domain)
  1. Template hardening pipeline:
  • Create base template in management domain
  • Apply CIS benchmark hardening (automated via Packer or Ansible)
  • Validate with Aria Operations compliance scan
  • Publish to Content Library
  • Workload domains consume hardened template
  • Update cadence: Monthly (patch Tuesday + 7 days testing)
  1. VM customization specs:
  • Per-domain customization: Different AD OUs, DNS servers, NTP sources
  • WD-1: Prod AD OU, production DNS (172.16.200.10)
  • WD-2: Dev AD OU, dev DNS (172.16.200.20)
  • Prevents dev VMs from joining production AD accidentally
Step 4

Create design decision D-010 — Workload Domain Architecture:

Decision: 2 workload domains (Production + Dev/Test) with separate transport zones and vSAN datastores

Alternatives:
A) Single workload domain — Rejected: SOC 2 audit scope includes dev/test; compliance cost and risk increase
B) 3 workload domains (Tier-1 + Tier-2 + Dev) — Rejected: Tier-1 and Tier-2 managed by same ops team; additional vCenter license not justified without PCI requirement
C) 2 domains with shared transport zone — Rejected: NSX blast radius spans both domains; separate TZ provides defense-in-depth

Cost analysis:

Item1 Domain2 Domains3 Domains
vCenter instances123
vCenter license cost$X$2X$3X
Operational complexityLowMediumHigh
Compliance audit scopeAll VMsProduction onlyTier-1 only
Upgrade independenceNoneFullFull

Net: 2 domains reduces compliance audit scope by ~37% (220 dev VMs excluded) while adding only 1 additional vCenter license.

Validation Gate

Check: Cross-domain design covers network isolation, storage boundaries, content sharing, and design decision with cost analysis

Expected: Transport zone strategy justified, storage isolation with shared services designed, Content Library architecture documented, design decision with 3+ alternatives and cost comparison

Common Errors

Sharing NSX transport zones between compliance-scoped and non-compliance domains — auditors will include the entire TZ in scope
Not having a Content Library strategy — without it, templates are managed independently per domain, leading to drift
Forgetting that HCI Mesh adds latency — never use remote datastores for latency-sensitive Tier-1 workloads
Not accounting for the additional vCenter license cost when proposing multiple domains — cost must be justified by compliance or operational benefit

Task 4 Lifecycle Management & Upgrade Sequencing

VCF upgrades are sequential and dependency-ordered. VCDX panelists test whether candidates know the official 19-step upgrade sequence (KB 390634) and can design maintenance windows accordingly.

Design the upgrade and patching strategy for a multi-domain VCF environment with proper sequencing and rollback procedures.

Step 1

Document VCF 9.0 upgrade sequence (per Broadcom KB 390634):

The official upgrade order within each domain:

  1. SDDC Manager
  2. VCF Operations Suite Lifecycle Manager
  3. VCF Operations
  4. VCF Operations for Logs
  5. VCF Automation
  6. VCF Automation Orchestrator
  7. VCF Operations Suite (remaining)
  8. vCenter Server (Management Domain)
  9. NSX Manager (Management Domain)
  10. ESXi hosts (Management Domain) — rolling, one host at a time
  11. vSAN (Management Domain) — disk format upgrade if applicable
  12. vCenter Server (Workload Domain 1)
  13. NSX Manager (Workload Domain 1)
  14. ESXi hosts (Workload Domain 1) — rolling
  15. vSAN (Workload Domain 1)
  16. vCenter Server (Workload Domain 2)
  17. NSX Manager (Workload Domain 2)
  18. ESXi hosts (Workload Domain 2) — rolling
  19. vSAN (Workload Domain 2)

Critical rule: Management domain MUST upgrade before ANY workload domain.

Step 2

Design maintenance window schedule:

Assumptions:

  • Host rolling upgrade: ~45 minutes per host (evacuate VMs, install, reboot, verify)
  • vCenter upgrade: ~30 minutes (VAMI upgrade + verification)
  • NSX Manager upgrade: ~60 minutes (3-node rolling upgrade)
  • Maximum maintenance window: 8 hours

Upgrade plan:

Week 1 — Management Domain:

  • Friday 22:00 - Saturday 06:00 (8-hour window)
- Sequence: SDDC Manager → Aria Suite → vCenter → NSX → ESXi (4 hosts × 45 min = 3 hrs) → vSAN
  • Estimated duration: 6.5 hours
  • Rollback: Snapshot before each component upgrade; revert if verification fails

Week 2 — Workload Domain 1 (Production):

  • Saturday 22:00 - Sunday 10:00 (12-hour window, extended for production)
- Sequence: vCenter → NSX → ESXi (13 hosts × 45 min = 9.75 hrs) → vSAN
  • Estimated duration: 11.5 hours
  • Risk: Long window for 13 hosts; consider splitting into 2 weekends if allowed

Week 3 — Workload Domain 2 (Dev/Test):

  • Wednesday 22:00 - Thursday 04:00 (6-hour window)
- Sequence: vCenter → NSX → ESXi (4 hosts × 45 min = 3 hrs) → vSAN
  • Estimated duration: 4.5 hours
  • Lower risk: Dev/test can tolerate business-hours maintenance if needed
Step 3

Design pre-upgrade validation checklist:

  1. Health checks:
  • [ ] vSAN Health: All green (no degraded components, no resync in progress)
  • [ ] NSX Health: All 3 Manager nodes healthy, DFW rules applied
  • [ ] vCenter Health: VAMI health status green, disk space >30% free
  • [ ] SDDC Manager: Bundle download complete, compatibility verified
  • [ ] Cluster health: No hosts in maintenance mode, HA admission control satisfied
  1. Backup verification:
  • [ ] SDDC Manager backup: Taken <4 hours before upgrade start
  • [ ] vCenter backup: File-level backup verified
  • [ ] NSX backup: Configuration backup exported and verified
  • [ ] VM snapshots: Management VMs snapshotted (delete after upgrade success)
  1. Compatibility:
  • [ ] VCF compatibility matrix: Target version compatible with all components
  • [ ] HCL check: ESXi target version supports all host hardware (CPU, NIC, HBA)
  • [ ] Third-party compatibility: Backup agents, monitoring agents verified
  • [ ] License entitlement: Updated licenses uploaded to SDDC Manager
  1. Communication:
  • [ ] Change management approval received
  • [ ] Stakeholders notified (24-hour advance notice)
  • [ ] On-call team briefed on rollback procedure
  • [ ] VMware Support SR pre-opened (Severity 2, reference to upgrade)
Step 4

Design rollback and failure handling:

  1. Rollback strategy per component:
  • SDDC Manager: Restore from file-level backup (30 min)
  • vCenter: Restore from file-level backup or revert VM snapshot (15-30 min)
  • NSX Manager: Restore from backup (30-60 min)
  • ESXi host: Roll back failed host only; boot from previous ESXi version if available
  • vSAN disk format: CANNOT be rolled back (one-way upgrade); ensure all hosts upgraded before triggering
  1. Failure decision tree:
   - Component upgrade fails → Revert to snapshot → Investigate → Retry or escalate
   - Verification fails post-upgrade → Assess impact → Rollback if data plane affected
   - Host rolling upgrade fails on 1 host → Skip host, continue remaining → Return to failed host last
   - vSAN health degrades during upgrade → STOP → Wait for resync → Resume when healthy
  1. Post-upgrade verification:
  • [ ] All management services accessible (vCenter, NSX, SDDC Manager)
  • [ ] DFW rules intact and enforcing (Traceflow test)
  • [ ] vSAN health all green
  • [ ] Workload VMs running and accessible
  • [ ] Monitoring operational (Aria dashboards showing data)
  • [ ] Backup jobs running successfully
  1. Post-upgrade cleanup:
  • Delete VM snapshots (within 24 hours — snapshots grow and degrade performance)
  • Confirm file-level backups taken with new version
  • Update asset inventory with new component versions
  • Close change management record

Validation Gate

Check: Lifecycle management design covers upgrade sequence, maintenance windows, pre-checks, and rollback procedures

Expected: 19-step upgrade sequence documented, maintenance windows scheduled with duration estimates, pre-upgrade checklist with 15+ items, rollback strategy per component, post-upgrade verification

Common Errors

Attempting to upgrade workload domain before management domain — VCF enforces this dependency; upgrade will fail
Triggering vSAN disk format upgrade before ALL hosts in the cluster are upgraded — mixed-version cluster with new disk format causes data accessibility issues
Not pre-opening a VMware Support SR — if the upgrade fails at 3 AM, waiting for SR creation wastes critical maintenance window time
Leaving VM snapshots after upgrade — snapshots grow rapidly and degrade vSAN performance; delete within 24 hours

Final Validation

Complete multi-workload domain design with isolation, management domain protection, cross-domain architecture, and lifecycle management

✓ Workload domain strategy justified with cost analysis → Domain count selected with alternatives evaluated, cluster mapping documented

✓ Management domain sized and protected → Component sizing table, backup strategy, recovery procedures with RTO per scenario

✓ Cross-domain isolation verified → Separate transport zones, storage isolation, controlled cross-domain pathways

✓ Lifecycle management complete → 19-step sequence, maintenance windows, pre-checks, rollback, post-upgrade verification

Cleanup / Restore

• Save all domain architecture documentation and upgrade procedures

• If using Holodeck: Remove test workload domains and revert to management-only configuration

Design Reflection (VCDX)

Multi-workload domain design tests architectural judgment. Too few domains = insufficient isolation. Too many = unnecessary cost and complexity. VCDX panelists look for: (1) Each domain boundary traces to a requirement (compliance, lifecycle, governance). (2) Management domain is treated as the most critical asset — backup/recovery procedure is thorough. (3) Upgrade sequencing follows the official 19-step order. (4) Cost analysis shows the trade-off of adding domains vs. value of isolation. The right answer is rarely the most complex one.

Requirements

  • SOC 2 compliance scope limited to production workloads
  • Independent upgrade schedules for production and dev/test
  • Centralized template management across all domains
  • Management domain RTO <4 hours for full recovery

Constraints

  • Each workload domain requires a dedicated vCenter (licensing cost)
  • Management domain must upgrade before workload domains (VCF dependency)
  • SDDC Manager has no native HA — single VM, backup-dependent recovery
  • VCF 9.0 supports maximum 15 vCenters in Enhanced Linked Mode

Assumptions

  • 2 workload domains sufficient (no PCI-DSS requirement for Tier-1 isolation)
  • 8-hour maintenance window available on weekends for upgrades
  • External SFTP server available for management backups
  • Single operations team manages all domains (no multi-tenancy requirement)

Risks

  • Management domain failure impacts visibility and control of all workload domains
  • SDDC Manager backup corruption discovered only during recovery attempt
  • Upgrade fails mid-sequence leaving mixed-version environment
  • vSAN disk format upgrade is irreversible — if hosts need rollback, disk format is incompatible

Self-Assessment Discussion Prompts

  1. If the customer later requires PCI-DSS isolation, how would you add a third workload domain without downtime?
  2. What if the customer wants to run Kubernetes (TKG) — does it go in an existing domain or a new one?
  3. How would you handle a scenario where the SDDC Manager backup is corrupted and unrecoverable?
  4. What is the impact on upgrade sequencing if two workload domains share the same NSX transport zone?

Extensions

VCF Multi-Domain with Tanzu Kubernetes Grid

Design a dedicated workload domain for TKG workloads. Address the vSphere Supervisor cluster configuration, harbor registry placement, storage class mapping to vSAN policies, and NSX networking for Kubernetes services. Document how TKG domain lifecycle differs from traditional VM domains.

VCF Multi-Tenancy with Workload Domains

Design a multi-tenant VCF environment where each tenant gets an isolated workload domain with separate SSO, billing, and SLA tracking. Address shared management domain governance, tenant onboarding automation via SDDC Manager API, and cross-tenant network isolation.

VCF Upgrade Automation with SDDC Manager API

Automate the 19-step upgrade sequence using SDDC Manager REST API. Create a script that orchestrates pre-checks, triggers upgrades per component, monitors progress, and sends notifications. Include automatic rollback triggers based on health check failures.

⚠ Known Pitfalls (from Community KB)

Creating workload domains purely for organizational separation without compliance or lifecycle requirements — each additional domain adds a vCenter license and operational overhead without proportional benefit
Storing management domain backups on the management domain vSAN — a complete domain failure makes backups inaccessible; always use an external backup target
Attempting to upgrade workload domains before the management domain — VCF enforces this dependency chain; trying to skip it will fail and may leave the environment in an inconsistent state
Sharing NSX transport zones between compliance-scoped and non-compliance workload domains — auditors consider the entire transport zone in scope, negating the isolation benefit of separate domains

References

  • VMware VCF 9.0 Administration Guide — Workload Domain Management
  • Broadcom KB 390634 — VCF Upgrade Sequence and Component Order
  • VMware VCF 9.0 Planning and Preparation Guide — Management Domain Sizing
  • vCenter Server 9.0 — Enhanced Linked Mode architecture
  • VMware VCF SDDC Manager API Reference — Domain and Cluster Operations
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.