Academy/VCP-VCF 9.0 Administrator (2V0-17.25)/Multi-Instance Fleet Operations (VCF 9.0)
This lab targets VCF 9.0

Multi-Instance Fleet Operations (VCF 9.0)

VCF 9.0Advancedvcp-foundationarchitectvcdx⏱ 180 min

VCF 9.0 introduces Fleet Manager as the new operational paradigm for multi-instance management. This lab focuses on fleet-level workflows: instance registration, policy propagation, cross-instance reporting, and fleet-wide failover scenarios. Contrast with VCF 5.2 where instances were essentially independent and required manual coordination.

Objectives

  • Understand Fleet Manager architecture and its role as the unified control plane for multiple VCF instances
  • Deploy a multi-instance fleet: register and inventory-sync two or more VCF instances under a single Fleet Manager
  • Configure and propagate fleet-level policies (certificates, password policies, DFW rules, patching schedules) across instances
  • Execute cross-instance policy rollout and validate per-instance compliance and application
  • Conduct fleet-wide reporting and compliance monitoring using Fleet Manager observability
  • Design and test fleet failover scenarios: instance failure, Fleet Manager failure, and multi-instance recovery
  • Articulate operational differences between VCF 5.2 (instance-centric) and VCF 9.0 (fleet-centric) architectures

Prerequisites

Two or more independently deployed VCF 9.0 instances in separate sites (or separate clusters), each with a complete management domain and at least one workload domain. Fleet Manager deployed and accessible. All instances and Fleet Manager have network connectivity and DNS resolution. Minimum 4 ESXi hosts per instance, 512 GB RAM, sufficient storage for multi-instance lab.

Prior labs: holodeck-02, vcp-admin-09

Required skills:

  • VCF 9.0 instance deployment and management domain configuration
  • Fleet Manager lifecycle management and policy administration
  • NSX distributed firewall (DFW) rules and cross-instance propagation
  • VCF Operations observability and compliance reporting
  • Understanding of federation and fleet-level orchestration concepts
  • SSH CLI access to appliances for validation
  • vSphere Client and NSX Manager UI navigation

Lab Environment

Multi-site VCF 9.0 fleet: (1) Fleet Manager VM (single central point) managing 2+ instances, (2) Instance A (Primary Site) with management domain + 2 workload domains, (3) Instance B (Secondary Site) with management domain + 2 workload domains, (4) NSX Manager deployed within each instance, (5) vCenter Server per instance, (6) VCF Operations monitoring all instances. Cross-instance routing configured via Tier-0 gateway peering or physical network routing to enable east-west traffic between instances.

graph TB
  FM[Fleet Manager<br/>10.0.0.50]
  VCFOPS[VCF Operations<br/>10.0.0.51]
  INST1[Instance A<br/>Primary Site<br/>10.1.0.0/16]
  INST2[Instance B<br/>Secondary Site<br/>10.2.0.0/16]
  VC1[vCenter A<br/>10.1.0.6]
  NSX1[NSX Manager A<br/>10.1.0.10]
  WLD1A[Workload Domain A1<br/>3 hosts]
  WLD1B[Workload Domain A2<br/>3 hosts]
  VC2[vCenter B<br/>10.2.0.6]
  NSX2[NSX Manager B<br/>10.2.0.10]
  WLD2A[Workload Domain B1<br/>3 hosts]
  WLD2B[Workload Domain B2<br/>3 hosts]
  INST1 --> VC1
  INST1 --> NSX1
  VC1 --> WLD1A
  VC1 --> WLD1B
  NSX1 --> WLD1A
  NSX1 --> WLD1B
  INST2 --> VC2
  INST2 --> NSX2
  VC2 --> WLD2A
  VC2 --> WLD2B
  NSX2 --> WLD2A
  NSX2 --> WLD2B
  FM --> INST1
  FM --> INST2
  FM --> VCFOPS
  VCFOPS --> INST1
  VCFOPS --> INST2
  WLD1A -.cross-instance-routing.-> WLD2A

IP Addressing

NetworkPurposeVLAN
10.0.0.0/24Fleet management network (Fleet Manager, VCF Operations, management appliances)VLAN 1644
10.1.0.0/16Instance A (management + workload)VLAN 1650-1659
10.2.0.0/16Instance B (management + workload)VLAN 1660-1669
192.168.1.0/24Out-of-band management (backup, NTP, syslog)VLAN 100

Credentials

SystemUsernamePassword
Fleet ManageradminFleet Manager admin password
vCenter Instance Aadministrator@vsphere.localInstance A SSO password
vCenter Instance Badministrator@vsphere.localInstance B SSO password
NSX Manager Instance AadminInstance A NSX admin password
NSX Manager Instance BadminInstance B NSX admin password
VCF Operationsadministrator@vsphere.localVCF Operations SSO password

Tasks

Task 1 Map fleet topology and verify instance autonomy and federation

Manageability

Fleet Manager is a new operational paradigm in VCF 9.0. Unlike SDDC Manager (VCF 5.2) which was mandatory for all management, Fleet Manager in 9.0 is an orchestrator of autonomous instances. This task clarifies the architecture and proves instance independence.

Step 1

Log into Fleet Manager UI (https://<fleet-manager-ip>). Navigate to Inventory > Managed Instances. Document each instance: Name, Type (Management or Workload), FQDN, Status, vCenter IP, NSX Manager IP, Management Domain hosts, Workload Domains count.

Spreadsheet with all instances and their component IPs
Screenshot the Fleet Manager inventory page. This is your fleet topology artifact for VCDX documentation.
Step 2

Verify instance autonomy. In Instance A, log into vCenter directly (bypass Fleet Manager). Navigate to Administration > System Configuration. Verify: (1) vCenter is fully functional without Fleet Manager connectivity, (2) All workload domains are visible and healthy, (3) VMs can be provisioned and managed independently.

vCenter Instance A operates independently. All management functions available without Fleet Manager.
In VCF 5.2, SDDC Manager was mandatory for many day-2 operations (patching, certificates, cluster expansion). In VCF 9.0, each instance is autonomous. This is a key architectural shift. Document this as a VCDX-critical design principle: 'Federation, not subordination.'
Step 3

Temporarily isolate Instance B from Fleet Manager. On the physical network, block traffic from Instance B management domain to Fleet Manager IP (10.0.0.50). This can be done via firewall rule or by disconnecting a physical network cable from the Instance B management switch (if lab topology permits). Verify: Instance B continues to operate autonomously — vCenter and NSX remain functional, workload VMs continue running.

Instance B operates without Fleet Manager connectivity. No workload disruption.
This demonstrates instance autonomy — a key difference from VCF 5.2 where SDDC Manager loss would impact operations.
Step 4

Restore Instance B connectivity to Fleet Manager. Return network connectivity. In Fleet Manager UI, navigate to Inventory > Managed Instances. Verify Instance B status changes from 'Unreachable' to 'Connected' within 2-5 minutes.

Instance B reconnects and resumes reporting to Fleet Manager
The reconnection delay demonstrates Fleet Manager's eventual consistency model — instances operate independently but reconcile state when connectivity is restored.
Step 5

Create a fleet topology diagram. Map: (1) Fleet Manager at the center, (2) Instances A and B as peer nodes, (3) Each instance's management domain and workload domains, (4) Cross-instance networking (if configured). Use Mermaid graph syntax or draw in Visio/Lucidchart. Include trust boundaries and control plane vs data plane separation.

Topology diagram showing multi-instance fleet architecture
Contrast this to VCF 5.2 where SDDC Manager was the central hub and instances were subordinates. In VCF 9.0, instances are peers coordinated by Fleet Manager for policy and lifecycle, but each remains independently viable.

Validation Gate

Check: Instance spreadsheet complete, instance autonomy verified (vCenter/NSX work standalone), isolation/reconnection test successful, topology diagram created

Expected: You understand the fleet model: autonomous instances + federation layer (Fleet Manager) coordinating policies and lifecycle

Common Errors

Instance shows 'status: unreachable' in Fleet Manager but is actually operational
Cause: Network connectivity broken but instance vCenter/NSX services are still running
Fix: Verify network: ping instance management IP from Fleet Manager. Check firewall rules allow port 443 bidirectional. Instance is still operational; reconnection will succeed once network is restored.
vCenter becomes inaccessible when Fleet Manager is offline
Cause: Misconfiguration — vCenter is pointing to Fleet Manager for some critical service (NTP, DNS, licensing) that is not truly required
Fix: Verify vCenter NTP and DNS settings point to infrastructure NTP/DNS servers, not Fleet Manager. Verify licensing is not contingent on Fleet Manager connectivity.
Cannot articulate the difference between VCF 5.2 (SDDC Manager mandatory) and VCF 9.0 (Fleet Manager optional for operations)
Cause: Insufficient understanding of the architectural evolution
Fix: Read Broadcom's VCF 9.0 architecture overview. Key insight: SDDC Manager was removed because it became a bottleneck for multi-instance management. Fleet Manager addresses this by being an orchestration layer, not a operational dependency.

Task 2 Register instances with Fleet Manager and establish inventory sync

Manageability

Instance registration is the first step in fleet management. This task covers the workflow of bringing instances under Fleet Manager control and validating inventory consistency.

Step 1

In Fleet Manager UI, navigate to Lifecycle > Add Instance. Click 'Register New Instance'. Enter Instance B details: FQDN of Instance B management domain vCenter, admin credentials, select instance type (Workload Domain), confirm SSL certificate.

Instance registration wizard initiates discovery of Instance B infrastructure
Instance must already be fully deployed and operational before registration. If registration fails with 'Connection refused', verify: (1) vCenter FQDN resolves, (2) Port 443 is open from Fleet Manager to vCenter, (3) Credentials are correct.
Step 2

Complete instance registration. Fleet Manager discovers: management domain, vCenter instance, NSX Manager, workload domains, ESXi hosts, storage, networking. This may take 5-15 minutes. Monitor progress via Fleet Manager UI.

Instance B fully discovered. Fleet Manager displays: 4 workload domains, 12 ESXi hosts, 15 datastores, NSX segments, DFW rules
Screenshot the inventory discovery page. This is the point where Fleet Manager has assumed management of Instance B.
Step 3

Validate inventory consistency. In Fleet Manager UI, compare Instance A and B inventories side-by-side. Verify: (1) Both instances have same number of workload domains (should be 2 each if lab is symmetric), (2) Host count per domain is similar, (3) Storage and network configurations are visible for both.

Inventory displays for both instances with consistent structure
If inventories are asymmetric (e.g., Instance A has 2 workload domains, Instance B has 3), document this as a fleet imbalance. VCDX panelists may ask: How would you manage capacity or failover with asymmetric instances?
Step 4

Enable continuous inventory sync. In Fleet Manager > Settings > Inventory Sync, set sync frequency to 'Every 15 minutes' (or 'Continuous' if available). This ensures Fleet Manager inventory is fresh and detects changes quickly.

Inventory sync configured and active. Next sync time is shown.
Continuous inventory sync is important for compliance monitoring and policy application — if a workload domain is added and sync is manual, the new domain may not inherit fleet policies until next sync.
Step 5

Validate cross-instance visibility. In Fleet Manager > Analytics/Reporting, check dashboard: total VMs count should be sum of Instance A + Instance B VMs. Total hosts should be sum of both instances. This validates that inventory sync is working and data is aggregated correctly.

Analytics dashboard shows aggregated metrics across both instances
This cross-instance observability is a new capability in VCF 9.0. In VCF 5.2, you would need separate SDDC Manager instances or external tools to correlate metrics across instances.

Validation Gate

Check: Instance B registered with Fleet Manager, inventory fully discovered, inventory sync enabled and active, analytics dashboard shows aggregated data for both instances

Expected: Both instances are now under Fleet Manager management with real-time inventory visibility

Common Errors

Instance registration fails: 'Invalid credentials' or 'Authentication failed'
Cause: vCenter credentials are incorrect or user account lacks required permissions
Fix: In Instance B vCenter, verify the admin user account is enabled and has Administrator role on vCenter. Test login: https://<instance-b-vcenter> with the same credentials. Re-run registration with confirmed credentials.
Inventory sync shows 0 hosts or incomplete discovery
Cause: NSX Manager or vCenter licensing not activated, or certain components not responding
Fix: Verify in Instance B vCenter: all ESXi hosts are connected and licensed. In NSX Manager: transport nodes are installed. Run pre-flight checks in vCenter: Administration > System Configuration > Components.
Analytics dashboard shows Instance A metrics but not Instance B
Cause: VCF Operations has not yet collected metrics from Instance B, or Instance B is not registered with VCF Operations
Fix: Wait 30 minutes for first metrics collection. In VCF Operations, verify Instance B is listed under Monitored Instances. Check: VCF Operations > Administration > Data Sources > Instance B (should show 'Connected').

Task 3 Configure and propagate fleet-level policies across instances

Manageability

Fleet-level policies are the key differentiator in VCF 9.0. This task covers creating policies that apply consistently across instances — a capability SDDC Manager never had.

Step 1

In Fleet Manager UI, navigate to Policies > Certificate Policy. Create a new fleet-level SSL certificate policy: Name: 'Fleet-CA-Cert-2024', Certificate File: upload a CA certificate (or generate a self-signed cert for lab), Apply Scope: 'All Instances'. Click Create.

Certificate policy created and shows 'Pending Application' status
Certificate policies are sensitive — in production, you would use your organization's CA certificate. For lab, a self-signed cert is acceptable but would trigger warnings in browsers.
Step 2

In Fleet Manager UI, navigate to Policies > Password Policy. Create fleet-level password policy: Name: 'Fleet-Password-Policy-2024', Minimum Length: 12, Require Uppercase: Yes, Require Digits: Yes, Require Special Chars: Yes, Rotation Interval: 90 days, Apply Scope: 'All Instances'. Click Create.

Password policy created. Shows which instances are targeted.
This policy will enforce consistent password standards across all instances — reducing the risk of weak passwords in a distributed fleet.
Step 3

In Fleet Manager UI, navigate to Policies > Patching Policy. Create fleet-level patching schedule: Name: 'Monthly-Patch-2nd-Tuesday', Frequency: Monthly, Day: 'Second Tuesday of month', Time Window: '02:00-06:00 UTC', Rollback on Failure: Yes, Apply Scope: 'All Instances'. Click Create.

Patching policy created with enforcement schedule
Patching policy is especially important in multi-instance fleet to ensure consistency. If Instance A patches on Day 1 and Instance B on Day 15, compatibility issues may arise. Fleet-level scheduling prevents this.
Step 4

Create a fleet-level NSX firewall (DFW) policy. In Fleet Manager > Policies > Security > Distributed Firewall, create a new rule: Name: 'Fleet-Block-Untrusted-DNS', Source: 'Any', Destination: 'DNS servers (UDP 53)', Action: 'Deny', Apply Scope: 'All Instances'. This rule will be pushed to NSX Managers in both instances.

DFW rule created and showing 'Pending Application' for Instance A and Instance B
This is a powerful feature — a single DFW rule propagated consistently across multi-site NSX deployment. In VCF 5.2, you would need to configure DFW rules manually in each NSX instance.
Step 5

Apply all pending policies. In Fleet Manager > Policies, click 'Apply All' or apply each policy individually. For each policy, select 'Apply Now' and confirm. Fleet Manager will push policies to each instance. Monitor application status.

Policies move from 'Pending' to 'Applied' status. Per-instance application log shows which instance received the policy and when.
Policy application is not instantaneous. Expect 5-10 minutes for policies to propagate to both instances and be activated. Large deployments may take longer.
Step 6

Validate policy application in individual instances. (1) Instance A: Log into vCenter > Administration > System Configuration > Security Settings. Verify password policy is active. (2) Instance B: Repeat verification. (3) For DFW: Log into NSX Manager A > System > Firewall > Policies. Verify the 'Fleet-Block-Untrusted-DNS' rule is present. Repeat for NSX B.

All policies visible in both instances with consistent names and settings
Take screenshots of the policies in each instance. This documents that propagation succeeded and provides evidence for VCDX design discussions.
Step 7

Test policy modification and re-application. In Fleet Manager, edit the Password Policy: change Rotation Interval from 90 to 60 days. Click Update. Then Apply. Monitor the application process. Verify in both instances that the new setting (60 days) is reflected.

Policy modification flows to both instances and updates are visible in instance configurations
This demonstrates the closed-loop: create policy -> apply to fleet -> modify -> re-apply. In production, implement change control: approval before policy apply, notification after application, rollback plan if issues arise.

Validation Gate

Check: All four policies (certificates, password, patching, DFW) created and applied successfully to both instances. Per-instance validation shows policies are active and consistent.

Expected: Fleet-level policy propagation working. Instances have consistent security and operational posture.

Common Errors

Policy shows 'Applied' in Fleet Manager but is not visible in Instance B vCenter/NSX
Cause: Policy propagation is asynchronous and still in-flight, or instance agent has not yet pulled the policy
Fix: Wait 10-15 minutes. In Instance B vCenter > Administration > Services, check that 'Fleet Manager Agent' is running. Force a sync: in Fleet Manager > Instance B > Settings > Sync Now.
Policy application fails with 'Incompatible with target instance'
Cause: Policy is configured for a feature not available in target instance (e.g., advanced DFW rule in Instance B which is at older NSX version)
Fix: Check instance versions: Instance A on NSX 9.0.2, Instance B on NSX 9.0.1. Simplify policy to be compatible with lowest NSX version, or split policy: apply to Instance A only.
DFW rule applied but not blocking untrusted DNS — traffic still flows
Cause: DFW rule was created as a logging rule (action: log, not deny) or is placed lower in rule order than a permit rule
Fix: In NSX Manager, check rule details: Action must be 'Deny'. Check rule order: 'Block-Untrusted-DNS' must be above any broader permit rules.

Task 4 Conduct fleet-wide reporting and compliance monitoring

Manageability

With multiple instances, unified observability becomes critical. This task uses Fleet Manager and VCF Operations to generate fleet-wide reports and verify compliance.

Step 1

In Fleet Manager UI, navigate to Analytics > Fleet Overview. Generate a fleet-wide summary report showing: total instances, total hosts, total VMs, total storage capacity, average cluster CPU/memory utilization across both instances, DFW rules count per instance.

Dashboard showing aggregated metrics for the entire fleet
Screenshot this dashboard — it's a powerful artifact showing fleet-level observability that VCF 5.2 did not have.
Step 2

In VCF Operations UI, navigate to Analytics > Compliance. Generate compliance report for the fleet: Verify (1) All instances are on supported versions of VCF/NSX, (2) All instances have patching policy applied, (3) All instances have password policy enforced, (4) All instances are backing up (if backup was configured in task vcp-admin-09). Generate a report and export as PDF.

Compliance report shows status of all compliance checks across both instances
If any instances show non-compliant status, document the specific non-compliance (e.g., 'Instance B missing patching policy') and steps to remediate.
Step 3

In Fleet Manager UI, navigate to Analytics > Policy Compliance. Create a custom compliance dashboard showing: Certificate Policy status (applied to all instances?), Password Policy status (enforced?), Patching Policy status (last execution time per instance). This provides at-a-glance verification that policies are consistently applied.

Policy compliance dashboard with per-instance status
Use this dashboard as your 'source of truth' for fleet compliance state. In production, export this dashboard weekly for compliance documentation.
Step 4

Generate cross-instance workload distribution report. In Fleet Manager > Analytics > Workloads, show: Total VMs across fleet, VMs per instance, VMs per workload domain, VM distribution by guest OS type, VM resource consumption (CPU, memory, storage) aggregated and per-instance.

Workload distribution report showing how VMs are spread across the fleet
This report is valuable for capacity planning: if one instance has 70% of VMs and the other has 30%, you have an imbalance that may affect performance or disaster recovery RTO.
Step 5

Generate security posture report. In NSX Manager (via Fleet Manager UI if available), query: Total DFW rules across fleet, rules per instance, rules in 'logging' vs 'enforcing' state, rules with high severity (block all traffic) vs low severity (allow specific traffic).

Security posture report showing DFW rule consistency and coverage across instances
If one instance has 100 rules and the other has 20, that's an asymmetry. Document it and determine if it's intentional (different security postures per site) or a gap (one site is under-protected).
Step 6

Export all reports to PDF or CSV. Save in a 'Fleet Reports' folder for archival and audit purposes. Create a report manifest: date generated, report type, instances covered, key findings, action items.

Fleet reports package with manifest file
In production, automate weekly report generation and storage. Use this as evidence of fleet monitoring and compliance for audits.

Validation Gate

Check: Fleet-wide analytics dashboard accessible, compliance report generated and shows pass/fail for all checks, policy compliance dashboard shows status per instance, workload distribution and security posture reports generated

Expected: Unified fleet observability established. You can answer: 'What is the compliance status of the fleet?' and 'Are policies consistently applied across instances?'

Common Errors

Compliance report shows Instance B as 'status: unknown' or 'last_check: never'
Cause: VCF Operations has not yet pulled compliance data from Instance B, or inventory sync is incomplete
Fix: Wait 30 minutes for first compliance check. In VCF Operations > Administration > Data Sources, verify Instance B is connected. Force sync: click Instance B > Refresh.
Workload distribution report shows VMs but not workload distribution by instance
Cause: Report configuration not filtered by instance, or instances are not properly tagged in VCF Operations
Fix: In Fleet Manager > Inventory, verify each instance is correctly labeled and tagged. Re-run report with 'Group By Instance' filter.
DFW rule count report shows 0 rules for Instance B
Cause: NSX integration not complete, or DFW rules are not being pulled from NSX API
Fix: In Fleet Manager > Integration > NSX, verify Instance B NSX Manager is connected. Test API: curl -k https://<nsx-b-ip>/api/v1/security-policies | grep rule-count

Task 5 Design and test fleet failover scenarios: instance failure and recovery

Availability

Multi-instance fleets introduce new failure modes. This task walks through recovery scenarios and validates that fleet design supports acceptable RTO/RPO.

Step 1

Document the fleet's critical dependencies. Create a failure mode matrix: (1) Fleet Manager failure — which operations are impacted?, (2) Instance A failure — can Instance B continue independently?, (3) Instance A vCenter failure — can workload VMs be managed?, (4) Cross-instance connectivity failure — can instances operate independently but unsynchronized?. For each scenario, estimate RTO and RPO.

Failure mode matrix with RTO/RPO estimates and impact analysis
This matrix is VCDX gold. Panelists will probe: 'What happens if Fleet Manager fails?' Your answer should be: 'Instances continue operating independently. Fleet Manager recovery takes <1 hour (rebuild appliance from backup). Workload continuity is not impacted.'
Step 2

Simulate Instance A failure: power off all ESXi hosts in Instance A workload domain (but leave management domain running). Workload VMs in Instance A will power off. Verify: (1) Fleet Manager detects Instance A as degraded, (2) Instance B continues operating normally with no impact, (3) Fleet Manager shows Instance A status as 'Partially Unavailable' or similar.

Fleet Manager detects Instance A degradation. Instance B unaffected.
This demonstrates instance isolation — a failure in one instance does not cascade to the other. This is a key design principle in VCF 9.0 multi-instance architecture.
Step 3

Recover Instance A: power on ESXi hosts in workload domain. Verify: (1) vCenter re-discovers hosts, (2) VMs re-register, (3) vSAN re-forms, (4) NSX transport nodes re-join fabric. This may take 10-20 minutes. Monitor via vCenter and NSX Manager.

Instance A cluster recovers and resumes normal operation
This demonstrates the autonomy of instances within the fleet — recovery is local and does not require Fleet Manager involvement.
Step 4

Simulate cross-instance network failure: block network traffic between Instance A and Instance B (firewall rule or cable disconnect). Verify: (1) Instances cannot communicate, (2) Workload VMs in Instance A cannot reach workloads in Instance B, (3) Each instance continues operating with no knowledge of the other, (4) Fleet Manager may show instances as 'network partitioned' or similar.

Cross-instance traffic blocked. Each instance operates independently.
In production, this scenario (network partition) is a split-brain risk. If not handled carefully, you could end up with two independent 'fleets' that diverge in configuration. Fleet Manager should have a split-brain detection mechanism — verify it's in place.
Step 5

Restore cross-instance connectivity. Verify: (1) Network traffic restored between instances, (2) Fleet Manager detects healing, (3) Policy sync occurs — if either instance was modified while partitioned, policies are re-synchronized, (4) No data loss or configuration inconsistency.

Cross-instance connectivity restored. Fleet converges to consistent state.
The convergence after split-brain is automatic if both instances have recorded changes to shared policies. However, if only one instance was partitioned, that instance should adopt the policies of the connected instance (write-through, not last-write-wins).
Step 6

Simulate Fleet Manager failure: power off the Fleet Manager VM. Verify: (1) Both instances continue operating, (2) Workload VMs in both instances unaffected, (3) No new policies can be applied (Fleet Manager is the policy control plane), (4) Instances may show 'disconnected from Fleet Manager' status.

Instances operate normally without Fleet Manager
This is the most critical test. If Fleet Manager loss cascades to instance failure, the fleet design is flawed. In VCF 9.0, it should not cascade.
Step 7

Recover Fleet Manager: power on the Fleet Manager VM. Verify: (1) Fleet Manager boots and re-initializes, (2) Instances reconnect, (3) Fleet Manager pulls fresh inventory from both instances and detects any changes that occurred during disconnection, (4) Policies are re-applied if needed, (5) Compliance checks run to ensure fleet is in expected state.

Fleet Manager recovers and reconciles fleet state
This demonstrates eventual consistency — instances record local changes, Fleet Manager reconciles on reconnection, conflicts are resolved via pre-defined rules (e.g., 'fleet policy wins over local policy').

Validation Gate

Check: All six failure scenarios tested: (1) Instance A failure and recovery, (2) Cross-instance network failure and restoration, (3) Fleet Manager failure and recovery. For each, verify instances remain operational and fleet converges to consistent state.

Expected: Fleet is resilient to single-instance failure, network partitions, and Fleet Manager loss. Instance autonomy is proven. Recovery procedures are validated.

Common Errors

Instance A failure cascades to Instance B — workload VMs in Instance B stop functioning
Cause: Workload routing or NSX configuration incorrectly depends on Instance A resources
Fix: Review Instance B NSX configuration: Tier-0 gateway should not route through Instance A. If cross-instance routing is needed, verify it uses direct Instance B-to-B network, not through Instance A.
Fleet Manager failure causes instances to failover to read-only mode or shutdown
Cause: Instances are incorrectly configured to depend on Fleet Manager for operational decisions
Fix: Verify instance configuration: Fleet Manager should be orchestration-only, not operational. If instances are querying Fleet Manager for cluster topology or policy validation, reconfigure to use local cache/replicas.
Cross-instance network restored but instances show conflicting policies
Cause: During partition, one instance applied a policy change that the other did not. Reconciliation failed.
Fix: In Fleet Manager, navigate to Policies > Conflict Resolution. Review conflicts and manually resolve by re-applying correct policy. Document the conflicting change and implement split-brain prevention (e.g., prevent policy changes during partition).
Instance recovery is slow (30+ minutes) after power-on
Cause: vSAN re-forming, NSX VTEP re-establishment, or vCenter database consistency checks are taking time
Fix: Expected behavior for large clusters. Monitor via vCenter > Cluster > Status to see detailed progress. Ensure vSAN is healthy before declaring instance fully recovered. In production, RTO for instance recovery is typically 20-40 minutes.

Final Validation

A multi-instance VCF 9.0 fleet has been deployed, registered with Fleet Manager, and validated for autonomous operation, policy propagation, observability, and failover resilience. Fleet-level policies (certificates, password, patching, DFW rules) are consistently applied across instances. Analytics and compliance reporting show unified fleet visibility. Failure scenarios (instance loss, network partition, Fleet Manager loss) have been tested and recovery procedures validated. The fleet demonstrates the key architectural advantages of VCF 9.0 over VCF 5.2: autonomous instances + federation control plane, rather than monolithic SDDC Manager dependency.

✓ Fleet topology documented with all instances, management domains, workload domains, and cross-instance routing → Multi-instance fleet architecture fully mapped and understood

✓ Instance autonomy proven: instances operate independently without Fleet Manager → Instances continue functioning when Fleet Manager is offline

✓ Inventory sync configured and active for both instances → Fleet Manager real-time visibility of both instances with continuous sync

✓ Fleet-level policies created and applied: certificates, password, patching, DFW → All four policy types visible in both instances with consistent configuration

✓ Fleet-wide compliance report generated → Compliance dashboard shows status across both instances

✓ Failure scenarios tested: instance failure, cross-instance network partition, Fleet Manager loss → All failure modes tested; instances resilient and recoverable; fleet eventually consistent

✓ Documentation: failure mode matrix, RTO/RPO estimates, topology diagram, policy compliance matrix → VCDX-ready design documentation completed

Cleanup / Restore

Snapshot: vcp-admin-10-post-lab

• Remove all test policies created during the lab (or keep if desired for subsequent labs)

• Revert any powered-off instances to powered-on state

• Restore cross-instance network connectivity (if partitioned for testing)

• Verify Fleet Manager is operational and all instances reconnected

• Take a post-lab snapshot of the fleet for future reference

• Archive all reports and documentation generated during the lab

Design Reflection (VCDX)

The shift from VCF 5.2 SDDC Manager to VCF 9.0 Fleet Manager represents a fundamental architectural change in multi-instance management philosophy. SDDC Manager was a single operational hub — all day-2 operations (patching, certificates, expansion, failover) flowed through it, making it a bottleneck and single point of failure. Fleet Manager is a coordination plane that enables instance autonomy while providing fleet-level policy and observability. A VCDX panelist will ask: 'Why this change?

What problem does it solve?' Your answer should reference scalability (supporting 10+ instances), blast radius reduction (instance failure doesn't cascade), and operational flexibility (instances can be regional or independent). They will also probe: 'What are the trade-offs?' Higher complexity (distributed state reconciliation), eventual consistency (not strong consistency), and more sophisticated observability requirements.

Be prepared to discuss: (1) how you would design a 20-instance global fleet with regional autonomy, (2) how you would ensure policy consistency in a split-brain scenario, (3) what RPO/RTO targets are realistic for multi-region failover, (4) how you would monitor and alert on fleet-wide compliance drift, and (5) the role of automation in managing operational complexity of a distributed fleet.

Requirements

  • Fleet Manager must coordinate policies and lifecycle operations across 2+ instances
  • All instances must maintain autonomous operation regardless of Fleet Manager availability
  • Policies created at fleet level must propagate consistently to all instances
  • Fleet-wide reporting and compliance monitoring must provide single-pane-of-glass visibility
  • Failure of one instance must not impact other instances or Fleet Manager
  • Fleet must support cross-instance workload communication and data center extension scenarios
  • Policy changes must be reversible and support rollback if an instance experiences incompatibility

Constraints

  • Instance autonomy limits Fleet Manager's ability to enforce synchronous consistency — eventual consistency is the model
  • Network latency and partitions between instances require local caching of policies and configuration
  • NSX and vCenter versions may differ across instances, limiting unified policy scope (e.g., advanced DFW rules may not apply to older NSX versions)
  • Workload mobility across instances is limited — VMs cannot live-migrate across instance boundaries (separate vCenter domains)
  • Policy propagation is asynchronous and may take minutes — near-real-time synchronization is not achievable
  • Instances in different geographic regions may have different security, compliance, or regulatory requirements — one-size-fits-all fleet policies may be infeasible

Assumptions

  • Instances are deployed and fully operational before being registered with Fleet Manager
  • Network connectivity between Fleet Manager and all instances is stable (not prone to frequent partitions)
  • Instance administrators accept eventual consistency for fleet policies — strong consistency is not a hard requirement
  • All instances are managed by a single organization (no multi-tenant scenarios in this lab)
  • Instances are in same or similar configuration (symmetric fleet) — asymmetric instances may require special handling
  • Fleet Manager is not needed for day-2 operations like VM management or workload patching — only for policy and lifecycle

Risks

  • Policy divergence across instances: if network partition occurs and both instances accept policy changes, reconciliation may be impossible. MITIGATION: implement policy versioning and conflict detection; prefer 'fleet policy wins' for conflicts.
  • Split-brain scenario: if fleet network partitions, one instance may assume itself is the fleet and begin authoritative policy application while the other does the same. MITIGATION: implement quorum or leader election; detect and alert on partition; require manual intervention to resolve.
  • Fleet Manager single point of failure: if Fleet Manager is lost and cannot be recovered from backup, instances are orphaned. MITIGATION: backup Fleet Manager regularly, document manual recovery, maintain out-of-band instance access (SSH) for emergency operations.
  • Workload disruption during cross-instance network failure: if instances are connected and one loses connectivity, workloads with cross-instance dependencies (distributed DB, cluster service) may experience split-brain or partial outage. MITIGATION: design workloads to be instance-local; if cross-instance communication is needed, use application-level consensus (e.g., Zookeeper, Raft) rather than relying on network topology.
  • Compliance drift: if instances are modified locally (outside of Fleet Manager policies), compliance reports may become inaccurate. MITIGATION: implement audit logging of all configuration changes, regular compliance scans, automated remediation of drift.
  • Scaling beyond 10-20 instances: Fleet Manager may become a bottleneck for inventory sync and policy propagation. MITIGATION: implement tiered fleet (sub-fleet managers) or hierarchical policy propagation; profile and optimize inventory sync queries.

Self-Assessment Discussion Prompts

  1. In VCF 5.2, SDDC Manager was mandatory for all management operations. In VCF 9.0, Fleet Manager is orchestration-only and instances are autonomous. What is the business advantage of this change? What are the operational challenges?
  2. Design a global VCF 9.0 fleet with 10 instances across 5 continents. How would you structure the fleet? Would you have one Fleet Manager for all instances, or multiple Fleet Managers (one per region)? Justify your choice.
  3. Describe a scenario where instance autonomy saves you from catastrophic failure. How would the same scenario play out in VCF 5.2 with SDDC Manager?
  4. Fleet Manager policies are eventually consistent — they propagate asynchronously across instances. What are the implications for security policy (e.g., DFW rules)? How long can an instance operate with 'old' policies before it violates compliance?
  5. If you have a split-brain scenario (fleet partition, both partitions think they are authoritative), how would you resolve it? What would be your conflict resolution strategy?
  6. Workload migration across instances in VCF 9.0 is not possible (no live vMotion between instances). How does this constraint affect your design of a multi-region fleet? Would you ever want to migrate workloads across instances?
  7. Monitoring and alerting on 10 instances is complex. What KPIs would you surface on a fleet dashboard to detect early signs of trouble (e.g., policy divergence, compliance drift, instance degradation)?

Extensions

Design a Global Fleet with Regional Autonomy

Extend the lab to 3+ regions (North America, Europe, Asia-Pacific). Each region has its own instance pair, managed by a single Fleet Manager. Design: (1) regional instance grouping, (2) region-specific policies (e.g., GDPR compliance for Europe, local encryption standards for Asia), (3) cross-region failover scenarios, (4) network latency and bandwidth considerations for policy sync. Document how you would handle policy conflicts between regions.

harder

Implement Hierarchical Fleet Management

Instead of a single Fleet Manager, design a multi-tier fleet: (1) Central Fleet Manager coordinates 2 regional Fleet Managers, (2) Each regional Fleet Manager manages 2-3 instances. Implement: (a) policy inheritance (parent policies propagate to children), (b) policy override (regional Fleet Manager can customize parent policy), (c) conflict resolution between regional and central policies. This models very large fleets (100+ instances).

expert

Test and Document Disaster Recovery at Fleet Level

Design a DR procedure for a multi-instance fleet: (1) simultaneous failure of Instance A and Fleet Manager, (2) recovery sequence (restore Fleet Manager first, then instances), (3) RTO/RPO targets and validation. Execute the DR procedure in lab and document: time to recover, data loss (if any), consistency checks, post-recovery validation. This mirrors production DR testing and is critical for VCDX credibility.

expert

Implement Workload Affinity and Anti-Affinity Across Fleet

Configure VM placement policies at fleet level: (1) critical workloads must not run in the same instance (anti-affinity for disaster isolation), (2) related workloads must run in same instance (affinity for latency), (3) workloads with geographic compliance needs must run in specific regions. Test how Fleet Manager (via vCenter) enforces these policies during VM provisioning. Explore limitations: can you guarantee affinity across instances, or only within instance?

harder

⚠ Known Pitfalls (from Community KB)

Instances Treated as Subordinates Instead of Autonomous Peers COMMON
Problem: Operational model assumes instances are managed by Fleet Manager for all day-2 operations (SDDC Manager legacy thinking). When Fleet Manager is lost, instances are assumed to fail.
Resolution: Reframe mental model: instances are autonomous, Fleet Manager is orchestration-only. Day-2 operations (VM provisioning, patching, cluster expansion) happen locally within each instance. Fleet Manager coordinates fleet-level policies and lifecycle. Test instance autonomy: verify instances operate without Fleet Manager.
Policy Propagation Expected to be Synchronous and Atomic COMMON
Problem: Expected that policy application completes instantly across all instances. If one instance fails to apply policy, assumption that fleet is now inconsistent.
Resolution: Understand eventual consistency: policy propagation is asynchronous and per-instance. If one instance fails to apply policy, it will retry. Monitor policy sync status, but don't require 100% instant consistency. In critical cases (security policies), implement verification checks before and after propagation.
Cross-Instance Workload Mobility Expected (Live vMotion Between Instances) RESOLVED
Problem: Assumption that VMs can live-migrate across instance boundaries (separate vCenter domains). Not possible in VCF 9.0 architecture.
Resolution: Workload mobility is within-instance only. Cross-instance workload movement requires shutdown, export, and reimport. Design fleets with workload locality in mind — don't assume you can rebalance VMs dynamically across instances.
Fleet Manager Loss Results in Full Fleet Outage CRITICAL_IF_MISDESIGNED
Problem: If Fleet Manager is not properly designed as optional orchestration layer, instances may fail when Fleet Manager is lost.
Resolution: Test instance autonomy thoroughly. If instances fail without Fleet Manager, reconfigure them to operate independently. Fleet Manager should be operational convenience, not operational requirement. See Task 5 for failure scenario testing.
Policy Conflict in Split-Brain Scenario Not Detected or Resolved LIKELY_RESOLVED
Problem: During network partition, both fleet partitions apply policies. When partition heals, conflicting policies are silently accepted, leading to inconsistency.
Resolution: Implement split-brain detection and conflict resolution in Fleet Manager. When partition heals, Fleet Manager must detect conflicts (e.g., 'certificate policy version 3 vs version 5') and resolve them per pre-defined strategy (e.g., 'later version wins', 'fleet policy wins', 'manual intervention required').

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.