Academy/VCP-VCF 9.0 Administrator (2V0-17.25)/VCF 9.0 Backup & Restore — SDDC Manager Absence Paradigm
This lab targets VCF 9.0

VCF 9.0 Backup & Restore — SDDC Manager Absence Paradigm

VCF 9.0Advancedvcp-foundationarchitectvcdx⏱ 150 min

VCF 9.0 introduces Fleet Manager as the primary management layer, fundamentally changing backup/restore architecture compared to VCF 5.2 SDDC Manager. This lab focuses on VCF 9.0.x backup strategies for multi-instance fleets without SDDC Manager dependency.

Objectives

  • Articulate VCF 9.0 backup architecture: what must be backed up (Fleet Manager, VCF Operations, vCenter, NSX), and why the absence of SDDC Manager changes the approach
  • Configure file-based backup targets (SFTP) for all critical management components in a fleet topology
  • Execute and validate a complete backup cycle including Fleet Manager, VCF Operations, NSX Manager, and vCenter Server
  • Perform a catastrophic recovery: rebuild Fleet Manager appliance from backup and restore full fleet state
  • Conduct DR testing: validate post-restore inventory consistency, policy application, and workload continuity across instances
  • Compare VCF 9.0 fleet-based backup to VCF 5.2 SDDC Manager backup and document operational differences

Prerequisites

Multi-instance VCF 9.0 lab environment deployed and operational. Fleet Manager active and managing 2+ instances. All management appliances (vCenter, NSX, VCF Operations) healthy with network connectivity to backup target SFTP server. Minimum 500 GB available storage on backup target.

Prior labs: holodeck-02

Required skills:

  • VCF 9.0 architecture: Fleet Manager, instances, management domains
  • SFTP/file-based backup concepts and SSH key authentication
  • vCenter Server backup and restore procedures
  • NSX Manager backup and restore procedures
  • SSH CLI access to appliances
  • VCF Operations (formerly VCF Manager) console navigation
  • Understanding of DR testing vs disaster recovery

Lab Environment

Multi-instance VCF 9.0 fleet with: (1) Fleet Manager VM managing instances, (2) Instance A (primary site) with management domain + workload domain, (3) Instance B (secondary site) with management domain + workload domain, (4) NSX Manager deployed within each instance, (5) vCenter Server deployed within each instance, (6) VCF Operations (embedded or standalone) monitoring both instances. SFTP backup server accessible from all management appliances.

graph TB
  FM[Fleet Manager 10.0.0.50]
  SFTP[SFTP Backup Server<br/>192.168.1.100]
  INST1[Instance A<br/>Site Primary]
  INST2[Instance B<br/>Site Secondary]
  VOP[VCF Operations<br/>10.0.0.51]
  VC1[vCenter<br/>Instance A]
  NSX1[NSX Manager<br/>Instance A]
  VC2[vCenter<br/>Instance B]
  NSX2[NSX Manager<br/>Instance B]
  FM --> SFTP
  VOP --> SFTP
  INST1 --> VC1
  INST1 --> NSX1
  INST2 --> VC2
  INST2 --> NSX2
  VC1 --> SFTP
  NSX1 --> SFTP
  VC2 --> SFTP
  NSX2 --> SFTP

IP Addressing

NetworkPurposeVLAN
10.0.0.0/24Management network (Fleet Manager, VCF Operations, instance management domains)VLAN 1644
10.1.0.0/16Instance A workload network (cross-instance routing)VLAN 1650+
10.2.0.0/16Instance B workload network (cross-instance routing)VLAN 1650+
192.168.1.0/24Backup SFTP server and out-of-band managementVLAN 100

Credentials

SystemUsernamePassword
Fleet ManageradminFleet Manager admin password set during deployment
VCF Operationsadministrator@vsphere.localSame as Instance A vCenter SSO password
vCenter Server (Instance A)administrator@vsphere.localInstance A SSO password
vCenter Server (Instance B)administrator@vsphere.localInstance B SSO password
NSX Manager (Instance A)adminInstance A NSX admin password
NSX Manager (Instance B)adminInstance B NSX admin password
SFTP Backup Servervcf-backupSFTP user credential for file deposit

Tasks

Task 1 Map VCF 9.0 backup architecture and compare to VCF 5.2 SDDC Manager model

Recoverability

In VCF 9.0, the backup surface has changed fundamentally. SDDC Manager no longer exists. Fleet Manager and VCF Operations have replaced it. Understanding what to back up and in what order is a VCDX-critical skill.

Step 1

Log into Fleet Manager UI (https://<fleet-manager-ip>). Navigate to Lifecycle > Backup & Recovery to view the backup topology diagram.

Backup topology showing Fleet Manager as root, connected to VCF Operations, and all managed instances below. Visual dependency graph shows backup order constraints.
Screenshot this diagram. It is your first VCDX design artifact showing backup dependencies.
Step 2

In Fleet Manager UI, navigate to Inventory > Managed Instances. Document the FQDN, IP address, and instance type (management/workload) for each instance. Create a spreadsheet with columns: Instance Name, Role, FQDN, Management IP, vCenter IP, NSX Manager IP, Backup Priority.

Spreadsheet with all instances mapped to their components
Incorrect instance-to-component mapping is the #1 cause of incomplete backups. Each instance can have separate vCenter and NSX managers, or shared fleet-level ones. Verify by examining each instance's configuration page.
Step 3

Compare to VCF 5.2: if you have access to a VCF 5.2 lab environment (or documentation), document how SDDC Manager backup differed. Key differences: (a) SDDC Manager was a single point for all instance state, (b) vCenter and NSX backup were optional, (c) no fleet-level policy backup. For VCDX, be prepared to articulate WHY VCF 9.0 moved to this distributed model (hint: multi-instance scalability, regional independence, reduced blast radius).

Comparison table with 5.2 vs 9.0 backup architecture differences
This comparison is gold for VCDX discussion. Panelists will ask: What was the architectural problem VCF 9.0 solved? Your answer should reference the SDDC Manager bottleneck and the need for multi-instance management.
Step 4

Review the VCF 9.0 backup requirements documentation: what are the RPO (Recovery Point Objective) and RTO (Recovery Time Objective) implications for each component? Document RTO for: (a) Fleet Manager alone, (b) Instance vCenter, (c) Instance NSX, (d) Full fleet recovery. Note that Fleet Manager recovery < 1 hour (appliance rebuild), but instance recovery depends on backup freshness.

RTO/RPO matrix for each backup component
RPO of 24 hours (daily backups) is typical, but RTO varies: Fleet Manager ~30 min, vCenter ~60 min, NSX ~90 min. Longer RTO due to appliance rebuild and state replication.
Step 5

Document the critical data that MUST be backed up (Fleet Manager configuration, instance registration, policy definitions) vs optional data (vCenter event logs, NSX flow logs). This distinction determines backup size and retention strategy.

Critical vs Optional backup data matrix
Critical data: Fleet Manager database, instance registration keys, SSL certificates, policy templates. Optional: logs, telemetry, old snapshots. This drives your 'lean backup' strategy for resource-constrained environments.

Validation Gate

Check: Comparison table complete with all instances mapped + RTO/RPO documented + critical data distinction made

Expected: You can articulate the backup topology and explain why each component is backed up independently (unlike SDDC Manager which centralized all state)

Common Errors

Backup topology diagram shows instances not connected to Fleet Manager
Cause: Instance registration incomplete — instance exists in infrastructure but not yet adopted by Fleet Manager
Fix: In Fleet Manager, navigate to Lifecycle > Add Instance and complete the registration workflow. Instance must register before backup can include it.
📋 KB: VCF 9.0 Instance Registration Guide (Broadcom docs)
Uncertainty about whether vCenter/NSX should be backed up separately or only through Fleet Manager
Cause: Misunderstanding the layered backup model — Fleet Manager backs up metadata, but vCenter and NSX require separate application-level backups
Fix: Fleet Manager backup ≠ vCenter backup. Each appliance has its own backup procedure. Fleet Manager backup is a fleet-level orchestration layer that can TRIGGER individual appliance backups, but does not replace them.
📋 KB: VCF 9.0 Backup Architecture (Broadcom docs § 3.2)
Cannot articulate difference between VCF 5.2 and 9.0 backup models
Cause: Lack of VCF 5.2 exposure — the architectural shift is subtle and requires comparison
Fix: Read the VCF 5.2 SDDC Manager Backup Guide (cited in references). Focus on: SDDC Manager was a monolithic backup target; VCF 9.0 Fleet Manager is an orchestrator that coordinates appliance-level backups.

Task 2 Configure file-based backup targets (SFTP) for all management components

Recoverability

In production, you must configure backup targets for Fleet Manager, VCF Operations, vCenter, and NSX independently. This task walks through the configuration of each.

Step 1

Prepare the SFTP backup server. Verify connectivity from Fleet Manager: ssh vcf-backup@192.168.1.100. Create backup directories: mkdir -p /backups/fleet-manager /backups/vcf-operations /backups/instance-a-vcenter /backups/instance-a-nsx /backups/instance-b-vcenter /backups/instance-b-nsx. Set permissions: chmod 755 /backups && chmod 755 /backups/*

SFTP server accessible, backup directories created and readable
Use SSH keys instead of passwords for production. Generate an SSH keypair on Fleet Manager and add the public key to ~vcf-backup/.ssh/authorized_keys on the backup server.
Step 2

In Fleet Manager UI, navigate to Lifecycle > Backup & Recovery > Backup Configuration. Add a new backup target: Protocol: SFTP, Host: 192.168.1.100, Port: 22, Username: vcf-backup, Password/Key: [use SSH key method], Remote Path: /backups/fleet-manager. Click 'Validate Connection'.

Validation succeeds: 'Connection to SFTP server verified. Write permissions confirmed.'
If validation fails with 'Permission denied', the SFTP user lacks write permissions or the path doesn't exist. SSH to the server and verify: ls -ld /backups/fleet-manager && su - vcf-backup -c 'touch /backups/fleet-manager/test.txt'
Step 3

Configure VCF Operations backup. Log into VCF Operations UI (https://<vcf-operations-ip>). Navigate to Settings > Backup & Restore. Add backup target: SFTP, Host: 192.168.1.100, Path: /backups/vcf-operations. Validate.

VCF Operations reports 'Backup target configured and validated'
VCF Operations backup includes monitoring database, analytics, and compliance snapshots. Typically 20-50 GB per backup.
Step 4

Configure vCenter backup for Instance A. Log into vCenter (Instance A) UI. Navigate to Administration > Backup & Restore. Add backup location: Protocol: SFTP, Server: 192.168.1.100, Username: vcf-backup, Directory: /backups/instance-a-vcenter. Test the connection.

vCenter reports 'Backup server connected and authenticated'
vCenter backup includes the vCenter database, certificates, and configuration. Size: 30-100 GB. Ensure backup server has sufficient space and the SFTP user can write at least 150 GB across all targets.
Step 5

Configure NSX Manager backup for Instance A. Log into NSX Manager (Instance A) UI. Navigate to System > Backup & Restore > Backups. Configure backup: Protocol: SFTP, Server: 192.168.1.100, Username: vcf-backup, Directory: /backups/instance-a-nsx. Click 'Test Backup'.

NSX Manager creates a test backup file and reports success
NSX backup includes the NSX database (configurations, segments, DFW rules), certificates, and user database. Size: 5-20 GB per backup.
Step 6

Repeat Steps 4 and 5 for Instance B (vCenter and NSX of Instance B), using paths /backups/instance-b-vcenter and /backups/instance-b-nsx respectively.

All backup targets configured: Fleet Manager, VCF Operations, Instance A (vCenter + NSX), Instance B (vCenter + NSX)
Verify all targets are reachable: in Fleet Manager Backup Configuration, click 'Test All Targets'. All six should report success.

Validation Gate

Check: Fleet Manager UI shows 6 backup targets all in 'Connected' state. SSH to SFTP server: ls -la /backups/ shows all directories with vcf-backup ownership.

Expected: All management components have validated SFTP backup targets. Backup server has write access to all directories.

Common Errors

SFTP connection fails: 'Connection refused' or 'No route to host'
Cause: Backup server unreachable — network routing, firewall, or SFTP service down
Fix: Verify: (1) ping 192.168.1.100 from Fleet Manager, (2) SSH into backup server and run: systemctl status ssh, (3) Verify firewall: iptables -L | grep 22
SFTP connection succeeds but validation fails: 'Permission denied' on write
Cause: SFTP user cannot write to remote path — directory permissions or ownership issue
Fix: SSH to backup server as root: chmod 777 /backups/<path> OR chown vcf-backup:vcf-backup /backups/<path>. Then re-validate from Fleet Manager.
vCenter backup configuration appears in UI but backup never executes
Cause: vCenter service can reach SFTP server but internal backup scheduler is disabled or erroring silently
Fix: In vCenter, check Configuration > System > Backup & Restore > Backup Schedule. Verify schedule is enabled and next execution time is shown. Check system logs: /var/log/vmware/backup/
NSX Manager backup target configured but subsequent backups fail with 'Insufficient space'
Cause: SFTP server disk full — multiple large backups consuming space
Fix: On backup server: df -h /backups. Implement retention policy: keep only last 7 daily backups. Script: find /backups -name '*.tgz' -mtime +7 -delete

Task 3 Execute a complete backup cycle: Fleet Manager, VCF Operations, vCenter, and NSX

Recoverability

Backing up is straightforward if targets are configured. The critical learning is understanding the dependency order, backup size, and duration.

Step 1

Start with Fleet Manager backup (has fewest dependencies). In Fleet Manager UI, navigate to Lifecycle > Backup & Recovery > Backup Now. Select Backup Scope: 'Full' (includes instance registration data, fleet policies, certificates). Click 'Initiate Backup'. Monitor progress.

Backup progresses through stages: (1) Preparing metadata, (2) Creating archive, (3) Uploading to SFTP. Final status: 'Backup completed: 2024-04-17-FM-001, Size: 3.2 GB, Duration: 12 minutes'
Fleet Manager backup must complete before instance backups begin. If instance backups run during Fleet Manager backup, there is a risk of state inconsistency. In production, script these backups with 30-minute delays between each component.
Step 2

Initiate VCF Operations backup. In VCF Operations UI, navigate to Backup & Restore > Backup Now. Select backup level: 'Full' (includes database, analytics, compliance snapshots). Click Start. Monitor progress.

Backup progresses. Status: 'Backup completed: 2024-04-17-VCFOPS-001, Size: 28 GB, Duration: 45 minutes'
VCF Operations backup is larger than Fleet Manager because it includes monitoring metrics and historical data. Allow 45-90 minutes for full backup.
Step 3

Initiate vCenter backup for Instance A. In vCenter UI, navigate to Administration > Backup & Restore > Backup Now. Backup type: 'Full' (includes database, events, tasks, configuration). Click Start.

vCenter begins backup. Status page shows: 'Backup in progress: 2024-04-17-VC-A-001'. Monitor via Administration > System Configuration > Services > Backup Status.
vCenter backup is a heavy operation. The vCenter service briefly pauses write operations during backup. In production, schedule vCenter backups during maintenance windows (off-hours). Expect 60-120 minutes.
Step 4

Initiate NSX Manager backup for Instance A (do not wait for vCenter to finish; these run in parallel). In NSX Manager UI, navigate to System > Backup & Restore > Backups > Create Backup. File name prefix: 'nsx-a-2024-04-17'. Click Create.

NSX creates backup file and uploads to SFTP. Status: 'Backup successful. File: nsx-a-2024-04-17-001.backup, Size: 12 GB'
NSX backup is quicker than vCenter (5-10 minutes for backup creation + upload).
Step 5

Once Instance A backups are complete (vCenter and NSX), initiate backups for Instance B. Repeat Steps 3-4 for Instance B vCenter and NSX, using directory paths and hostnames specific to Instance B.

All Instance B backups initiated and progress monitored
In production, use a backup orchestration script (PowerShell or bash) to automate this sequence with proper error handling and notifications.
Step 6

Monitor the SFTP backup server to confirm all backups have been deposited. SSH: ls -lah /backups/ and df -h /backups. Document total backup size, number of files, and oldest vs newest backup timestamps.

Backup server shows all backup files with timestamps. Example: /backups/fleet-manager/2024-04-17-FM-001.tgz (3.2 GB), /backups/vcf-operations/2024-04-17-VCFOPS-001.tgz (28 GB), etc.
Total backup size is typically 80-150 GB for a two-instance fleet with full backups. If your server has less free space, implement incremental backup or retention pruning.
Step 7

Document the backup manifest. Create a file /backups/MANIFEST-2024-04-17.txt with entries for each backup: Component, Timestamp, Filename, Size, Checksum (md5sum), Notes (e.g., 'Full backup before storage expansion'). This becomes your recovery procedure reference.

MANIFEST file created with all backup metadata
In production, generate this manifest programmatically and store a copy off-site or in an external repo (git). The manifest is critical for recovery — without it, you don't know which backup corresponds to which point in time or which components were included.

Validation Gate

Check: All six backups completed and verified on SFTP server. MANIFEST file created with checksums. Backup server has sufficient remaining space (> 100 GB free).

Expected: Complete fleet backup cycle finished. All components backed up. Ready for restore testing.

Common Errors

Backup fails midway: 'SFTP upload interrupted, connection timeout'
Cause: Network interruption during large file upload (vCenter/VCF Operations backups are 20-50 GB). Resume/retry mechanism not triggered.
Fix: Fleet Manager and vCenter have built-in retry logic. Check UI — they should show 'Backup in progress, retrying upload'. If stuck, click 'Retry' or restart the backup.
Backup server disk full after second or third backup cycle
Cause: Retention policy not enforced — old backups accumulating
Fix: Implement retention: keep last 7 daily + last 4 weekly + last 12 monthly. Script: find /backups -name '*.tgz' -mtime +7 -delete. In production, integrate this into a cron job.
Backup reports success but file on SFTP server is 0 bytes or corrupted
Cause: SFTP write succeeded but file was not properly flushed or transferred was incomplete
Fix: Check file integrity: md5sum /backups/<file>. If checksum is missing or file size is wrong, re-initiate backup. Verify disk space on backup server is not full (df -h).
vCenter backup timeout: 'Backup did not complete within 4 hours'
Cause: Large vCenter database or insufficient disk I/O on backup server
Fix: Increase backup timeout in vCenter: Administration > Backup & Restore > Backup Configuration > Advanced > Timeout (default 4 hours). Alternatively, split vCenter backup into incremental backups.

Task 4 Perform disaster recovery: rebuild Fleet Manager from backup and restore instance state

Recoverability

Backups are only valuable if you can restore them. This task simulates a catastrophic Fleet Manager failure and walks through the recovery procedure.

Step 1

Create a 'disaster' snapshot of the current state. In the physical ESXi host, take a snapshot of the Fleet Manager VM: Right-click Fleet Manager VM > Snapshots > Take Snapshot, Name: 'pre-disaster-recovery-test', Description: 'Snapshot before DR test — revert here if recovery fails'.

Snapshot created successfully
This snapshot is your safety net. If the recovery procedure goes wrong, you can revert and retry.
Step 2

Simulate catastrophic Fleet Manager failure: In vSphere Client, power off the Fleet Manager VM. Wait 30 seconds. Then delete the Fleet Manager VM's virtual disk. This simulates a storage failure where the appliance is unrecoverable.

Fleet Manager VM powered off and disk deleted
After this step, Fleet Manager is gone. Instances can still operate independently (they are autonomous in VCF 9.0), but fleet-level operations are impossible. Proceed with recovery in the next steps.
Step 3

Deploy a new Fleet Manager appliance OVA (using the same OVA as the original deployment). Complete the initial deployment wizard: hostname, IP address (use the same IP as the original Fleet Manager: 10.0.0.50), domain, NTP, DNS. Do NOT join it to any existing configuration — this is a blank appliance.

New Fleet Manager appliance deployed and running with base configuration, no instances registered
The new appliance has no knowledge of the previous instances or fleet state. This is expected — the state will be restored from backup in the next steps.
Step 4

Restore Fleet Manager from backup. Log into the new Fleet Manager UI as admin. Navigate to Lifecycle > Backup & Recovery > Restore. Select backup source: SFTP, Server: 192.168.1.100, Path: /backups/fleet-manager. Select the latest backup file (2024-04-17-FM-001.tgz). Click Restore.

Fleet Manager begins restore process. UI shows: 'Restore in progress: extracting backup, validating database, restoring configuration'. Process takes 15-30 minutes.
During restore, Fleet Manager service will restart. Do not interrupt. Web UI will be briefly unavailable. Monitor via SSH: tail -f /var/log/vmware/fleet-manager/backup-restore.log
Step 5

Wait for Fleet Manager restore to complete. Web UI returns. Log in as admin. Verify: (1) navigate to Inventory > Managed Instances — both Instance A and Instance B should be listed and should show 'status: registered', (2) Navigate to Lifecycle > Policies — all fleet policies (certificates, passwords, patching schedules) should be restored.

Fleet Manager shows all instances and policies as they were at backup time
Instances may temporarily show 'status: unreachable' until they re-establish communication with Fleet Manager (usually within 5 minutes). This is normal.
Step 6

Restore vCenter for Instance A. Log into the new Fleet Manager. Navigate to Lifecycle > Backup & Restore > Restore Instance. Select Instance A. Choose Component: vCenter. Backup source: SFTP, Directory: /backups/instance-a-vcenter. Select the latest vCenter backup. Click Restore.

Fleet Manager orchestrates the vCenter restore: (1) Detects the existing Instance A vCenter VM, (2) Powers it off, (3) Restores from backup, (4) Powers it back on, (5) Validates connectivity. Process takes 60-120 minutes depending on vCenter database size.
While vCenter is being restored, Instance A workload domain will be unavailable (no vCenter = no VM management). This is acceptable in a DR scenario. In production, consider restoring to a standby vCenter first before cutover.
Step 7

Restore NSX Manager for Instance A. In Fleet Manager, navigate to Lifecycle > Backup & Restore > Restore Instance. Select Instance A. Choose Component: NSX Manager. Backup source: /backups/instance-a-nsx. Select latest NSX backup. Click Restore.

Fleet Manager restores NSX: (1) Detects existing NSX Manager VM, (2) Powers it off, (3) Restores from backup, (4) Powers it back on, (5) Validates all transport nodes are still connected. Process takes 45-90 minutes.
NSX restore includes the network configuration: segments, firewall rules, load balancers, VTEPs. If transport nodes are not on the same network as the original deployment, NSX restore may fail. Ensure network topology has not changed.
Step 8

Repeat Steps 6-7 for Instance B (vCenter and NSX). This is done in parallel with Instance A restore or sequentially, depending on capacity.

Both instances fully restored from backups. All components show 'status: healthy'
Total recovery time for a full two-instance fleet: 3-4 hours (Fleet Manager 15-30 min + 2x vCenter 60-120 min each in parallel + 2x NSX 45-90 min each in parallel). In production, pre-stage standby appliances to reduce RTO further.

Validation Gate

Check: Fleet Manager: (1) All instances registered, (2) All policies present. vCenter A & B: (1) Inventory shows all workload domains and VMs, (2) vSAN cluster healthy, (3) Datastore capacity matches original. NSX A & B: (1) All transport nodes connected, (2) All segments present, (3) DFW rule count matches original.

Expected: Full fleet restored to the state at backup time. All management and workload components operational. Inventory and policies consistent.

Common Errors

Fleet Manager restore hangs at 'Validating database integrity' for 30+ minutes
Cause: Large Fleet Manager database or slow SFTP connection causes restore to timeout
Fix: Monitor via SSH: ps aux | grep backup. If process is running, let it continue. If stuck, check: (a) SFTP server reachability: ping 192.168.1.100, (b) Disk space: df -h / (must have >50 GB free)
Instance A shows 'status: unreachable' after Fleet Manager restore completes
Cause: Instance management network routing or DNS broken during disaster. Instance can reach its own vCenter/NSX but cannot reach Fleet Manager.
Fix: Verify network connectivity from Instance A management domain to Fleet Manager 10.0.0.50. SSH to an ESXi host in Instance A: ping 10.0.0.50. If unreachable, check: (1) physical network routing, (2) NSX distributed firewall rules blocking management traffic
vCenter restore fails: 'Database restore failed, VC_DB_ERROR'
Cause: vCenter database corruption in backup file or database version mismatch between backup and new appliance
Fix: Verify backup file integrity: md5sum /backups/instance-a-vcenter/*.tgz. If checksum matches original MANIFEST, backup is not corrupted. Check vCenter logs: /var/log/vmware/vc-db-restore.log for specific errors.
NSX restore completes but transport nodes show 'status: down'
Cause: NSX segment configuration restored but network VLAN or routing has changed since backup
Fix: In NSX Manager, navigate to System > Fabric > Nodes > Host Transport Nodes. For each down node, click 'Resolve' — NSX will re-validate connectivity. If transport nodes remain down, verify: (1) Network VLAN still exists, (2) ESXi hosts still have VMkernel on management network

Task 5 Conduct DR testing: validate inventory consistency, policy application, and workload continuity

Recoverability

After restore, you must validate that the recovered fleet is truly equivalent to the original. This is where DR testing separates the professionals from the lucky.

Step 1

Inventory consistency check. In Fleet Manager, navigate to Inventory > Workload Domains. For Instance A and B, verify: (1) Number of clusters matches original (should be 2+ per instance: 1 management, 1+ workload), (2) Number of hosts per cluster matches original, (3) Storage capacity matches (vSAN capacity should be identical). Create a checklist.

Checklist completed. Example: Instance A Workload Domain has 4 hosts, vSAN capacity 200 GB (matches original). Instance A Management Domain has 3 hosts, vSAN capacity 150 GB (matches original).
Screenshot both the original state (from your pre-disaster snapshot if you reverted and took a screen capture earlier) and the restored state to document consistency.
Step 2

Policy application test. In Fleet Manager, navigate to Lifecycle > Policies. Verify: (1) Certificate policy is active and both instances show 'certificates applied: yes', (2) Patching policy is scheduled and last execution timestamp is consistent with original, (3) Password policy shows 'next rotation: [date]' matching the original schedule.

All policies applied and synchronized across instances
If policies show different states between Instance A and B, or if 'last applied' timestamps are significantly different from pre-disaster state, there may be a policy sync issue. Navigate to each instance's configuration and manually re-apply the policy.
Step 3

Workload continuity test. Log into vCenter (Instance A). Navigate to Inventory > VMs. Count the total number of VMs in the workload domain. Compare to the pre-disaster count (record this from your pre-disaster snapshot or test logs). Verify: (1) VM count is identical, (2) Each VM's power state matches expected state (should all be powered on if they were running at backup time), (3) Storage assignments (datastore names) are consistent.

VM inventory matches pre-disaster state. Example: 'Workload domain has 24 VMs, all powered on, distributed across 3 datastores with expected capacity allocation'
Take a screenshot of the VM inventory for documentation. This is a key artifact for VCDX showing that recovery was successful and complete.
Step 4

NSX network validation. Log into NSX Manager (Instance A). Navigate to Networking > Segments. Verify: (1) Number of segments matches original (should be 3-5 typical segments: management, workload, edge, etc.), (2) Each segment shows 'status: active' and 'transport nodes: connected', (3) Tier-0 and Tier-1 gateways are present and routing is established.

NSX network configuration is consistent and operational. Example: '4 segments present (mgmt, web, app, db), all active, 4 transport nodes connected per segment'
If any segment shows 'transport nodes: disconnected', run the 'Resolve' command in NSX. If issues persist, check: (1) ESXi hosts have NSX VIBs installed (esxcli software vib list | grep nsx), (2) Network connectivity to NSX Manager from hosts (ping mgmt IP)
Step 5

Cross-instance policy sync test. Create a new DFW rule in Instance A: Source = web segment, Destination = app segment, Action = allow. Save. Navigate to Fleet Manager > Policies > Firewall Policy. Verify: (1) Rule appears in fleet policy, (2) DFW policy is marked as 'applied to Instance B', (3) Log into Instance B NSX Manager and verify the same rule is present (it may be auto-propagated if fleet sync is enabled, or require manual application).

New policy created and applied to both instances. Rule is visible in both Instance A and B NSX Managers.
This tests the fleet-level policy propagation — a capability that SDDC Manager in VCF 5.2 never had. Document this as a VCF 9.0 architectural advancement for your VCDX discussion.
Step 6

Execute a workload test. In Instance A workload domain, identify a test VM (or create one if not present). Connect to a workload segment. Attempt to ping a resource in another segment (e.g., app segment). This validates end-to-end network connectivity post-restore.

Ping succeeds, demonstrating network routing and firewall rules are working correctly post-restore
If ping fails, check: (1) NSX segments have correct CIDR assignments, (2) Tier-0 router is active and advertising routes, (3) Default route on test VM points to Tier-0 gateway IP
Step 7

Revert to the pre-disaster snapshot. In vSphere Client, navigate to the Fleet Manager VM. Right-click > Snapshots > Revert to Snapshot 'pre-disaster-recovery-test'. This returns the environment to the original state before the disaster.

Fleet Manager VM reverted to pre-disaster state. All instances return to original configuration.
After this revert, your original Fleet Manager, vCenter, and NSX instances are back online as they were. The backup and restore cycle has been validated successfully.

Validation Gate

Check: All seven validation steps completed and documented: (1) Inventory checklist passed, (2) Policies applied correctly, (3) VM count and state match pre-disaster, (4) NSX segments operational, (5) Cross-instance policy test passed, (6) Workload connectivity test passed, (7) Revert to snapshot successful.

Expected: DR testing complete. Fleet is proven recoverable to a consistent state. All components validated. Environment reverted to original state.

Common Errors

VM count in Instance A matches pre-disaster but several VMs show 'status: unknown' or gray icon
Cause: VM objects were restored but vCenter connectivity to ESXi hosts is incomplete or delayed
Fix: Wait 10-15 minutes for vCenter to re-establish host connections and re-inventory. In vSphere Client, right-click cluster > Refresh. If status persists, SSH to an ESXi host: esxcfg-nas -l (verify NFS mounts are accessible) and ping vCenter IP.
NSX segment shows 'transport nodes: 2/4 connected' — missing nodes
Cause: NSX VIB on some ESXi hosts failed to initialize post-restore or host is isolated
Fix: In NSX Manager, select the affected host transport node and click 'Resolve'. NSX will reinstall the VIB. If still down, SSH to the host and check: systemctl status nsx-firewall
New DFW rule created in Instance A is not visible in Instance B after 30 minutes
Cause: Fleet policy sync is not enabled or is configured to manual-apply mode
Fix: In Fleet Manager Policies, verify 'Auto-propagate changes' is enabled. If disabled, manually apply the policy to Instance B.
Workload VM ping test fails: 'Destination host unreachable'
Cause: Default route on VM is incorrect or NSX Tier-0 is not advertising the route to workload segment
Fix: In the test VM, check: route -n (Linux) or route print (Windows). Verify default gateway is the NSX Tier-0 IP. In NSX Manager, verify Tier-0 is in 'Active' state and advertising the workload segment subnet.
Snapshot revert fails: 'Cannot revert snapshot — VM is not in snapshot tree'
Cause: Snapshot was deleted or corrupted during disaster recovery testing
Fix: If snapshot no longer exists, restore from a prior snapshot or redeploy the environment from backup. For future tests, create a fresh snapshot before each DR test cycle.

Final Validation

The VCF 9.0 fleet backup and recovery workflow is validated end-to-end. All critical management components (Fleet Manager, vCenter, NSX) have been backed up, and a catastrophic failure scenario has been simulated and successfully recovered. The recovered fleet is consistent with the original, with all instances, VMs, policies, and network configurations restored to their pre-disaster state. DR testing demonstrates that the backup strategy is sound and the recovery procedures are repeatable.

✓ Backup targets configured for all 6 components (Fleet Manager, VCF Operations, 2x vCenter, 2x NSX) → All targets show 'Connected' status in Fleet Manager UI

✓ Complete backup cycle executed successfully → Backup MANIFEST file created with 6 backup files, total size 80-150 GB, all checksums validated

✓ Fleet Manager disaster recovery: rebuild from backup successful → Restored Fleet Manager shows all instances registered and policies applied

✓ Instance A recovery: vCenter and NSX restored from backup → vCenter shows original VM inventory, NSX shows original segments and connectivity

✓ Instance B recovery: vCenter and NSX restored from backup → Same consistency checks as Instance A

✓ DR validation tests passed: inventory, policies, VMs, network, cross-instance sync → All seven validation tests completed and documented; environment reverted to pre-disaster state

Cleanup / Restore

Snapshot: vcp-admin-09-post-lab

• Delete disaster recovery test snapshots (keep only clean pre-lab and post-lab snapshots)

• Archive backup MANIFEST and selected backup files to external storage for retention

• Document backup procedures in your lab notebook: backup sequence, backup duration, restore sequence, restore duration, RTO/RPO by component

• Revert to vcp-admin-09-post-lab snapshot for subsequent labs

Design Reflection (VCDX)

A VCDX panelist examining VCF 9.0 backup design will probe deeply: Why has the backup architecture shifted from SDDC Manager centralization (VCF 5.2) to fleet-level orchestration (VCF 9.0)? The answer is multi-instance scalability and regional independence — each instance can be backed up and recovered independently, reducing blast radius and enabling geographically distributed fleets. Panelists will ask: What is the RPO/RTO trade-off? In a 24-hour RPO with 2-4 hour RTO, what are the implications for a customer with strict availability requirements?
How would you design backup/recovery for a 10-instance fleet? Can you articulate why VCF 9.0 moved away from a single SDDC Manager as the backup focal point? Be prepared to discuss: (1) the architectural bottleneck of SDDC Manager as a single point of management (and backup), (2) how fleet-level policy propagation enables consistent recovery across instances, (3) the operational complexity introduced by distributed backups (more targets, more failure points, more coordination), and (4) how to mitigate that complexity with orchestration and automation.

Requirements

  • All critical management components must be backed up: Fleet Manager, vCenter, NSX Manager, and VCF Operations
  • Backup targets must be reachable from all management appliances with authentication and write permissions
  • RPO must be achievable within operational constraints (typically 24 hours for daily backups)
  • RTO for full fleet recovery must be acceptable for business continuity (< 4 hours typical)
  • Backup files must be validated for integrity (checksums) and stored securely
  • Recovery procedures must be documented, tested, and repeatable without manual intervention

Constraints

  • Backup window must not impact production VCF operations — backups should be scheduled outside business hours
  • Network bandwidth limits may force sequential or staggered backups (e.g., vCenter + NSX cannot run simultaneously on a slow WAN)
  • Backup storage is limited — retention policies must balance recoverability (keep multiple backups) with cost (delete old backups)
  • Single backup server is a single point of failure — in production, replicate backups to secondary location or cloud storage
  • Recovery of large components (vCenter, NSX) requires maintenance window — instances are unavailable during recovery
  • Appliance rebuild (Fleet Manager) requires redeployment from OVA, which adds 30 minutes to RTO

Assumptions

  • Backup server has sufficient capacity (500 GB minimum, ideally 1-2 TB for rolling retention)
  • Network connectivity to backup server is stable and has sufficient bandwidth (at least 50 Mbps for SFTP)
  • VCF appliances (Fleet Manager, vCenter, NSX) are deployed with redundancy within each site (HA clusters where applicable)
  • Instances are independently viable — they can operate without Fleet Manager coordination (autonomy design)
  • Disaster is not universal — at least one instance remains partially operational to bootstrap recovery
  • Operator has access to out-of-band management (SSH to appliances) if UI connectivity is lost

Risks

  • SFTP backup server failure: no backups can be created or restored. MITIGATION: replicate backups to secondary server or cloud storage (S3, GCS). Monitor server health with alerts.
  • Backup corruption: backup file appears valid but contains corrupted data, making recovery impossible. MITIGATION: test recovery procedures monthly. Validate file checksums against MANIFEST. Verify backup file integrity: tar -tzf backup.tgz
  • Network isolation: backup server unreachable due to firewall or routing change. MITIGATION: maintain DNS/IP records of backup server in change management. Test connectivity regularly.
  • Incomplete backup: backup completes but component state was mid-transaction at backup time. MITIGATION: use application-level quiescing (vCenter snapshots, NSX consistency checks) before backup. Implement pre-backup validation.
  • Recovery failure: appliance rebuild succeeds but data restore fails due to version mismatch or database corruption. MITIGATION: test recovery with test appliances in parallel environment. Maintain detailed recovery runbooks.
  • Long RTO: recovery of multi-instance fleet takes 4+ hours, violating SLA. MITIGATION: design backup strategy with incremental/differential backups. Pre-stage standby appliances. Implement async replication for critical databases.

Self-Assessment Discussion Prompts

  1. In VCF 9.0, backup of Fleet Manager is separate from backup of managed instances. Why did Broadcom design it this way instead of a monolithic SDDC Manager-like backup (as in VCF 5.2)? What are the trade-offs?
  2. If your backup server has a catastrophic failure and you lose all backups, what is your recovery path? Discuss RTO implications and mitigation strategies.
  3. Design a multi-region VCF 9.0 fleet with 5 instances across 3 geographic regions. How would you architect backup and recovery to minimize RTO for a region-wide disaster while keeping cost reasonable?
  4. VCF 9.0 enables cross-instance policy propagation via Fleet Manager. How does this change the backup requirements compared to VCF 5.2 where instances were largely independent?
  5. If vCenter backup takes 4 hours and your backup window is 8 hours, you can only back up 2 vCenter instances per night. How would you design a staggered backup strategy for a 10-instance fleet to ensure daily backups of all components?
  6. Recovery of Fleet Manager also requires rediscovery of instances and policy re-sync. Describe the dependencies and sequencing: which components must be recovered first, and why?

Extensions

Implement Incremental Backup Strategy for Faster Recovery

Instead of full backups daily, implement a strategy: full backup once per week, incremental backups daily. This reduces backup window and storage, but complicates restore (you must replay multiple incremental backup chains). Document the trade-offs: RPO improves (smaller incremental backups = less data loss if backup fails mid-cycle), but RTO may increase (must apply multiple incremental backups during restore). Test this strategy with a 2-week retention policy.

harder

Design Off-Site Backup Replication

Extend the lab: after each backup completes, replicate the backup files to an off-site location (simulate using a second SFTP server on a different subnet, or use S3/GCS if you have cloud access). Implement automatic replication using rsync or s3sync. Test recovery from the off-site backup to ensure the replicated backups are valid and can be used for DR at a remote location.

harder

Automate Backup Orchestration and Validation

Write a PowerShell script or Python script that orchestrates the complete backup cycle: (1) Check SFTP connectivity, (2) Initiate backups in proper sequence, (3) Monitor progress and alert on failures, (4) Validate backup file integrity, (5) Update MANIFEST file with checksums, (6) Prune old backups per retention policy. This mirrors production backup automation and demonstrates maturity for VCDX.

harder

Test Recovery of a Single vCenter Without Fleet Manager

Advanced scenario: what if Fleet Manager is permanently lost and unrecoverable? Can you restore vCenter directly without Fleet Manager orchestration? Manually restore vCenter from backup using vCenter's native backup/restore tools (not via Fleet Manager). This tests your understanding of appliance-level restore independence and is a critical failsafe skill.

expert

⚠ Known Pitfalls (from Community KB)

Incomplete Backup: Missing Instance Registration COMMON
Problem: Fleet Manager backup completes but does not include instance registration metadata. When restored, Fleet Manager has no knowledge of instances.
Resolution: Verify backup scope in Fleet Manager: Backup Now dialog must have 'Full' scope selected, which includes instance registration data. Check backup content: tar -tzf backup.tgz | grep instance-registry
Backup Server Disk Full After Second Cycle COMMON
Problem: Retention policy not enforced; old backups accumulate and consume all disk space
Resolution: Implement retention: find /backups -name '*.tgz' -mtime +7 -delete. Document retention policy before production rollout.
vCenter Restore Timeout COMMON
Problem: vCenter backup file is 50+ GB, SFTP upload or restore takes longer than default timeout (4 hours)
Resolution: Increase timeout in vCenter Administration > Backup & Restore > Configuration > Advanced > Backup Timeout. Set to 8 hours or implement incremental backups to reduce size.
NSX Transport Nodes Disconnected Post-Restore RESOLVED
Problem: NSX backup restored but transport nodes show 'down' — network VLAN changed since backup
Resolution: Verify network VLAN is same as backup time. If changed, update NSX segment VLAN or revert network to original configuration.
Fleet Manager Restore Fails: Database Incompatibility LIKELY_RESOLVED
Problem: Fleet Manager 9.0.0 backup cannot be restored to Fleet Manager 9.0.2 due to database schema changes
Resolution: Ensure backup and appliance versions match. If upgrading, upgrade Fleet Manager before restoring old backups, or restore to same version then upgrade after.

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.