VCF 9.0 Backup & Restore — SDDC Manager Absence Paradigm
Objectives
- Articulate VCF 9.0 backup architecture: what must be backed up (Fleet Manager, VCF Operations, vCenter, NSX), and why the absence of SDDC Manager changes the approach
- Configure file-based backup targets (SFTP) for all critical management components in a fleet topology
- Execute and validate a complete backup cycle including Fleet Manager, VCF Operations, NSX Manager, and vCenter Server
- Perform a catastrophic recovery: rebuild Fleet Manager appliance from backup and restore full fleet state
- Conduct DR testing: validate post-restore inventory consistency, policy application, and workload continuity across instances
- Compare VCF 9.0 fleet-based backup to VCF 5.2 SDDC Manager backup and document operational differences
Prerequisites
Multi-instance VCF 9.0 lab environment deployed and operational. Fleet Manager active and managing 2+ instances. All management appliances (vCenter, NSX, VCF Operations) healthy with network connectivity to backup target SFTP server. Minimum 500 GB available storage on backup target.
Prior labs: holodeck-02
Required skills:
- VCF 9.0 architecture: Fleet Manager, instances, management domains
- SFTP/file-based backup concepts and SSH key authentication
- vCenter Server backup and restore procedures
- NSX Manager backup and restore procedures
- SSH CLI access to appliances
- VCF Operations (formerly VCF Manager) console navigation
- Understanding of DR testing vs disaster recovery
Lab Environment
Multi-instance VCF 9.0 fleet with: (1) Fleet Manager VM managing instances, (2) Instance A (primary site) with management domain + workload domain, (3) Instance B (secondary site) with management domain + workload domain, (4) NSX Manager deployed within each instance, (5) vCenter Server deployed within each instance, (6) VCF Operations (embedded or standalone) monitoring both instances. SFTP backup server accessible from all management appliances.
graph TB FM[Fleet Manager 10.0.0.50] SFTP[SFTP Backup Server<br/>192.168.1.100] INST1[Instance A<br/>Site Primary] INST2[Instance B<br/>Site Secondary] VOP[VCF Operations<br/>10.0.0.51] VC1[vCenter<br/>Instance A] NSX1[NSX Manager<br/>Instance A] VC2[vCenter<br/>Instance B] NSX2[NSX Manager<br/>Instance B] FM --> SFTP VOP --> SFTP INST1 --> VC1 INST1 --> NSX1 INST2 --> VC2 INST2 --> NSX2 VC1 --> SFTP NSX1 --> SFTP VC2 --> SFTP NSX2 --> SFTP
IP Addressing
| Network | Purpose | VLAN |
|---|---|---|
10.0.0.0/24 | Management network (Fleet Manager, VCF Operations, instance management domains) | VLAN 1644 |
10.1.0.0/16 | Instance A workload network (cross-instance routing) | VLAN 1650+ |
10.2.0.0/16 | Instance B workload network (cross-instance routing) | VLAN 1650+ |
192.168.1.0/24 | Backup SFTP server and out-of-band management | VLAN 100 |
Credentials
| System | Username | Password |
|---|---|---|
| Fleet Manager | admin | Fleet Manager admin password set during deployment |
| VCF Operations | administrator@vsphere.local | Same as Instance A vCenter SSO password |
| vCenter Server (Instance A) | administrator@vsphere.local | Instance A SSO password |
| vCenter Server (Instance B) | administrator@vsphere.local | Instance B SSO password |
| NSX Manager (Instance A) | admin | Instance A NSX admin password |
| NSX Manager (Instance B) | admin | Instance B NSX admin password |
| SFTP Backup Server | vcf-backup | SFTP user credential for file deposit |
Tasks
Task 1 Map VCF 9.0 backup architecture and compare to VCF 5.2 SDDC Manager model
RecoverabilityIn VCF 9.0, the backup surface has changed fundamentally. SDDC Manager no longer exists. Fleet Manager and VCF Operations have replaced it. Understanding what to back up and in what order is a VCDX-critical skill.
Log into Fleet Manager UI (https://<fleet-manager-ip>). Navigate to Lifecycle > Backup & Recovery to view the backup topology diagram.
In Fleet Manager UI, navigate to Inventory > Managed Instances. Document the FQDN, IP address, and instance type (management/workload) for each instance. Create a spreadsheet with columns: Instance Name, Role, FQDN, Management IP, vCenter IP, NSX Manager IP, Backup Priority.
Compare to VCF 5.2: if you have access to a VCF 5.2 lab environment (or documentation), document how SDDC Manager backup differed. Key differences: (a) SDDC Manager was a single point for all instance state, (b) vCenter and NSX backup were optional, (c) no fleet-level policy backup. For VCDX, be prepared to articulate WHY VCF 9.0 moved to this distributed model (hint: multi-instance scalability, regional independence, reduced blast radius).
Review the VCF 9.0 backup requirements documentation: what are the RPO (Recovery Point Objective) and RTO (Recovery Time Objective) implications for each component? Document RTO for: (a) Fleet Manager alone, (b) Instance vCenter, (c) Instance NSX, (d) Full fleet recovery. Note that Fleet Manager recovery < 1 hour (appliance rebuild), but instance recovery depends on backup freshness.
Document the critical data that MUST be backed up (Fleet Manager configuration, instance registration, policy definitions) vs optional data (vCenter event logs, NSX flow logs). This distinction determines backup size and retention strategy.
Validation Gate
Check: Comparison table complete with all instances mapped + RTO/RPO documented + critical data distinction made
Expected: You can articulate the backup topology and explain why each component is backed up independently (unlike SDDC Manager which centralized all state)
Common Errors
Task 2 Configure file-based backup targets (SFTP) for all management components
RecoverabilityIn production, you must configure backup targets for Fleet Manager, VCF Operations, vCenter, and NSX independently. This task walks through the configuration of each.
Prepare the SFTP backup server. Verify connectivity from Fleet Manager: ssh vcf-backup@192.168.1.100. Create backup directories: mkdir -p /backups/fleet-manager /backups/vcf-operations /backups/instance-a-vcenter /backups/instance-a-nsx /backups/instance-b-vcenter /backups/instance-b-nsx. Set permissions: chmod 755 /backups && chmod 755 /backups/*
In Fleet Manager UI, navigate to Lifecycle > Backup & Recovery > Backup Configuration. Add a new backup target: Protocol: SFTP, Host: 192.168.1.100, Port: 22, Username: vcf-backup, Password/Key: [use SSH key method], Remote Path: /backups/fleet-manager. Click 'Validate Connection'.
Configure VCF Operations backup. Log into VCF Operations UI (https://<vcf-operations-ip>). Navigate to Settings > Backup & Restore. Add backup target: SFTP, Host: 192.168.1.100, Path: /backups/vcf-operations. Validate.
Configure vCenter backup for Instance A. Log into vCenter (Instance A) UI. Navigate to Administration > Backup & Restore. Add backup location: Protocol: SFTP, Server: 192.168.1.100, Username: vcf-backup, Directory: /backups/instance-a-vcenter. Test the connection.
Configure NSX Manager backup for Instance A. Log into NSX Manager (Instance A) UI. Navigate to System > Backup & Restore > Backups. Configure backup: Protocol: SFTP, Server: 192.168.1.100, Username: vcf-backup, Directory: /backups/instance-a-nsx. Click 'Test Backup'.
Repeat Steps 4 and 5 for Instance B (vCenter and NSX of Instance B), using paths /backups/instance-b-vcenter and /backups/instance-b-nsx respectively.
Validation Gate
Check: Fleet Manager UI shows 6 backup targets all in 'Connected' state. SSH to SFTP server: ls -la /backups/ shows all directories with vcf-backup ownership.
Expected: All management components have validated SFTP backup targets. Backup server has write access to all directories.
Common Errors
Task 3 Execute a complete backup cycle: Fleet Manager, VCF Operations, vCenter, and NSX
RecoverabilityBacking up is straightforward if targets are configured. The critical learning is understanding the dependency order, backup size, and duration.
Start with Fleet Manager backup (has fewest dependencies). In Fleet Manager UI, navigate to Lifecycle > Backup & Recovery > Backup Now. Select Backup Scope: 'Full' (includes instance registration data, fleet policies, certificates). Click 'Initiate Backup'. Monitor progress.
Initiate VCF Operations backup. In VCF Operations UI, navigate to Backup & Restore > Backup Now. Select backup level: 'Full' (includes database, analytics, compliance snapshots). Click Start. Monitor progress.
Initiate vCenter backup for Instance A. In vCenter UI, navigate to Administration > Backup & Restore > Backup Now. Backup type: 'Full' (includes database, events, tasks, configuration). Click Start.
Initiate NSX Manager backup for Instance A (do not wait for vCenter to finish; these run in parallel). In NSX Manager UI, navigate to System > Backup & Restore > Backups > Create Backup. File name prefix: 'nsx-a-2024-04-17'. Click Create.
Once Instance A backups are complete (vCenter and NSX), initiate backups for Instance B. Repeat Steps 3-4 for Instance B vCenter and NSX, using directory paths and hostnames specific to Instance B.
Monitor the SFTP backup server to confirm all backups have been deposited. SSH: ls -lah /backups/ and df -h /backups. Document total backup size, number of files, and oldest vs newest backup timestamps.
Document the backup manifest. Create a file /backups/MANIFEST-2024-04-17.txt with entries for each backup: Component, Timestamp, Filename, Size, Checksum (md5sum), Notes (e.g., 'Full backup before storage expansion'). This becomes your recovery procedure reference.
Validation Gate
Check: All six backups completed and verified on SFTP server. MANIFEST file created with checksums. Backup server has sufficient remaining space (> 100 GB free).
Expected: Complete fleet backup cycle finished. All components backed up. Ready for restore testing.
Common Errors
Task 4 Perform disaster recovery: rebuild Fleet Manager from backup and restore instance state
RecoverabilityBackups are only valuable if you can restore them. This task simulates a catastrophic Fleet Manager failure and walks through the recovery procedure.
Create a 'disaster' snapshot of the current state. In the physical ESXi host, take a snapshot of the Fleet Manager VM: Right-click Fleet Manager VM > Snapshots > Take Snapshot, Name: 'pre-disaster-recovery-test', Description: 'Snapshot before DR test — revert here if recovery fails'.
Simulate catastrophic Fleet Manager failure: In vSphere Client, power off the Fleet Manager VM. Wait 30 seconds. Then delete the Fleet Manager VM's virtual disk. This simulates a storage failure where the appliance is unrecoverable.
Deploy a new Fleet Manager appliance OVA (using the same OVA as the original deployment). Complete the initial deployment wizard: hostname, IP address (use the same IP as the original Fleet Manager: 10.0.0.50), domain, NTP, DNS. Do NOT join it to any existing configuration — this is a blank appliance.
Restore Fleet Manager from backup. Log into the new Fleet Manager UI as admin. Navigate to Lifecycle > Backup & Recovery > Restore. Select backup source: SFTP, Server: 192.168.1.100, Path: /backups/fleet-manager. Select the latest backup file (2024-04-17-FM-001.tgz). Click Restore.
Wait for Fleet Manager restore to complete. Web UI returns. Log in as admin. Verify: (1) navigate to Inventory > Managed Instances — both Instance A and Instance B should be listed and should show 'status: registered', (2) Navigate to Lifecycle > Policies — all fleet policies (certificates, passwords, patching schedules) should be restored.
Restore vCenter for Instance A. Log into the new Fleet Manager. Navigate to Lifecycle > Backup & Restore > Restore Instance. Select Instance A. Choose Component: vCenter. Backup source: SFTP, Directory: /backups/instance-a-vcenter. Select the latest vCenter backup. Click Restore.
Restore NSX Manager for Instance A. In Fleet Manager, navigate to Lifecycle > Backup & Restore > Restore Instance. Select Instance A. Choose Component: NSX Manager. Backup source: /backups/instance-a-nsx. Select latest NSX backup. Click Restore.
Repeat Steps 6-7 for Instance B (vCenter and NSX). This is done in parallel with Instance A restore or sequentially, depending on capacity.
Validation Gate
Check: Fleet Manager: (1) All instances registered, (2) All policies present. vCenter A & B: (1) Inventory shows all workload domains and VMs, (2) vSAN cluster healthy, (3) Datastore capacity matches original. NSX A & B: (1) All transport nodes connected, (2) All segments present, (3) DFW rule count matches original.
Expected: Full fleet restored to the state at backup time. All management and workload components operational. Inventory and policies consistent.
Common Errors
Task 5 Conduct DR testing: validate inventory consistency, policy application, and workload continuity
RecoverabilityAfter restore, you must validate that the recovered fleet is truly equivalent to the original. This is where DR testing separates the professionals from the lucky.
Inventory consistency check. In Fleet Manager, navigate to Inventory > Workload Domains. For Instance A and B, verify: (1) Number of clusters matches original (should be 2+ per instance: 1 management, 1+ workload), (2) Number of hosts per cluster matches original, (3) Storage capacity matches (vSAN capacity should be identical). Create a checklist.
Policy application test. In Fleet Manager, navigate to Lifecycle > Policies. Verify: (1) Certificate policy is active and both instances show 'certificates applied: yes', (2) Patching policy is scheduled and last execution timestamp is consistent with original, (3) Password policy shows 'next rotation: [date]' matching the original schedule.
Workload continuity test. Log into vCenter (Instance A). Navigate to Inventory > VMs. Count the total number of VMs in the workload domain. Compare to the pre-disaster count (record this from your pre-disaster snapshot or test logs). Verify: (1) VM count is identical, (2) Each VM's power state matches expected state (should all be powered on if they were running at backup time), (3) Storage assignments (datastore names) are consistent.
NSX network validation. Log into NSX Manager (Instance A). Navigate to Networking > Segments. Verify: (1) Number of segments matches original (should be 3-5 typical segments: management, workload, edge, etc.), (2) Each segment shows 'status: active' and 'transport nodes: connected', (3) Tier-0 and Tier-1 gateways are present and routing is established.
Cross-instance policy sync test. Create a new DFW rule in Instance A: Source = web segment, Destination = app segment, Action = allow. Save. Navigate to Fleet Manager > Policies > Firewall Policy. Verify: (1) Rule appears in fleet policy, (2) DFW policy is marked as 'applied to Instance B', (3) Log into Instance B NSX Manager and verify the same rule is present (it may be auto-propagated if fleet sync is enabled, or require manual application).
Execute a workload test. In Instance A workload domain, identify a test VM (or create one if not present). Connect to a workload segment. Attempt to ping a resource in another segment (e.g., app segment). This validates end-to-end network connectivity post-restore.
Revert to the pre-disaster snapshot. In vSphere Client, navigate to the Fleet Manager VM. Right-click > Snapshots > Revert to Snapshot 'pre-disaster-recovery-test'. This returns the environment to the original state before the disaster.
Validation Gate
Check: All seven validation steps completed and documented: (1) Inventory checklist passed, (2) Policies applied correctly, (3) VM count and state match pre-disaster, (4) NSX segments operational, (5) Cross-instance policy test passed, (6) Workload connectivity test passed, (7) Revert to snapshot successful.
Expected: DR testing complete. Fleet is proven recoverable to a consistent state. All components validated. Environment reverted to original state.
Common Errors
Final Validation
The VCF 9.0 fleet backup and recovery workflow is validated end-to-end. All critical management components (Fleet Manager, vCenter, NSX) have been backed up, and a catastrophic failure scenario has been simulated and successfully recovered. The recovered fleet is consistent with the original, with all instances, VMs, policies, and network configurations restored to their pre-disaster state. DR testing demonstrates that the backup strategy is sound and the recovery procedures are repeatable.
✓ Backup targets configured for all 6 components (Fleet Manager, VCF Operations, 2x vCenter, 2x NSX) → All targets show 'Connected' status in Fleet Manager UI
✓ Complete backup cycle executed successfully → Backup MANIFEST file created with 6 backup files, total size 80-150 GB, all checksums validated
✓ Fleet Manager disaster recovery: rebuild from backup successful → Restored Fleet Manager shows all instances registered and policies applied
✓ Instance A recovery: vCenter and NSX restored from backup → vCenter shows original VM inventory, NSX shows original segments and connectivity
✓ Instance B recovery: vCenter and NSX restored from backup → Same consistency checks as Instance A
✓ DR validation tests passed: inventory, policies, VMs, network, cross-instance sync → All seven validation tests completed and documented; environment reverted to pre-disaster state
Cleanup / Restore
Snapshot: vcp-admin-09-post-lab
• Delete disaster recovery test snapshots (keep only clean pre-lab and post-lab snapshots)
• Archive backup MANIFEST and selected backup files to external storage for retention
• Document backup procedures in your lab notebook: backup sequence, backup duration, restore sequence, restore duration, RTO/RPO by component
• Revert to vcp-admin-09-post-lab snapshot for subsequent labs
Design Reflection (VCDX)
A VCDX panelist examining VCF 9.0 backup design will probe deeply: Why has the backup architecture shifted from SDDC Manager centralization (VCF 5.2) to fleet-level orchestration (VCF 9.0)? The answer is multi-instance scalability and regional independence — each instance can be backed up and recovered independently, reducing blast radius and enabling geographically distributed fleets. Panelists will ask: What is the RPO/RTO trade-off? In a 24-hour RPO with 2-4 hour RTO, what are the implications for a customer with strict availability requirements?
How would you design backup/recovery for a 10-instance fleet? Can you articulate why VCF 9.0 moved away from a single SDDC Manager as the backup focal point? Be prepared to discuss: (1) the architectural bottleneck of SDDC Manager as a single point of management (and backup), (2) how fleet-level policy propagation enables consistent recovery across instances, (3) the operational complexity introduced by distributed backups (more targets, more failure points, more coordination), and (4) how to mitigate that complexity with orchestration and automation.
Requirements
- All critical management components must be backed up: Fleet Manager, vCenter, NSX Manager, and VCF Operations
- Backup targets must be reachable from all management appliances with authentication and write permissions
- RPO must be achievable within operational constraints (typically 24 hours for daily backups)
- RTO for full fleet recovery must be acceptable for business continuity (< 4 hours typical)
- Backup files must be validated for integrity (checksums) and stored securely
- Recovery procedures must be documented, tested, and repeatable without manual intervention
Constraints
- Backup window must not impact production VCF operations — backups should be scheduled outside business hours
- Network bandwidth limits may force sequential or staggered backups (e.g., vCenter + NSX cannot run simultaneously on a slow WAN)
- Backup storage is limited — retention policies must balance recoverability (keep multiple backups) with cost (delete old backups)
- Single backup server is a single point of failure — in production, replicate backups to secondary location or cloud storage
- Recovery of large components (vCenter, NSX) requires maintenance window — instances are unavailable during recovery
- Appliance rebuild (Fleet Manager) requires redeployment from OVA, which adds 30 minutes to RTO
Assumptions
- Backup server has sufficient capacity (500 GB minimum, ideally 1-2 TB for rolling retention)
- Network connectivity to backup server is stable and has sufficient bandwidth (at least 50 Mbps for SFTP)
- VCF appliances (Fleet Manager, vCenter, NSX) are deployed with redundancy within each site (HA clusters where applicable)
- Instances are independently viable — they can operate without Fleet Manager coordination (autonomy design)
- Disaster is not universal — at least one instance remains partially operational to bootstrap recovery
- Operator has access to out-of-band management (SSH to appliances) if UI connectivity is lost
Risks
- SFTP backup server failure: no backups can be created or restored. MITIGATION: replicate backups to secondary server or cloud storage (S3, GCS). Monitor server health with alerts.
- Backup corruption: backup file appears valid but contains corrupted data, making recovery impossible. MITIGATION: test recovery procedures monthly. Validate file checksums against MANIFEST. Verify backup file integrity: tar -tzf backup.tgz
- Network isolation: backup server unreachable due to firewall or routing change. MITIGATION: maintain DNS/IP records of backup server in change management. Test connectivity regularly.
- Incomplete backup: backup completes but component state was mid-transaction at backup time. MITIGATION: use application-level quiescing (vCenter snapshots, NSX consistency checks) before backup. Implement pre-backup validation.
- Recovery failure: appliance rebuild succeeds but data restore fails due to version mismatch or database corruption. MITIGATION: test recovery with test appliances in parallel environment. Maintain detailed recovery runbooks.
- Long RTO: recovery of multi-instance fleet takes 4+ hours, violating SLA. MITIGATION: design backup strategy with incremental/differential backups. Pre-stage standby appliances. Implement async replication for critical databases.
Self-Assessment Discussion Prompts
- In VCF 9.0, backup of Fleet Manager is separate from backup of managed instances. Why did Broadcom design it this way instead of a monolithic SDDC Manager-like backup (as in VCF 5.2)? What are the trade-offs?
- If your backup server has a catastrophic failure and you lose all backups, what is your recovery path? Discuss RTO implications and mitigation strategies.
- Design a multi-region VCF 9.0 fleet with 5 instances across 3 geographic regions. How would you architect backup and recovery to minimize RTO for a region-wide disaster while keeping cost reasonable?
- VCF 9.0 enables cross-instance policy propagation via Fleet Manager. How does this change the backup requirements compared to VCF 5.2 where instances were largely independent?
- If vCenter backup takes 4 hours and your backup window is 8 hours, you can only back up 2 vCenter instances per night. How would you design a staggered backup strategy for a 10-instance fleet to ensure daily backups of all components?
- Recovery of Fleet Manager also requires rediscovery of instances and policy re-sync. Describe the dependencies and sequencing: which components must be recovered first, and why?
Extensions
Implement Incremental Backup Strategy for Faster Recovery
Instead of full backups daily, implement a strategy: full backup once per week, incremental backups daily. This reduces backup window and storage, but complicates restore (you must replay multiple incremental backup chains). Document the trade-offs: RPO improves (smaller incremental backups = less data loss if backup fails mid-cycle), but RTO may increase (must apply multiple incremental backups during restore). Test this strategy with a 2-week retention policy.
harderDesign Off-Site Backup Replication
Extend the lab: after each backup completes, replicate the backup files to an off-site location (simulate using a second SFTP server on a different subnet, or use S3/GCS if you have cloud access). Implement automatic replication using rsync or s3sync. Test recovery from the off-site backup to ensure the replicated backups are valid and can be used for DR at a remote location.
harderAutomate Backup Orchestration and Validation
Write a PowerShell script or Python script that orchestrates the complete backup cycle: (1) Check SFTP connectivity, (2) Initiate backups in proper sequence, (3) Monitor progress and alert on failures, (4) Validate backup file integrity, (5) Update MANIFEST file with checksums, (6) Prune old backups per retention policy. This mirrors production backup automation and demonstrates maturity for VCDX.
harderTest Recovery of a Single vCenter Without Fleet Manager
Advanced scenario: what if Fleet Manager is permanently lost and unrecoverable? Can you restore vCenter directly without Fleet Manager orchestration? Manually restore vCenter from backup using vCenter's native backup/restore tools (not via Fleet Manager). This tests your understanding of appliance-level restore independence and is a critical failsafe skill.
expert⚠ Known Pitfalls (from Community KB)
References
- VMware Cloud Foundation 9.0 Backup and Recovery GuideTier 1 — Official
Official Broadcom documentation for VCF 9.0 backup/restore procedures, RPO/RTO guidelines, and component-specific backup steps - VMware Cloud Foundation 5.2 SDDC Manager Backup Guide (for historical comparison)Tier 1 — Official
VCF 5.2 documentation showing the monolithic SDDC Manager backup model that has been replaced in VCF 9.0 by fleet-level orchestration - vCenter Server 8.0 Backup and Recovery ProceduresTier 1 — Official
Deep dive into vCenter appliance backup/restore, VCSA backup options, and post-restore validation - NSX Manager 9.0 Backup and RestoreTier 1 — Official
NSX-specific backup procedures including SFTP configuration, backup testing, and recovery best practices - Cormac Hogan — VCF 9.0 Backup Topology and Fleet OperationsTier 3 — Expert Blog
Expert blog covering VCF 9.0 fleet-level backup architecture, failure scenarios, and operational maturity practices - William Lam — Automating VCF 9.0 Backup OperationsTier 3 — Expert Blog
Automation and scripting approaches for VCF 9.0 backup orchestration, reducing manual intervention