Site Protection and Disaster Recovery for VMware Cloud Foundation
This is the VCF 9.0 evolution of the VCF 5.2 Site Protection and Disaster Recovery validated solution. Key changes from VCF 5.2 → 9.0 include: consolidation of VMware Live Site Recovery and vSphere Replication into the unified VMware Live Recovery (VLR) appliance, native VCF 9.0 workload domain integration, updated support for vSAN ESA at the recovery site, tighter NSX integration for recovery networks, and new Planning Workbook templates aligned with the VCF 9.0 topology.
Key Components: VMware Live Recovery, vSphere Replication, NSX, SDDC Manager, vCenter, VCF Operations
Purpose: Instructions for adapting a dual-instance SDDC on top of VCF to provide disaster recovery of SDDC management components using VMware Live Recovery for VCF Automation, VCF Operations, and VCF Operations fleet management.
32 design decisions
| DD-ID | Decision | Quality |
|---|---|---|
| SPR-VLR-CFG-001 | Deploy VLR instance in protected and recovery VCF instances. | Recoverability |
Decision: Deploy VLR instance in protected and recovery VCF instances. Rationale: Allows orchestrated recovery of VCF mgmt components in another VCF instance. Implication: No significant trade-offs identified for this decision. Component: VMware Live Recovery | ||
| SPR-VLR-CFG-002 | Deploy each VLR in the management domain. | Manageability |
Decision: Deploy each VLR in the management domain. Rationale: Consistent deployment model for all mgmt applications. Implication: No significant trade-offs identified for this decision. Component: VMware Live Recovery | ||
| SPR-VLR-CFG-003 | Deploy VLR with highest scale level supported. | Manageability |
Decision: Deploy VLR with highest scale level supported. Rationale: Highest protection capacity for mgmt components. Implication: No significant trade-offs identified for this decision. Component: VMware Live Recovery | ||
| SPR-VLR-CFG-004 | Use vSphere Replication in VLR as VM replication protection. | Manageability |
Decision: Use vSphere Replication in VLR as VM replication protection. Rationale: Flexibility in storage/vendor; minimizes SRA compatibility overhead. Implication: All mgmt components must be in same cluster; reduced total VM count vs storage-based replication. Component: VMware Live Recovery | ||
| SPR-VLR-CFG-005 | Use VLR to automate recovery of: VCF Operations, VCF Operations fleet mgmt, VCF Operations for logs, | ManageabilityRecoverability |
Decision: Use VLR to automate recovery of: VCF Operations, VCF Operations fleet mgmt, VCF Operations for logs, VCF Operations for networks. Rationale: Automated runbook; RTO <= 4 hours. Implication: No significant trade-offs identified for this decision. Component: VMware Live Recovery | ||
| SPR-VLR-CFG-006 | Configure PIT instances keeping 3 copies over 24-hour period. | Recoverability |
Decision: Configure PIT instances keeping 3 copies over 24-hour period. Rationale: Application integrity after DR recovery event. Implication: More PIT = more disk usage on vSAN datastore. Component: VMware Live Recovery | ||
| SPR-VLR-NET-001 | Place VLR on the management network. | Manageability |
Decision: Place VLR on the management network. Rationale: Co-locate with VCF components. Implication: No significant trade-offs identified for this decision. Component: VMware Live Recovery | ||
| SPR-VLR-NET-002 | Allocate static IP to VLR instances. | Manageability |
Decision: Allocate static IP to VLR instances. Rationale: Removes DHCP constraints/risks on management networks. Implication: Precise IP management. Component: VMware Live Recovery | ||
| SPR-VLR-NET-003 | Configure A + PTR DNS records for each VLR. | ManageabilitySecurity |
Decision: Configure A + PTR DNS records for each VLR. Rationale: FQDN accessibility. Implication: DNS infrastructure required; manage DNS records; firewalls allow DNS. Component: VMware Live Recovery | ||
| SPR-VLR-NET-004 | Configure VLR to use NTP servers (not VMTools). | Security |
Decision: Configure VLR to use NTP servers (not VMTools). Rationale: Accurate time sync; prevent mismatch. Implication: NTP services required; firewalls allow NTP. Component: VMware Live Recovery | ||
| SPR-VLR-NET-005 | Use dedicated vDS portgroup for VLR traffic. | Security |
Decision: Use dedicated vDS portgroup for VLR traffic. Rationale: Isolates VLR traffic from other system traffic. Implication: Manually create portgroup. Component: VMware Live Recovery | ||
| SPR-VLR-NET-006 | Use dedicated VLAN for VLR traffic. | Security |
Decision: Use dedicated VLAN for VLR traffic. Rationale: Traffic isolation. Implication: Manually create VLAN. Component: VMware Live Recovery | ||
| SPR-VLR-NET-007 | Use dedicated VMkernel interface on each ESX for VLR traffic. | Security |
Decision: Use dedicated VMkernel interface on each ESX for VLR traffic. Rationale: Traffic isolation. Implication: Manually create VMkernel interface. Component: VMware Live Recovery | ||
| SPR-VLR-NET-008 | Assign static IP to dedicated VMkernel for VLR traffic. | Manageability |
Decision: Assign static IP to dedicated VMkernel for VLR traffic. Rationale: Removes DHCP risks. Implication: Precise IP management. Component: VMware Live Recovery | ||
| SPR-VLR-LCM-001 | Lifecycle manage VLR via native appliance tools. | Manageability |
Decision: Lifecycle manage VLR via native appliance tools. Rationale: VLR not managed by SDDC Manager. Implication: Deployment, patching, updates, upgrades without native automation. Component: VMware Live Recovery | ||
| SPR-VLR-SEC-001 | Configure service account in each vCenter for app-to-app communication VLR > vSphere, in SSO Adminis | Manageability |
Decision: Configure service account in each vCenter for app-to-app communication VLR > vSphere, in SSO Administrators group. Rationale: VLSR DR orchestration + site pairing; VR replication; accountability; limited blast radius. Implication: Maintain service account lifecycle outside VCF. Component: VMware Live Recovery | ||
| SPR-VLR-SEC-003 | Configure password expiration policy for VLR appliance. | Manageability |
Decision: Configure password expiration policy for VLR appliance. Rationale: Align with org/compliance policies (local users only). Implication: Manage via appliance console or SSH. Component: VMware Live Recovery | ||
| SPR-VLR-SEC-004 | Configure password complexity policy for VLR appliance. | Manageability |
Decision: Configure password complexity policy for VLR appliance. Rationale: Align with org/compliance policies (local users only). Implication: Manage via console or SSH. Component: VMware Live Recovery | ||
| SPR-VLR-SEC-005 | Configure account lockout policy for VLR appliance. | Manageability |
Decision: Configure account lockout policy for VLR appliance. Rationale: Align with org/compliance policies (local users only). Implication: Manage via console or SSH. Component: VMware Live Recovery | ||
| SPR-VLR-SEC-006 | Change VLR root password on recurring/event-initiated schedule. | Manageability |
Decision: Change VLR root password on recurring/event-initiated schedule. Rationale: Root passwords never expire by default. Implication: Routine password changes required. Component: VMware Live Recovery | ||
| SPR-VLR-SEC-007 | Change VLR admin password on recurring/event-initiated schedule. | Manageability |
Decision: Change VLR admin password on recurring/event-initiated schedule. Rationale: Admin passwords never expire by default. Implication: Routine password changes required. Component: VMware Live Recovery | ||
| SPR-VLR-SEC-008 | Replace default self-signed cert with CA-signed in each VLR. | Security |
Decision: Replace default self-signed cert with CA-signed in each VLR. Rationale: All externally facing Web UI + cross-product communication encrypted. Implication: Must have PKI access. Component: VMware Live Recovery | ||
| SPR-VLR-RP-001 | Use prioritized startup order for VCF Operations + VCF Operations FM appliance nodes. | ManageabilityRecoverability |
Decision: Use prioritized startup order for VCF Operations + VCF Operations FM appliance nodes. Rationale: Individual VCF Ops nodes started in order for operational monitoring. Implication: Maintain customized recovery plan if node count increases. Component: VMware Live Recovery | ||
| SPR-VLR-RP-002 | Use prioritized startup order for VCF Operations for logs nodes. | ManageabilityRecoverability |
Decision: Use prioritized startup order for VCF Operations for logs nodes. Rationale: Cloud automation services restored; restored within 4h. Implication: VMware Tools on each VCF Ops Logs node. Component: VMware Live Recovery | ||
| SPR-VLR-RP-003 | Use prioritized startup order for VCF Operations for networks nodes. | ManageabilityRecoverability |
Decision: Use prioritized startup order for VCF Operations for networks nodes. Rationale: Cloud automation services restored; restored within 4h. Implication: VMware Tools on each VCF Ops Networks node. Component: VMware Live Recovery | ||
| SPR-VLR-RP-005 | Do NOT run test recovery of mgmt components. | PerformanceRecoverabilitySecurity |
Decision: Do NOT run test recovery of mgmt components. Rationale: Some protected apps (e.g., clustered) don't function correctly in isolated test. Implication: Cannot test DR; perform planned migrations to validate instead. Component: VMware Live Recovery | ||
| SPR-DNS-NET-001 | With multiple VCF instances, configure DNS settings for each protected component to use DNS servers | Availability |
Decision: With multiple VCF instances, configure DNS settings for each protected component to use DNS servers across all VCF instances. Rationale: Higher DNS availability and resilience during DR/migration. Implication: As you scale from single to multiple instances, update DNS settings on each protected component. Component: DNS | ||
| SPR-NTP-NET-001 | With multiple VCF instances, configure NTP settings across all instances. | Availability |
Decision: With multiple VCF instances, configure NTP settings across all instances. Rationale: Higher NTP availability during DR/migration. Implication: Update NTP settings as instances scale. Component: NTP | ||
| SPR-VCFO-CFG-001 | Install VLR management pack for VCF Operations. | Manageability |
Decision: Install VLR management pack for VCF Operations. Rationale: Establishes communication between VCF Ops and VLR endpoints. Implication: Manual management pack installation. Component: VCFO | ||
| SPR-VCFO-CFG-002 | Configure VLR endpoints to use VCF Operations collector/collector group for the VCF Instance. | Manageability |
Decision: Configure VLR endpoints to use VCF Operations collector/collector group for the VCF Instance. Rationale: Offloads data collection from analytics cluster. Implication: No significant trade-offs identified for this decision. Component: VCFO | ||
| SPR-VLR-LOG-001 | When using VLR, install VCF Operations for logs agent on VLR appliance. | Manageability |
Decision: When using VLR, install VCF Operations for logs agent on VLR appliance. Rationale: Simplifies configuration via agent-based log forwarding. Implication: Must configure VCF Ops for logs as aggregation target. Component: VMware Live Recovery | ||
| SPR-VLR-LOG-002 | Configure VCF Operations for logs as central log aggregation target for VLR. | Manageability |
Decision: Configure VCF Operations for logs as central log aggregation target for VLR. Rationale: Ensures transmission of logs to centralized system. Implication: Configuration required per VCF component. Component: VMware Live Recovery | ||
Prerequisites
- VCF management domain operational
- DNS/NTP configured
- CA infrastructure in place
Implementation Procedure
Implementation
Environment: VCF version in support matrix; per Before You Apply; Site Protection and Disaster Recovery tab of Planning Workbook; VCF healthy.
DNS: Required DNS entries in forward and reverse zones
Network: IP mobility between VCF instances; jumbo frames; L3 routing; max 150ms latency; bandwidth sized with VLR Calculator
Software: VLR .iso mounted
License: VLR license with quantity per design
- Active Directory: AD Domain Controllers, service accounts, security groups
- CA: CA available
- Reconfigure DNS NTP On VCF Management Components
VCF Operations Fleet Mgmt Appliance
DNSSSH to fleet mgmt node as root. Edit /etc/systemd/network/10-eth0.network - add DNS entry. Edit /etc/systemd/resolved.conf - modify Domains entry (space separated).
NTPVCF Ops interface > Fleet Management > Lifecycle > Settings > NTP Servers > Add NTP Server or Edit Server Selection
VCF Operations Nodes
DNSSSH to primary node as root. Edit /etc/systemd/network/10-eth0.network and /etc/systemd/resolved.conf (same as fleet mgmt). Repeat for all cluster nodes + VCF Ops Logs + VCF Ops Networks.
NTPVCF Ops > Administration > Control Panel > Cluster Management > Actions > Network Time Protocol Settings > Add NTP
VCF Operations For Logs
NTPSSH to primary logs node as root. Edit /etc/ntp.conf - add 'server ntp.lax.rainpole.io iburst'. Repeat for each.
VCF Operations For Networks
NTPSSH to primary networks node as support user. sudo vi /etc/ntp.conf - add server. Repeat for each.
VCF Automation
DNSTake backup first (per VCF Fleet Management). VCF Ops > Fleet Management > Lifecycle > VCF Management > Components > automation > ellipsis > Update DNS Configuration > Add Server, Change Priority, Run Precheck, Finish.
NTPSame path > Update NTP Configuration > Add, Priority, Precheck, Finish
VCF Identity Broker
DNSTake backup first. VCF Ops > Fleet Management > Lifecycle > VCF Management > Components > identity broker > ellipsis > Update DNS Configuration
NTPSame > Update NTP Configuration
Deploy VLR Appliance
- MethodDeploy OVF template in vSphere Client for mgmt domain
- PlacementDefault mgmt cluster
- StorageMgmt domain vSAN datastore
- NetworkManagement network
Day-2 Operations Tasks
Operations
As neededOperational Verification
As neededEndpointProtected vCenter > VMware Live Site Recovery
As neededCertificate Management
As neededIn multi-instance environment, after generating signed certs for VLR, replace and update on connected components to maintain secure connection
As neededPassword Management
As neededExpiration Policy
As neededMethodSSH to VLR as admin, switch to root
As neededRepeat Foradmin user and recovery instance
As neededCommands
As neededMonitoring Points
- Verify Pairing
- CheckVMware Live Site Recovery + vSphere Replication status = OK (green)
- Power on VLR via vCenter. Verify VR + VLSR operational after startup.
- Monitoring And Alerting
- Validate connection, accept cert, verify Collecting status
- Checklist
- Activation AssessmentVerify DR is required (e.g., extended instance outage); plan for business continuity
- When prompted, verify LB VIP responds to ping, Dismiss
- Monitor analytics cluster node status (Not Running/Offline)
- Monitor all collectors until Not Running/Offline
Troubleshooting
Likely Panelist Questions
Q: Why did you choose this architecture?
See design decisions for rationale
Failure Scenarios
Trade-off Analysis
Trade-Offs Analysis
Chosen:
Justification:
Explain IP mobility strategies (NSX Federation vs stretched L2) and their trade-offs
Chosen:
Justification:
Quiz — Site Protection & DR
- 1 hour
- 4 hours
- 8 hours
- 24 hours
- 15 minutes
- 1 hour
- 4 hours
- 24 hours
- VCF Identity Broker + VCF Automation
- VCF Operations stack (Ops, Ops FM, Ops Logs, Ops Networks)
- SDDC Manager + NSX
- vCenter + SDDC Manager
- 1
- 3 over 24 hours
- 10
- 200
- Static routing only
- DHCP
- NSX Federation with stretched overlay segments OR stretched L2 networks
- BGP without overlay
- VCF Operations
- VCF Operations for logs
- VCF Automation
- VCF Operations for networks
- Unsupported
- Some protected apps (clustered) don't function correctly in isolated test network
- Too expensive
- Breaks replication
- Delete and recreate
- Protected -> Secondary, Recovery -> Primary
- Both Primary
- No change needed
- All components simultaneously
- Identity Broker/Automation, then Ops Networks, then Ops Logs, then take Ops cluster offline and shut down
- Random order
- VCF Ops first
- Nothing
- Take cluster offline via admin UI System status
- Force shutdown
- Snapshot each node
- 50 ms
- 100 ms
- 150 ms
- 300 ms
- Nothing special
- Force-provisioning
- Maximum redundancy
- Disable SPBM
- Read-Only
- No Access
- Administrators
- Custom
- Disaster recovery
- Planned migration
- Test recovery
- Reprotect
- To update dashboards
- To enable day-2 operations (power on/off VMs) after extended stay in recovery - so VCF Ops FM knows new VM location
- Required by license
- To enable backups
Flashcards — Site Protection & DR
Labs
Lab 1: Deploy VLR and Pair Across Two VCF Instances with NSX Federation
Deploy VLR appliances in both VCF instances, reconfigure DNS/NTP across instances, set up NSX Federation Global Manager active-standby, and pair VLR via a cross-instance site pair.
Starting State: Two VCF 9.0 instances operational; NSX Federation Global Manager deployed in each (Active-Standby); DNS/NTP/CA services in place; VLR .iso mounted.
Lab 2: Configure VCF Operations Recovery Plan and Run Planned Migration
Configure VM folder/network/resource mappings, set up VR protection for VCF Operations stack with RPO 15min + 3 PIT, create a recovery plan with prioritized startup, then run a planned migration and reprotect.
Starting State: Lab 1 complete. VCF Operations cluster deployed in protected instance. AD DC + DNS/NTP/CA + SMTP + Syslog services available at both sites. NSX Edge cluster configured in both instances for North-South routing. NSX load balancer deployed/configured in both.
Lab 3: Perform VCF Identity Broker and VCF Automation Recovery via Backup/Restore
Take backups of VCF Identity Broker and VCF Automation, simulate DR by powering off in protected, deploy new instances in recovery, and restore from backup.
Starting State: VCF Identity Broker + VCF Automation deployed in protected VCF instance. VCF Fleet Management configured with SFTP backup (HA between sites or replicated). Known-good backups exist.