Academy/VVS/Site Protection & DR
This solution targets VCF 9.0

Site Protection and Disaster Recovery for VMware Cloud Foundation

VCF 9.0architectvcdxautomationstoragePages 83-178

This is the VCF 9.0 evolution of the VCF 5.2 Site Protection and Disaster Recovery validated solution. Key changes from VCF 5.2 → 9.0 include: consolidation of VMware Live Site Recovery and vSphere Replication into the unified VMware Live Recovery (VLR) appliance, native VCF 9.0 workload domain integration, updated support for vSAN ESA at the recovery site, tighter NSX integration for recovery networks, and new Planning Workbook templates aligned with the VCF 9.0 topology.

Key Components: VMware Live Recovery, vSphere Replication, NSX, SDDC Manager, vCenter, VCF Operations

Purpose: Instructions for adapting a dual-instance SDDC on top of VCF to provide disaster recovery of SDDC management components using VMware Live Recovery for VCF Automation, VCF Operations, and VCF Operations fleet management.

Design Decisions
Implementation
Operations
VCDX Defense
Quiz (15)
Flashcards (15)

32 design decisions

DD-IDDecisionQuality
SPR-VLR-CFG-001Deploy VLR instance in protected and recovery VCF instances.Recoverability

Decision: Deploy VLR instance in protected and recovery VCF instances.

Rationale: Allows orchestrated recovery of VCF mgmt components in another VCF instance.

Implication: No significant trade-offs identified for this decision.

Component: VMware Live Recovery

SPR-VLR-CFG-002Deploy each VLR in the management domain.Manageability

Decision: Deploy each VLR in the management domain.

Rationale: Consistent deployment model for all mgmt applications.

Implication: No significant trade-offs identified for this decision.

Component: VMware Live Recovery

SPR-VLR-CFG-003Deploy VLR with highest scale level supported.Manageability

Decision: Deploy VLR with highest scale level supported.

Rationale: Highest protection capacity for mgmt components.

Implication: No significant trade-offs identified for this decision.

Component: VMware Live Recovery

SPR-VLR-CFG-004Use vSphere Replication in VLR as VM replication protection.Manageability

Decision: Use vSphere Replication in VLR as VM replication protection.

Rationale: Flexibility in storage/vendor; minimizes SRA compatibility overhead.

Implication: All mgmt components must be in same cluster; reduced total VM count vs storage-based replication.

Component: VMware Live Recovery

SPR-VLR-CFG-005Use VLR to automate recovery of: VCF Operations, VCF Operations fleet mgmt, VCF Operations for logs,ManageabilityRecoverability

Decision: Use VLR to automate recovery of: VCF Operations, VCF Operations fleet mgmt, VCF Operations for logs, VCF Operations for networks.

Rationale: Automated runbook; RTO <= 4 hours.

Implication: No significant trade-offs identified for this decision.

Component: VMware Live Recovery

SPR-VLR-CFG-006Configure PIT instances keeping 3 copies over 24-hour period.Recoverability

Decision: Configure PIT instances keeping 3 copies over 24-hour period.

Rationale: Application integrity after DR recovery event.

Implication: More PIT = more disk usage on vSAN datastore.

Component: VMware Live Recovery

SPR-VLR-NET-001Place VLR on the management network.Manageability

Decision: Place VLR on the management network.

Rationale: Co-locate with VCF components.

Implication: No significant trade-offs identified for this decision.

Component: VMware Live Recovery

SPR-VLR-NET-002Allocate static IP to VLR instances.Manageability

Decision: Allocate static IP to VLR instances.

Rationale: Removes DHCP constraints/risks on management networks.

Implication: Precise IP management.

Component: VMware Live Recovery

SPR-VLR-NET-003Configure A + PTR DNS records for each VLR.ManageabilitySecurity

Decision: Configure A + PTR DNS records for each VLR.

Rationale: FQDN accessibility.

Implication: DNS infrastructure required; manage DNS records; firewalls allow DNS.

Component: VMware Live Recovery

SPR-VLR-NET-004Configure VLR to use NTP servers (not VMTools).Security

Decision: Configure VLR to use NTP servers (not VMTools).

Rationale: Accurate time sync; prevent mismatch.

Implication: NTP services required; firewalls allow NTP.

Component: VMware Live Recovery

SPR-VLR-NET-005Use dedicated vDS portgroup for VLR traffic.Security

Decision: Use dedicated vDS portgroup for VLR traffic.

Rationale: Isolates VLR traffic from other system traffic.

Implication: Manually create portgroup.

Component: VMware Live Recovery

SPR-VLR-NET-006Use dedicated VLAN for VLR traffic.Security

Decision: Use dedicated VLAN for VLR traffic.

Rationale: Traffic isolation.

Implication: Manually create VLAN.

Component: VMware Live Recovery

SPR-VLR-NET-007Use dedicated VMkernel interface on each ESX for VLR traffic.Security

Decision: Use dedicated VMkernel interface on each ESX for VLR traffic.

Rationale: Traffic isolation.

Implication: Manually create VMkernel interface.

Component: VMware Live Recovery

SPR-VLR-NET-008Assign static IP to dedicated VMkernel for VLR traffic.Manageability

Decision: Assign static IP to dedicated VMkernel for VLR traffic.

Rationale: Removes DHCP risks.

Implication: Precise IP management.

Component: VMware Live Recovery

SPR-VLR-LCM-001Lifecycle manage VLR via native appliance tools.Manageability

Decision: Lifecycle manage VLR via native appliance tools.

Rationale: VLR not managed by SDDC Manager.

Implication: Deployment, patching, updates, upgrades without native automation.

Component: VMware Live Recovery

SPR-VLR-SEC-001Configure service account in each vCenter for app-to-app communication VLR > vSphere, in SSO AdminisManageability

Decision: Configure service account in each vCenter for app-to-app communication VLR > vSphere, in SSO Administrators group.

Rationale: VLSR DR orchestration + site pairing; VR replication; accountability; limited blast radius.

Implication: Maintain service account lifecycle outside VCF.

Component: VMware Live Recovery

SPR-VLR-SEC-003Configure password expiration policy for VLR appliance.Manageability

Decision: Configure password expiration policy for VLR appliance.

Rationale: Align with org/compliance policies (local users only).

Implication: Manage via appliance console or SSH.

Component: VMware Live Recovery

SPR-VLR-SEC-004Configure password complexity policy for VLR appliance.Manageability

Decision: Configure password complexity policy for VLR appliance.

Rationale: Align with org/compliance policies (local users only).

Implication: Manage via console or SSH.

Component: VMware Live Recovery

SPR-VLR-SEC-005Configure account lockout policy for VLR appliance.Manageability

Decision: Configure account lockout policy for VLR appliance.

Rationale: Align with org/compliance policies (local users only).

Implication: Manage via console or SSH.

Component: VMware Live Recovery

SPR-VLR-SEC-006Change VLR root password on recurring/event-initiated schedule.Manageability

Decision: Change VLR root password on recurring/event-initiated schedule.

Rationale: Root passwords never expire by default.

Implication: Routine password changes required.

Component: VMware Live Recovery

SPR-VLR-SEC-007Change VLR admin password on recurring/event-initiated schedule.Manageability

Decision: Change VLR admin password on recurring/event-initiated schedule.

Rationale: Admin passwords never expire by default.

Implication: Routine password changes required.

Component: VMware Live Recovery

SPR-VLR-SEC-008Replace default self-signed cert with CA-signed in each VLR.Security

Decision: Replace default self-signed cert with CA-signed in each VLR.

Rationale: All externally facing Web UI + cross-product communication encrypted.

Implication: Must have PKI access.

Component: VMware Live Recovery

SPR-VLR-RP-001Use prioritized startup order for VCF Operations + VCF Operations FM appliance nodes.ManageabilityRecoverability

Decision: Use prioritized startup order for VCF Operations + VCF Operations FM appliance nodes.

Rationale: Individual VCF Ops nodes started in order for operational monitoring.

Implication: Maintain customized recovery plan if node count increases.

Component: VMware Live Recovery

SPR-VLR-RP-002Use prioritized startup order for VCF Operations for logs nodes.ManageabilityRecoverability

Decision: Use prioritized startup order for VCF Operations for logs nodes.

Rationale: Cloud automation services restored; restored within 4h.

Implication: VMware Tools on each VCF Ops Logs node.

Component: VMware Live Recovery

SPR-VLR-RP-003Use prioritized startup order for VCF Operations for networks nodes.ManageabilityRecoverability

Decision: Use prioritized startup order for VCF Operations for networks nodes.

Rationale: Cloud automation services restored; restored within 4h.

Implication: VMware Tools on each VCF Ops Networks node.

Component: VMware Live Recovery

SPR-VLR-RP-005Do NOT run test recovery of mgmt components.PerformanceRecoverabilitySecurity

Decision: Do NOT run test recovery of mgmt components.

Rationale: Some protected apps (e.g., clustered) don't function correctly in isolated test.

Implication: Cannot test DR; perform planned migrations to validate instead.

Component: VMware Live Recovery

SPR-DNS-NET-001With multiple VCF instances, configure DNS settings for each protected component to use DNS servers Availability

Decision: With multiple VCF instances, configure DNS settings for each protected component to use DNS servers across all VCF instances.

Rationale: Higher DNS availability and resilience during DR/migration.

Implication: As you scale from single to multiple instances, update DNS settings on each protected component.

Component: DNS

SPR-NTP-NET-001With multiple VCF instances, configure NTP settings across all instances.Availability

Decision: With multiple VCF instances, configure NTP settings across all instances.

Rationale: Higher NTP availability during DR/migration.

Implication: Update NTP settings as instances scale.

Component: NTP

SPR-VCFO-CFG-001Install VLR management pack for VCF Operations.Manageability

Decision: Install VLR management pack for VCF Operations.

Rationale: Establishes communication between VCF Ops and VLR endpoints.

Implication: Manual management pack installation.

Component: VCFO

SPR-VCFO-CFG-002Configure VLR endpoints to use VCF Operations collector/collector group for the VCF Instance.Manageability

Decision: Configure VLR endpoints to use VCF Operations collector/collector group for the VCF Instance.

Rationale: Offloads data collection from analytics cluster.

Implication: No significant trade-offs identified for this decision.

Component: VCFO

SPR-VLR-LOG-001When using VLR, install VCF Operations for logs agent on VLR appliance.Manageability

Decision: When using VLR, install VCF Operations for logs agent on VLR appliance.

Rationale: Simplifies configuration via agent-based log forwarding.

Implication: Must configure VCF Ops for logs as aggregation target.

Component: VMware Live Recovery

SPR-VLR-LOG-002Configure VCF Operations for logs as central log aggregation target for VLR.Manageability

Decision: Configure VCF Operations for logs as central log aggregation target for VLR.

Rationale: Ensures transmission of logs to centralized system.

Implication: Configuration required per VCF component.

Component: VMware Live Recovery

Prerequisites

  • VCF management domain operational
  • DNS/NTP configured
  • CA infrastructure in place

Implementation Procedure

Implementation

Environment: VCF version in support matrix; per Before You Apply; Site Protection and Disaster Recovery tab of Planning Workbook; VCF healthy.

DNS: Required DNS entries in forward and reverse zones

Network: IP mobility between VCF instances; jumbo frames; L3 routing; max 150ms latency; bandwidth sized with VLR Calculator

Software: VLR .iso mounted

License: VLR license with quantity per design

  • Active Directory: AD Domain Controllers, service accounts, security groups
  • CA: CA available
  • Reconfigure DNS NTP On VCF Management Components

VCF Operations Fleet Mgmt Appliance

DNSSSH to fleet mgmt node as root. Edit /etc/systemd/network/10-eth0.network - add DNS entry. Edit /etc/systemd/resolved.conf - modify Domains entry (space separated).
NTPVCF Ops interface > Fleet Management > Lifecycle > Settings > NTP Servers > Add NTP Server or Edit Server Selection

VCF Operations Nodes

DNSSSH to primary node as root. Edit /etc/systemd/network/10-eth0.network and /etc/systemd/resolved.conf (same as fleet mgmt). Repeat for all cluster nodes + VCF Ops Logs + VCF Ops Networks.
NTPVCF Ops > Administration > Control Panel > Cluster Management > Actions > Network Time Protocol Settings > Add NTP

VCF Operations For Logs

NTPSSH to primary logs node as root. Edit /etc/ntp.conf - add 'server ntp.lax.rainpole.io iburst'. Repeat for each.

VCF Operations For Networks

NTPSSH to primary networks node as support user. sudo vi /etc/ntp.conf - add server. Repeat for each.

VCF Automation

DNSTake backup first (per VCF Fleet Management). VCF Ops > Fleet Management > Lifecycle > VCF Management > Components > automation > ellipsis > Update DNS Configuration > Add Server, Change Priority, Run Precheck, Finish.
NTPSame path > Update NTP Configuration > Add, Priority, Precheck, Finish

VCF Identity Broker

DNSTake backup first. VCF Ops > Fleet Management > Lifecycle > VCF Management > Components > identity broker > ellipsis > Update DNS Configuration
NTPSame > Update NTP Configuration

Deploy VLR Appliance

  • MethodDeploy OVF template in vSphere Client for mgmt domain
  • PlacementDefault mgmt cluster
  • StorageMgmt domain vSAN datastore
  • NetworkManagement network

Day-2 Operations Tasks

Operations

As needed

Operational Verification

As needed

EndpointProtected vCenter > VMware Live Site Recovery

As needed

Certificate Management

As needed

In multi-instance environment, after generating signed certs for VLR, replace and update on connected components to maintain secure connection

As needed

Password Management

As needed

Expiration Policy

As needed

MethodSSH to VLR as admin, switch to root

As needed

Repeat Foradmin user and recovery instance

As needed

Commands

As needed

Monitoring Points

  • Verify Pairing
  • CheckVMware Live Site Recovery + vSphere Replication status = OK (green)
  • Power on VLR via vCenter. Verify VR + VLSR operational after startup.
  • Monitoring And Alerting
  • Validate connection, accept cert, verify Collecting status
  • Checklist
  • Activation AssessmentVerify DR is required (e.g., extended instance outage); plan for business continuity
  • When prompted, verify LB VIP responds to ping, Dismiss
  • Monitor analytics cluster node status (Not Running/Offline)
  • Monitor all collectors until Not Running/Offline

Troubleshooting

TroubleshootingEnsure network connectivity between vCenter instances and VLR appliance
Cause:
Fix:
Backupcp -p /etc/security/faillock.conf ...back
Cause:
Fix:
Stop VLR from 5480 UI > Summary > Stop. Paired sites expectedly show connection errors.
Cause:
Fix:
Failover Of SDDC Management Applications
Cause:
Fix:
Applications With Failover Support
Cause:
Fix:

Likely Panelist Questions

Q: Why did you choose this architecture?

See design decisions for rationale

Failure Scenarios

Why reconfigure DNS/NTP to include both instances - avoids outage of protected site DNS/NTP breaking recovered components
Impact:
Mitigation:
Why NSX load balancer reconfiguration (Tier-1 active-standby) during failover
Impact:
Mitigation:
15 min RPO = low data loss but high replication load; 24h RPO = less load but more data loss
Impact:
Mitigation:
Failure to reconfigure DNS on protected components before DR (split-brain name resolution)
Impact:
Mitigation:
Not reconfiguring Tier-1 Gateway for dynamic routing during failover (routing fails)
Impact:
Mitigation:

Trade-off Analysis

Trade-Offs Analysis

Chosen:

Justification:

Explain IP mobility strategies (NSX Federation vs stretched L2) and their trade-offs

Chosen:

Justification:

Quiz — Site Protection & DR

0/15
Q1
What RTO target does this design achieve for the VCF management components?
  • 1 hour
  • 4 hours
  • 8 hours
  • 24 hours
VLR automates management component recovery with an RTO target of 4 hours or less (SPR-VLR-CFG-005).
Q2
What is the RPO configured on management virtual machine policies?
  • 15 minutes
  • 1 hour
  • 4 hours
  • 24 hours
RPO of 15 minutes is configured - ensures recovered mgmt apps contain data except any changes in last 15 minutes before DR event (SPR-VLR-CFG-005-b).
Q3
Which VCF management components use VMware Live Recovery replication (vs backup/restore)?
  • VCF Identity Broker + VCF Automation
  • VCF Operations stack (Ops, Ops FM, Ops Logs, Ops Networks)
  • SDDC Manager + NSX
  • vCenter + SDDC Manager
VLR replicates VCF Operations, VCF Operations fleet management appliance, VCF Operations for logs, and VCF Operations for networks. VCF Identity Broker + VCF Automation use backup/restore. Collectors are site-specific and NOT protected.
Q4
How many PIT instances are configured per management VM?
  • 1
  • 3 over 24 hours
  • 10
  • 200
3 PIT copies over 24-hour period - ensures application integrity for mgmt components after DR event (SPR-VLR-CFG-006).
Q5
Which IP mobility mechanism is preferred for this design?
  • Static routing only
  • DHCP
  • NSX Federation with stretched overlay segments OR stretched L2 networks
  • BGP without overlay
IP mobility can be achieved via NSX Federation global overlay networks (preferred for programmability) or stretched layer 2 networks using physical fabric.
Q6
Which VCF management component canNOT be recovered with VLR replication?
  • VCF Operations
  • VCF Operations for logs
  • VCF Automation
  • VCF Operations for networks
VCF Automation and VCF Identity Broker are recovered via backup/restore - not VLR replication.
Q7
Why should test recovery plans NOT be run for management components?
  • Unsupported
  • Some protected apps (clustered) don't function correctly in isolated test network
  • Too expensive
  • Breaks replication
Management components include clustered apps (like VCF Operations cluster) that cannot function correctly in VLR's isolated test network (SPR-VLR-RP-005). Validate via planned migration instead.
Q8
Which NSX Tier-1 gateway change is required during failover?
  • Delete and recreate
  • Protected -> Secondary, Recovery -> Primary
  • Both Primary
  • No change needed
Reconfigure the cross-instance Tier-1 gateway: Protected VCF -> Secondary, Recovery VCF -> Primary. This provides dynamic routing from recovery instance.
Q9
What shutdown order is required before planned migration?
  • All components simultaneously
  • Identity Broker/Automation, then Ops Networks, then Ops Logs, then take Ops cluster offline and shut down
  • Random order
  • VCF Ops first
Shut down sequence: VCF Identity Broker + VCF Automation, then VCF Operations for Networks, VCF Operations for Logs, then VCF Operations (take cluster offline first), then VCF Operations FM.
Q10
What is required for VCF Operations cluster before VM shutdown?
  • Nothing
  • Take cluster offline via admin UI System status
  • Force shutdown
  • Snapshot each node
Before shutting down VCF Operations VMs, you must Take cluster offline via the primary node admin UI (https://<primary>/admin > System status). Wait until all nodes are Offline.
Q11
What is the maximum allowed latency between protected and recovery VCF instances?
  • 50 ms
  • 100 ms
  • 150 ms
  • 300 ms
Max 150 ms latency between instances - critical for VLR replication.
Q12
In multi-AZ deployments, what storage policy setting might be required during DR failback?
  • Nothing special
  • Force-provisioning
  • Maximum redundancy
  • Disable SPBM
If vSAN witness appliance is unavailable (e.g., recovery VCF instance outage during failback), turn on force-provisioning in the storage policy to ensure VCF Automation + VCF Operations VMs can be provisioned.
Q13
What role must the VLR service account be a member of in vCenter SSO?
  • Read-Only
  • No Access
  • Administrators
  • Custom
Service account must be member of vCenter SSO Administrators group - VLSR for DR orchestration + site pairing, VR for site-to-site replication (SPR-VLR-SEC-001).
Q14
What is the recovery type for planned migration (failback) in VLR?
  • Disaster recovery
  • Planned migration
  • Test recovery
  • Reprotect
For failback or non-disaster move, select Recovery type: Planned migration. Disaster recovery is for unplanned DR scenarios.
Q15
Why is VCF Operations fleet management inventory sync required after DR?
  • To update dashboards
  • To enable day-2 operations (power on/off VMs) after extended stay in recovery - so VCF Ops FM knows new VM location
  • Required by license
  • To enable backups
After VMs reside in recovery for extended time, VCF Ops FM inventory must be synchronized so day-2 operations (power on/off) work with the new VM location.

Flashcards — Site Protection & DR

Card 1 of 15
Components protected via VLR replication?
VCF Operations cluster, VCF Operations fleet management appliance, VCF Operations for logs, VCF Operations for networks. Collectors are site-specific, NOT protected.

Labs

Lab 1: Deploy VLR and Pair Across Two VCF Instances with NSX Federation

Deploy VLR appliances in both VCF instances, reconfigure DNS/NTP across instances, set up NSX Federation Global Manager active-standby, and pair VLR via a cross-instance site pair.

Starting State: Two VCF 9.0 instances operational; NSX Federation Global Manager deployed in each (Active-Standby); DNS/NTP/CA services in place; VLR .iso mounted.

Lab 2: Configure VCF Operations Recovery Plan and Run Planned Migration

Configure VM folder/network/resource mappings, set up VR protection for VCF Operations stack with RPO 15min + 3 PIT, create a recovery plan with prioritized startup, then run a planned migration and reprotect.

Starting State: Lab 1 complete. VCF Operations cluster deployed in protected instance. AD DC + DNS/NTP/CA + SMTP + Syslog services available at both sites. NSX Edge cluster configured in both instances for North-South routing. NSX load balancer deployed/configured in both.

Lab 3: Perform VCF Identity Broker and VCF Automation Recovery via Backup/Restore

Take backups of VCF Identity Broker and VCF Automation, simulate DR by powering off in protected, deploy new instances in recovery, and restore from backup.

Starting State: VCF Identity Broker + VCF Automation deployed in protected VCF instance. VCF Fleet Management configured with SFTP backup (HA between sites or replicated). Known-good backups exist.

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.