Site Protection and Disaster Recovery for VMware Cloud Foundation
The solution uses VMware Site Recovery Manager (SRM) and vSphere Replication to orchestrate protection and recovery of SDDC management components between two VMware Cloud Foundation instances. Each VCF instance runs an SRM virtual appliance with embedded vPostgreSQL database integrated with management-domain vCenter Server, plus a vSphere Replication appliance. NSX Federation + cross-instance overlay segment provides IP mobility, allowing Aria Suite Lifecycle, Workspace ONE Access, Aria Operations, and Aria Automation to fail over without changing IP addresses.
Key Components: vSphere Replication, Site Recovery Manager, NSX, Aria Operations, Aria Automation, SDDC Manager, vCenter, VCF Operations, Workspace ONE
Replication: vSphere Replication — hypervisor-based, no storage adapter required.
25 design decisions
| DD-ID | Decision | Quality |
|---|---|---|
| SPR-SRM-RP-005 | s (20) | Recoverability |
Decision: s (20) Rationale: Orchestrates recovery. Implication: No significant trade-offs identified for this decision. Component: Site Recovery Manager | ||
| SPR-SRM-CFG-001 | Deploy SRM as a virtual appliance. | Recoverability |
Decision: Deploy SRM as a virtual appliance. Rationale: Orchestrates recovery. Implication: No significant trade-offs identified for this decision. Component: Site Recovery Manager | ||
| SPR-SRM-CFG-002 | Deploy SRM in the management domain. | Manageability |
Decision: Deploy SRM in the management domain. Rationale: Consistent deployment. Implication: No significant trade-offs identified for this decision. Component: Site Recovery Manager | ||
| SPR-SRM-CFG-003 | Use Light SRM deployment type. | AvailabilityRecoverability |
Decision: Use Light SRM deployment type. Rationale: Highest availability for defined mgmt components (10 VMs, 3 protection groups, 3 recovery plans). Implication: No significant trade-offs identified for this decision. Component: Site Recovery Manager | ||
| SPR-SRM-CFG-004 | Use vSphere Replication as replication method (not array-based). | Manageability |
Decision: Use vSphere Replication as replication method (not array-based). Rationale: Flexibility in storage vendor; no SRA maintenance overhead. Implication: All mgmt components must be in same cluster; fewer VMs supported vs array-based. Component: Site Recovery Manager | ||
| SPR-SRM-CFG-005 | Automate recovery of Aria Suite Lifecycle, WS1 Access, Aria Operations analytics, Aria Automation. | ManageabilityRecoverability |
Decision: Automate recovery of Aria Suite Lifecycle, WS1 Access, Aria Operations analytics, Aria Automation. Rationale: Automated run book; 4-hour RTO. Implication: No significant trade-offs identified for this decision. Component: Site Recovery Manager | ||
| SPR-VR-CFG-001 | Deploy vSphere Replication in the vCenter it will register with. | Manageability |
Decision: Deploy vSphere Replication in the vCenter it will register with. Rationale: Cert thumbprint discovered during OVF deployment. Implication: No significant trade-offs identified for this decision. Component: vSphere Replication | ||
| SPR-VR-CFG-002 | Deploy vSphere Replication with 4 vCPU size. | Manageability |
Decision: Deploy vSphere Replication with 4 vCPU size. Rationale: Supports expected VM count from the four protected components. Implication: No significant trade-offs identified for this decision. Component: vSphere Replication | ||
| SPR-VR-CFG-003 | Do NOT activate guest OS quiescing. | Manageability |
Decision: Do NOT activate guest OS quiescing. Rationale: Not all mgmt VMs support it; could cause outage. Implication: Replicas are crash-consistent. Component: vSphere Replication | ||
| SPR-VR-CFG-004 | Activate network compression. | Manageability |
Decision: Activate network compression. Rationale: Reduced network footprint, less buffer memory. Implication: More CPU on source as more VMs are protected. Component: vSphere Replication | ||
| SPR-VR-CFG-005 | Configure 15-minute RPO. | Recoverability |
Decision: Configure 15-minute RPO. Rationale: Ensures recovered app has all data except changes in last 15 min. Implication: Up to 15 min data loss on recovery. Component: vSphere Replication | ||
| SPR-VR-CFG-006 | Configure 3 point-in-time instances over 24 hours. | Manageability |
Decision: Configure 3 point-in-time instances over 24 hours. Rationale: Application integrity after compromise. Implication: Increases vSAN disk usage. Component: vSphere Replication | ||
| SPR-SRM-NET-001 | Place SRM and vSphere Replication on management VLAN. | Manageability |
Decision: Place SRM and vSphere Replication on management VLAN. Rationale: Same network as VCF components. Implication: No significant trade-offs identified for this decision. Component: Site Recovery Manager | ||
| SPR-VR-NET-001 | Place SRM and vSphere Replication on management VLAN. | Manageability |
Decision: Place SRM and vSphere Replication on management VLAN. Rationale: Same network as VCF components. Implication: No significant trade-offs identified for this decision. Component: vSphere Replication | ||
| SPR-SRM-NET-002 | Static IPs. | Manageability |
Decision: Static IPs. Rationale: Removes DHCP risks. Implication: IP management discipline. Component: Site Recovery Manager | ||
| SPR-VR-NET-002 | Static IPs. | Manageability |
Decision: Static IPs. Rationale: Removes DHCP risks. Implication: IP management discipline. Component: vSphere Replication | ||
| SPR-SRM-NET-003 | Forward/reverse DNS records. | Security |
Decision: Forward/reverse DNS records. Rationale: FQDN accessibility. Implication: Maintain DNS records; allow DNS through firewalls. Component: Site Recovery Manager | ||
| SPR-VR-NET-003 | Forward/reverse DNS records. | Security |
Decision: Forward/reverse DNS records. Rationale: FQDN accessibility. Implication: Maintain DNS records; allow DNS through firewalls. Component: vSphere Replication | ||
| SPR-SRM-NET-004 | Configure NTP, not VMTools sync. | AvailabilitySecurity |
Decision: Configure NTP, not VMTools sync. Rationale: Accurate time sync; prevents mismatch. Implication: NTP HA and firewall allow. Component: Site Recovery Manager | ||
| SPR-VR-NET-004 | Configure NTP, not VMTools sync. | AvailabilitySecurity |
Decision: Configure NTP, not VMTools sync. Rationale: Accurate time sync; prevents mismatch. Implication: NTP HA and firewall allow. Component: vSphere Replication | ||
| SPR-SRM-LCM-001 | LCM via native appliance tools, not SDDC Manager. | Manageability |
Decision: LCM via native appliance tools, not SDDC Manager. Rationale: Not integrated with SDDC Manager LCM. Implication: Manual LCM. Component: Site Recovery Manager | ||
| SPR-VR-LCM-001 | LCM via native appliance tools, not SDDC Manager. | Manageability |
Decision: LCM via native appliance tools, not SDDC Manager. Rationale: Not integrated with SDDC Manager LCM. Implication: Manual LCM. Component: vSphere Replication | ||
| SPR-SRM-RP-001 | Prioritized startup order: Aria Suite Lifecycle + WS1 Access → Operations analytics → Operations rem | ManageabilityRecoverability |
Decision: Prioritized startup order: Aria Suite Lifecycle + WS1 Access → Operations analytics → Operations remote collectors → Aria Automation. Rationale: Dependency ordering; RTO target. Implication: Maintain customized plans; VMware Tools required on each VM. Component: Site Recovery Manager | ||
| SPR-DNS-NET-001 | Configure DNS settings for each protected component to use DNS servers across both VCF instances. | Manageability |
Decision: Configure DNS settings for each protected component to use DNS servers across both VCF instances. Rationale: Resolve DNS during planned migration or DR. Implication: Update DNS settings as you scale from single to multi-instance. Component: DNS | ||
| SPR-NTP-NET-001 | Configure NTP settings for each protected component to use NTP servers across both VCF instances. | Manageability |
Decision: Configure NTP settings for each protected component to use NTP servers across both VCF instances. Rationale: Resolve NTP during migration/DR. Implication: Update NTP settings as you scale. Component: NTP | ||
Prerequisites
- VCF healthy per Support Matrix (protected and recovery instances); same VCF version.
- Captured Site Protection and DR tab of Planning & Preparation Workbook.
- NSX Federation deployed; cross-instance overlay segment promoted via NSX Global Manager.
- Standalone Tier-1 gateway in recovery region for LB failover.
- Mgmt cluster in recovery region with matching compute/storage capacity.
- External services available in recovery: AD DCs, DNS, NTP, SMTP, Syslog.
- SRM and vSphere Replication OVAs downloaded.
- Implementation Methods
- Powershell
Implementation Procedure
Implementation
PowerValidatedSolutions module — end-to-end deployment and configuration of SRM and vSphere Replication.
UI
Per VCF version: Deploy vSphere Replication in protected and recovery mgmt vCenters (use vCenter OVF deploy). Deploy SRM in protected and recovery; pair sites. Configure placeholder datastore mappings, network mappings, folder mappings. Create protection groups for each management component grouping. Create recovery plans with prioritized startup order + prompts. Customize recovery plan for Aria Suite Lifecycle + WS1 Access (includes VMware Aria Suite Lifecycle PowerOn workflow + vidmPowerOnSkip retry). Create anti-affinity rules for placeholder VMs. Create VM groups with restart order on recovery site. If applicable, back up/restore OVF properties of protected VMs.
External Services / Integration Points
Active Directory
DNS
NTP
SMTP
Syslog
Configuration Values
RPO15 minutes
PIT3 copies over 24 hours
Network CompressionEnabled
Guest QuiescingDisabled
Protection Groups3 groups covering Aria Suite Lifecycle + WS1 Access, Operations analytics, Aria Automation
RTO≤ 4 hours
Additional Instance
Not applicable — solution always spans two VCF instances (primary + recovery).
Day-2 Operations Tasks
Operations
As neededPersonas
As neededNameRole
As neededCloud AdminSRM Administrator; recovery plan authoring and execution
As neededAuditorRead-only access to recovery plans and history
As neededOperational Verification
As neededCertificate Management
As neededAfter generating CA-signed certs for SRM and vSphere Replication, update certs on connected management components (paired SRM, vCenter registration) t
As neededPassword Management
As neededPolicies
As neededMonitoring Points
- Site Reliability EngineerRecovery plan execution and monitoring
- Verify SRM pairing across VCF instances — Site Recovery page shows OK (green check) for both sites.
- Verify vSphere Replication pairing — Site Recovery page shows green tick for VR subsection on both sites.
- MonitoringAria Operations consumes SRM via non-native VMware Site Recovery Manager management pack (not packaged with Aria Operations).
- Failover Checklist
- LogisticsFacilities/personnel available; SRM available in recovery; replication status healthy; NSX state (Edge nodes, overlay IPs, LB config) verifie
- Post FailoverUpdate DNS for VAOL; reconfigure Operations logging target; inventory sync in Aria Suite Lifecycle; verify each component.
- ProcedureIn SRM recovery plan: right-click → Reprotect → confirm irreversible. If Reprotect Interrupted, use Force cleanup option. After Ready state,
Troubleshooting
Likely Panelist Questions
Q: Why did you choose this architecture?
See design decisions for rationale
Failure Scenarios
Trade-off Analysis
Trade-Offs Analysis
Chosen:
Justification:
Quiz — Site Protection & DR
- Array-based replication only
- vSphere Replication
- Both simultaneously
- vSAN stretched cluster
- 1 hour
- 4 hours or less
- 8 hours
- 24 hours
- 5 minutes
- 15 minutes
- 1 hour
- 24 hours
- 1 snapshot / 1h
- 3 snapshots / 24h
- 5 snapshots / 12h
- 10 snapshots / 7d
- VCF Operations analytics nodes
- Aria Automation
- VCF Operations Cloud Proxies
- WS1 Access
- Too expensive
- Protected apps use NSX LB and DNS unavailable in isolated test network
- SRM doesn't support test
- Not supported on VCF
- Light (1000 VMs)
- Standard (5000 VMs)
- Extra Large (10,000 VMs)
- Custom
- Cross-instance NSX overlay segment
- Local-instance NSX overlay segment
- Management VLAN
- VMkernel VMotion network
- Aria Automation → WS1 Access → Operations
- Aria Suite Lifecycle → WS1 Access → Operations analytics → Operations remote collectors → Aria Automation
- Operations → Aria Automation → WS1 Access
- WS1 Access → Aria Automation → Aria Suite Lifecycle
- adminResetSkip
- vidmPowerOnSkip
- srmSkipDNS
- vraSkipK8s
- Not all mgmt VMs support quiescing; could cause outage
- It would increase RPO
- Performance impact on ESXi
- VCF forbids it
- SDDC Manager
- Aria Suite Lifecycle
- Native appliance tools (manual)
- vCenter Server
- Nothing
- Reprotect VMs and re-point logs to recovery VAOL
- Redeploy all Operations nodes
- Uninstall Cloud Proxies
- Always
- For failback DR where recovery site's vSAN witness is unavailable in multi-AZ
- For test recoveries
- Never
- Immediately reprotect
- Run recovery plans in 'Recovery Required' state → planned migration → reprotect
- Reinstall SRM always
- No reprotect needed
Flashcards — Site Protection & DR
Labs
Pair SRM + vSphere Replication across VCF instances
Deploy SRM and vSphere Replication in both instances, pair sites, validate.
Starting State: Two VCF instances operational with mgmt domains; NSX Federation configured.
Create protection groups and recovery plans with prioritized startup
Build 3 protection groups and 3 recovery plans covering Aria Suite Lifecycle + WS1 Access, Operations analytics, Aria Automation.
Starting State: SRM/VR paired; Aria Suite products deployed.
Planned migration + reprotect drill
Fail over all mgmt apps to recovery site and then reprotect to prepare for failback.
Starting State: Recovery plans Ready; management apps healthy on protected site.