Academy/VVS/Site Protection & DR
This solution targets VCF 5.2

Site Protection and Disaster Recovery for VMware Cloud Foundation

VCF 5.2architectvcdxautomationstoragePages 656-793

The solution uses VMware Site Recovery Manager (SRM) and vSphere Replication to orchestrate protection and recovery of SDDC management components between two VMware Cloud Foundation instances. Each VCF instance runs an SRM virtual appliance with embedded vPostgreSQL database integrated with management-domain vCenter Server, plus a vSphere Replication appliance. NSX Federation + cross-instance overlay segment provides IP mobility, allowing Aria Suite Lifecycle, Workspace ONE Access, Aria Operations, and Aria Automation to fail over without changing IP addresses.

Key Components: vSphere Replication, Site Recovery Manager, NSX, Aria Operations, Aria Automation, SDDC Manager, vCenter, VCF Operations, Workspace ONE

Replication: vSphere Replication — hypervisor-based, no storage adapter required.

Design Decisions
Implementation
Operations
VCDX Defense
Quiz (15)
Flashcards (15)

25 design decisions

DD-IDDecisionQuality
SPR-SRM-RP-005s (20)Recoverability

Decision: s (20)

Rationale: Orchestrates recovery.

Implication: No significant trade-offs identified for this decision.

Component: Site Recovery Manager

SPR-SRM-CFG-001Deploy SRM as a virtual appliance.Recoverability

Decision: Deploy SRM as a virtual appliance.

Rationale: Orchestrates recovery.

Implication: No significant trade-offs identified for this decision.

Component: Site Recovery Manager

SPR-SRM-CFG-002Deploy SRM in the management domain.Manageability

Decision: Deploy SRM in the management domain.

Rationale: Consistent deployment.

Implication: No significant trade-offs identified for this decision.

Component: Site Recovery Manager

SPR-SRM-CFG-003Use Light SRM deployment type.AvailabilityRecoverability

Decision: Use Light SRM deployment type.

Rationale: Highest availability for defined mgmt components (10 VMs, 3 protection groups, 3 recovery plans).

Implication: No significant trade-offs identified for this decision.

Component: Site Recovery Manager

SPR-SRM-CFG-004Use vSphere Replication as replication method (not array-based).Manageability

Decision: Use vSphere Replication as replication method (not array-based).

Rationale: Flexibility in storage vendor; no SRA maintenance overhead.

Implication: All mgmt components must be in same cluster; fewer VMs supported vs array-based.

Component: Site Recovery Manager

SPR-SRM-CFG-005Automate recovery of Aria Suite Lifecycle, WS1 Access, Aria Operations analytics, Aria Automation.ManageabilityRecoverability

Decision: Automate recovery of Aria Suite Lifecycle, WS1 Access, Aria Operations analytics, Aria Automation.

Rationale: Automated run book; 4-hour RTO.

Implication: No significant trade-offs identified for this decision.

Component: Site Recovery Manager

SPR-VR-CFG-001Deploy vSphere Replication in the vCenter it will register with.Manageability

Decision: Deploy vSphere Replication in the vCenter it will register with.

Rationale: Cert thumbprint discovered during OVF deployment.

Implication: No significant trade-offs identified for this decision.

Component: vSphere Replication

SPR-VR-CFG-002Deploy vSphere Replication with 4 vCPU size.Manageability

Decision: Deploy vSphere Replication with 4 vCPU size.

Rationale: Supports expected VM count from the four protected components.

Implication: No significant trade-offs identified for this decision.

Component: vSphere Replication

SPR-VR-CFG-003Do NOT activate guest OS quiescing.Manageability

Decision: Do NOT activate guest OS quiescing.

Rationale: Not all mgmt VMs support it; could cause outage.

Implication: Replicas are crash-consistent.

Component: vSphere Replication

SPR-VR-CFG-004Activate network compression.Manageability

Decision: Activate network compression.

Rationale: Reduced network footprint, less buffer memory.

Implication: More CPU on source as more VMs are protected.

Component: vSphere Replication

SPR-VR-CFG-005Configure 15-minute RPO.Recoverability

Decision: Configure 15-minute RPO.

Rationale: Ensures recovered app has all data except changes in last 15 min.

Implication: Up to 15 min data loss on recovery.

Component: vSphere Replication

SPR-VR-CFG-006Configure 3 point-in-time instances over 24 hours.Manageability

Decision: Configure 3 point-in-time instances over 24 hours.

Rationale: Application integrity after compromise.

Implication: Increases vSAN disk usage.

Component: vSphere Replication

SPR-SRM-NET-001Place SRM and vSphere Replication on management VLAN.Manageability

Decision: Place SRM and vSphere Replication on management VLAN.

Rationale: Same network as VCF components.

Implication: No significant trade-offs identified for this decision.

Component: Site Recovery Manager

SPR-VR-NET-001Place SRM and vSphere Replication on management VLAN.Manageability

Decision: Place SRM and vSphere Replication on management VLAN.

Rationale: Same network as VCF components.

Implication: No significant trade-offs identified for this decision.

Component: vSphere Replication

SPR-SRM-NET-002Static IPs.Manageability

Decision: Static IPs.

Rationale: Removes DHCP risks.

Implication: IP management discipline.

Component: Site Recovery Manager

SPR-VR-NET-002Static IPs.Manageability

Decision: Static IPs.

Rationale: Removes DHCP risks.

Implication: IP management discipline.

Component: vSphere Replication

SPR-SRM-NET-003Forward/reverse DNS records.Security

Decision: Forward/reverse DNS records.

Rationale: FQDN accessibility.

Implication: Maintain DNS records; allow DNS through firewalls.

Component: Site Recovery Manager

SPR-VR-NET-003Forward/reverse DNS records.Security

Decision: Forward/reverse DNS records.

Rationale: FQDN accessibility.

Implication: Maintain DNS records; allow DNS through firewalls.

Component: vSphere Replication

SPR-SRM-NET-004Configure NTP, not VMTools sync.AvailabilitySecurity

Decision: Configure NTP, not VMTools sync.

Rationale: Accurate time sync; prevents mismatch.

Implication: NTP HA and firewall allow.

Component: Site Recovery Manager

SPR-VR-NET-004Configure NTP, not VMTools sync.AvailabilitySecurity

Decision: Configure NTP, not VMTools sync.

Rationale: Accurate time sync; prevents mismatch.

Implication: NTP HA and firewall allow.

Component: vSphere Replication

SPR-SRM-LCM-001LCM via native appliance tools, not SDDC Manager.Manageability

Decision: LCM via native appliance tools, not SDDC Manager.

Rationale: Not integrated with SDDC Manager LCM.

Implication: Manual LCM.

Component: Site Recovery Manager

SPR-VR-LCM-001LCM via native appliance tools, not SDDC Manager.Manageability

Decision: LCM via native appliance tools, not SDDC Manager.

Rationale: Not integrated with SDDC Manager LCM.

Implication: Manual LCM.

Component: vSphere Replication

SPR-SRM-RP-001Prioritized startup order: Aria Suite Lifecycle + WS1 Access → Operations analytics → Operations remManageabilityRecoverability

Decision: Prioritized startup order: Aria Suite Lifecycle + WS1 Access → Operations analytics → Operations remote collectors → Aria Automation.

Rationale: Dependency ordering; RTO target.

Implication: Maintain customized plans; VMware Tools required on each VM.

Component: Site Recovery Manager

SPR-DNS-NET-001Configure DNS settings for each protected component to use DNS servers across both VCF instances.Manageability

Decision: Configure DNS settings for each protected component to use DNS servers across both VCF instances.

Rationale: Resolve DNS during planned migration or DR.

Implication: Update DNS settings as you scale from single to multi-instance.

Component: DNS

SPR-NTP-NET-001Configure NTP settings for each protected component to use NTP servers across both VCF instances.Manageability

Decision: Configure NTP settings for each protected component to use NTP servers across both VCF instances.

Rationale: Resolve NTP during migration/DR.

Implication: Update NTP settings as you scale.

Component: NTP

Prerequisites

  • VCF healthy per Support Matrix (protected and recovery instances); same VCF version.
  • Captured Site Protection and DR tab of Planning & Preparation Workbook.
  • NSX Federation deployed; cross-instance overlay segment promoted via NSX Global Manager.
  • Standalone Tier-1 gateway in recovery region for LB failover.
  • Mgmt cluster in recovery region with matching compute/storage capacity.
  • External services available in recovery: AD DCs, DNS, NTP, SMTP, Syslog.
  • SRM and vSphere Replication OVAs downloaded.
  • Implementation Methods
  • Powershell

Implementation Procedure

Implementation

PowerValidatedSolutions module — end-to-end deployment and configuration of SRM and vSphere Replication.

UI

Per VCF version: Deploy vSphere Replication in protected and recovery mgmt vCenters (use vCenter OVF deploy). Deploy SRM in protected and recovery; pair sites. Configure placeholder datastore mappings, network mappings, folder mappings. Create protection groups for each management component grouping. Create recovery plans with prioritized startup order + prompts. Customize recovery plan for Aria Suite Lifecycle + WS1 Access (includes VMware Aria Suite Lifecycle PowerOn workflow + vidmPowerOnSkip retry). Create anti-affinity rules for placeholder VMs. Create VM groups with restart order on recovery site. If applicable, back up/restore OVF properties of protected VMs.

External Services / Integration Points

Active Directory

DNS

NTP

SMTP

Syslog

Configuration Values

RPO15 minutes

PIT3 copies over 24 hours

Network CompressionEnabled

Guest QuiescingDisabled

Protection Groups3 groups covering Aria Suite Lifecycle + WS1 Access, Operations analytics, Aria Automation
RTO≤ 4 hours

Additional Instance

Not applicable — solution always spans two VCF instances (primary + recovery).

Day-2 Operations Tasks

Operations

As needed

Personas

As needed

NameRole

As needed

Cloud AdminSRM Administrator; recovery plan authoring and execution

As needed

AuditorRead-only access to recovery plans and history

As needed

Operational Verification

As needed

Certificate Management

As needed

After generating CA-signed certs for SRM and vSphere Replication, update certs on connected management components (paired SRM, vCenter registration) t

As needed

Password Management

As needed

Policies

As needed

Monitoring Points

  • Site Reliability EngineerRecovery plan execution and monitoring
  • Verify SRM pairing across VCF instances — Site Recovery page shows OK (green check) for both sites.
  • Verify vSphere Replication pairing — Site Recovery page shows green tick for VR subsection on both sites.
  • MonitoringAria Operations consumes SRM via non-native VMware Site Recovery Manager management pack (not packaged with Aria Operations).
  • Failover Checklist
  • LogisticsFacilities/personnel available; SRM available in recovery; replication status healthy; NSX state (Edge nodes, overlay IPs, LB config) verifie
  • Post FailoverUpdate DNS for VAOL; reconfigure Operations logging target; inventory sync in Aria Suite Lifecycle; verify each component.
  • ProcedureIn SRM recovery plan: right-click → Reprotect → confirm irreversible. If Reprotect Interrupted, use Force cleanup option. After Ready state,

Troubleshooting

ActivationConfirm DR/failback required (e.g., extended outage, scheduled maintenance, disaster).
Cause:
Fix:
Multi AZFor failback DR where recovery VCF instance remains unavailable, vSAN witness may be unavailable — use force-provisioning in storage policy.
Cause:
Fix:
Post FailoverRedirect logs to VAOL in recovery site; complete post-recovery assessment.
Cause:
Fix:
ExecutionRun recovery plans as Planned Migration type: first Aria Suite Lifecycle + WS1 Access (with retry using vidmPowerOnSkip=true after VMware Ari
Cause:
Fix:
NSX LB Failover
Cause:
Fix:

Likely Panelist Questions

Q: Why did you choose this architecture?

See design decisions for rationale

Failure Scenarios

NSX LB failover via standalone Tier-1 attach/detach on cross-instance segment.
Impact:
Mitigation:
Aria Suite Lifecycle Power On workflow requires vidmPowerOnSkip=true retry after failover (WS1 Access VMs relocated).
Impact:
Mitigation:
Cloud Proxies not failed over — must remain instance-local (breaks centralized monitoring during DR).
Impact:
Mitigation:
Forgetting to update DNS/NTP to cross-instance settings before failover (SPR-DNS-NET-001/SPR-NTP-NET-001).
Impact:
Mitigation:
Aria Suite Lifecycle: Power On/Off workflows with vidmPowerOnSkip=true retry for post-failover WS1 Access.
Impact:
Mitigation:

Trade-off Analysis

Trade-Offs Analysis

Chosen:

Justification:

Quiz — Site Protection & DR

0/15
Q1
Which replication technology is used in this design?
  • Array-based replication only
  • vSphere Replication
  • Both simultaneously
  • vSAN stretched cluster
SPR-SRM-CFG-004: vSphere Replication for flexibility across vendors; no SRA maintenance.
Q2
What is the RTO target?
  • 1 hour
  • 4 hours or less
  • 8 hours
  • 24 hours
SPR-SRM-CFG-005 and recovery plan design target RTO ≤ 4 hours.
Q3
What is the configured RPO?
  • 5 minutes
  • 15 minutes
  • 1 hour
  • 24 hours
SPR-VR-CFG-005: 15 minutes.
Q4
How many Point-in-Time snapshots are retained, and over what period?
  • 1 snapshot / 1h
  • 3 snapshots / 24h
  • 5 snapshots / 12h
  • 10 snapshots / 7d
SPR-VR-CFG-006: 3 PITs over 24 hours.
Q5
Which component is NOT failed over?
  • VCF Operations analytics nodes
  • Aria Automation
  • VCF Operations Cloud Proxies
  • WS1 Access
Cloud Proxies are instance-local and remain in each VCF instance; only analytics nodes are failed over.
Q6
Why are recovery plan tests NOT run?
  • Too expensive
  • Protected apps use NSX LB and DNS unavailable in isolated test network
  • SRM doesn't support test
  • Not supported on VCF
SPR-SRM-RP-005: NSX LB can't come online in isolated test network; DNS unavailable.
Q7
Which SRM deployment size is used?
  • Light (1000 VMs)
  • Standard (5000 VMs)
  • Extra Large (10,000 VMs)
  • Custom
SPR-SRM-CFG-003: Light — 2 vCPU, 8 GB RAM, 1000 protected VMs.
Q8
Where are SRM and vSphere Replication placed?
  • Cross-instance NSX overlay segment
  • Local-instance NSX overlay segment
  • Management VLAN
  • VMkernel VMotion network
SPR-SRM-NET-001/SPR-VR-NET-001: Management VLAN, not NSX overlay.
Q9
What is the startup priority order?
  • Aria Automation → WS1 Access → Operations
  • Aria Suite Lifecycle → WS1 Access → Operations analytics → Operations remote collectors → Aria Automation
  • Operations → Aria Automation → WS1 Access
  • WS1 Access → Aria Automation → Aria Suite Lifecycle
SPR-SRM-RP-001..004: ASL+WS1 Access first for LCM/auth, then Operations, then Aria Automation.
Q10
What flag must be set to true when retrying WS1 Access Power On after failover?
  • adminResetSkip
  • vidmPowerOnSkip
  • srmSkipDNS
  • vraSkipK8s
After failover, Aria Suite Lifecycle can't locate WS1 Access VMs; retry with vidmPowerOnSkip=true.
Q11
Why is guest OS quiescing disabled?
  • Not all mgmt VMs support quiescing; could cause outage
  • It would increase RPO
  • Performance impact on ESXi
  • VCF forbids it
SPR-VR-CFG-003: quiescing could cause outages and not all mgmt VMs support it; crash-consistent replicas are acceptable.
Q12
Who performs LCM of SRM and vSphere Replication?
  • SDDC Manager
  • Aria Suite Lifecycle
  • Native appliance tools (manual)
  • vCenter Server
SPR-SRM-LCM-001/SPR-VR-LCM-001: Not integrated with SDDC Manager; use native tools.
Q13
What must be done after DR if Cloud Foundation Operations comes back up?
  • Nothing
  • Reprotect VMs and re-point logs to recovery VAOL
  • Redeploy all Operations nodes
  • Uninstall Cloud Proxies
Post-failover: update DNS for VAOL, reconfigure logging target, then reprotect.
Q14
When must force-provisioning be enabled in vSAN storage policy?
  • Always
  • For failback DR where recovery site's vSAN witness is unavailable in multi-AZ
  • For test recoveries
  • Never
Failover checklist notes: if recovery VCF instance remains unavailable in failback DR, witness may be unavailable; force-provisioning ensures VM provisioning succeeds.
Q15
After DR, what is the reprotect prerequisite sequence?
  • Immediately reprotect
  • Run recovery plans in 'Recovery Required' state → planned migration → reprotect
  • Reinstall SRM always
  • No reprotect needed
Reprotect prerequisites: after DR, run Recovery Required plans, then planned migration, then reprotect.

Flashcards — Site Protection & DR

Card 1 of 15
Replication technology choice?
vSphere Replication (not array-based) — no SRA dependency; heterogeneous storage.

Labs

Pair SRM + vSphere Replication across VCF instances

Deploy SRM and vSphere Replication in both instances, pair sites, validate.

Starting State: Two VCF instances operational with mgmt domains; NSX Federation configured.

Create protection groups and recovery plans with prioritized startup

Build 3 protection groups and 3 recovery plans covering Aria Suite Lifecycle + WS1 Access, Operations analytics, Aria Automation.

Starting State: SRM/VR paired; Aria Suite products deployed.

Planned migration + reprotect drill

Fail over all mgmt apps to recovery site and then reprotect to prepare for failback.

Starting State: Recovery plans Ready; management apps healthy on protected site.

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.