Multi-Site Design Decision Framework
Objectives
- Evaluate and select multi-site architecture patterns based on RTO/RPO requirements
- Design failover workflows for tiered applications with different SLAs
- Select appropriate replication technology for each workload tier
- Create a multi-site decision framework that maps requirements to patterns
- Document failback procedures and data consistency verification
Prerequisites
Access to VCF 9.0, SRM, HCX, and vSphere Replication documentation
Prior labs: vcp-architect-01, vcp-architect-04
Required skills:
- DR and BC concepts (RTO, RPO, failover, failback)
- vSAN stretched cluster fundamentals
- SRM/vSphere Replication basics
- HCX migration capabilities
Lab Environment
Design exercise — compare multi-site patterns for different requirements scenarios
Tasks
Task 1 Multi-Site Pattern Evaluation Matrix
Create a comprehensive decision framework that maps business requirements (RTO/RPO/cost/distance) to the optimal multi-site architecture pattern.
Document all VCF-native multi-site patterns:
- vSAN Stretched Cluster (Active-Active Synchronous):
- RPO: 0 (synchronous replication)
- RTO: Seconds (automatic HA restart)
- Max distance: ~100km (≤5ms RTT)
- Cost: High (2× hosts, dedicated inter-site link)
- Complexity: High (witness, partition handling)
- Best for: Zero data loss, automatic failover
- VMware Site Recovery (SRM) with vSphere Replication:
- RPO: 5 minutes to 24 hours (configurable)
- RTO: 15-60 minutes (automated recovery plan execution)
- Max distance: Unlimited (asynchronous)
- Cost: Medium (fewer DR hosts, standard WAN)
- Complexity: Medium (recovery plans, testing)
- Best for: Cost-effective DR with acceptable data loss
- VMware Site Recovery with Array-Based Replication:
- RPO: Near-zero to minutes (depends on array)
- RTO: 15-60 minutes (SRM recovery plan)
- Max distance: Array-dependent
- Cost: High (array licenses, replication)
- Complexity: Medium-High (array integration)
- Best for: Existing array investment, low RPO requirement
- HCX Disaster Recovery:
- RPO: 5 minutes to 24 hours (similar to vSphere Replication)
- RTO: Minutes (automated failover)
- Max distance: Unlimited (supports cloud DR)
- Cost: Medium (HCX license)
- Complexity: Low-Medium (simplified mobility)
- Best for: Cloud-based DR, heterogeneous environments
- Pilot Light:
- RPO: Hours (backup-based, last backup point)
- RTO: 2-8 hours (power up infrastructure + restore)
- Max distance: Unlimited
- Cost: Low (minimal DR infrastructure, powered off)
- Complexity: Low (standard backup/restore)
- Best for: Lowest cost DR for non-critical workloads
- Backup-Only (Cold DR):
- RPO: 24 hours (daily backup)
- RTO: 8-24 hours (provision + restore)
- Max distance: Unlimited
- Cost: Minimal (backup storage only)
- Complexity: Low
- Best for: Archival, regulatory backup requirement
Create the decision framework matrix:
| Requirement | Stretched | SRM/vSR | SRM/Array | HCX DR | Pilot Light | Backup |
|---|---|---|---|---|---|---|
| RPO = 0 | ✅ | ❌ | ⚠️ | ❌ | ❌ | ❌ |
| RPO < 15 min | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ |
| RPO < 4 hours | ✅ | ✅ | ✅ | ✅ | ⚠️ | ❌ |
| RTO < 1 min | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
| RTO < 30 min | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ |
| RTO < 4 hours | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ |
| Distance > 100km | ❌ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Cloud DR target | ❌ | ⚠️ | ❌ | ✅ | ✅ | ✅ |
| Lowest cost | ❌ | ⚠️ | ❌ | ⚠️ | ✅ | ✅ |
| Auto failover | ✅ | ⚠️ | ⚠️ | ⚠️ | ❌ | ❌ |
| Single vendor (VCF) | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ |
Legend: ✅ = Fully supported, ⚠️ = Partial/with limitations, ❌ = Not supported
How to use: Map customer requirements → check columns → select pattern with most ✅ for their constraints.
Design tiered DR strategy for 600-VM environment:
Principle: Not all VMs need the same DR level. Match DR pattern to business criticality.
| Tier | VMs | RPO Required | RTO Required | DR Pattern | Justification |
|---|---|---|---|---|---|
| Tier-1 | 80 | 0 | <1 min | Stretched Cluster | Zero data loss, auto failover |
| Tier-2 | 300 | <15 min | <30 min | SRM + vSphere Replication | Cost-effective, automated recovery |
| Tier-3 | 220 | <24 hours | <4 hours | Pilot Light (backup) | Lowest cost, non-critical |
Cost optimization:
- Only Tier-1 (80 VMs) needs expensive stretched cluster (16 hosts × 2 sites)
- Tier-2 uses SRM with 4-6 DR hosts (not 1:1 ratio — accept degraded performance during DR)
- Tier-3 uses backup only — no dedicated DR hosts, restore to Tier-2 DR hosts if needed
Total DR infrastructure:
- Primary site: 21 hosts (management + production + dev)
- DR site: 12 hosts (Tier-1 stretched: 8 hosts shared with primary, Tier-2 DR: 4 hosts)
- Witness site: 1 appliance
Document inter-site latency impact on pattern selection:
Latency measurement guide:
- Tool: iPerf3 for bandwidth, ping with timestamps for RTT
- Measure at multiple times of day (RTT varies with network load)
- Measure for sustained period (24+ hours) to capture worst case
| Latency (RTT) | Distance | Viable Patterns |
|---|---|---|
| ≤1ms | <30km | All patterns (stretched cluster optimal performance) |
| 1-5ms | 30-100km | All patterns (stretched cluster functional, some write penalty) |
| 5-10ms | 100-300km | NOT stretched cluster; SRM/HCX/Pilot Light/Backup |
| 10-50ms | 300-2000km | SRM/HCX/Pilot Light/Backup |
| >50ms | Intercontinental | HCX/Pilot Light/Backup (SRM possible but slow failover) |
Design decision rule:
- If RPO=0 AND RTT ≤5ms → Stretched Cluster - If RPO=0 AND RTT >5ms → Cannot achieve; negotiate RPO relaxation - If RPO >0 AND RTT any → SRM/HCX/Pilot Light based on RTO and cost
Validation Gate
Check: Multi-site decision framework created with all patterns evaluated
Expected: 6 patterns documented with RPO/RTO/cost/distance characteristics, decision matrix with requirement mapping, tiered DR strategy for 600-VM environment, latency-based pattern selection guide
Common Errors
Task 2 Failover Workflow Design for Tiered Applications
Design detailed failover and failback workflows for each DR tier, including application startup sequencing and data consistency verification.
Design Tier-1 failover workflow (Stretched Cluster — automatic):
- Trigger: Site A failure detected by vSAN and HA
- vSAN detects host unreachability (5 heartbeat misses × 1s = 5s)
- HA declares hosts as failed (30s default isolation timeout)
- Total detection time: ~35 seconds
- Automatic actions (no human intervention):
- vSAN: Witness votes for Site B (preferred site was A, but A is down)
- HA: Restarts Tier-1 VMs on Site B hosts (priority order)
- DRS: Distributes restarted VMs across available Site B hosts
- DFW: Rules follow VMs (pushed to new host kernel modules)
- Restart priority order:
a. Database VMs (highest) — data consistency depends on clean startup
b. Application VMs (high) — reconnect to databases after DB is up
c. Web VMs (medium) — front-end reconnects to app tier
- Post-failover validation (manual, by on-call team):
- [ ] All 80 Tier-1 VMs running on Site B
- [ ] Database consistency check (application-specific: MySQL → check InnoDB recovery log)
- [ ] Application health checks passing (HTTP 200 from health endpoints)
- [ ] DNS: If using GSLB, verify Site B endpoints are active
- [ ] Monitoring: Aria Operations shows all VMs on Site B, no lingering alerts
- [ ] Communication: Notify stakeholders of site failure and successful failover
Design Tier-2 failover workflow (SRM — semi-automated):
- Trigger: Site A declared unavailable by operations team
- NOT automatic — requires human decision to invoke SRM recovery plan
- Reason: Avoiding false failover due to transient network issue
- Pre-failover checklist:
- [ ] Verify Site A is truly unavailable (out-of-band verification via IPMI/iLO)
- [ ] Verify last successful replication timestamp (RPO check)
- [ ] Notify stakeholders: planned failover in progress
- [ ] Confirm DR site capacity can handle Tier-2 workload
- SRM Recovery Plan execution:
a. SRM powers off placeholder VMs at DR site (if any)
b. SRM reconfigures replicated VMs:
- Update IP addresses (if DR uses different subnet) or rely on NSX L2 extension
- Mount replicated VMDKs
- Adjust VM settings (if different hardware profile at DR)
c. Power on VMs in priority groups:
- Group 1: Infrastructure services (AD, DNS, monitoring)
- Group 2: Database VMs (delay: 120s for service stabilization)
- Group 3: Application VMs (delay: 60s)
- Group 4: Web/front-end VMs (delay: 30s)
d. Run recovery plan custom scripts:
- Post-power-on: Verify database connectivity
- Post-recovery: Update load balancer VIP to DR site
- Expected timeline:
- Decision to invoke: 5-15 minutes (human assessment)
- SRM plan execution: 15-30 minutes (depends on VM count and startup delays)
- Total RTO: 20-45 minutes for 300 Tier-2 VMs
Design failback workflow (return to primary site):
Failback is more complex than failover because data has been modified at the DR site.
Tier-1 (Stretched Cluster) failback:
- Repair/restore Site A hosts
2. Add hosts back to cluster (SDDC Manager → Commission Host) 3. vSAN detects stale components on Site A 4. vSAN resync: Copies changed data from Site B → Site A
- Duration: Proportional to data change rate × time since failure
- Example: 10 TB changed over 24 hours, 10 Gbps link = ~2.5 hours
- After resync complete: DRS migrates VMs back to Site A (per site affinity rules)
- Verification: All components in compliance, site affinity satisfied
Tier-2 (SRM) failback:
- Ensure primary site restored and healthy
2. SRM reprotect: Reverse replication direction (DR → Primary) - SRM creates placeholder VMs at primary site - Replication begins DR → Primary (initial sync may take hours) 3. Wait for replication to be in sync 4. Execute SRM planned failback (not a forced failover — clean cutover) a. Shut down VMs at DR site b. Final replication sync (delta only) c. Power on VMs at primary site with original network config d. Verify application health 5. Re-protect: Restore replication Primary → DR
- Test: Run SRM test recovery to verify DR readiness restored
Tier-3 (Pilot Light) failback:
- Restore VMs from backup at primary site
- Verify data from last backup is acceptable
- Power on VMs at primary site
- Update monitoring to track primary site VMs
Design DR testing strategy:
- SRM recovery plan testing:
- Frequency: Quarterly
- Method: SRM test recovery (non-disruptive — creates test VMs at DR site in isolated network)
- Duration: 2-4 hours
- Validation: Application health checks, database connectivity, user acceptance
- Output: Test report documenting RTO achieved, failures, and remediation
- Stretched cluster failover testing:
- Frequency: Semi-annual
- Method: Planned maintenance failover (move all VMs to Site B via DRS)
- Duration: 4-6 hours (including failback)
- Validation: Zero data loss, VM accessibility throughout, performance baseline comparison
- Cannot do: Simulate actual site failure without planned outage (use Holodeck instead)
- Backup restore testing:
- Frequency: Monthly
- Method: Restore random sample of VMs from backup to isolated network
- Validation: VM boots, application starts, data is recent and consistent
- Output: Restore success/failure report, backup tool health
- DR test calendar:
| Month | Test Type | Scope | Estimated Duration |
|---|---|---|---|
| Jan | SRM Test Recovery | All Tier-2 VMs | 3 hours |
| Feb | Backup Restore Spot Check | 10 random VMs | 2 hours |
| Mar | SRM Test Recovery | Subset (50 VMs) | 1.5 hours |
| Apr | Stretched Cluster Planned Failover | All Tier-1 VMs | 5 hours |
| May | Backup Restore Spot Check | 10 random VMs | 2 hours |
| Jun | SRM Test Recovery + Full Failback Test | All Tier-2 VMs | 6 hours |
| (repeat pattern H2) |
Validation Gate
Check: Failover, failback, and DR testing workflows documented for all tiers
Expected: Tier-1 automatic failover with 35s detection + HA restart, Tier-2 SRM workflow with 20-45 min RTO, failback procedures for all tiers, quarterly DR test calendar
Common Errors
Task 3 Replication Technology Selection & Configuration
Select and configure the appropriate replication technology for each workload tier with bandwidth sizing and RPO verification.
Configure vSphere Replication for Tier-2 workloads:
vSphere Replication architecture:
- vSphere Replication Appliance (VRA): Deployed at both sites
- Replication: VM-level, asynchronous, host-based (no array dependency)
- RPO: 5 minutes minimum, configurable per VM
- Protocol: TCP/31031 (encrypted by default in 9.0)
- Replication data flow: ESXi host → VRA (source) → VRA (target) → target datastore
Configuration for Tier-2 (300 VMs):
- RPO: 15 minutes (balance between data protection and bandwidth)
- Replication schedule: Continuous (not scheduled — delta-based)
- Bandwidth calculation:
- Average daily change rate: 5% of VM storage per day
- Tier-2 storage: 300 VMs × 150 GB = 45 TB
- Daily change: 45 TB × 5% = 2.25 TB/day
- Hourly: 2.25 TB / 24 = 93.75 GB/hour
- Bandwidth needed: 93.75 GB × 8 / 3600 = ~208 Mbps sustained
- With overhead (protocol + encryption): ~250 Mbps
- Available WAN: 1 Gbps → sufficient with headroom
- Seed copy: For initial replication of 45 TB
- Over 1 Gbps WAN: 45 TB × 8 / 1000 Mbps = ~100 hours (4+ days)
- Alternative: Ship physical disk with initial copy, then switch to network replication
Design SRM recovery plan configuration:
SRM recovery plan structure:
- Protection groups:
- PG-Tier2-Databases: All Tier-2 database VMs (30 VMs)
- PG-Tier2-AppServers: All Tier-2 application VMs (180 VMs)
- PG-Tier2-WebServers: All Tier-2 web servers (90 VMs)
- Recovery plan: RP-Tier2-Full-Recovery
- Priority Group 1: PG-Tier2-Databases
- Pre-power-on steps: Check target datastore capacity
- Power-on delay: 0 seconds
- Post-power-on steps: Wait for SQL/MySQL service to respond (timeout: 300s)
- Priority Group 2: PG-Tier2-AppServers
- Pre-power-on steps: None
- Power-on delay: 120 seconds (wait for databases)
- Post-power-on steps: HTTP health check on port 8080 (timeout: 120s)
- Priority Group 3: PG-Tier2-WebServers
- Pre-power-on steps: None
- Power-on delay: 60 seconds (wait for app servers)
- Post-power-on steps: HTTP health check on port 443 (timeout: 60s)
- Network mapping:
- Option A: Same IP addresses at DR (if using NSX L2 extension or stretched VLANs)
- Option B: IP customization (if DR uses different subnets)
- SRM handles: IP re-addressing, default gateway, DNS server update
- Not handled: Application-level connection strings (requires DNS update or config management)
- Design decision: Use NSX L2 extension for critical VMs (no IP change needed)
Compare replication technologies for edge cases:
| Feature | vSphere Replication | vSAN Stretched | HCX DR | Array Replication |
|---|---|---|---|---|
| Granularity | Per-VM | Per-cluster | Per-VM | Per-LUN/volume |
| RPO minimum | 5 minutes | 0 (synchronous) | 5 minutes | Near-zero |
| Network requirement | TCP (any WAN) | ≤5ms RTT | TCP (any WAN) | Array-specific |
| Storage dependency | None (host-based) | vSAN only | None | Specific array |
| VMs per appliance | ~500 | N/A | ~1,000 per HCX Mgr | N/A |
| Bandwidth efficient | Delta-based | Full write mirroring | Delta-based | Array-level dedup |
| VCF integration | Native (SRM) | Native (SDDC Mgr) | VCF add-on | Plugin/SRA |
| Crash consistency | Per-VM | Per-cluster | Per-VM | Per-LUN |
| App consistency | With VSS/quiesce | Not native | With VSS/quiesce | With agent |
Decision rule:
- If vSAN and RPO=0: Stretched cluster
- If vSAN and RPO>0: vSphere Replication (native, no array dependency)
- If external storage: Array replication OR vSphere Replication
- If cloud DR target: HCX DR (best cloud mobility)
- If cross-platform (VMware → non-VMware): HCX (heterogeneous support)
Design RPO verification and monitoring:
- vSphere Replication monitoring:
- Dashboard: Aria Operations → vSphere Replication management pack
- Key metrics:
- RPO violation count (VMs exceeding configured RPO)
- Replication lag (time since last successful sync)
- Transfer rate (MB/s per VM)
- Alerts:
- RPO breach >2× configured value → Warning
- RPO breach >4× configured value → Critical
- Replication error (connection lost) → Critical- Stretched cluster monitoring:
- vSAN Health → Stretched Cluster section
- Inter-site latency trending
- Resync progress (during recovery)
- Component health across sites
- Compliance reporting:
- Monthly RPO compliance report:
- % of time each VM met its RPO target
- Number of RPO violations per VM
- Root cause of violations (bandwidth saturation, host issues)
- Target: 99.9% RPO compliance for Tier-2 VMs
- Bandwidth monitoring:
- Track actual replication bandwidth vs provisioned
- If actual approaches 80% of available: trigger capacity alert
- Trending: Project when bandwidth will be insufficient (growth-adjusted)
Validation Gate
Check: Replication technology selected and configured for each tier with bandwidth sizing and RPO monitoring
Expected: vSphere Replication configured for Tier-2 with bandwidth math, SRM recovery plan with priority groups and startup delays, technology comparison table, RPO monitoring with alert thresholds
Common Errors
Task 4 Multi-Site Design Summary & VCDX Defense
Consolidate the multi-site design into a defensible architecture summary with design decisions and VCDX panel preparation.
Create design decision D-011 — Multi-Site DR Strategy:
Decision: Tiered DR with stretched cluster (Tier-1), SRM (Tier-2), backup (Tier-3)
Alternatives:
A) Stretched cluster for ALL VMs:
- RPO=0 for everything
- Cost: Extreme (600+ VMs × 2 sites, 30+ hosts per site)
- Rejected: Cost prohibitive; Tier-2/3 do not require RPO=0
B) SRM for ALL VMs:
- RPO=5-15 min for everything
- Cost: Moderate (DR hosts need not be 1:1)
- Rejected: Tier-1 requires RPO=0; SRM cannot provide this
C) HCX DR for ALL VMs:
- Similar to SRM capability set
- Pro: Simplified management, cloud DR option
- Con: HCX DR is newer, less mature DR testing framework than SRM
- Rejected: SRM has more robust recovery plan testing and compliance reporting
D) Backup-only for ALL VMs:
- RPO=24 hours, RTO=8-24 hours
- Cost: Minimal
- Rejected: Unacceptable for Tier-1 and Tier-2 SLAs
Justification: Tiered approach optimizes cost by matching DR investment to business criticality.
Cost savings: ~45% reduction vs all-stretched, with Tier-1 still at RPO=0.
Prepare VCDX defense responses:
Challenge 1: 'Your Tier-2 has RPO=15 min. What if the customer says they need RPO=5 min?'
Response: vSphere Replication supports RPO=5 min. Change is configuration-only. However, bandwidth impact increases 3×: 250 Mbps → ~750 Mbps. Verify WAN capacity before committing. If WAN insufficient, either upgrade link or negotiate RPO=10 min as compromise.
Challenge 2: 'What is the total cost of the multi-site design?'
Response: Primary: 21 hosts + DR: 12 hosts + Witness: 1 appliance = 33 hosts total + 1 Gbps WAN + 100 Gbps inter-site dark fiber. Licensing: 33 × 64 cores × per-core cost. Compare to single-site: 21 hosts — DR adds ~57% to host count but protects $X million in application availability. Business case: If Tier-1 outage costs $Y/hour, stretched cluster pays for itself after Z hours of prevented downtime.
Challenge 3: 'How do you handle a DR test for Tier-1 stretched cluster?'
Response: Cannot simulate site failure without actual outage. Alternative: Planned failover — DRS evacuates all Tier-1 VMs to Site B, verify application health, then fail back. Additionally, test partition scenarios on Holodeck to validate failure matrix without production risk.
Challenge 4: 'What if the customer adds a regulatory requirement for a third geographic region?'
Response: vSAN stretched cluster supports 2 data sites only. For 3-region DR: Deploy primary at Region A, stretched cluster to Region B (same country/jurisdiction), asynchronous SRM replication to Region C (different jurisdiction). This gives RPO=0 within jurisdiction and RPO=15min to remote region.
Create the multi-site architecture summary:
- Sites:
- Primary (Site A): 21 hosts — Management (4) + Production (13) + Dev/Test (4)
- DR (Site B): 12 hosts — Stretched mirror for Tier-1 (8, shared with Site A count) + DR for Tier-2 (4)
- Witness (Site C): 1 Medium OVA appliance
- DR pattern mapping:
| Tier | VMs | DR Pattern | RPO | RTO | DR Hosts | Replication Technology |
|---|---|---|---|---|---|---|
| 1 | 80 | Stretched Cluster | 0 | <1 min | 8 (shared) | vSAN synchronous |
| 2 | 300 | SRM | 15 min | 30 min | 4 | vSphere Replication |
| 3 | 220 | Pilot Light | 24 hours | 4 hours | 0 (shared) | Backup restore |
- Network:
- Inter-site (A↔B): 100 Gbps dark fiber (stretched cluster + replication) - WAN (A→B for SRM): 1 Gbps MPLS (replication traffic) - Witness (C↔A/B): 1 Gbps MPLS
- Testing schedule:
- Monthly: Backup restore spot check (10 VMs)
- Quarterly: SRM test recovery (all Tier-2)
- Semi-annual: Stretched cluster planned failover
- Monitoring:
- Aria Operations dashboards for replication health, RPO compliance, inter-site latency
- Alert thresholds documented and configured
Final design validation checklist:
- [ ] Every VM has an assigned DR tier with defined RPO/RTO
- [ ] DR infrastructure sized for the workload it must protect
- [ ] Bandwidth calculated and sufficient for replication load
- [ ] Failover procedure documented for each tier with startup order
- [ ] Failback procedure documented (including SRM reprotect)
- [ ] DR testing calendar established with test types and frequency
- [ ] RPO monitoring configured with alert thresholds
- [ ] Cost comparison documented (tiered vs uniform DR)
- [ ] Design decision register entry with alternatives and justification
- [ ] VCDX defense responses prepared for 4+ challenge questions
- [ ] Architecture diagram shows all sites, connectivity, and data flow
- [ ] Runbook: On-call procedure for site failure detection → failover invocation → verification
Validation Gate
Check: Multi-site design consolidated with design decision, defense responses, architecture summary, and validation checklist
Expected: Design decision with 4+ alternatives, 4+ defense responses, complete architecture summary with site/pattern/network details, 12+ item validation checklist
Common Errors
Final Validation
Complete multi-site design decision framework with pattern selection, failover workflows, replication configuration, and VCDX-ready documentation
✓ All multi-site patterns evaluated → 6 patterns documented with RPO/RTO/cost/distance characteristics
✓ Tiered DR strategy designed → Each tier has assigned DR pattern with infrastructure sizing
✓ Failover/failback workflows complete → Step-by-step procedures for all tiers, including application-level recovery
✓ Replication configured and monitored → Bandwidth sized, RPO monitoring configured, DR testing scheduled
Cleanup / Restore
• Save all DR design documentation, failover procedures, and test calendar
• If using Holodeck: Remove SRM configuration and revert to base snapshot
Design Reflection (VCDX)
Multi-site design is the capstone of VCF architecture. VCDX panelists assess: (1) Do you know ALL the patterns? Not just the one you chose — you must articulate why each alternative was rejected. (2) Is the DR strategy tiered? Protecting everything equally shows poor cost optimization. (3) Can you execute? Failover procedures must be specific — startup order, database consistency checks, DNS updates — not just 'VMs restart.' (4) Do you test? An untested DR plan is theatre. Quarterly tests with documented results prove operational readiness.
Requirements
- RPO=0 for Tier-1 (80 VMs), RPO<15min for Tier-2 (300 VMs), RPO<24hrs for Tier-3 (220 VMs)
- Automatic failover for Tier-1 (no human intervention needed)
- DR test capability without production impact
- Full failback procedure with data consistency verification
Constraints
- Inter-site RTT ≤5ms for stretched cluster (limits distance to ~100km)
- WAN bandwidth limited to 1 Gbps for asynchronous replication
- DR budget is 40% of primary infrastructure budget
- Quarterly DR testing window of 6 hours maximum
Assumptions
- Dark fiber available between primary and DR sites
- DR site has sufficient power and cooling for full failover load
- Business accepts different RPO/RTO per tier
- vSphere Replication can sustain 300 VMs on single appliance pair
Risks
- DR site capacity insufficient during peak failover — not all Tier-2/3 VMs may start
- Replication lag exceeds RPO during high-change periods (month-end processing)
- SRM recovery plan drift — plan not updated when new VMs are added to production
- Failback takes longer than expected due to large data resync volume
Self-Assessment Discussion Prompts
- What if the customer's insurance requires annual DR test with full failover and 72-hour operation at DR site?
- How would you modify the design if the customer wants DR in AWS/Azure instead of a physical DR site?
- What happens to Tier-3 VMs during a Tier-1 stretched cluster failover — do they compete for resources at Site B?
- If you could only pick ONE DR technology for all tiers, which would it be and why?
Extensions
Cloud-Based DR with HCX and VMware Cloud
Design a DR solution using HCX to replicate workloads to VMware Cloud on AWS. Document the HCX network extension setup, replication scheduling, cloud compute sizing, and cost model (reserved vs on-demand instances). Compare TCO of cloud DR vs physical DR site over 3-year lifecycle.
Ransomware Recovery Architecture
Extend the multi-site design with ransomware-specific recovery capabilities. Design isolated recovery environment (IRE), immutable backups with air-gap, recovery validation in sandbox, and clean-room restore procedure. Document how stretched cluster, SRM, and backup integrate into the ransomware response plan.
Automated DR Orchestration with Aria Automation
Build an automated DR orchestration workflow using Aria Automation that monitors replication health, automatically triggers failover based on configurable thresholds, and executes post-failover validation scripts. Include Slack/Teams notification integration for status updates.
⚠ Known Pitfalls (from Community KB)
References
- VMware Site Recovery Manager 9.0 Administration Guide
- vSphere Replication 9.0 Administration Guide
- HCX Documentation — Disaster Recovery chapter
- vSAN 9.0 Stretched Cluster Guide
- VMware VCF 9.0 Multi-Site Architecture Reference