Academy/VCP-VCF 9.0 Architect (2V0-13.25)/Disaster Recovery Architecture Design
This lab targets VCF 9.0

Disaster Recovery Architecture Design

VCF 9.0Advancedvcp-architect⏱ 90 min

Comprehensive DR architecture design covering RTO/RPO classification, multi-tier topology design, failover validation, and RCAR-based decision documentation

Objectives

  • Classify workloads by RTO/RPO tiers and map to VCF DR capabilities
  • Design multi-tier disaster recovery topology with stretched clusters, SRM, and backup
  • Validate DR design against failure scenarios and RPO compliance
  • Document DR architecture decisions using RCAR framework

Prerequisites

VCF 9.0 lab environment deployed with 2+ vSphere clusters, NSX, vSAN storage, and SRM installed

Prior labs: vcp-architect-04: vSphere Clustering and Stretch Design, vcp-architect-05: NSX Architecture and Segmentation

Required skills:

  • vSphere 7.x/8.x cluster administration
  • Storage replication technologies (vSAN, array-based sync/async)
  • vSphere Replication and Site Recovery Manager
  • NSX cross-site networking and BGP failover
  • Bandwidth planning for synchronous replication
  • Business Impact Analysis and tier definition
  • Failover testing and runbook creation

Lab Environment

Two-site stretched VCF environment: Site-A (primary) with 4-node vSAN cluster + 3-node management cluster, Site-B (secondary) with matching compute cluster, witness host at Site-C or local storage, dedicated 10Gbps replication links with 1ms latency (simulated), NSX Edge Cluster at each site for cross-site routing and BGP failover, SRM 8.x and vSphere Replication 8.x installed and licensed at both sites

Tasks

Task 1 Classify Workloads by RTO/RPO and Map to VCF DR Capabilities

recoverability

Conduct a thorough Business Impact Analysis for the three test workloads, classify them into RPO/RTO tiers (Tier-1: RTO<5min RPO=0, Tier-2: RTO=30min RPO=1hr, Tier-3: RTO=4hrs RPO=24hrs), and document how each tier maps to specific VCF disaster recovery technologies. This foundational task ensures technical decisions are aligned with business requirements.

Step 1

Conduct a Business Impact Analysis for the three sample workloads. Document the Recovery Time Objective (RTO) and Recovery Point Objective (RPO) for the Trading Platform, Risk Analytics, and Reporting System. For each workload, record: business criticality (critical/high/medium/low), acceptable downtime in minutes, acceptable data loss in minutes/hours, and dependencies on other systems. Create a BIA spreadsheet with columns for workload name, tier assignment, RTO, RPO, max tolerable downtime, and business justification.

Completed BIA spreadsheet with all three workloads classified into tiers: Trading Platform (Tier-1), Risk Analytics (Tier-2), Reporting System (Tier-3). Each row includes RTO/RPO values, criticality level, and brief business driver. Example: Trading Platform | Tier-1 | RTO:5min | RPO:0min | Critical (no transaction loss)
Validate RTO/RPO numbers against VMware sizing guides. For Tier-1 workloads with RPO=0, synchronous replication bandwidth should not exceed 30% of link capacity. For Trading Platform with 500GB and 5min RTO, estimate required replication throughput: plan for 5-10% change rate per day, minimum 50-100 MB/s sync replication bandwidth.
Step 2

Map each workload tier to VCF DR capabilities. For Tier-1 (Trading Platform), document why vSAN stretched cluster with synchronous replication is required. For Tier-2 (Risk Analytics), explain why Site Recovery Manager with vSphere Replication and asynchronous replication suits the 30min RTO and 1hr RPO. For Tier-3 (Reporting System), justify why VADP-based backup with 24hr RPO and 4hr RTO is appropriate. Create a mapping document that shows: Tier, RTO/RPO, selected DR technology, and technology rationale.

DR Technology Mapping Document with 3 rows (one per tier) including: Tier-1→Stretched Cluster (synchronous vSAN replication, RPO achieves 0), Tier-2→SRM + vSphere Replication (async, RPO<1hr), Tier-3→VADP Backup (snapshot-based, RPO 24hrs). Include deployment rationale: stretched cluster requires 1ms latency (achievable), SRM suitable for warm standby with 30min failover, backup suitable for low-priority workloads.
Do not recommend synchronous replication for Tier-2 or Tier-3—the 1ms latency requirement and bandwidth overhead are not justified by their RTO/RPO. Conversely, do not recommend backup-only for Tier-1 trading workload; an RPO of hours is unacceptable for zero-loss transactions. Ensure RTO also factors in failover time: stretched cluster failover ~1min, SRM recovery plan ~15-30min, backup restore ~2-4hrs.
Step 3

Document VCF component prerequisites and licensing for each tier. For Tier-1, specify vSAN Enterprise license (required for stretched cluster), NSX Enterprise for cross-site BGP failover, and vSphere Replication Enterprise. For Tier-2, confirm SRM Standard or Enterprise, vSphere Replication Standard, and NSX standard for network remapping. For Tier-3, ensure vSphere Replication or VADP license, backup storage licensing (on-premises or cloud). Create a License and Component Checklist.

License Checklist table with columns: Tier, DR Technology, Required Licenses (product/edition), VCF Components (vSAN stretch, SRM, vSphere Replication, VADP), Notes. Example row: Tier-1 | Stretched Cluster | vSAN Enterprise, NSX Enterprise, vSphere Replication Enterprise | vSAN Cluster Stretch, BGP on NSX Edges | Requires 1ms site-to-site latency
Cross-reference VMware's VCF 9.0 Bill of Materials and support matrix to confirm all DR components are available in your lab VCF environment. If vSAN Enterprise is unavailable, document alternative: iSCSI arrays with third-party replication (e.g., Compellent, Pure Storage with async replication) for Tier-2.

Validation Gate

Check: All three workloads assigned to appropriate RTO/RPO tiers with documented BIA, DR technology mapping aligned to tier requirements, and VCF license prerequisites verified

Expected: BIA spreadsheet + mapping document + license checklist completed and reviewed

Common Errors

DR technology mapping shows 'asynchronous vSphere Replication' for Trading Platform, contradicting zero RPO requirement
Cause: Misunderstanding that async replication inherently has RTO of minutes/hours and RPO of 15min-hours depending on replication lag; cannot achieve RPO=0
Fix: Correct the mapping: Tier-1 must use synchronous replication only (stretched vSAN cluster or array-based sync). Verify replication lag in vSphere Replication monitor shows <5 seconds for Tier-2; if lag exceeds RPO window, increase replication link bandwidth or reduce workload change rate.
Tier-1 Trading Platform assigned to stretched cluster but bandwidth estimate is '50MB/s', which is insufficient for a 500GB workset with daily change rate
Cause: Failed to calculate required replication bandwidth: (Workload size × Daily Change Rate) / Seconds per Day = required sustained bandwidth. Typical financial systems see 5-10% daily change, so (500GB × 7.5%) / 86400s ≈ 434MB/s needed, far exceeding 50MB/s
Fix: Recalculate using realistic change rates. For Trading Platform, plan minimum 500MB/s dedicated replication link. If unavailable, reduce synchronous scope (replicate only hot data) or recommend multi-site active-active architecture instead of stretched cluster.
Tier-1 and Tier-2 mapping documents mention 'network failover' but do not address NSX Edge Cluster state replication, BGP convergence time, or stateful service failover (Firewall, NAT)
Cause: Focused only on storage replication, ignoring that NSX Control Plane and Edge stateful services are not automatically replicated; manual failover or time-consuming BGP reconvergence required
Fix: Add NSX Tier-1 Logical Router and Edge Cluster failover steps to all tiers. For Tier-1, document BGP failover convergence time (typically 30-180 seconds); if unacceptable, design active-active NSX on both sites with ECMP routing. For Tier-2, plan for manual BGP weight adjustments during SRM failover (documented in recovery plan).

Task 2 Design the DR Topology for Each Tier

availability

Create detailed architectural designs for Tier-1 (stretched cluster), Tier-2 (SRM + vSphere Replication), and Tier-3 (backup-based DR). Each design should specify replication technology, network topology including NSX considerations, failover mechanisms, and data flow paths. This task translates tier requirements into concrete infrastructure blueprints.

Step 1

Design the Tier-1 stretched cluster topology. Define the 4-node vSAN cluster split across Site-A (2 nodes) and Site-B (2 nodes) with a 1-node witness host at Site-C (or logically isolated on-prem). Document the placement: hosts A1, A2 at Site-A datastore 1; hosts B1, B2 at Site-B datastore 2; witness at Site-C (witness appliance ~40 GB footprint; stores metadata/quorum only, no VM data). Specify vSAN cluster configuration: stripes=1, replicas=2 (for 4 hosts, this ensures each site has full copy), failures to tolerate (FTT)=1. Include network topology: dedicated 10Gbps replication link between Site-A and Site-B for vSAN traffic, dedicated NSX overlay network for VM communication, separate management VLAN at each site with redundant NSX Edges for cross-site routing. Create a Visio or ASCII diagram showing host placement, replication links, and network zones.

Topology diagram (Visio or detailed ASCII art) showing: Site-A with nodes A1, A2 and local datastore; Site-B with nodes B1, B2 and local datastore; Site-C with witness; replication link (10Gbps, colored to indicate vSAN traffic); NSX Edges at each site with BGP peering; VM workload VLAN and management VLAN clearly separated. Include legend: solid line = sync replication, dashed line = management, double line = BGP adjacency. Example ASCII snippet: [Site-A: 2 nodes] ====10Gbps===== [Site-B: 2 nodes] / \ (Witness at Site-C)
Verify that the 1ms latency requirement is achievable in your lab (measure with ping, aim for <5ms including encryption overhead). If latency exceeds 5ms, stretched cluster will suffer from high latency penalties; consider recommending async SRM replication for Tier-1 as fallback, though this compromises RPO=0 goal.
Step 2

Design the Tier-2 SRM + vSphere Replication topology. Create two separate vSphere clusters: primary at Site-A (4 nodes, local vSAN storage) and secondary at Site-B (2-4 nodes, local vSAN storage, initially powered down or in standby). Document vSphere Replication configuration: Recovery Point Objective = 1 hour (set replication interval to 60 minutes or smaller for granularity), Recovery Verification Test (RVT) scheduled daily, replication network isolated to dedicated VLAN to avoid impacting production traffic. Design the SRM recovery plan: specify which VMs from Site-A will replicate to Site-B, network mapping configuration (IP address remapping from Site-A subnet 10.0.0.0/24 to Site-B subnet 10.1.0.0/24), and NSX Tier-1 router migration (power off Source Router at Site-A, activate Destination Router at Site-B, update BGP weight on destination Edge to attract traffic). Include failback plan: resync Site-A from Site-B, reverse network mappings, reactivate Site-A primary systems. Create SRM recovery plan diagram and failover playbook.

SRM Topology Diagram with: Primary Site-A cluster (4 nodes, local storage), Secondary Site-B cluster (2-4 standby nodes), vSphere Replication traffic on dedicated 1Gbps async link, SRM pair connected over TCP 443 (site pairing); vSphere Replication traffic on TCP 31031/32032, recovery plan showing VM list, network mapping (Site-A 10.0.0.0/24 → Site-B 10.1.0.0/24), NSX Router failover steps. Include failover duration estimate: 5min (SRM startup) + 5min (network failover) + 10min (application startup) = 20min total, within 30min RTO. Failback playbook with reverse network mapping and resync steps.
Do not assume automatic application startup after SRM failover; many applications require manual startup or dependency orchestration. Document expected application startup sequence in recovery plan. Also, vSphere Replication lag may reach 10-15 minutes during peak load, risking RPO violation; include monitoring alert for replication lag >60min.
Step 3

Design the Tier-3 backup-based DR topology. Specify VADP backup policy: full backup on weekends, incremental backups daily, scheduled at 2 AM to avoid business hours. Define backup retention: 7-day incremental chain for point-in-time recovery, 4-week full backup retention (for 4hr RTO requirement, restore from most recent backup). Document backup storage: 6 TB local or cloud storage for 2.5TB Reporting System with compression ratio 2:1, dual-site backup copies (Site-A primary, Site-B secondary for disaster recovery of backup infrastructure itself). Define recovery procedures: restore RTO 4 hours assumes 1.5 hours for backup retrieval + 1 hour for storage re-initialization + 1.5 hours for application startup and data consistency checks. Include backup verification testing: monthly full restore to isolated cluster, document recovery time and data integrity validation. Create backup topology diagram and RTO calculation worksheet.

Backup Topology Diagram showing: Site-A with backup appliance (e.g., Commvault, Veeam), 6TB backup storage, daily incremental + weekly full backup schedule, backup copies replicated to Site-B storage (async, daily). RTO Calculation: full restore from backup on Site-B = retrieval 1.5hr + initialize storage 1hr + app startup 1.5hr = 4hr, meeting Tier-3 RTO. RPO Document: most recent daily backup = max 24hr RPO (acceptable for non-time-sensitive reporting). Include monthly recovery drill schedule and responsible team.
Backup restore performance depends on storage IO; in lab with limited bandwidth, estimate restore time as (2.5TB × compression_ratio) / link_bandwidth. If backup link is 1Gbps (125 MB/s), restore time ≈ 2.5TB × 2 / 125MB = 40 hours (unacceptable). Plan either higher bandwidth or multi-site backup copies to minimize restore latency.
Step 4

Consolidate NSX cross-site networking requirements for all tiers. Document NSX Tier-1 Logical Router configuration: distributed routing at Site-A and Site-B with BGP peering over Edge Cluster uplinks (BGP weight 100 at primary, 200 at secondary). Specify BGP failover convergence: typical 30-180 seconds depending on router configuration. Design failover behavior for each tier: Tier-1 stretched cluster failover is near-instantaneous (vSAN quorum-based), BGP weight change can be simultaneous; Tier-2 SRM failover requires manual BGP weight adjustment as part of recovery plan (documented step); Tier-3 backup restore means manual BGP re-pointing to Site-B backup infrastructure. Document DNS/DHCP handling: specify whether Site-B has read-only or writable DNS replicas, whether DHCP relay agents exist at both sites. Create an NSX Cross-Site Networking checklist.

NSX Networking Diagram: Tier-1 Logical Router with BGP peers at Site-A (weight 100) and Site-B (weight 200), failover path showing BFD convergence (typically 50ms) + BGP update propagation (typically 30-60s). Checklist table with rows: Tier-1 (BGP auto-failover on vSAN stretch), Tier-2 (manual BGP weight in SRM plan), Tier-3 (manual BGP re-point in backup restore SOP). DNS/DHCP: confirm Site-B has DNS forwarders pointing to Site-A primary, DHCP relay agents on Site-B Edges for client requests. Example entry: Tier-1 | BGP auto-failover via BFD (50ms detection) + weight change (30s propagation) | Total failover time <2min | Acceptable for 5min RTO
If NSX cannot automatically detect failure (e.g., replication link down but BGP peers still up), configure BFD (Bidirectional Forwarding Detection) with 300ms hello interval, 3-second multiplier for 1-second failure detection. This improves failover speed from 30+ seconds (slow timer) to <2 seconds for Tier-1.

Validation Gate

Check: All three tier topologies designed with replication technology, network paths, failover mechanisms, and RTO/RPO calculations documented

Expected: Tier-1 stretched cluster topology diagram, Tier-2 SRM recovery plan diagram with failover playbook, Tier-3 backup topology with RTO calculation, NSX cross-site networking checklist

Common Errors

Topology diagram shows 10Gbps link but fails to account for latency impact: actual replication throughput drops under load due to RTO (round-trip time) windowing in vSAN replication protocol
Cause: Assumed 10Gbps link = 10Gbps usable replication bandwidth, ignoring that RAID-like encoding (stripes, replicas) and ACK-wait cycles reduce effective throughput; also, mixing vSAN, management, and other traffic on same link reduces available bandwidth
Fix: Recalculate available bandwidth: 10Gbps link × 0.8 (protocol overhead) × 0.5 (simultaneous bidirectional replication) = 4Gbps effective, sufficient for ~100 MB/s sustained replication. Alternatively, dedicate separate 10Gbps link exclusively for vSAN replication traffic and measure actual latency; if it exceeds 10ms, upgrade to 40Gbps or reduce workload change rate.
Tier-1 topology diagram shows 2 nodes at Site-A and 2 nodes at Site-B but no witness host mentioned; if replication link fails, either site could believe itself is majority and continue serving VMs, causing data corruption
Cause: Overlooked witness host requirement for even-number-of-nodes stretched clusters; vSAN must have quorum (>50% nodes) to accept writes, so with 4 nodes (2 + 2 split), cluster cannot determine majority without witness
Fix: Add witness host at neutral Site-C (witness appliance ~40 GB footprint; holds only witness components for quorum), or place witness in same site as one node if network partition is less likely (e.g., Site-C local colo). Document witness role: quorum only, no data storage, powered always-on.
SRM recovery plan shows VM network mapping and reactivation but does not mention NSX Tier-0/Tier-1 routers, firewall policies, or NAT rules; failover succeeds but traffic is blocked by default deny policies
Cause: Assumed NSX policies replicate automatically with vSphere Replication, but NSX Control Plane state (policies, routing configs) is not replicated; must be manually failed over or pre-configured identically on Site-B
Fix: Add NSX failover steps to SRM recovery plan: (1) verify NSX Tier-1 router at Site-B has identical policies as Site-A, (2) activate Site-B router (power on or state change), (3) update BGP weight on Site-B Edge to attract inbound traffic, (4) test outbound connectivity before declaring failover complete. For automated failover, use NSX Federation (if available in VCF 9.0) to replicate policies across sites.
Tier-3 RTO calculation shows 4 hours but assumes unrealistic backup retrieval speed (e.g., assumes 10Gbps throughput); actual lab tests show 40+ hour restore time with limited bandwidth
Cause: Did not account for lab network constraints: backup link shared with other traffic, backup storage is HDD-based (slow IO), network latency adds overhead
Fix: Run actual backup restore test with 2.5TB backup to Site-B, measure wall-clock time including all steps: retrieve backup metadata (5-10min), stream backup data (depends on link speed: 1Gbps = 5.5 hours for 2.5TB), initialize storage (1-2 hours), validate data integrity (30-60min). If total exceeds 4hr RTO, either upgrade to faster backup storage, increase link bandwidth, or accept longer RTO for Tier-3.

Task 3 Validate DR Design Against Failure Scenarios

recoverability

Execute practical failover tests and RPO compliance checks to verify that the DR design actually meets business requirements. This task moves from theory (design documents) to validation (tested procedures), ensuring that when a real disaster occurs, the recovery process works as documented.

Step 1

Simulate a host failure in the Tier-1 stretched cluster. Gracefully power off one node at Site-A (e.g., host A2) or ungracefully disconnect it from the network (simulate network partition by disabling vSAN heartbeat). Monitor vSAN cluster status: expected behavior is immediate failover (within 60 seconds) to remaining majority (A1 + B1 + B2 + witness = quorum). Verify that all Trading Platform VMs continue running without interruption or data loss; use vSAN observer or vsphere CLI (vsan.health -a) to confirm cluster state. Document the Recovery Time from host failure to full cluster availability: should be <30 seconds. Document any dropped packets or storage IO stalls using vSphere performance monitoring (vSAN IO latency, VM packet loss).

Lab test log documenting: (1) timestamp of host failure injection, (2) vSAN cluster status immediately after failure (should show 1 host down, cluster healthy with quorum), (3) VM uptime verification (ping or SSH to Trading Platform VM, confirm continuous operation), (4) recovery time measurement (time from failure to first IO request succeeding = <30 sec), (5) data integrity check (query Trading Platform database for committed transaction count, verify no transactions lost). Example log: '14:23:45 A2 powered off -> 14:23:58 A2 missing in vSAN, quorum verified (4 nodes) -> 14:24:02 Trading VM IO resuming, latency spike 5ms (acceptable).'
Use vSAN Observer VM to get detailed replication metrics during failure; look for Snapshot Create Delay and Write Queuing Depth spikes during convergence. For Tier-1, acceptable threshold is <1 second of write delay; if spike exceeds 5 seconds, your change rate or replication latency may exceed architecture limits.
Step 2
Test SRM failover for Tier-2 Risk Analytics workload. Create a network partition between Site-A and Site-B (simulate replication link failure) and monitor vSphere Replication lag. Execute the SRM recovery plan for Risk Analytics VMs: (1) open SRM, navigate to Protection Groups, select Risk Analytics group, verify replication status (should show 'Ready'), (2) perform a Recovery Plan Test (non-disruptive) first: SRM creates temporary copies of VMs at Site-B without affecting Site-A, verify connectivity by pinging Site-B Recovery VMs, check that network mapping (10.0.0.0/24 → 10.1.0.0/24) is working, confirm database consistency (connect to Risk Analytics DB on Site-B test copy, verify record count matches Site-A), (3) clean up test copy, (4) execute actual failover: click 'Failover', confirm all VMs are migrated to Site-B, update BGP weight to redirect traffic to Site-B Edge (documented in playbook), verify access to Risk Analytics application (should be available within 20-30 minutes from failover start). Document the actual failover time and any manual interventions required.
SRM Failover Test Report including: (1) initial replication lag measurement (should be <60min for 1hr RPO compliance), (2) Test Failover results: all VMs powered on at Site-B, network mapping confirmed (ping test from recovered VM to site-local gateway), database consistency check (query record count), Recovery Plan Test completed with 0 errors, (3) Test Failover cleanup completed, (4) Planned Failover results (if approved by test scenario): failover start time, all VMs offline at Site-A, powered on at Site-B, BGP update propagation time measured, application access verified (login to Risk Analytics UI), total failover time from initiation to full availability recorded (target <30min). Example: 'Test Failover: 0 VMs failed, lag=45min (<60min OK), Site-B VMs online, DB query count=500,000 (matches Site-A).'
Do not perform failover without a verified Test Failover first; SRM test mode allows you to catch configuration issues without disrupting production. Common issues: network mapping missing IP gateway, resulting in Site-B VMs unable to reach default gateway; NSX firewall policies not applied to Site-B, blocking traffic. If test fails, do not proceed to actual failover until issues are resolved.
Step 3

Validate RPO compliance through replication lag monitoring. For Tier-1 and Tier-2, enable continuous replication lag monitoring. For Tier-1 (stretched vSAN), check the vSAN Replication Witness Latency metric (should be <5ms for healthy sync), and check the Replication Network Throughput to ensure sustained replication of workload changes (500GB workload with 5-10% daily change = 25-50GB/day change, should fit within replication window of 24 hours). For Tier-2 (vSphere Replication), open vSphere Replication dashboards, check the Replication Lag metric for each VM (target <60min for 1hr RPO). Simulate a peak load period: run a synthetic workload (e.g., IOmeter) on Tier-2 VMs that increases write IO by 50% for 1 hour, measure replication lag during peak (lag should not exceed 90min, otherwise RPO is violated). Document the test results and identify any conditions where RPO is at risk.

RPO Compliance Report with: (1) Tier-1 vSAN replication health metrics: witness latency <5ms (sample from vSAN Observer), replication throughput <4Gbps (expected for 500GB Trading workload with 5% daily change = 25GB/day ≈ 300MB/s average), replication lag consistently <1ms. (2) Tier-2 vSphere Replication lag baseline: recorded hourly lag measurements over 24hr period, lag typically 5-30min (good), peak lag during business hours 40-60min (acceptable). (3) Load test results: during synthetic 50% IO increase, lag measured at 65-85min (within acceptable range for 1hr RPO), lag recovered to baseline <30min within 30 minutes after load ends. (4) Conclusion: 'RPO requirements met under tested conditions. Risk areas: Tier-1 RPO of 0 is vulnerable to simultaneous multi-host failure (unlikely) and replication link failure (mitigated by witness quorum). Tier-2 RPO of 1hr is met with typical lag <60min, but peak periods approach limit; monitor and alert if lag exceeds 75min.'
Set replication lag alert thresholds at 75% of RPO target (e.g., 45min alert threshold for 60min RPO); this gives your team 15 minutes to address latency issues before RPO is violated. Common causes of lag spikes: storage contention (high IO from other VMs), network congestion (check QoS settings), burst writes from backups or batch jobs. Add storage IO and network bandwidth monitoring alongside replication lag.
Step 4

Document recovery runbook for each tier based on validated test results. For Tier-1, create a 'Stretched Cluster Failure Response' runbook: (1) detect cluster failure (monitor quorum), (2) if majority at Site-A, continue operations (no failover needed), (3) if majority at Site-B, check Site-A status (network partition vs. hardware failure), (4) if Site-A hosts are permanently lost, declare disaster, contact storage vendor to confirm vSAN data consistency, verify all critical VMs are running at Site-B, document incident. For Tier-2, create 'SRM Risk Analytics Failover Runbook': prerequisites (verify replication lag <60min, test failover succeeds), steps (contact incident commander, execute test failover, notify Site-B team to prepare, execute planned failover, update DNS/BGP, verify application access, begin recovery operations). For Tier-3, create 'Backup Restore Runbook': (1) identify backup point (most recent daily backup), (2) request backup restore from backup admin, (3) allow 4-6 hours for restore (verify actual time from validation test), (4) validate data integrity, (5) notify business stakeholders. Include contact info, escalation path, and success criteria for each runbook.

Three recovery runbooks (Tier-1, Tier-2, Tier-3) in standardized format: [Runbook Title] [Objective] [Scope] [Prerequisites] [Decision Tree / Steps] [Success Criteria] [Rollback] [Contacts]. Example Tier-2 snippet: 'Objective: Failover Risk Analytics to Site-B within 30min. Scope: All Risk Analytics VMs (10 VMs, 1TB data). Prerequisites: Verify replication lag <60min (check vSphere Replication UI), perform test failover with 0 failures. Steps: (1) Incident commander declared, (2) SRM admin logs into recovery site SRM, (3) executes Planned Failover for Risk Analytics group (5min), (4) Network admin updates BGP weight at Site-B Edge (2min), (5) Verify application access (app loads <3sec), (6) update status page. Success: Risk Analytics UI accessible, queries return <5sec, no data errors. Rollback: revert BGP, execute failback when Site-A restored.'
Runbooks should be tested at least quarterly; include a 'last updated' date and 'last tested' date to track currency. Assign ownership (who maintains the runbook) and review cycle (e.g., annually or after infrastructure changes). Distribute runbooks to all on-call personnel and include in incident response playbooks.

Validation Gate

Check: Host failure in Tier-1 tested with documented recovery time; SRM failover for Tier-2 tested with successful test failover and documented failover procedure; RPO compliance validated with replication lag monitoring; recovery runbooks created and validated

Expected: Host failure test log + SRM failover report + RPO compliance report + three validated recovery runbooks

Common Errors

SRM Planned Failover initiated without prior Test Failover; halfway through, VMs fail to power on at Site-B due to missing network mapping or resource constraints, causing data access disruption
Cause: Assumed network and resource configuration is correct without validation; network mapping typos or firewall policy misconfigurations go undetected
Fix: Always execute SRM Test Failover first (non-disruptive), verify all aspects: VM power-on at Site-B, network connectivity, data access. Only after test succeeds, proceed to planned failover. SRM test mode creates temporary VM copies; if test fails, analyze errors in SRM UI (Error Report) and fix configuration before attempting actual failover.
Replication lag gradually increases from 30min to 2 hours without alerting; discovery happens after failover when Site-B data is found to be stale by >1 hour, violating Tier-2 RPO
Cause: No proactive monitoring of replication lag; team only checks lag reactively when failover is needed
Fix: Implement continuous replication lag monitoring with alerting at 75% threshold (45min for 60min RPO). Use vSphere Replication metric 'Replication Lag' and set critical alert at >60min. Common causes: storage bottleneck (increase storage IO priority), network congestion (verify QoS), backup window interference (schedule backups outside peak hours).
Ungraceful host shutdown (pulling power) instead of graceful off triggers vSAN host failure detection (30 seconds wait for heartbeat timeout), during which write IOs are blocked; test lasts 5+ minutes instead of expected 30 seconds due to slow recovery
Cause: Simulated failure too aggressively; vSAN double-checks before assuming host is truly down to avoid false-positives from network glitches
Fix: For more realistic testing, use vSphere HA to force a host failure: select host in vCenter, click 'HA Restart VM' (graceful shutdown of host), which allows vSAN to detect failure in <10 seconds. Alternatively, disable vSAN heartbeat on one host (vSAN iSCSI target configuration) to simulate partial network failure; this triggers faster detection (5-10 seconds).

Task 4 Document DR Architecture Decisions Using RCAR Framework

manageability

Create formal decision documents (DD-001 through DD-003) for each tier's DR architecture using the VMware RCAR template (Requirements, Constraints, Assumptions, Risks). This task ensures that design decisions are explicitly documented with tradeoff analysis, risk acknowledgment, and stakeholder approval, meeting enterprise architecture governance standards.

Step 1

Create Decision Document DD-001 for Tier-1 Stretched Cluster DR. Structure the document with sections: (1) Business Context: Trading Platform must have zero data loss (RPO=0) and <5min recovery (RTO<5min); transactions are critical and any loss is unacceptable. (2) Requirements (R): must support synchronous replication (RPO=0), must tolerate 1 host failure without loss of quorum (requires witness), must meet 1ms latency SLA between sites, must support 500GB workload with 5-10% daily change rate (25-50GB/day replication traffic). (3) Constraints (C): bandwidth limitation (10Gbps link available), latency requirement (<1ms for sync replication, achievable only with dedicated link and fiber optic connectivity), licensing (requires vSAN Enterprise), staffing (requires 2 dedicated SREs trained on vSAN stretch cluster troubleshooting). (4) Assumptions (A): Site-to-site latency will remain <1ms (assume fiber link stability, no network overload), witness host will remain available 99.99% uptime (host placement at neutral Site-C or HA-protected), workload change rate will not exceed 10% daily (if trading volume increases >10% change/day, architecture must be revisited). (5) Risks (R) with IMPACT and MITIGATION: Witness host failure (IMPACT: split-brain, VMs may diverge on two sites; MITIGATION: Host HA on witness, daily failover tests, automated alerting on witness heartbeat loss); Replication link failure (IMPACT: if latency spike >5ms, replication throughput drops, queuing delays increase, RPO drifts; MITIGATION: dedicated 10Gbps link, bandwidth reservation QoS, lag monitoring alert at 5min, failover to Site-B if lag approaches 5min); Workload change rate exceeds capacity (IMPACT: replication lag accumulates, RPO violations; MITIGATION: cap trading volume in DB max row insert rate, monitor change rate continuously, scale up to 40Gbps link if needed). Include alternatives analysis: (1) Multi-site active-active with conflict resolution (rejected due to licensing cost for multiple clusters, complexity of distributed transaction handling), (2) SRM + vSphere Replication for Tier-1 (rejected: async replication cannot meet RPO=0). Conclude with recommendation: 'Implement stretched vSAN cluster for Tier-1 based on RPO=0 requirement, pending confirmation that <1ms latency is achievable.'

Decision Document DD-001 (1-2 pages) with all RCAR sections filled. Explicit Requirements list with 4+ requirements, Constraints list with 3+ constraints (bandwidth, latency, licensing, staffing), Assumptions with 3+ assumptions, Risks section with 4+ risks each containing IMPACT statement and MITIGATION approach. Alternatives section comparing 3 options with pros/cons. Example Risk entry: 'Replication link failure causes lag >5min (IMPACT: RPO violation, potential data loss on failover). MITIGATION: QoS reservation, redundant links, automated lag monitoring alert at 4min, automatic failover to Site-B if replication unavailable for >1min.'
Use a standard decision document template (available from VMware Consulting or ITIL frameworks); ensure all stakeholders (storage team, networking team, app owner) review and sign off. RCAR is particularly important for Tier-1 because high cost (enterprise licensing) and high risk (split-brain, data loss potential) justify deep analysis.
Step 2

Create Decision Document DD-002 for Tier-2 SRM + vSphere Replication DR. Structure document: (1) Business Context: Risk Analytics workload has 30min RTO and 1hr RPO; data loss of up to 1 hour is acceptable, but >30min recovery would impact daily risk reporting. (2) Requirements: support asynchronous replication (RPO up to 1 hour acceptable), support failover to standby cluster in <30min, include network remapping for IP address translation between sites, must include NSX policy failover for firewall/NAT rules, must support Tier-2 application startup on Site-B (dependency ordering, app readiness checks). (3) Constraints: SRM licensing limits number of VMs (check VCF edition), replication network link limited to 1Gbps (shared with other traffic, not dedicated like Tier-1), standby cluster at Site-B has fewer nodes (2-4) to reduce cost, recovery plan complexity grows with number of VMs. (4) Assumptions: vSphere Replication lag will remain <60min under normal operations (workload change rate <5GB/hour), NSX policies are identical on both sites (pre-configured, not replicated), application startup is idempotent (can be safely restarted at Site-B after failover), database consistency is maintained during async replication (app-level transactions are ACID, no application-specific replication needed). (5) Risks: Replication lag exceeds 60min during business hours (IMPACT: RPO violation, Site-B data stale; MITIGATION: monitor lag, alert at 45min, consider upgrade to dedicated replication link if peak hours consistently breach threshold), NSX policy drift between Site-A and Site-B (IMPACT: failover succeeds but firewall blocks traffic; MITIGATION: version control NSX policies, automated policy sync or annual audit to ensure Sites are identical), application dependency ordering failures (IMPACT: database starts before file services, transactions fail; MITIGATION: define startup dependency sequence in recovery plan, test monthly), human error during SRM failover (IMPACT: wrong recovery plan executed or step skipped; MITIGATION: SRM runbook with checklist, two-person approval for planned failover). Alternatives: (1) Stretch cluster for Tier-2 (rejected: 1ms latency not critical for 30min RTO, cost of sync replication infrastructure not justified), (2) Backup-only for Tier-2 (rejected: 1-4hr restore time cannot meet 30min RTO). Conclude: 'Implement SRM + vSphere Replication for Tier-2 based on RTO/RPO balance and cost efficiency.'

Decision Document DD-002 (2 pages) with RCAR sections. Requirements include async replication, failover <30min, network remapping, NSX policy failover, app startup orchestration. Constraints include SRM licensing limits, 1Gbps replication link, reduced standby cluster size. Assumptions include lag <60min, NSX policies pre-synchronized, app idempotence, database ACID compliance. Risks section with 4+ risks including lag violation, NSX policy drift, app dependency failures, human error. Alternative analysis showing stretch cluster and backup rejected. Recommendation to proceed with SRM.
For SRM, pay special attention to network remapping and NSX policy sync; these are common failure points in real failovers. Include a network mapping verification test in the decision doc (verify IP range 10.0.0.0/24 maps to 10.1.0.0/24, confirm gateway 10.1.0.1 is reachable, ping test from recovered VM succeeds).
Step 3

Create Decision Document DD-003 for Tier-3 Backup-Based DR. Structure document: (1) Business Context: Reporting System is non-critical, used for historical analysis and not time-sensitive; 4hr RTO and 24hr RPO are acceptable for business continuity. (2) Requirements: backup system must support daily incremental snapshots, weekly full backups, retention of 7-day incremental chain and 4-week full backups (covering up to 28 days of recovery points), backup must be stored off-site or on secondary storage at Site-B, restore process must be documented and tested quarterly to ensure 4hr RTO is achievable. (3) Constraints: backup storage is shared infrastructure (limited to 6TB for 2.5TB data with 2:1 compression), backup link is 1Gbps (shared with replication and other traffic, limiting restore speed), backup restore depends on manual intervention (no automated recovery plan, requires human to initiate restore and validate), staff resources for backup management are limited (1 FTE). (4) Assumptions: data does not need to be point-in-time consistent (application can accept 24hr RPO, no transactions are lost, just day-old snapshots), restore from backup does not require parallel restore paths (serial restore acceptable within 4hr window), storage administrators can reliably execute restore procedures without error (assumption on staff competency). (5) Risks: Backup storage failure (IMPACT: cannot restore Reporting System after disaster, hours/days of data unavailable; MITIGATION: dual-site backup copies, monthly restore validation test, alert if backup copy fails), Restore exceeds 4hr RTO (IMPACT: Reporting System unavailable for extended period; MITIGATION: test actual restore time with full 2.5TB workload monthly, if test shows >4hr, reduce retention window or upgrade link bandwidth), RPO violation from missed backups (IMPACT: if backup job fails and is not noticed, next available backup is >24hr old; MITIGATION: backup job monitoring, alert if backup fails, daily verification that backup completed successfully). Alternatives: (1) Tier-2 SRM for Reporting System (rejected: cost of continuous replication not justified for non-critical workload, 30min RTO is overkill), (2) Snapshot-based replication (rejected: snapshot cost and storage overhead higher than backup approach for this use case). Conclude: 'Implement VADP backup-based DR for Tier-3 based on RTO/RPO requirements and cost efficiency.'

Decision Document DD-003 (1.5-2 pages) with RCAR sections. Requirements include daily incremental + weekly full backup, 7-day incremental retention, 4-week full backup retention, off-site backup copies, quarterly restore testing. Constraints include shared 6TB storage, 1Gbps link, manual restore process. Assumptions include non-critical data, 24hr RPO acceptable, point-in-time consistency not required, staff competency. Risks include backup storage failure, restore exceeding RTO, missed backups. Alternatives rejected (SRM too costly, snapshots too expensive). Recommendation to proceed with backup-based DR.
For Tier-3, emphasize the cost/benefit tradeoff: backup approach costs ~30% of SRM approach but trades off RTO (4 hours vs 30 min). For non-critical workloads, this is acceptable and recommended.
Step 4

Create an Executive Summary document consolidating DD-001, DD-002, DD-003 into a single 1-page overview for stakeholder approval. Summary should include: (1) Overall DR architecture for all 3 tiers in a single table (Tier | RTO/RPO | Technology | Key Features | Cost Estimate), (2) Risk Register: top 5 risks across all tiers with IMPACT, MITIGATION, and Residual Risk (after mitigation), (3) Critical Success Factors: assumptions that must hold for DR to work (e.g., <1ms latency for Tier-1, NSX policies synchronized for Tier-2, backup link available for Tier-3), (4) Implementation Schedule: Tier-1 (months 1-3, stretch cluster build-out), Tier-2 (months 2-4, SRM configuration and testing), Tier-3 (months 1-2, backup policy configuration), (5) Approval signatures: CTO/VP Eng, Storage Lead, Networking Lead, Backup Admin, App Owner (Risk Analytics and Reporting), (6) Success Metrics: all 3 tiers tested and validated, zero RPO violations during monthly monitoring, failover runbooks approved and team trained, recovery RTO measured within 10% of design estimate. Include a visio diagram showing all three tiers and data flow paths for quick reference.

Executive Summary document (1-2 pages) with: (1) Architecture summary table: Tier-1 | RTO<5min RPO=0 | Stretched vSAN Cluster | Sync replication, witness, 10Gbps link | $500K (licensing + infrastructure); Tier-2 | RTO=30min RPO=1hr | SRM + vSphere Replication | Async replication, network remapping, NSX failover | $150K; Tier-3 | RTO=4hrs RPO=24hrs | VADP Backup | Daily incremental, weekly full, off-site copies | $50K. (2) Risk Register table with 5 top risks: Witness failure (Impact: split-brain, Mitigation: HA, Residual: medium), Lag violation (Impact: RPO miss, Mitigation: monitoring + failover, Residual: low), SRM policy drift (Impact: failover blocked, Mitigation: annual audit, Residual: medium), backup storage failure (Impact: no restore, Mitigation: dual copies, Residual: low), human error (Impact: wrong failover, Mitigation: runbook + approval, Residual: medium). (3) CSF section listing assumptions. (4) Timeline Gantt-style chart. (5) Signature blocks for stakeholders. (6) Success Metrics and KPIs.
Keep the Executive Summary actionable and business-focused; avoid technical jargon where possible. Highlight total DR budget and timeline for approval. Use a Visio diagram for visual impact: show three tiers as boxes, replication arrows between Site-A and Site-B, witness placement, NSX control plane. This diagram will be useful for presentations to executives and for onboarding new team members.

Validation Gate

Check: All RCAR decisions (DD-001, DD-002, DD-003) completed with detailed Requirements, Constraints, Assumptions, Risks; Executive Summary created with risk register and implementation timeline; decision documents reviewed by all stakeholders

Expected: Three decision documents + Executive Summary + signed stakeholder approval

Common Errors

Risk section lists 'Replication link failure' but does not specify impact (what happens?), mitigation steps (how to prevent or respond?), or residual risk (what remains after mitigation?)
Cause: Rushed RCAR creation without deep analysis; team treated risks as checkbox rather than critical planning input
Fix: Expand each risk with explicit IMPACT statement ('Replication link down for 1hr causes lag to spike from 30min to 2hrs, violating RPO'), MITIGATION approach ('Deploy redundant 10Gbps links, automatic failover to secondary link, monitor link status every 60 seconds'), and RESIDUAL RISK ('After mitigation, risk of dual-link failure is low probability but high impact; accepted because infrastructure cost and complexity of triple-redundancy is not justified for Tier-2').
DD-001 approved by architecture team but storage team later objects that stretched cluster requires too much vSAN licensing; decision must be revisited
Cause: Architecture team created RCAR in isolation; storage team not involved in decision-making and constraints discovery
Fix: Include all stakeholders in RCAR creation: conduct a working session with storage, networking, app owner, and security; solicit input on constraints, assumptions, and risks from each perspective. Use the RCAR as a decision-making forum, not a post-hoc documentation task. Require sign-off from all stakeholders before finalizing decision.
DD-001 assumes 'Site-to-site latency will remain <1ms' without actual measurement; during implementation, discovered that latency is 5-10ms, invalidating stretched cluster architecture
Cause: Assumed network characteristics without baselining; did not test actual latency before committing to design
Fix: Before finalizing RCAR, validate all critical assumptions: measure actual site-to-site latency (ping and sustained traffic test), verify bandwidth available for replication (iperf test on replication link), confirm workload change rate (run transaction monitor for 1-2 weeks, collect actual write IO stats). If measurements contradict assumptions, revise design (e.g., switch from stretched cluster to SRM if latency is 5ms instead of <1ms).

Final Validation

All four DR architecture tasks completed: workload classification with tier mapping, topology design for each tier, failover validation with test results, and RCAR-based decision documentation

✓ Task 1: BIA spreadsheet, DR technology mapping, and license checklist completed → All three workloads classified into appropriate tiers with RTO/RPO requirements and technology selection justified

✓ Task 2: Tier-1 topology diagram (stretched cluster), Tier-2 topology with SRM recovery plan, Tier-3 backup topology, NSX cross-site networking checklist completed → Three detailed topology diagrams with replication paths, failover mechanisms, and RTO/RPO calculations documented

✓ Task 3: Host failure test log, SRM failover test report, RPO compliance validation, recovery runbooks completed → Practical validation of failover procedures with measured recovery times and documented runbooks for incident response

✓ Task 4: DD-001 (Tier-1), DD-002 (Tier-2), DD-003 (Tier-3), Executive Summary with risk register and implementation timeline, stakeholder approval signatures → Formal decision documentation with RCAR framework, alternatives analysis, risk mitigation, and executive-level approval

Cleanup / Restore

• Revert Tier-1 stretched cluster to stable state (power on any offline hosts, verify cluster quorum)

• Remove SRM test failover temporary VMs at Site-B

• Remove synthetic workload generators used for RPO validation testing

• Archive all lab test logs and validation reports to shared drive or wiki for future reference

• Document any infrastructure changes made during lab (e.g., new monitoring alerts, bandwidth reservations) for handoff to operations team

Design Reflection (VCDX)

VCDX panelists will probe the candidate's understanding of tradeoff analysis between RPO/RTO requirements and technology cost/complexity.

Expect questions: 'Why stretched cluster for Tier-1 and not SRM?' (answer: RPO=0 requires sync replication, SRM is async), 'What happens if witness host fails?' (answer: quorum logic, failover to majority site, show understanding of split-brain prevention), 'How do you ensure NSX policies are consistent across sites during SRM failover?' (answer: policy versioning, pre-configuration, optional NSX Federation), 'What is the financial impact of a 1-hour RPO violation for Risk Analytics?' (answer: depends on business context, in finance likely high cost of decisions made on stale data).

Panelists will also assess communication skills: can the candidate explain complex DR architecture to non-technical stakeholders using the Executive Summary? Can they articulate risks and mitigations without hand-waving?

Requirements

  • Trading Platform (Tier-1) must achieve RPO=0 (zero data loss) and RTO<5 minutes
  • Risk Analytics (Tier-2) must achieve RPO=1 hour and RTO=30 minutes
  • Reporting System (Tier-3) must achieve RPO=24 hours and RTO=4 hours
  • All tiers must support failover between two geographically separated sites (Site-A primary, Site-B secondary)
  • DR architecture must integrate with NSX for cross-site networking, including policy and routing failover
  • All RTO/RPO requirements must be validated through practical testing before production deployment
  • Recovery procedures must be documented in runbooks and tested at least quarterly
  • DR investment must be justified through RCAR analysis with alternatives comparison for each tier

Constraints

  • Bandwidth limitation: 10Gbps dedicated link for Tier-1, 1Gbps shared link for Tier-2/Tier-3 replication and other traffic
  • Latency requirement: Tier-1 stretched cluster requires <1ms site-to-site latency (achievable only with dedicated fiber, geographically proximate sites within 100km)
  • Licensing: vSAN Enterprise for stretched cluster (Tier-1), SRM Standard/Enterprise for Tier-2, VADP for Tier-3; cost must fit IT budget
  • Staffing: limited to 2 SREs for ongoing DR operations; complex architectures (stretched cluster) require specialized training
  • Storage capacity: 6TB backup storage for 2.5TB Tier-3 workload with 2:1 compression (limited expansion headroom)
  • Site-B resources: secondary cluster sized for minimum viable service (may be smaller than Site-A to reduce cost); impacts failover performance and application startup time
  • Change management: any DR topology change requires testing in lab, validation through failover test, stakeholder approval (slow process, limits architecture flexibility)

Assumptions

  • Site-to-site latency will remain <1ms under normal operations (requires stable network, no congestion); if latency spikes, stretched cluster efficiency degrades and may violate RPO
  • Witness host for stretched cluster will remain available 99.99% of the time; failure of witness triggers quorum recalculation and potential service disruption
  • Workload change rates will not exceed baseline estimates (Trading Platform 5-10% daily, Risk Analytics 5% daily); if change rate increases, replication lag grows and RPO may be violated
  • NSX policies (firewall rules, NAT, routing) are pre-configured identically on both sites and will remain synchronized; drift between sites (e.g., new rule added at Site-A but not Site-B) will break failover
  • Application workloads are idempotent and can be safely restarted at Site-B after failover without side effects (true for stateless apps, may not be true for apps with local state or singleton processes)
  • Recovery procedures documented in runbooks are accurate and will be correctly followed under stress (assumes good runbook quality and staff training)
  • Infrastructure changes (network, storage, hypervisor) will not invalidate DR design; any change requires RCAR re-evaluation

Risks

  • Risk: Witness host failure in Tier-1 stretched cluster — Impact: If witness host becomes unavailable (network partition, hardware failure, power loss), stretched cluster loses quorum decision-making. With 2+2 host split between sites, cluster cannot determine which site has majority; both sites may attempt to accept writes, causing split-brain and data corruption. — Mitigation: Deploy witness host on HA-protected infrastructure (e.g., local vSAN cluster with separate power/network feeds), implement automated heartbeat monitoring with alert if witness is unreachable for >60 seconds, execute quarterly failover test (fail witness, verify cluster remains functional with remaining quorum), maintain runbook with manual quorum recovery steps if needed. — Residual_risk: Low; after mitigation, witness host failure is unlikely to cause data loss if detected and responded to quickly. Residual risk is that manual intervention may be required during witness recovery, causing brief service interruption. Acceptable risk level.
  • Risk: Replication link failure or latency spike causing RPO violation — Impact: Tier-1 stretched cluster replication depends on <1ms latency and sustained throughput; if link fails or latency increases to 5ms+, replication queues grow, lag accumulates, RPO=0 target cannot be met. Tier-2 vSphere Replication lag may exceed 60min during peak load if link is congested, violating RPO. Tier-3 backup restore time may exceed 4hr RTO if backup link is unavailable. — Mitigation: Design dual replication links with active-active load balancing (if supported by switch/array) or active-standby automatic failover; implement QoS to reserve bandwidth for replication traffic; monitor replication link bandwidth and latency continuously with alerts at 80% utilization or latency >2ms; perform monthly link failover tests to verify automatic failover works; maintain low-cost fallback procedures (e.g., manual fail to Site-B if lag approaches RPO limit for Tier-2). — Residual_risk: Medium; despite mitigation, link failures can happen unexpectedly. Residual risk is reduced to acceptable level if monitoring and alerting are in place and team is trained to respond quickly. For Tier-1, if lag breaches RPO window, failover to Site-B is the only recovery option.
  • Risk: NSX policy drift between Site-A and Site-B, causing firewall/routing failure during failover — Impact: SRM failover succeeds and VMs power on at Site-B, but network policies (firewall, NAT, routing) are not applied correctly because they differ from Site-A. Result: legitimate traffic is blocked by default-deny firewall, or traffic is routed incorrectly. Application appears down from external perspective even though VMs are running. — Mitigation: Pre-configure identical NSX Tier-1 routers and security policies on both sites before implementing DR; use version control (GitHub) to track NSX policy changes; implement policy comparison tool (NSX API or third-party tool) to audit policy consistency monthly; include NSX policy verification as a step in SRM test failover (verify firewall rules, NAT entries, routes); document NSX failover steps in recovery runbook (e.g., activate Site-B router, confirm traffic flow); consider NSX Federation (if available in VCF 9.0) for automatic policy replication. — Residual_risk: Medium-low; policy drift is a common failover issue but can be prevented with good change management. Residual risk is that new policies added at Site-A are not immediately applied to Site-B, requiring manual sync. Acceptable if monthly audit ensures drift is detected and corrected before failover is needed.
  • Risk: Backup restore exceeds 4-hour RTO, causing Tier-3 reporting system to be unavailable beyond acceptable window — Impact: Tier-3 Reporting System backup restore from 2.5TB data with limited 1Gbps link and HDD-based storage takes 6+ hours, exceeding 4hr RTO. Business stakeholders cannot access reporting for 6+ hours, impacting strategic decision-making and daily reporting obligations. — Mitigation: Validate actual backup restore time with full-size 2.5TB workload in quarterly lab test; if test shows restore exceeds 4 hours, identify bottleneck (link bandwidth, storage IO, metadata retrieval) and remediate (upgrade link to 10Gbps, add SSD-based backup storage, parallelize restore across multiple backup appliances); pre-stage backup copies at Site-B to avoid retrieval latency; document actual measured RTO in recovery runbook rather than assuming 4 hours; set RTO expectations with business stakeholders based on validated times. — Residual_risk: Medium; backup restore time depends on many variables (link bandwidth, storage speed, workload characteristics) that may be unpredictable. After mitigation, residual risk is that restore may take longer than expected if storage or link is congested at time of disaster. Acceptable if backup restore time is monitored and communicated to stakeholders.
  • Risk: Human error during SRM failover or recovery runbook execution, causing wrong VMs to be failed over, network mapping errors, or recovery steps skipped — Impact: During crisis, team member executing recovery runbook makes typo (e.g., selects wrong recovery plan in SRM, fails over wrong VM group), or skips validation step, resulting in partial failover or data inconsistency. Application is partially recovered or inaccessible, extending RTO and potentially causing customer impact. — Mitigation: Create detailed, step-by-step recovery runbooks with screenshots and decision trees for different failure scenarios; implement two-person approval process for planned failovers (incident commander approves, separate team member executes); conduct quarterly runbook training and failover drills with full team; use SRM test failover mode extensively before any planned failover; add pre-flight checks to runbook (verify backup is complete, check replication lag, confirm no in-flight transactions); document all steps with expected outputs (e.g., 'After step 3, confirm 10 VMs show 'Ready' status in SRM UI'). — Residual_risk: Low-medium; after mitigation, human error risk is reduced through process discipline and training. Residual risk remains that under high stress, team may deviate from runbook or miss steps. Acceptable if runbooks are actively maintained and team is well-trained.

Self-Assessment Discussion Prompts

  1. If the site-to-site latency is measured at 5ms instead of the assumed <1ms, what changes to the Tier-1 stretched cluster design would you recommend? Would you still use stretched cluster or switch to SRM for Tier-1? What is the business impact of the change?
  2. Describe a scenario where Tier-1 stretched cluster RPO=0 is violated despite no hardware failure. What conditions could cause replication lag to accumulate? How would you detect and respond?
  3. In the Tier-2 SRM failover runbook, NSX policies must be applied at Site-B before application traffic reaches Site-B routers. How would you orchestrate this in the recovery plan? What automation or manual steps are needed?
  4. For Tier-3 backup-based DR, the 4-hour RTO assumes backup retrieval time of 1.5 hours. What factors could cause backup retrieval to take longer? How would you design the backup infrastructure to minimize retrieval time without breaking budget constraints?
  5. If both Site-A and Site-B fail simultaneously (rare but possible), how would you recover Tier-1 and Tier-2 workloads? What is your strategy for this scenario and what are the time and data loss implications?
  6. The RCAR document assumes NSX policies are pre-configured identically on both sites. In practice, how would you enforce and audit this assumption? What tools or processes would help maintain policy synchronization over time?

Extensions

Extend Tier-1 to Three-Site Active-Active Cluster with Distributed Replication

Instead of two-site stretched cluster with witness, design a three-site vSAN cluster where all sites actively serve workloads (active-active-active) with distributed replication across three sites. Each site has a 2-node cluster (6 nodes total), no dedicated witness needed (odd number of sites = automatic quorum). Requires evaluation of inter-site replication topology: star pattern (all sites replicate to Site-A hub) vs. mesh pattern (all sites replicate to each other) vs. hierarchical pattern (Site-A primary, Site-B secondary, Site-C backup). Model the trade-offs: star pattern reduces replication overhead but creates single point of failure (Site-A replication hub); mesh pattern maximizes redundancy but increases replication traffic 3x. Document recovery time if one entire site is lost (should be <30 seconds due to automatic failover to majority sites).

+2 (advanced to expert)

Design RPO=0 for Tier-2 Using Array-Based Synchronous Replication

Challenge the assumption that Tier-2 requires RPO=1hr; investigate whether array-based synchronous replication (e.g., Pure Storage ActiveCluster, Dell/EMC SRDF, NetApp MetroCluster) could be used to achieve RPO=0 for Tier-2 Risk Analytics without upgrading to full stretched cluster architecture. Compare cost, complexity, and failover performance of array-based sync replication vs. stretched vSAN cluster. Document network latency requirements (typically <5ms for array replication vs. <1ms for vSAN stretched cluster). Evaluate whether synchronous array replication imposes write latency penalties that would degrade application performance. Conclude whether this extension is economically justified compared to current SRM design.

+1 (advanced, requires deep storage knowledge)

Implement Automated Failover Using vSAN Stretched Cluster with NSX AVI Load Balancer Global Service Load Balancing (GSLB)

Design automatic failover without manual intervention by combining vSAN stretched cluster (automatic host failure detection) with NSX AVI Global Load Balancer for DNS-based failover. When Site-A cluster quorum is lost, AVI GSLB automatically updates DNS to point to Site-B IP addresses, redirecting client traffic. Define the failover sequence: vSAN detects quorum loss at Site-B (Site-A hosts down), AVI GSLB detects Site-A datacenters unreachable (health check failure), AVI updates DNS to Site-B addresses, clients re-resolve DNS and traffic flows to Site-B. Model failover time including DNS update latency (typically 5-30 seconds) and client re-connection time. Evaluate whether fully automated failover (no manual runbook steps) is appropriate for Tier-1 trading workload or if human oversight is still needed.

+2 (advanced, requires load balancer and DNS knowledge)

Evaluate VMware Cloud Disaster Recovery (VCDR) as Alternative to On-Premises DR

Instead of building on-premises two-site DR, evaluate VMware's cloud-native disaster recovery service (VCDR, formerly Zerto) as an alternative for Tier-2 and Tier-3 workloads. VCDR offers cloud-based standby environment, continuous asynchronous replication, and point-in-time recovery. Compare VCDR cost/performance/complexity vs. on-premises SRM approach: VCDR eliminates need for secondary on-premises cluster (cost savings), provides frequent recovery point objectives (minutes rather than hours), but introduces cloud provider dependency and potential multi-tenancy security concerns. Model failover workflow in VCDR (activate cloud VMs in VCDR-managed environment) vs. SRM (activate VMs in on-premises Site-B cluster). Determine appropriate tier for cloud DR: is Tier-3 reporting system a good fit for cloud DR? Would Tier-1 trading platform ever be cloud DR candidate (likely not due to latency and data residency concerns)?

+1 (advanced, requires cloud services knowledge)

⚠ Known Pitfalls (from Community KB)

Designing DR for peak capacity instead of average capacity, resulting in over-provisioned secondary cluster and inflated costs
Problem: Team designs Tier-2 secondary cluster to handle 100% of primary cluster capacity (all 10 Risk Analytics VMs running simultaneously at full load). In reality, failover scenarios rarely need full parallel capacity; if primary site is down, users will be redirected to secondary site, but some workloads may be deprioritized or run in degraded mode. Recommendation: design secondary cluster to handle 70-80% of primary capacity during peak, accept that some users may experience slower performance during failover.
Resolution: Conduct a detailed failover impact analysis for each tier: estimate how many concurrent users will be active during disaster (may be lower than peak due to time-of-day or geographic distribution), determine which workloads are critical vs. deferrable, right-size secondary cluster based on critical workload capacity, not peak capacity. Document assumptions in RCAR.
Assuming RPO/RTO requirements are fixed in stone, without reviewing business impact of relaxing requirements to save cost
Problem: Team commits to RPO=0 for Tier-1 trading workload based on initial requirements, not realizing that cost of stretched cluster ($500K infrastructure + $50K annual licensing) is 10x the cost of SRM approach ($50K infra + $10K licensing). Follow-up analysis shows that accepting RPO=15min (instead of 0) would reduce cost by 50% and still meet business needs (15min of re-entry is acceptable for trading platform). Recommendation: challenge requirements during RCAR phase, run cost-benefit analysis, present options to stakeholders.
Resolution: During Task 1 (BIA), do not accept RTO/RPO requirements at face value. Interview business stakeholders to understand what happens if RPO is missed by 10% or 20%. For trading workload, ask 'What is the cost of re-entering 15 minutes of transactions vs. cost of infrastructure to achieve RPO=0?' Use answers to validate or adjust requirements before committing to architecture.
Building elaborate DR architecture that is never tested until actual disaster occurs, ensuring it will fail when needed most
Problem: Tier-1 stretched cluster is designed and deployed but never tested during normal operations. When actual disaster occurs (Site-A power loss), team discovers that stretched cluster has split-brain condition because witness host also lost power; VMs continue running at both sites with diverged disk states, causing data corruption. Root cause: witness host was not protected by UPS, deployment deviation from design document went undetected.
Resolution: Establish quarterly failover testing cadence (minimum 4 times per year, one per quarter). For Tier-1, test at least: (1) single host failure (graceful and ungraceful), (2) entire site failure (network partition simulation), (3) witness host failure. Document test results; failures must be root-caused and fixed before test is closed. Assign test ownership and accountability.
Relying on RPO/RTO estimates without validating that infrastructure can actually deliver those metrics under real load
Problem: Team designed Tier-2 SRM replication with estimated RPO=1hour based on workload change rate of 5GB/day. In production, Risk Analytics workload experiences 50GB/day change rate during business hours (10x higher than estimate due to new trading volume), causing replication lag to spike to 10 hours, violating RPO by 9x. Team did not detect violation until RTO event occurred; by then, Site-B data was stale and partially unusable.
Resolution: During Task 3 (validation), create a synthetic workload that matches estimated change rates and run sustained load test for 24 hours. Measure actual replication lag during load test. If measured lag exceeds 75% of RPO target, either reduce workload change rate assumptions or upgrade replication infrastructure before going to production. Document validated RPO/RTO metrics in runbooks, not estimated metrics.
Creating recovery runbooks that are too high-level, missing critical details that cause failures during actual execution
Problem: Tier-2 SRM recovery runbook states 'Failover Risk Analytics VMs to Site-B' as a single step. During actual failover, operator spends 30 minutes searching for correct recovery plan in SRM UI, selects wrong plan by mistake, partially fails over (5 of 10 VMs), then has to re-do failover. Total failover time becomes 45 minutes, exceeding 30-minute RTO.
Resolution: Recovery runbooks must include screenshots, exact UI paths, checkbox-style steps, and decision trees. Example step: 'Open vCenter Web Client → VMware Site Recovery Manager → select 'Recovery' tab → click 'Protection Groups' → select 'Risk Analytics (PG-2)' → verify status shows 'Ready' → click 'Recover' button (do not click 'Test').' Include expected output for each step ('After clicking Recover, protection group status changes to 'Recovering'; wait for all 10 VMs to show 'Powered On' status'). Test runbook with team members who are not familiar with SRM to identify gaps.

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.