Disaster Recovery Architecture Design
Objectives
- Classify workloads by RTO/RPO tiers and map to VCF DR capabilities
- Design multi-tier disaster recovery topology with stretched clusters, SRM, and backup
- Validate DR design against failure scenarios and RPO compliance
- Document DR architecture decisions using RCAR framework
Prerequisites
VCF 9.0 lab environment deployed with 2+ vSphere clusters, NSX, vSAN storage, and SRM installed
Prior labs: vcp-architect-04: vSphere Clustering and Stretch Design, vcp-architect-05: NSX Architecture and Segmentation
Required skills:
- vSphere 7.x/8.x cluster administration
- Storage replication technologies (vSAN, array-based sync/async)
- vSphere Replication and Site Recovery Manager
- NSX cross-site networking and BGP failover
- Bandwidth planning for synchronous replication
- Business Impact Analysis and tier definition
- Failover testing and runbook creation
Lab Environment
Two-site stretched VCF environment: Site-A (primary) with 4-node vSAN cluster + 3-node management cluster, Site-B (secondary) with matching compute cluster, witness host at Site-C or local storage, dedicated 10Gbps replication links with 1ms latency (simulated), NSX Edge Cluster at each site for cross-site routing and BGP failover, SRM 8.x and vSphere Replication 8.x installed and licensed at both sites
Tasks
Task 1 Classify Workloads by RTO/RPO and Map to VCF DR Capabilities
recoverabilityConduct a thorough Business Impact Analysis for the three test workloads, classify them into RPO/RTO tiers (Tier-1: RTO<5min RPO=0, Tier-2: RTO=30min RPO=1hr, Tier-3: RTO=4hrs RPO=24hrs), and document how each tier maps to specific VCF disaster recovery technologies. This foundational task ensures technical decisions are aligned with business requirements.
Conduct a Business Impact Analysis for the three sample workloads. Document the Recovery Time Objective (RTO) and Recovery Point Objective (RPO) for the Trading Platform, Risk Analytics, and Reporting System. For each workload, record: business criticality (critical/high/medium/low), acceptable downtime in minutes, acceptable data loss in minutes/hours, and dependencies on other systems. Create a BIA spreadsheet with columns for workload name, tier assignment, RTO, RPO, max tolerable downtime, and business justification.
Map each workload tier to VCF DR capabilities. For Tier-1 (Trading Platform), document why vSAN stretched cluster with synchronous replication is required. For Tier-2 (Risk Analytics), explain why Site Recovery Manager with vSphere Replication and asynchronous replication suits the 30min RTO and 1hr RPO. For Tier-3 (Reporting System), justify why VADP-based backup with 24hr RPO and 4hr RTO is appropriate. Create a mapping document that shows: Tier, RTO/RPO, selected DR technology, and technology rationale.
Document VCF component prerequisites and licensing for each tier. For Tier-1, specify vSAN Enterprise license (required for stretched cluster), NSX Enterprise for cross-site BGP failover, and vSphere Replication Enterprise. For Tier-2, confirm SRM Standard or Enterprise, vSphere Replication Standard, and NSX standard for network remapping. For Tier-3, ensure vSphere Replication or VADP license, backup storage licensing (on-premises or cloud). Create a License and Component Checklist.
Validation Gate
Check: All three workloads assigned to appropriate RTO/RPO tiers with documented BIA, DR technology mapping aligned to tier requirements, and VCF license prerequisites verified
Expected: BIA spreadsheet + mapping document + license checklist completed and reviewed
Common Errors
Task 2 Design the DR Topology for Each Tier
availabilityCreate detailed architectural designs for Tier-1 (stretched cluster), Tier-2 (SRM + vSphere Replication), and Tier-3 (backup-based DR). Each design should specify replication technology, network topology including NSX considerations, failover mechanisms, and data flow paths. This task translates tier requirements into concrete infrastructure blueprints.
Design the Tier-1 stretched cluster topology. Define the 4-node vSAN cluster split across Site-A (2 nodes) and Site-B (2 nodes) with a 1-node witness host at Site-C (or logically isolated on-prem). Document the placement: hosts A1, A2 at Site-A datastore 1; hosts B1, B2 at Site-B datastore 2; witness at Site-C (witness appliance ~40 GB footprint; stores metadata/quorum only, no VM data). Specify vSAN cluster configuration: stripes=1, replicas=2 (for 4 hosts, this ensures each site has full copy), failures to tolerate (FTT)=1. Include network topology: dedicated 10Gbps replication link between Site-A and Site-B for vSAN traffic, dedicated NSX overlay network for VM communication, separate management VLAN at each site with redundant NSX Edges for cross-site routing. Create a Visio or ASCII diagram showing host placement, replication links, and network zones.
Design the Tier-2 SRM + vSphere Replication topology. Create two separate vSphere clusters: primary at Site-A (4 nodes, local vSAN storage) and secondary at Site-B (2-4 nodes, local vSAN storage, initially powered down or in standby). Document vSphere Replication configuration: Recovery Point Objective = 1 hour (set replication interval to 60 minutes or smaller for granularity), Recovery Verification Test (RVT) scheduled daily, replication network isolated to dedicated VLAN to avoid impacting production traffic. Design the SRM recovery plan: specify which VMs from Site-A will replicate to Site-B, network mapping configuration (IP address remapping from Site-A subnet 10.0.0.0/24 to Site-B subnet 10.1.0.0/24), and NSX Tier-1 router migration (power off Source Router at Site-A, activate Destination Router at Site-B, update BGP weight on destination Edge to attract traffic). Include failback plan: resync Site-A from Site-B, reverse network mappings, reactivate Site-A primary systems. Create SRM recovery plan diagram and failover playbook.
Design the Tier-3 backup-based DR topology. Specify VADP backup policy: full backup on weekends, incremental backups daily, scheduled at 2 AM to avoid business hours. Define backup retention: 7-day incremental chain for point-in-time recovery, 4-week full backup retention (for 4hr RTO requirement, restore from most recent backup). Document backup storage: 6 TB local or cloud storage for 2.5TB Reporting System with compression ratio 2:1, dual-site backup copies (Site-A primary, Site-B secondary for disaster recovery of backup infrastructure itself). Define recovery procedures: restore RTO 4 hours assumes 1.5 hours for backup retrieval + 1 hour for storage re-initialization + 1.5 hours for application startup and data consistency checks. Include backup verification testing: monthly full restore to isolated cluster, document recovery time and data integrity validation. Create backup topology diagram and RTO calculation worksheet.
Consolidate NSX cross-site networking requirements for all tiers. Document NSX Tier-1 Logical Router configuration: distributed routing at Site-A and Site-B with BGP peering over Edge Cluster uplinks (BGP weight 100 at primary, 200 at secondary). Specify BGP failover convergence: typical 30-180 seconds depending on router configuration. Design failover behavior for each tier: Tier-1 stretched cluster failover is near-instantaneous (vSAN quorum-based), BGP weight change can be simultaneous; Tier-2 SRM failover requires manual BGP weight adjustment as part of recovery plan (documented step); Tier-3 backup restore means manual BGP re-pointing to Site-B backup infrastructure. Document DNS/DHCP handling: specify whether Site-B has read-only or writable DNS replicas, whether DHCP relay agents exist at both sites. Create an NSX Cross-Site Networking checklist.
Validation Gate
Check: All three tier topologies designed with replication technology, network paths, failover mechanisms, and RTO/RPO calculations documented
Expected: Tier-1 stretched cluster topology diagram, Tier-2 SRM recovery plan diagram with failover playbook, Tier-3 backup topology with RTO calculation, NSX cross-site networking checklist
Common Errors
Task 3 Validate DR Design Against Failure Scenarios
recoverabilityExecute practical failover tests and RPO compliance checks to verify that the DR design actually meets business requirements. This task moves from theory (design documents) to validation (tested procedures), ensuring that when a real disaster occurs, the recovery process works as documented.
Simulate a host failure in the Tier-1 stretched cluster. Gracefully power off one node at Site-A (e.g., host A2) or ungracefully disconnect it from the network (simulate network partition by disabling vSAN heartbeat). Monitor vSAN cluster status: expected behavior is immediate failover (within 60 seconds) to remaining majority (A1 + B1 + B2 + witness = quorum). Verify that all Trading Platform VMs continue running without interruption or data loss; use vSAN observer or vsphere CLI (vsan.health -a) to confirm cluster state. Document the Recovery Time from host failure to full cluster availability: should be <30 seconds. Document any dropped packets or storage IO stalls using vSphere performance monitoring (vSAN IO latency, VM packet loss).
Test SRM failover for Tier-2 Risk Analytics workload. Create a network partition between Site-A and Site-B (simulate replication link failure) and monitor vSphere Replication lag. Execute the SRM recovery plan for Risk Analytics VMs: (1) open SRM, navigate to Protection Groups, select Risk Analytics group, verify replication status (should show 'Ready'), (2) perform a Recovery Plan Test (non-disruptive) first: SRM creates temporary copies of VMs at Site-B without affecting Site-A, verify connectivity by pinging Site-B Recovery VMs, check that network mapping (10.0.0.0/24 → 10.1.0.0/24) is working, confirm database consistency (connect to Risk Analytics DB on Site-B test copy, verify record count matches Site-A), (3) clean up test copy, (4) execute actual failover: click 'Failover', confirm all VMs are migrated to Site-B, update BGP weight to redirect traffic to Site-B Edge (documented in playbook), verify access to Risk Analytics application (should be available within 20-30 minutes from failover start). Document the actual failover time and any manual interventions required.
Validate RPO compliance through replication lag monitoring. For Tier-1 and Tier-2, enable continuous replication lag monitoring. For Tier-1 (stretched vSAN), check the vSAN Replication Witness Latency metric (should be <5ms for healthy sync), and check the Replication Network Throughput to ensure sustained replication of workload changes (500GB workload with 5-10% daily change = 25-50GB/day change, should fit within replication window of 24 hours). For Tier-2 (vSphere Replication), open vSphere Replication dashboards, check the Replication Lag metric for each VM (target <60min for 1hr RPO). Simulate a peak load period: run a synthetic workload (e.g., IOmeter) on Tier-2 VMs that increases write IO by 50% for 1 hour, measure replication lag during peak (lag should not exceed 90min, otherwise RPO is violated). Document the test results and identify any conditions where RPO is at risk.
Document recovery runbook for each tier based on validated test results. For Tier-1, create a 'Stretched Cluster Failure Response' runbook: (1) detect cluster failure (monitor quorum), (2) if majority at Site-A, continue operations (no failover needed), (3) if majority at Site-B, check Site-A status (network partition vs. hardware failure), (4) if Site-A hosts are permanently lost, declare disaster, contact storage vendor to confirm vSAN data consistency, verify all critical VMs are running at Site-B, document incident. For Tier-2, create 'SRM Risk Analytics Failover Runbook': prerequisites (verify replication lag <60min, test failover succeeds), steps (contact incident commander, execute test failover, notify Site-B team to prepare, execute planned failover, update DNS/BGP, verify application access, begin recovery operations). For Tier-3, create 'Backup Restore Runbook': (1) identify backup point (most recent daily backup), (2) request backup restore from backup admin, (3) allow 4-6 hours for restore (verify actual time from validation test), (4) validate data integrity, (5) notify business stakeholders. Include contact info, escalation path, and success criteria for each runbook.
Validation Gate
Check: Host failure in Tier-1 tested with documented recovery time; SRM failover for Tier-2 tested with successful test failover and documented failover procedure; RPO compliance validated with replication lag monitoring; recovery runbooks created and validated
Expected: Host failure test log + SRM failover report + RPO compliance report + three validated recovery runbooks
Common Errors
Task 4 Document DR Architecture Decisions Using RCAR Framework
manageabilityCreate formal decision documents (DD-001 through DD-003) for each tier's DR architecture using the VMware RCAR template (Requirements, Constraints, Assumptions, Risks). This task ensures that design decisions are explicitly documented with tradeoff analysis, risk acknowledgment, and stakeholder approval, meeting enterprise architecture governance standards.
Create Decision Document DD-001 for Tier-1 Stretched Cluster DR. Structure the document with sections: (1) Business Context: Trading Platform must have zero data loss (RPO=0) and <5min recovery (RTO<5min); transactions are critical and any loss is unacceptable. (2) Requirements (R): must support synchronous replication (RPO=0), must tolerate 1 host failure without loss of quorum (requires witness), must meet 1ms latency SLA between sites, must support 500GB workload with 5-10% daily change rate (25-50GB/day replication traffic). (3) Constraints (C): bandwidth limitation (10Gbps link available), latency requirement (<1ms for sync replication, achievable only with dedicated link and fiber optic connectivity), licensing (requires vSAN Enterprise), staffing (requires 2 dedicated SREs trained on vSAN stretch cluster troubleshooting). (4) Assumptions (A): Site-to-site latency will remain <1ms (assume fiber link stability, no network overload), witness host will remain available 99.99% uptime (host placement at neutral Site-C or HA-protected), workload change rate will not exceed 10% daily (if trading volume increases >10% change/day, architecture must be revisited). (5) Risks (R) with IMPACT and MITIGATION: Witness host failure (IMPACT: split-brain, VMs may diverge on two sites; MITIGATION: Host HA on witness, daily failover tests, automated alerting on witness heartbeat loss); Replication link failure (IMPACT: if latency spike >5ms, replication throughput drops, queuing delays increase, RPO drifts; MITIGATION: dedicated 10Gbps link, bandwidth reservation QoS, lag monitoring alert at 5min, failover to Site-B if lag approaches 5min); Workload change rate exceeds capacity (IMPACT: replication lag accumulates, RPO violations; MITIGATION: cap trading volume in DB max row insert rate, monitor change rate continuously, scale up to 40Gbps link if needed). Include alternatives analysis: (1) Multi-site active-active with conflict resolution (rejected due to licensing cost for multiple clusters, complexity of distributed transaction handling), (2) SRM + vSphere Replication for Tier-1 (rejected: async replication cannot meet RPO=0). Conclude with recommendation: 'Implement stretched vSAN cluster for Tier-1 based on RPO=0 requirement, pending confirmation that <1ms latency is achievable.'
Create Decision Document DD-002 for Tier-2 SRM + vSphere Replication DR. Structure document: (1) Business Context: Risk Analytics workload has 30min RTO and 1hr RPO; data loss of up to 1 hour is acceptable, but >30min recovery would impact daily risk reporting. (2) Requirements: support asynchronous replication (RPO up to 1 hour acceptable), support failover to standby cluster in <30min, include network remapping for IP address translation between sites, must include NSX policy failover for firewall/NAT rules, must support Tier-2 application startup on Site-B (dependency ordering, app readiness checks). (3) Constraints: SRM licensing limits number of VMs (check VCF edition), replication network link limited to 1Gbps (shared with other traffic, not dedicated like Tier-1), standby cluster at Site-B has fewer nodes (2-4) to reduce cost, recovery plan complexity grows with number of VMs. (4) Assumptions: vSphere Replication lag will remain <60min under normal operations (workload change rate <5GB/hour), NSX policies are identical on both sites (pre-configured, not replicated), application startup is idempotent (can be safely restarted at Site-B after failover), database consistency is maintained during async replication (app-level transactions are ACID, no application-specific replication needed). (5) Risks: Replication lag exceeds 60min during business hours (IMPACT: RPO violation, Site-B data stale; MITIGATION: monitor lag, alert at 45min, consider upgrade to dedicated replication link if peak hours consistently breach threshold), NSX policy drift between Site-A and Site-B (IMPACT: failover succeeds but firewall blocks traffic; MITIGATION: version control NSX policies, automated policy sync or annual audit to ensure Sites are identical), application dependency ordering failures (IMPACT: database starts before file services, transactions fail; MITIGATION: define startup dependency sequence in recovery plan, test monthly), human error during SRM failover (IMPACT: wrong recovery plan executed or step skipped; MITIGATION: SRM runbook with checklist, two-person approval for planned failover). Alternatives: (1) Stretch cluster for Tier-2 (rejected: 1ms latency not critical for 30min RTO, cost of sync replication infrastructure not justified), (2) Backup-only for Tier-2 (rejected: 1-4hr restore time cannot meet 30min RTO). Conclude: 'Implement SRM + vSphere Replication for Tier-2 based on RTO/RPO balance and cost efficiency.'
Create Decision Document DD-003 for Tier-3 Backup-Based DR. Structure document: (1) Business Context: Reporting System is non-critical, used for historical analysis and not time-sensitive; 4hr RTO and 24hr RPO are acceptable for business continuity. (2) Requirements: backup system must support daily incremental snapshots, weekly full backups, retention of 7-day incremental chain and 4-week full backups (covering up to 28 days of recovery points), backup must be stored off-site or on secondary storage at Site-B, restore process must be documented and tested quarterly to ensure 4hr RTO is achievable. (3) Constraints: backup storage is shared infrastructure (limited to 6TB for 2.5TB data with 2:1 compression), backup link is 1Gbps (shared with replication and other traffic, limiting restore speed), backup restore depends on manual intervention (no automated recovery plan, requires human to initiate restore and validate), staff resources for backup management are limited (1 FTE). (4) Assumptions: data does not need to be point-in-time consistent (application can accept 24hr RPO, no transactions are lost, just day-old snapshots), restore from backup does not require parallel restore paths (serial restore acceptable within 4hr window), storage administrators can reliably execute restore procedures without error (assumption on staff competency). (5) Risks: Backup storage failure (IMPACT: cannot restore Reporting System after disaster, hours/days of data unavailable; MITIGATION: dual-site backup copies, monthly restore validation test, alert if backup copy fails), Restore exceeds 4hr RTO (IMPACT: Reporting System unavailable for extended period; MITIGATION: test actual restore time with full 2.5TB workload monthly, if test shows >4hr, reduce retention window or upgrade link bandwidth), RPO violation from missed backups (IMPACT: if backup job fails and is not noticed, next available backup is >24hr old; MITIGATION: backup job monitoring, alert if backup fails, daily verification that backup completed successfully). Alternatives: (1) Tier-2 SRM for Reporting System (rejected: cost of continuous replication not justified for non-critical workload, 30min RTO is overkill), (2) Snapshot-based replication (rejected: snapshot cost and storage overhead higher than backup approach for this use case). Conclude: 'Implement VADP backup-based DR for Tier-3 based on RTO/RPO requirements and cost efficiency.'
Create an Executive Summary document consolidating DD-001, DD-002, DD-003 into a single 1-page overview for stakeholder approval. Summary should include: (1) Overall DR architecture for all 3 tiers in a single table (Tier | RTO/RPO | Technology | Key Features | Cost Estimate), (2) Risk Register: top 5 risks across all tiers with IMPACT, MITIGATION, and Residual Risk (after mitigation), (3) Critical Success Factors: assumptions that must hold for DR to work (e.g., <1ms latency for Tier-1, NSX policies synchronized for Tier-2, backup link available for Tier-3), (4) Implementation Schedule: Tier-1 (months 1-3, stretch cluster build-out), Tier-2 (months 2-4, SRM configuration and testing), Tier-3 (months 1-2, backup policy configuration), (5) Approval signatures: CTO/VP Eng, Storage Lead, Networking Lead, Backup Admin, App Owner (Risk Analytics and Reporting), (6) Success Metrics: all 3 tiers tested and validated, zero RPO violations during monthly monitoring, failover runbooks approved and team trained, recovery RTO measured within 10% of design estimate. Include a visio diagram showing all three tiers and data flow paths for quick reference.
Validation Gate
Check: All RCAR decisions (DD-001, DD-002, DD-003) completed with detailed Requirements, Constraints, Assumptions, Risks; Executive Summary created with risk register and implementation timeline; decision documents reviewed by all stakeholders
Expected: Three decision documents + Executive Summary + signed stakeholder approval
Common Errors
Final Validation
All four DR architecture tasks completed: workload classification with tier mapping, topology design for each tier, failover validation with test results, and RCAR-based decision documentation
✓ Task 1: BIA spreadsheet, DR technology mapping, and license checklist completed → All three workloads classified into appropriate tiers with RTO/RPO requirements and technology selection justified
✓ Task 2: Tier-1 topology diagram (stretched cluster), Tier-2 topology with SRM recovery plan, Tier-3 backup topology, NSX cross-site networking checklist completed → Three detailed topology diagrams with replication paths, failover mechanisms, and RTO/RPO calculations documented
✓ Task 3: Host failure test log, SRM failover test report, RPO compliance validation, recovery runbooks completed → Practical validation of failover procedures with measured recovery times and documented runbooks for incident response
✓ Task 4: DD-001 (Tier-1), DD-002 (Tier-2), DD-003 (Tier-3), Executive Summary with risk register and implementation timeline, stakeholder approval signatures → Formal decision documentation with RCAR framework, alternatives analysis, risk mitigation, and executive-level approval
Cleanup / Restore
• Revert Tier-1 stretched cluster to stable state (power on any offline hosts, verify cluster quorum)
• Remove SRM test failover temporary VMs at Site-B
• Remove synthetic workload generators used for RPO validation testing
• Archive all lab test logs and validation reports to shared drive or wiki for future reference
• Document any infrastructure changes made during lab (e.g., new monitoring alerts, bandwidth reservations) for handoff to operations team
Design Reflection (VCDX)
VCDX panelists will probe the candidate's understanding of tradeoff analysis between RPO/RTO requirements and technology cost/complexity.
Expect questions: 'Why stretched cluster for Tier-1 and not SRM?' (answer: RPO=0 requires sync replication, SRM is async), 'What happens if witness host fails?' (answer: quorum logic, failover to majority site, show understanding of split-brain prevention), 'How do you ensure NSX policies are consistent across sites during SRM failover?' (answer: policy versioning, pre-configuration, optional NSX Federation), 'What is the financial impact of a 1-hour RPO violation for Risk Analytics?' (answer: depends on business context, in finance likely high cost of decisions made on stale data).
Panelists will also assess communication skills: can the candidate explain complex DR architecture to non-technical stakeholders using the Executive Summary? Can they articulate risks and mitigations without hand-waving?
Requirements
- Trading Platform (Tier-1) must achieve RPO=0 (zero data loss) and RTO<5 minutes
- Risk Analytics (Tier-2) must achieve RPO=1 hour and RTO=30 minutes
- Reporting System (Tier-3) must achieve RPO=24 hours and RTO=4 hours
- All tiers must support failover between two geographically separated sites (Site-A primary, Site-B secondary)
- DR architecture must integrate with NSX for cross-site networking, including policy and routing failover
- All RTO/RPO requirements must be validated through practical testing before production deployment
- Recovery procedures must be documented in runbooks and tested at least quarterly
- DR investment must be justified through RCAR analysis with alternatives comparison for each tier
Constraints
- Bandwidth limitation: 10Gbps dedicated link for Tier-1, 1Gbps shared link for Tier-2/Tier-3 replication and other traffic
- Latency requirement: Tier-1 stretched cluster requires <1ms site-to-site latency (achievable only with dedicated fiber, geographically proximate sites within 100km)
- Licensing: vSAN Enterprise for stretched cluster (Tier-1), SRM Standard/Enterprise for Tier-2, VADP for Tier-3; cost must fit IT budget
- Staffing: limited to 2 SREs for ongoing DR operations; complex architectures (stretched cluster) require specialized training
- Storage capacity: 6TB backup storage for 2.5TB Tier-3 workload with 2:1 compression (limited expansion headroom)
- Site-B resources: secondary cluster sized for minimum viable service (may be smaller than Site-A to reduce cost); impacts failover performance and application startup time
- Change management: any DR topology change requires testing in lab, validation through failover test, stakeholder approval (slow process, limits architecture flexibility)
Assumptions
- Site-to-site latency will remain <1ms under normal operations (requires stable network, no congestion); if latency spikes, stretched cluster efficiency degrades and may violate RPO
- Witness host for stretched cluster will remain available 99.99% of the time; failure of witness triggers quorum recalculation and potential service disruption
- Workload change rates will not exceed baseline estimates (Trading Platform 5-10% daily, Risk Analytics 5% daily); if change rate increases, replication lag grows and RPO may be violated
- NSX policies (firewall rules, NAT, routing) are pre-configured identically on both sites and will remain synchronized; drift between sites (e.g., new rule added at Site-A but not Site-B) will break failover
- Application workloads are idempotent and can be safely restarted at Site-B after failover without side effects (true for stateless apps, may not be true for apps with local state or singleton processes)
- Recovery procedures documented in runbooks are accurate and will be correctly followed under stress (assumes good runbook quality and staff training)
- Infrastructure changes (network, storage, hypervisor) will not invalidate DR design; any change requires RCAR re-evaluation
Risks
- Risk: Witness host failure in Tier-1 stretched cluster — Impact: If witness host becomes unavailable (network partition, hardware failure, power loss), stretched cluster loses quorum decision-making. With 2+2 host split between sites, cluster cannot determine which site has majority; both sites may attempt to accept writes, causing split-brain and data corruption. — Mitigation: Deploy witness host on HA-protected infrastructure (e.g., local vSAN cluster with separate power/network feeds), implement automated heartbeat monitoring with alert if witness is unreachable for >60 seconds, execute quarterly failover test (fail witness, verify cluster remains functional with remaining quorum), maintain runbook with manual quorum recovery steps if needed. — Residual_risk: Low; after mitigation, witness host failure is unlikely to cause data loss if detected and responded to quickly. Residual risk is that manual intervention may be required during witness recovery, causing brief service interruption. Acceptable risk level.
- Risk: Replication link failure or latency spike causing RPO violation — Impact: Tier-1 stretched cluster replication depends on <1ms latency and sustained throughput; if link fails or latency increases to 5ms+, replication queues grow, lag accumulates, RPO=0 target cannot be met. Tier-2 vSphere Replication lag may exceed 60min during peak load if link is congested, violating RPO. Tier-3 backup restore time may exceed 4hr RTO if backup link is unavailable. — Mitigation: Design dual replication links with active-active load balancing (if supported by switch/array) or active-standby automatic failover; implement QoS to reserve bandwidth for replication traffic; monitor replication link bandwidth and latency continuously with alerts at 80% utilization or latency >2ms; perform monthly link failover tests to verify automatic failover works; maintain low-cost fallback procedures (e.g., manual fail to Site-B if lag approaches RPO limit for Tier-2). — Residual_risk: Medium; despite mitigation, link failures can happen unexpectedly. Residual risk is reduced to acceptable level if monitoring and alerting are in place and team is trained to respond quickly. For Tier-1, if lag breaches RPO window, failover to Site-B is the only recovery option.
- Risk: NSX policy drift between Site-A and Site-B, causing firewall/routing failure during failover — Impact: SRM failover succeeds and VMs power on at Site-B, but network policies (firewall, NAT, routing) are not applied correctly because they differ from Site-A. Result: legitimate traffic is blocked by default-deny firewall, or traffic is routed incorrectly. Application appears down from external perspective even though VMs are running. — Mitigation: Pre-configure identical NSX Tier-1 routers and security policies on both sites before implementing DR; use version control (GitHub) to track NSX policy changes; implement policy comparison tool (NSX API or third-party tool) to audit policy consistency monthly; include NSX policy verification as a step in SRM test failover (verify firewall rules, NAT entries, routes); document NSX failover steps in recovery runbook (e.g., activate Site-B router, confirm traffic flow); consider NSX Federation (if available in VCF 9.0) for automatic policy replication. — Residual_risk: Medium-low; policy drift is a common failover issue but can be prevented with good change management. Residual risk is that new policies added at Site-A are not immediately applied to Site-B, requiring manual sync. Acceptable if monthly audit ensures drift is detected and corrected before failover is needed.
- Risk: Backup restore exceeds 4-hour RTO, causing Tier-3 reporting system to be unavailable beyond acceptable window — Impact: Tier-3 Reporting System backup restore from 2.5TB data with limited 1Gbps link and HDD-based storage takes 6+ hours, exceeding 4hr RTO. Business stakeholders cannot access reporting for 6+ hours, impacting strategic decision-making and daily reporting obligations. — Mitigation: Validate actual backup restore time with full-size 2.5TB workload in quarterly lab test; if test shows restore exceeds 4 hours, identify bottleneck (link bandwidth, storage IO, metadata retrieval) and remediate (upgrade link to 10Gbps, add SSD-based backup storage, parallelize restore across multiple backup appliances); pre-stage backup copies at Site-B to avoid retrieval latency; document actual measured RTO in recovery runbook rather than assuming 4 hours; set RTO expectations with business stakeholders based on validated times. — Residual_risk: Medium; backup restore time depends on many variables (link bandwidth, storage speed, workload characteristics) that may be unpredictable. After mitigation, residual risk is that restore may take longer than expected if storage or link is congested at time of disaster. Acceptable if backup restore time is monitored and communicated to stakeholders.
- Risk: Human error during SRM failover or recovery runbook execution, causing wrong VMs to be failed over, network mapping errors, or recovery steps skipped — Impact: During crisis, team member executing recovery runbook makes typo (e.g., selects wrong recovery plan in SRM, fails over wrong VM group), or skips validation step, resulting in partial failover or data inconsistency. Application is partially recovered or inaccessible, extending RTO and potentially causing customer impact. — Mitigation: Create detailed, step-by-step recovery runbooks with screenshots and decision trees for different failure scenarios; implement two-person approval process for planned failovers (incident commander approves, separate team member executes); conduct quarterly runbook training and failover drills with full team; use SRM test failover mode extensively before any planned failover; add pre-flight checks to runbook (verify backup is complete, check replication lag, confirm no in-flight transactions); document all steps with expected outputs (e.g., 'After step 3, confirm 10 VMs show 'Ready' status in SRM UI'). — Residual_risk: Low-medium; after mitigation, human error risk is reduced through process discipline and training. Residual risk remains that under high stress, team may deviate from runbook or miss steps. Acceptable if runbooks are actively maintained and team is well-trained.
Self-Assessment Discussion Prompts
- If the site-to-site latency is measured at 5ms instead of the assumed <1ms, what changes to the Tier-1 stretched cluster design would you recommend? Would you still use stretched cluster or switch to SRM for Tier-1? What is the business impact of the change?
- Describe a scenario where Tier-1 stretched cluster RPO=0 is violated despite no hardware failure. What conditions could cause replication lag to accumulate? How would you detect and respond?
- In the Tier-2 SRM failover runbook, NSX policies must be applied at Site-B before application traffic reaches Site-B routers. How would you orchestrate this in the recovery plan? What automation or manual steps are needed?
- For Tier-3 backup-based DR, the 4-hour RTO assumes backup retrieval time of 1.5 hours. What factors could cause backup retrieval to take longer? How would you design the backup infrastructure to minimize retrieval time without breaking budget constraints?
- If both Site-A and Site-B fail simultaneously (rare but possible), how would you recover Tier-1 and Tier-2 workloads? What is your strategy for this scenario and what are the time and data loss implications?
- The RCAR document assumes NSX policies are pre-configured identically on both sites. In practice, how would you enforce and audit this assumption? What tools or processes would help maintain policy synchronization over time?
Extensions
Extend Tier-1 to Three-Site Active-Active Cluster with Distributed Replication
Instead of two-site stretched cluster with witness, design a three-site vSAN cluster where all sites actively serve workloads (active-active-active) with distributed replication across three sites. Each site has a 2-node cluster (6 nodes total), no dedicated witness needed (odd number of sites = automatic quorum). Requires evaluation of inter-site replication topology: star pattern (all sites replicate to Site-A hub) vs. mesh pattern (all sites replicate to each other) vs. hierarchical pattern (Site-A primary, Site-B secondary, Site-C backup). Model the trade-offs: star pattern reduces replication overhead but creates single point of failure (Site-A replication hub); mesh pattern maximizes redundancy but increases replication traffic 3x. Document recovery time if one entire site is lost (should be <30 seconds due to automatic failover to majority sites).
+2 (advanced to expert)Design RPO=0 for Tier-2 Using Array-Based Synchronous Replication
Challenge the assumption that Tier-2 requires RPO=1hr; investigate whether array-based synchronous replication (e.g., Pure Storage ActiveCluster, Dell/EMC SRDF, NetApp MetroCluster) could be used to achieve RPO=0 for Tier-2 Risk Analytics without upgrading to full stretched cluster architecture. Compare cost, complexity, and failover performance of array-based sync replication vs. stretched vSAN cluster. Document network latency requirements (typically <5ms for array replication vs. <1ms for vSAN stretched cluster). Evaluate whether synchronous array replication imposes write latency penalties that would degrade application performance. Conclude whether this extension is economically justified compared to current SRM design.
+1 (advanced, requires deep storage knowledge)Implement Automated Failover Using vSAN Stretched Cluster with NSX AVI Load Balancer Global Service Load Balancing (GSLB)
Design automatic failover without manual intervention by combining vSAN stretched cluster (automatic host failure detection) with NSX AVI Global Load Balancer for DNS-based failover. When Site-A cluster quorum is lost, AVI GSLB automatically updates DNS to point to Site-B IP addresses, redirecting client traffic. Define the failover sequence: vSAN detects quorum loss at Site-B (Site-A hosts down), AVI GSLB detects Site-A datacenters unreachable (health check failure), AVI updates DNS to Site-B addresses, clients re-resolve DNS and traffic flows to Site-B. Model failover time including DNS update latency (typically 5-30 seconds) and client re-connection time. Evaluate whether fully automated failover (no manual runbook steps) is appropriate for Tier-1 trading workload or if human oversight is still needed.
+2 (advanced, requires load balancer and DNS knowledge)Evaluate VMware Cloud Disaster Recovery (VCDR) as Alternative to On-Premises DR
Instead of building on-premises two-site DR, evaluate VMware's cloud-native disaster recovery service (VCDR, formerly Zerto) as an alternative for Tier-2 and Tier-3 workloads. VCDR offers cloud-based standby environment, continuous asynchronous replication, and point-in-time recovery. Compare VCDR cost/performance/complexity vs. on-premises SRM approach: VCDR eliminates need for secondary on-premises cluster (cost savings), provides frequent recovery point objectives (minutes rather than hours), but introduces cloud provider dependency and potential multi-tenancy security concerns. Model failover workflow in VCDR (activate cloud VMs in VCDR-managed environment) vs. SRM (activate VMs in on-premises Site-B cluster). Determine appropriate tier for cloud DR: is Tier-3 reporting system a good fit for cloud DR? Would Tier-1 trading platform ever be cloud DR candidate (likely not due to latency and data residency concerns)?
+1 (advanced, requires cloud services knowledge)⚠ Known Pitfalls (from Community KB)
References
- vSAN Stretched Cluster Deployment GuideTier 1 — Official
- VMware Site Recovery Manager (SRM) Administration and DeploymentTier 1 — Official
- vSphere Replication Administration GuideTier 1 — Official
- VMware vCloud Foundation 9.0 Architecture and PlanningTier 1 — Official