On-Premises Ransomware Recovery for VMware Cloud Foundation
This solution addresses on-premises ransomware recovery using an Isolated Recovery Environment (IRE) at a secondary VCF 9.0 site. The VCF 5.2 validated solution library instead covers Cloud-Based Ransomware Recovery via VMware Live Cyber Recovery SaaS. Compared with the cloud variant, the VCF 9.0 on-premises pattern trades SaaS convenience for data sovereignty, gives customers full control of the recovery stack (VLR appliances, NSX isolation, vSAN ESA IRE), and introduces Python-based nsxtRwrIsolationPolicies.py automation plus native VCF Ops integration for entropy-based detection.
Key Components: VMware Live Recovery, vSphere Replication, vSAN Data Protection, NSX, vDefend, vCenter, VCF Operations, Carbon Black, CrowdStrike
Purpose: Detailed design, implementation, configuration, and operational guidance for recovering on-premises business critical workloads within an Isolated Recovery Environment (IRE) in a VMware Cloud Foundation platform, in the event of a ransomware attack in the protected instance.
33 design decisions
| DD-ID | Decision | Quality |
|---|---|---|
| RWR-VLR-CFG-001 | Deploy each VLR instance in the management domain. | Manageability |
Decision: Deploy each VLR instance in the management domain. Rationale: Consistent deployment model for all management applications. Implication: No significant trade-offs identified for this decision. Component: VMware Live Recovery | ||
| RWR-VLR-CFG-002 | Use vSphere Replication in VLR as protection method for VM replication. | Recoverability |
Decision: Use vSphere Replication in VLR as protection method for VM replication. Rationale: Storage agnostic; flexibility in storage selection; minimizes SRA compatibility overhead; Enhanced VR = direct host-to-host at target site. Implication: Requires TCP 32032 network connectivity from protected ESX hosts to recovery vSAN storage cluster ESX hosts and VLR appliance. Component: VMware Live Recovery | ||
| RWR-VLR-NET-001 | Place VLR instances on the management network. | Manageability |
Decision: Place VLR instances on the management network. Rationale: Co-locate with VCF components VLR must communicate with. Implication: No significant trade-offs identified for this decision. Component: VMware Live Recovery | ||
| RWR-VLR-NET-002 | Allocate and assign static IP to VLR instances. | Manageability |
Decision: Allocate and assign static IP to VLR instances. Rationale: Removes DHCP constraints/risks on management networks. Implication: Requires precise IP management. Component: VMware Live Recovery | ||
| RWR-VLR-NET-003 | Configure forward (A) and reverse (PTR) DNS records for each VLR. | Manageability |
Decision: Configure forward (A) and reverse (PTR) DNS records for each VLR. Rationale: VLR accessible via FQDN. Implication: DNS infrastructure must be available; must manage DNS records. Component: VMware Live Recovery | ||
| RWR-VLR-NET-004 | Configure VLR to use NTP servers (not VMTools). | Manageability |
Decision: Configure VLR to use NTP servers (not VMTools). Rationale: Ensures accurate time sync; prevents time mismatch between management components. Implication: NTP services must be available. Component: VMware Live Recovery | ||
| RWR-VLR-LCM-001 | Life cycle management of VLR via native appliance tools. | ManageabilityPerformance |
Decision: Life cycle management of VLR via native appliance tools. Rationale: VLR not managed by VCF Operations. Implication: Deployment, patching, updates, upgrades performed without native automation. Component: VMware Live Recovery | ||
| RWR-VLR-SEC-001 | Configure service account in each vCenter for app-to-app communication from VLR to vSphere, member o | Manageability |
Decision: Configure service account in each vCenter for app-to-app communication from VLR to vSphere, member of vCenter SSO Administrators group. Rationale: VLSR accesses vSphere with required permissions for DR orchestration + site pairing; VR accesses for replication; improved accountability; limits blast radius on compromise. Implication: Maintain se Component: VMware Live Recovery | ||
| RWR-VLR-SEC-002 | Configure password expiration policy for VLR appliance. | Manageability |
Decision: Configure password expiration policy for VLR appliance. Rationale: Align with org policies and compliance standards; only local users. Implication: Manage via appliance console or SSH. Component: VMware Live Recovery | ||
| RWR-VLR-SEC-003 | Configure password complexity policy for VLR appliance. | Manageability |
Decision: Configure password complexity policy for VLR appliance. Rationale: Align with org/compliance policies; local users only. Implication: Manage via console or SSH. Component: VMware Live Recovery | ||
| RWR-VLR-SEC-004 | Configure account lockout policy for VLR appliance. | Manageability |
Decision: Configure account lockout policy for VLR appliance. Rationale: Align with org/compliance policies; local users only. Implication: Manage via console or SSH. Component: VMware Live Recovery | ||
| RWR-VLR-SEC-005 | Change VLR root password on recurring/event-initiated schedule via console or SSH. | Performance |
Decision: Change VLR root password on recurring/event-initiated schedule via console or SSH. Rationale: Root passwords never expire by default. Implication: Must routinely perform password changes. Component: VMware Live Recovery | ||
| RWR-VLR-SEC-006 | Change VLR admin password on recurring/event-initiated schedule via console or SSH. | Performance |
Decision: Change VLR admin password on recurring/event-initiated schedule via console or SSH. Rationale: Admin passwords never expire by default. Implication: Must routinely perform password changes. Component: VMware Live Recovery | ||
| RWR-VLR-SEC-007 | Replace default self-signed certificate with CA-signed in each VLR. | Security |
Decision: Replace default self-signed certificate with CA-signed in each VLR. Rationale: All externally facing Web UI and cross-product communication encrypted. Implication: CA-signed cert acquisition may increase deployment preparation time. Component: VMware Live Recovery | ||
| RWR-IRE-CFG-001 | Configure isolated workload domain with 1x vSAN ESA storage cluster + 1x compute cluster with NSX Ed | PerformanceRecoverabilitySecurity |
Decision: Configure isolated workload domain with 1x vSAN ESA storage cluster + 1x compute cluster with NSX Edges and remote datastore from storage cluster. Rationale: vSAN ESA performance/efficiency/scalability via NVMe TLC flash; network-restricted environment disconnected from prod for safe ransomware recovery. Implication: Additional infr Component: IRE | ||
| RWR-IRE-CFG-002 | Configure Tier-1 gateway and IRE network segment in isolated workload domain NSX. | RecoverabilitySecurity |
Decision: Configure Tier-1 gateway and IRE network segment in isolated workload domain NSX. Rationale: Connect recovered test VMs. Implication: No significant trade-offs identified for this decision. Component: IRE | ||
| RWR-IRE-CFG-003 | Configure DHCP service for IRE segment in isolated workload domain NSX. | RecoverabilitySecurity |
Decision: Configure DHCP service for IRE segment in isolated workload domain NSX. Rationale: Recovered test VMs receive IP addresses and DNS settings from DHCP server. Implication: No significant trade-offs identified for this decision. Component: IRE | ||
| RWR-IRE-CFG-004 | Use Python script to create different network isolation levels on IRE. | RecoverabilitySecurity |
Decision: Use Python script to create different network isolation levels on IRE. Rationale: Network isolation prevents lateral malware spread; allows granular behavioral analysis at different phases of recovery. Implication: Separate vDefend DFW license required. Component: IRE | ||
| RWR-IRE-CFG-005 | Configure org firewall to allow outbound traffic to EDR from IRE. | RecoverabilitySecurity |
Decision: Configure org firewall to allow outbound traffic to EDR from IRE. Rationale: Recovered test VM must access EDR portal for sensor install + security analysis. Implication: Network configuration required. Component: IRE | ||
| RWR-IRE-CFG-006 | Configure separate DNS server on recovery instance. | ManageabilityRecoverability |
Decision: Configure separate DNS server on recovery instance. Rationale: IRE must use DNS distinct from protected instance DNS. Implication: Additional infrastructure + operational overhead. Component: IRE | ||
| RWR-IRE-CFG-007 | Configure NTP server different from protected instance. | Manageability |
Decision: Configure NTP server different from protected instance. Rationale: IRE Internet path should NOT follow production workload path. Implication: Open access to Internet-based NTP server. Component: IRE | ||
| RWR-IRE-CFG-008 | Isolate vSphere Replication network traffic from all other traffic in both instances; in recovery, i | RecoverabilitySecurity |
Decision: Isolate vSphere Replication network traffic from all other traffic in both instances; in recovery, isolate on vSAN storage cluster ESX hosts. Rationale: Security isolation via dedicated VR VLAN for replication traffic. Implication: Network configuration required. Component: IRE | ||
| VLR-VR-CFG-001 | Do NOT activate guest OS quiescing in VR policies. | Manageability |
Decision: Do NOT activate guest OS quiescing in VR policies. Rationale: Not all workloads support quiescing; may cause outage. Implication: Crash-consistent (not application-consistent) replicas. Component: vSphere Replication | ||
| VLR-VR-CFG-002 | Activate network compression on VR policies. | Performance |
Decision: Activate network compression on VR policies. Rationale: Reduces bandwidth for replication. Implication: More CPU at source and destination. Component: vSphere Replication | ||
| VLR-VR-CFG-003 | Configure PIT instances, keeping last 200 instances. | Recoverability |
Decision: Configure PIT instances, keeping last 200 instances. Rationale: Ransomware recovery needs comprehensive PIT history to find uncompromised snapshot; vSAN ESA recommended for values over 24. Implication: More retained instances = more disk usage on vSAN ESA datastore. Component: vSphere Replication | ||
| VLR-VR-CFG-004 | Configure RPO of 24 hours. | Recoverability |
Decision: Configure RPO of 24 hours. Rationale: Ransomware detection can lag for weeks/months; most recent replicas may also be infected. 200 instances * 24-hour RPO = up to 200 days history. Implication: Changes made after ransomware attack typically lost. Component: vSphere Replication | ||
| RWR-IRE-RP-001 | Use Python script to select replica from list and create test VM on IRE. | Recoverability |
Decision: Use Python script to select replica from list and create test VM on IRE. Rationale: No UI option to select required replica; ransomware recovery is delicate iterative process needing uncompromised snapshots. Implication: No significant trade-offs identified for this decision. Component: IRE | ||
| RWR-IRE-RP-002 | Use Python script to assign appropriate isolation level to test VM during analysis. | Security |
Decision: Use Python script to assign appropriate isolation level to test VM during analysis. Rationale: Isolation options limit access; Quarantined+Analysis permits EDR sensor install while denying general Internet access. Implication: No significant trade-offs identified for this decision. Component: IRE | ||
| RWR-IRE-RP-003 | Use EDR for security analysis on test VM. | ManageabilityRecoverabilitySecurity |
Decision: Use EDR for security analysis on test VM. Rationale: Recovered VM may contain malware; need to inspect OS + app vulnerabilities; monitor events. Implication: Separate EDR license; allow outbound to EDR; manually install sensor pre-analysis, uninstall post. Component: IRE | ||
| RWR-IRE-RP-004 | Use VMware Live Site Recovery to automate recovery from promoted (staged) image. | ManageabilityRecoverability |
Decision: Use VMware Live Site Recovery to automate recovery from promoted (staged) image. Rationale: Automated runbook for business critical workload recovery. Implication: No significant trade-offs identified for this decision. Component: IRE | ||
| VLR-SRM-RP-001 | Do NOT run test recovery of recovery plans. | Recoverability |
Decision: Do NOT run test recovery of recovery plans. Rationale: Changes made to remove malware are not persistent after test run. Implication: No significant trade-offs identified for this decision. Component: Site Recovery Manager | ||
| RWR-DNS-NET-001 | Configure DNS settings for each protected workload to use DNS servers across protected and recovery | Recoverability |
Decision: Configure DNS settings for each protected workload to use DNS servers across protected and recovery instances. Rationale: Enables DNS resolution when recovered test VMs powered on in IRE or during planned migration/DR. Implication: No significant trade-offs identified for this decision. Component: DNS | ||
| RWR-NTP-NET-001 | Configure NTP settings for each protected workload to use NTP servers across both instances. | Recoverability |
Decision: Configure NTP settings for each protected workload to use NTP servers across both instances. Rationale: Enables NTP when recovered test VMs powered on in IRE or during migration/DR. Implication: No significant trade-offs identified for this decision. Component: NTP | ||
Prerequisites
- VCF management domain operational
- DNS/NTP configured
- CA infrastructure in place
Implementation Procedure
Implementation
Environment: VCF version in support matrix; configured per guidance; capture all Ransomware Recovery tab parameters; VCF platforms healthy.
DNS: Required DNS entries in forward and reverse zones
Network: Protected <-> Recovery connectivity with jumbo frames, Layer 3 routing, max 150ms latency; bandwidth sized with VLR Calculator
Software: VLR .iso image mounted on vSphere Client machine
License: VCF license per design (protected + recovery), VMware vDefend Firewall license for IRE isolation
Active Directory: Domain Controllers available
CA: Microsoft CA available
- Workloads: Windows or Linux OS with latest updates + VMware Tools
- Deploy VLR Appliance
- MethodDeploy OVF template in vSphere Client
- PlacementDefault management cluster in management domain vCenter
- StorageManagement domain vSAN datastore
- NetworkManagement network
- Repeat ForRecovery instance using Planning Workbook for Recovery
- Vmdk Files
- VLR-9.0.3.0.XXXXXXXX_OVF10.ovf
- -database.vmdk
- -data.vmdk
- -support.vmdk
- -swap.vmdk
- -system.vmdk
- Customizations
- Enable SSHDSelected
- Host Network IP Address Familyipv4
- Host Network Modestatic
- Certificate Replacement
ProcessVLR appliance config UI (https://<vlr_fqdn>:5480) > Certificates
StepsGenerate CSR (Appliance certificate > Generate CSR), submit to Microsoft AD CS, upload signed cert back
Repeat ForRecovery instance
Day-2 Operations Tasks
Operations
As neededOperational Verification
As neededEndpointProtected vCenter > VMware Live Site Recovery
As neededPassword Management
As neededExpiration Policy
As neededMethodSSH to VLR as admin, switch to root
As neededRepeat Foradmin user and recovery instance
As neededCommands
As neededchage --maxdays <value> root
As neededchage --mindays <value> root
As neededMonitoring Points
- Verify Pairing
- CheckStatus of VLSR and vSphere Replication = OK (green checkmark)
- PreVerify all replications paused (VMware Live Site Recovery > site pair > View Details > Replications > Outgoing = Paused)
- Verify VR and VLCR operational
- VerifyNSX Manager > Security > Policy Management > Distributed Firewall > Application tab - verify ransomware recovery policies listed
- Failover Failback Checklist
- Verify 'Task finished successfully' and 'Created test vm with name <vm_name>-test-vm'
- Inspect OS/app vulnerabilities on EDR console; monitor events; manually patch + remove malware; behavioral analysis
- Change isolation levels to iterate (e.g., EXTERNAL_OUTBOUND) and monitor behavior on EDR
- WindowsDownload from falcon.crowdstrike.com > Host setup and management > Sensor downloads > Windows. Install FalconSensor_Windows.exe, enter Customer
Troubleshooting
Likely Panelist Questions
Q: Why did you choose this architecture?
See design decisions for rationale
Failure Scenarios
Trade-off Analysis
Why DISABLE guest OS quiescing - performance impact on high-I/O workloads (trade crash-consistent for app-consistent)
Chosen:
Justification:
Trade-Offs Analysis
Chosen:
Justification:
Service account with Admin role = simplicity vs. least-privilege principle
Chosen:
Justification:
Not running test recovery plans = accept recovery plan execution risk vs. persistent changes preservation
Chosen:
Justification:
Quiz — Ransomware Recovery
- 1 hour
- 4 hours
- 24 hours
- 72 hours
- Isolated
- Quarantined
- Quarantined+Analysis
- Open
- 50 ms
- 100 ms
- 150 ms
- 300 ms
- 443
- 902
- 32032
- 8443
- vSAN OSA (Original)
- vSAN ESA (Express Storage Architecture)
- VMFS
- NFS
- UI is deprecated
- There is no option in the UI to select a required replica
- Scripts are faster
- UI doesn't support NSX
- 50
- 100
- 200
- 500
- It is insecure
- Not all business critical workloads support it, and it may cause an outage on high-I/O VMs
- It requires additional licensing
- It is incompatible with vSAN
- 4 vCPU / 16 GB / 400 GB
- 8 vCPU / 24 GB / 826 GB
- 12 vCPU / 32 GB / 1 TB
- 16 vCPU / 64 GB / 2 TB
- They are not supported by VLR
- Changes made to remove malware are not persistent after test run
- They corrupt replicas
- Licensing restrictions
- > 0.3
- > 0.5
- > 0.7
- > 0.9
- Read-Only
- No Access
- Administrators
- Custom
- Planned migration
- Disaster recovery
- Test recovery
- Reprotect
- vcf-data-protection/examples/
- vcf-data-protection/examples/rwr-ref-arch/
- vcf-data-protection/examples/rwr-ref-arch/nsxtRwrIsolationPolicies
- vcf-data-protection/scripts/
- 30 days
- 90 days
- 200 days
- 365 days
Flashcards — Ransomware Recovery
Labs
Lab 1: Deploy and Pair VLR Across Protected and Recovery Instances
Deploy VMware Live Recovery appliances in both the protected and recovery VCF instances, register with vCenter SSO, replace certificates with CA-signed, and create a site pair.
Starting State: Two VCF 9.0 instances (protected + recovery) with management domains operational, DNS/NTP/CA in place, VLR .iso mounted, Planning and Preparation Workbook complete.
Lab 2: Configure IRE Network Isolation Levels via Python Script
Create Tier-1 gateway + IRE segment with DHCP in the isolated workload domain, deploy 8 network isolation levels via the Python script, and assign QUARANTINED_ANALYSIS policy to a test VM.
Starting State: Recovery VCF instance with isolated workload domain deployed (compute cluster + vSAN ESA storage cluster). vDefend Firewall license applied. Windows Jump Box with Python 3 and vcf-data-protection repo cloned.
Lab 3: Full Ransomware Recovery Workflow - Recover a Workload from a Clean Snapshot
Simulate a ransomware attack (using entropy analysis), create test VM from clean replica on IRE, install/uninstall EDR sensor, promote image, fail over via VLCR, then reprotect and fail back.
Starting State: Labs 1 and 2 complete. Protected workload replicating to recovery with RPO=24h + 200 PIT. VCF Ops with RWR management pack installed; dashboards imported. CrowdStrike Falcon or Carbon Black portal access.