Academy/VVS/Ransomware Recovery
This solution targets VCF 9.0

On-Premises Ransomware Recovery for VMware Cloud Foundation

VCF 9.0architectvcdxautomationstoragePages 7-82

This solution addresses on-premises ransomware recovery using an Isolated Recovery Environment (IRE) at a secondary VCF 9.0 site. The VCF 5.2 validated solution library instead covers Cloud-Based Ransomware Recovery via VMware Live Cyber Recovery SaaS. Compared with the cloud variant, the VCF 9.0 on-premises pattern trades SaaS convenience for data sovereignty, gives customers full control of the recovery stack (VLR appliances, NSX isolation, vSAN ESA IRE), and introduces Python-based nsxtRwrIsolationPolicies.py automation plus native VCF Ops integration for entropy-based detection.

Key Components: VMware Live Recovery, vSphere Replication, vSAN Data Protection, NSX, vDefend, vCenter, VCF Operations, Carbon Black, CrowdStrike

Purpose: Detailed design, implementation, configuration, and operational guidance for recovering on-premises business critical workloads within an Isolated Recovery Environment (IRE) in a VMware Cloud Foundation platform, in the event of a ransomware attack in the protected instance.

Design Decisions
Implementation
Operations
VCDX Defense
Quiz (15)
Flashcards (15)

33 design decisions

DD-IDDecisionQuality
RWR-VLR-CFG-001Deploy each VLR instance in the management domain.Manageability

Decision: Deploy each VLR instance in the management domain.

Rationale: Consistent deployment model for all management applications.

Implication: No significant trade-offs identified for this decision.

Component: VMware Live Recovery

RWR-VLR-CFG-002Use vSphere Replication in VLR as protection method for VM replication.Recoverability

Decision: Use vSphere Replication in VLR as protection method for VM replication.

Rationale: Storage agnostic; flexibility in storage selection; minimizes SRA compatibility overhead; Enhanced VR = direct host-to-host at target site.

Implication: Requires TCP 32032 network connectivity from protected ESX hosts to recovery vSAN storage cluster ESX hosts and VLR appliance.

Component: VMware Live Recovery

RWR-VLR-NET-001Place VLR instances on the management network.Manageability

Decision: Place VLR instances on the management network.

Rationale: Co-locate with VCF components VLR must communicate with.

Implication: No significant trade-offs identified for this decision.

Component: VMware Live Recovery

RWR-VLR-NET-002Allocate and assign static IP to VLR instances.Manageability

Decision: Allocate and assign static IP to VLR instances.

Rationale: Removes DHCP constraints/risks on management networks.

Implication: Requires precise IP management.

Component: VMware Live Recovery

RWR-VLR-NET-003Configure forward (A) and reverse (PTR) DNS records for each VLR.Manageability

Decision: Configure forward (A) and reverse (PTR) DNS records for each VLR.

Rationale: VLR accessible via FQDN.

Implication: DNS infrastructure must be available; must manage DNS records.

Component: VMware Live Recovery

RWR-VLR-NET-004Configure VLR to use NTP servers (not VMTools).Manageability

Decision: Configure VLR to use NTP servers (not VMTools).

Rationale: Ensures accurate time sync; prevents time mismatch between management components.

Implication: NTP services must be available.

Component: VMware Live Recovery

RWR-VLR-LCM-001Life cycle management of VLR via native appliance tools.ManageabilityPerformance

Decision: Life cycle management of VLR via native appliance tools.

Rationale: VLR not managed by VCF Operations.

Implication: Deployment, patching, updates, upgrades performed without native automation.

Component: VMware Live Recovery

RWR-VLR-SEC-001Configure service account in each vCenter for app-to-app communication from VLR to vSphere, member oManageability

Decision: Configure service account in each vCenter for app-to-app communication from VLR to vSphere, member of vCenter SSO Administrators group.

Rationale: VLSR accesses vSphere with required permissions for DR orchestration + site pairing; VR accesses for replication; improved accountability; limits blast radius on compromise.

Implication: Maintain se

Component: VMware Live Recovery

RWR-VLR-SEC-002Configure password expiration policy for VLR appliance.Manageability

Decision: Configure password expiration policy for VLR appliance.

Rationale: Align with org policies and compliance standards; only local users.

Implication: Manage via appliance console or SSH.

Component: VMware Live Recovery

RWR-VLR-SEC-003Configure password complexity policy for VLR appliance.Manageability

Decision: Configure password complexity policy for VLR appliance.

Rationale: Align with org/compliance policies; local users only.

Implication: Manage via console or SSH.

Component: VMware Live Recovery

RWR-VLR-SEC-004Configure account lockout policy for VLR appliance.Manageability

Decision: Configure account lockout policy for VLR appliance.

Rationale: Align with org/compliance policies; local users only.

Implication: Manage via console or SSH.

Component: VMware Live Recovery

RWR-VLR-SEC-005Change VLR root password on recurring/event-initiated schedule via console or SSH.Performance

Decision: Change VLR root password on recurring/event-initiated schedule via console or SSH.

Rationale: Root passwords never expire by default.

Implication: Must routinely perform password changes.

Component: VMware Live Recovery

RWR-VLR-SEC-006Change VLR admin password on recurring/event-initiated schedule via console or SSH.Performance

Decision: Change VLR admin password on recurring/event-initiated schedule via console or SSH.

Rationale: Admin passwords never expire by default.

Implication: Must routinely perform password changes.

Component: VMware Live Recovery

RWR-VLR-SEC-007Replace default self-signed certificate with CA-signed in each VLR.Security

Decision: Replace default self-signed certificate with CA-signed in each VLR.

Rationale: All externally facing Web UI and cross-product communication encrypted.

Implication: CA-signed cert acquisition may increase deployment preparation time.

Component: VMware Live Recovery

RWR-IRE-CFG-001Configure isolated workload domain with 1x vSAN ESA storage cluster + 1x compute cluster with NSX EdPerformanceRecoverabilitySecurity

Decision: Configure isolated workload domain with 1x vSAN ESA storage cluster + 1x compute cluster with NSX Edges and remote datastore from storage cluster.

Rationale: vSAN ESA performance/efficiency/scalability via NVMe TLC flash; network-restricted environment disconnected from prod for safe ransomware recovery.

Implication: Additional infr

Component: IRE

RWR-IRE-CFG-002Configure Tier-1 gateway and IRE network segment in isolated workload domain NSX.RecoverabilitySecurity

Decision: Configure Tier-1 gateway and IRE network segment in isolated workload domain NSX.

Rationale: Connect recovered test VMs.

Implication: No significant trade-offs identified for this decision.

Component: IRE

RWR-IRE-CFG-003Configure DHCP service for IRE segment in isolated workload domain NSX.RecoverabilitySecurity

Decision: Configure DHCP service for IRE segment in isolated workload domain NSX.

Rationale: Recovered test VMs receive IP addresses and DNS settings from DHCP server.

Implication: No significant trade-offs identified for this decision.

Component: IRE

RWR-IRE-CFG-004Use Python script to create different network isolation levels on IRE.RecoverabilitySecurity

Decision: Use Python script to create different network isolation levels on IRE.

Rationale: Network isolation prevents lateral malware spread; allows granular behavioral analysis at different phases of recovery.

Implication: Separate vDefend DFW license required.

Component: IRE

RWR-IRE-CFG-005Configure org firewall to allow outbound traffic to EDR from IRE.RecoverabilitySecurity

Decision: Configure org firewall to allow outbound traffic to EDR from IRE.

Rationale: Recovered test VM must access EDR portal for sensor install + security analysis.

Implication: Network configuration required.

Component: IRE

RWR-IRE-CFG-006Configure separate DNS server on recovery instance.ManageabilityRecoverability

Decision: Configure separate DNS server on recovery instance.

Rationale: IRE must use DNS distinct from protected instance DNS.

Implication: Additional infrastructure + operational overhead.

Component: IRE

RWR-IRE-CFG-007Configure NTP server different from protected instance.Manageability

Decision: Configure NTP server different from protected instance.

Rationale: IRE Internet path should NOT follow production workload path.

Implication: Open access to Internet-based NTP server.

Component: IRE

RWR-IRE-CFG-008Isolate vSphere Replication network traffic from all other traffic in both instances; in recovery, iRecoverabilitySecurity

Decision: Isolate vSphere Replication network traffic from all other traffic in both instances; in recovery, isolate on vSAN storage cluster ESX hosts.

Rationale: Security isolation via dedicated VR VLAN for replication traffic.

Implication: Network configuration required.

Component: IRE

VLR-VR-CFG-001Do NOT activate guest OS quiescing in VR policies.Manageability

Decision: Do NOT activate guest OS quiescing in VR policies.

Rationale: Not all workloads support quiescing; may cause outage.

Implication: Crash-consistent (not application-consistent) replicas.

Component: vSphere Replication

VLR-VR-CFG-002Activate network compression on VR policies.Performance

Decision: Activate network compression on VR policies.

Rationale: Reduces bandwidth for replication.

Implication: More CPU at source and destination.

Component: vSphere Replication

VLR-VR-CFG-003Configure PIT instances, keeping last 200 instances.Recoverability

Decision: Configure PIT instances, keeping last 200 instances.

Rationale: Ransomware recovery needs comprehensive PIT history to find uncompromised snapshot; vSAN ESA recommended for values over 24.

Implication: More retained instances = more disk usage on vSAN ESA datastore.

Component: vSphere Replication

VLR-VR-CFG-004Configure RPO of 24 hours.Recoverability

Decision: Configure RPO of 24 hours.

Rationale: Ransomware detection can lag for weeks/months; most recent replicas may also be infected. 200 instances * 24-hour RPO = up to 200 days history.

Implication: Changes made after ransomware attack typically lost.

Component: vSphere Replication

RWR-IRE-RP-001Use Python script to select replica from list and create test VM on IRE.Recoverability

Decision: Use Python script to select replica from list and create test VM on IRE.

Rationale: No UI option to select required replica; ransomware recovery is delicate iterative process needing uncompromised snapshots.

Implication: No significant trade-offs identified for this decision.

Component: IRE

RWR-IRE-RP-002Use Python script to assign appropriate isolation level to test VM during analysis.Security

Decision: Use Python script to assign appropriate isolation level to test VM during analysis.

Rationale: Isolation options limit access; Quarantined+Analysis permits EDR sensor install while denying general Internet access.

Implication: No significant trade-offs identified for this decision.

Component: IRE

RWR-IRE-RP-003Use EDR for security analysis on test VM.ManageabilityRecoverabilitySecurity

Decision: Use EDR for security analysis on test VM.

Rationale: Recovered VM may contain malware; need to inspect OS + app vulnerabilities; monitor events.

Implication: Separate EDR license; allow outbound to EDR; manually install sensor pre-analysis, uninstall post.

Component: IRE

RWR-IRE-RP-004Use VMware Live Site Recovery to automate recovery from promoted (staged) image.ManageabilityRecoverability

Decision: Use VMware Live Site Recovery to automate recovery from promoted (staged) image.

Rationale: Automated runbook for business critical workload recovery.

Implication: No significant trade-offs identified for this decision.

Component: IRE

VLR-SRM-RP-001Do NOT run test recovery of recovery plans.Recoverability

Decision: Do NOT run test recovery of recovery plans.

Rationale: Changes made to remove malware are not persistent after test run.

Implication: No significant trade-offs identified for this decision.

Component: Site Recovery Manager

RWR-DNS-NET-001Configure DNS settings for each protected workload to use DNS servers across protected and recovery Recoverability

Decision: Configure DNS settings for each protected workload to use DNS servers across protected and recovery instances.

Rationale: Enables DNS resolution when recovered test VMs powered on in IRE or during planned migration/DR.

Implication: No significant trade-offs identified for this decision.

Component: DNS

RWR-NTP-NET-001Configure NTP settings for each protected workload to use NTP servers across both instances.Recoverability

Decision: Configure NTP settings for each protected workload to use NTP servers across both instances.

Rationale: Enables NTP when recovered test VMs powered on in IRE or during migration/DR.

Implication: No significant trade-offs identified for this decision.

Component: NTP

Prerequisites

  • VCF management domain operational
  • DNS/NTP configured
  • CA infrastructure in place

Implementation Procedure

Implementation

Environment: VCF version in support matrix; configured per guidance; capture all Ransomware Recovery tab parameters; VCF platforms healthy.

DNS: Required DNS entries in forward and reverse zones

Network: Protected <-> Recovery connectivity with jumbo frames, Layer 3 routing, max 150ms latency; bandwidth sized with VLR Calculator

Software: VLR .iso image mounted on vSphere Client machine

License: VCF license per design (protected + recovery), VMware vDefend Firewall license for IRE isolation
Active Directory: Domain Controllers available

CA: Microsoft CA available

  • Workloads: Windows or Linux OS with latest updates + VMware Tools
  • Deploy VLR Appliance
  • MethodDeploy OVF template in vSphere Client
  • PlacementDefault management cluster in management domain vCenter
  • StorageManagement domain vSAN datastore
  • NetworkManagement network
  • Repeat ForRecovery instance using Planning Workbook for Recovery
  • Vmdk Files
  • VLR-9.0.3.0.XXXXXXXX_OVF10.ovf
  • -database.vmdk
  • -data.vmdk
  • -support.vmdk
  • -swap.vmdk
  • -system.vmdk
  • Customizations
  • Enable SSHDSelected
  • Host Network IP Address Familyipv4
  • Host Network Modestatic
  • Certificate Replacement

ProcessVLR appliance config UI (https://<vlr_fqdn>:5480) > Certificates

StepsGenerate CSR (Appliance certificate > Generate CSR), submit to Microsoft AD CS, upload signed cert back
Repeat ForRecovery instance

Day-2 Operations Tasks

Operations

As needed

Operational Verification

As needed

EndpointProtected vCenter > VMware Live Site Recovery

As needed

Password Management

As needed

Expiration Policy

As needed

MethodSSH to VLR as admin, switch to root

As needed

Repeat Foradmin user and recovery instance

As needed

Commands

As needed

chage --maxdays <value> root

As needed

chage --mindays <value> root

As needed

Monitoring Points

  • Verify Pairing
  • CheckStatus of VLSR and vSphere Replication = OK (green checkmark)
  • PreVerify all replications paused (VMware Live Site Recovery > site pair > View Details > Replications > Outgoing = Paused)
  • Verify VR and VLCR operational
  • VerifyNSX Manager > Security > Policy Management > Distributed Firewall > Application tab - verify ransomware recovery policies listed
  • Failover Failback Checklist
  • Verify 'Task finished successfully' and 'Created test vm with name <vm_name>-test-vm'
  • Inspect OS/app vulnerabilities on EDR console; monitor events; manually patch + remove malware; behavioral analysis
  • Change isolation levels to iterate (e.g., EXTERNAL_OUTBOUND) and monitor behavior on EDR
  • WindowsDownload from falcon.crowdstrike.com > Host setup and management > Sensor downloads > Windows. Install FalconSensor_Windows.exe, enter Customer

Troubleshooting

TroubleshootingEnsure network connectivity between vCenter instances and VLR appliance
Cause:
Fix:
Backupcp -p /etc/security/faillock.conf /etc/security/faillock.conf-`date +%F_%H:%M:%S`.back
Cause:
Fix:
sed -i -E 's/deny = [-]?[0-9]+/deny = <value>/g' /etc/security/faillock.conf
Cause:
Fix:
sed -i -E 's/root_unlock_time = [-]?[0-9]+/root_unlock_time = <value>/g' /etc/security/faillock.conf
Cause:
Fix:
sed -i -E 's/unlock_time = [-]?[0-9]+/unlock_time = <value>/g' /etc/security/faillock.conf
Cause:
Fix:

Likely Panelist Questions

Q: Why did you choose this architecture?

See design decisions for rationale

Failure Scenarios

Walk through failover (Disaster recovery) vs failback (Planned migration) recovery types
Impact:
Mitigation:

Trade-off Analysis

Why DISABLE guest OS quiescing - performance impact on high-I/O workloads (trade crash-consistent for app-consistent)

Chosen:

Justification:

Trade-Offs Analysis

Chosen:

Justification:

Service account with Admin role = simplicity vs. least-privilege principle

Chosen:

Justification:

Not running test recovery plans = accept recovery plan execution risk vs. persistent changes preservation

Chosen:

Justification:

Quiz — Ransomware Recovery

0/15
Q1
What is the recommended RPO for business critical workloads in this solution?
  • 1 hour
  • 4 hours
  • 24 hours
  • 72 hours
24-hour RPO is configured on business critical workload policies in VR. Detection lag for ransomware can be weeks/months - 200 PIT instances * 24h = 200 days rollback history (RWR-VLR-VR-CFG-004).
Q2
Which network isolation level permits EDR sensor communication but blocks general Internet access?
  • Isolated
  • Quarantined
  • Quarantined+Analysis
  • Open
Quarantined+Analysis allows DHCP/DNS/NTP + EDR (CrowdStrike/Carbon Black) cloud portal access; denies all other external, inbound, and east-west traffic.
Q3
What is the maximum supported latency between protected and recovery VCF instances?
  • 50 ms
  • 100 ms
  • 150 ms
  • 300 ms
Maximum supported latency between instances is 150 ms for vSphere Replication to work.
Q4
What is the required TCP port for Enhanced vSphere Replication host-to-host communication?
  • 443
  • 902
  • 32032
  • 8443
Enhanced VR goes directly host-to-host at target site, requiring TCP 32032 from protected ESX hosts to recovery vSAN storage cluster ESX hosts and VLR appliance (RWR-VLR-CFG-002).
Q5
What storage architecture is recommended for the IRE vSAN storage cluster?
  • vSAN OSA (Original)
  • vSAN ESA (Express Storage Architecture)
  • VMFS
  • NFS
vSAN ESA (Express Storage Architecture) leverages NVMe TLC flash devices for improved performance, efficiency, and scalability - required for snapshot-heavy ransomware recovery workloads (RWR-IRE-CFG-001).
Q6
Why should the Python script ransomware-recovery.py be used instead of the UI to select a replica for recovery?
  • UI is deprecated
  • There is no option in the UI to select a required replica
  • Scripts are faster
  • UI doesn't support NSX
There is no option in the VLR UI to select a specific replica from history. Ransomware recovery requires iteratively locating an uncompromised snapshot (RWR-IRE-RP-001).
Q7
What is the PIT instance count configured for business critical workload replication?
  • 50
  • 100
  • 200
  • 500
200 PIT instances retained, providing up to 200 days of point-in-time history. Recommended vSAN ESA datastore for values over 24 (RWR-VLR-VR-CFG-003).
Q8
Why is guest OS quiescing disabled in the replication policy?
  • It is insecure
  • Not all business critical workloads support it, and it may cause an outage on high-I/O VMs
  • It requires additional licensing
  • It is incompatible with vSAN
Quiescing improves reliability but impacts performance and may cause outages on high-I/O workloads. Replicas are crash-consistent rather than application-consistent (RWR-VLR-VR-CFG-001).
Q9
What is the sizing of the VMware Live Recovery appliance?
  • 4 vCPU / 16 GB / 400 GB
  • 8 vCPU / 24 GB / 826 GB
  • 12 vCPU / 32 GB / 1 TB
  • 16 vCPU / 64 GB / 2 TB
VLR appliance: 8 vCPUs, 24 GB memory, 826 GB disk capacity.
Q10
Why should test recovery plans NOT be run in this design?
  • They are not supported by VLR
  • Changes made to remove malware are not persistent after test run
  • They corrupt replicas
  • Licensing restrictions
Test recovery creates temporary VMs - any changes made during malware remediation are discarded on cleanup. Production recovery (not test) must be used (RWR-VLR-SRM-RP-001).
Q11
Which entropy rate threshold indicates likely encryption?
  • > 0.3
  • > 0.5
  • > 0.7
  • > 0.9
Entropy rate > 0.7 combined with abnormal prior CPU usage indicates likely ransomware encryption. Entropy = 1 / compression ratio; value close to 1 = encrypted.
Q12
What role must the VLR service account be a member of in vCenter SSO?
  • Read-Only
  • No Access
  • Administrators
  • Custom
Service account must be member of vCenter SSO Administrators group - VLSR needs permissions for DR orchestration + site pairing, VR needs permissions for site-to-site replication (RWR-VLR-SEC-001).
Q13
What recovery type should be selected in VLR for the initial failover in ransomware recovery?
  • Planned migration
  • Disaster recovery
  • Test recovery
  • Reprotect
Disaster recovery - because protected VMs under inspection are paused during ransomware recovery. Planned migration is used for failback after reprotect.
Q14
Where are the NSX isolation scripts located in the GitHub repository?
  • vcf-data-protection/examples/
  • vcf-data-protection/examples/rwr-ref-arch/
  • vcf-data-protection/examples/rwr-ref-arch/nsxtRwrIsolationPolicies
  • vcf-data-protection/scripts/
NSX isolation scripts (nsxtRwrIsolationPolicies.py, constants.py, nsxServiceClient.py) are in vcf-data-protection/examples/rwr-ref-arch/nsxtRwrIsolationPolicies directory.
Q15
What VCF Ops data retention period is configured for RWR dashboards?
  • 30 days
  • 90 days
  • 200 days
  • 365 days
Global Settings > Data Retention > Object History is set to 200 days to match the 200 PIT instance history for entropy trend analysis.

Flashcards — Ransomware Recovery

Card 1 of 15
What does VMware Live Recovery (VLR) contain?
VLR appliance = VMware Live Cyber Recovery + VMware Live Site Recovery + vSphere Replication + vSAN Data Protection services.

Labs

Lab 1: Deploy and Pair VLR Across Protected and Recovery Instances

Deploy VMware Live Recovery appliances in both the protected and recovery VCF instances, register with vCenter SSO, replace certificates with CA-signed, and create a site pair.

Starting State: Two VCF 9.0 instances (protected + recovery) with management domains operational, DNS/NTP/CA in place, VLR .iso mounted, Planning and Preparation Workbook complete.

Lab 2: Configure IRE Network Isolation Levels via Python Script

Create Tier-1 gateway + IRE segment with DHCP in the isolated workload domain, deploy 8 network isolation levels via the Python script, and assign QUARANTINED_ANALYSIS policy to a test VM.

Starting State: Recovery VCF instance with isolated workload domain deployed (compute cluster + vSAN ESA storage cluster). vDefend Firewall license applied. Windows Jump Box with Python 3 and vcf-data-protection repo cloned.

Lab 3: Full Ransomware Recovery Workflow - Recover a Workload from a Clean Snapshot

Simulate a ransomware attack (using entropy analysis), create test VM from clean replica on IRE, install/uninstall EDR sensor, promote image, fail over via VLCR, then reprotect and fail back.

Starting State: Labs 1 and 2 complete. Protected workload replicating to recovery with RPO=24h + 200 PIT. VCF Ops with RWR management pack installed; dashboards imported. CrowdStrike Falcon or Carbon Black portal access.

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.