This lab targets VCF 9.0

Alert Automation & Closed-Loop Remediation

VCF 9.0Intermediatevcp-foundation⏱ 75 min

Objectives

  • Configure multi-action alerts with REST/SMTP/SNMP notifications and integration with ticketing system.

Prerequisites

VCF lab environment deployed and operational

Lab Environment

Standard VCF lab environment for Cloud Operations 8.x Professional

Tasks

Task 1 Alert Automation & Closed-Loop Remediation

Configure multi-action alerts with REST/SMTP/SNMP notifications and integration with ticketing system.

Step 1

Define Symptom: Name: "Host Memory Utilization Critical"Condition: (memory|usage_percent > 90) AND (memory|swapused_kb > 1000000) for 10 minutesSeverity: CriticalWait cycle: 1 (fire on first evaluation)Cancel cycle: 2 (resolve after 2 cycles without trigger)Scope: ESXi Host resource kind

Step 2

Create Alert Definition: Name: "Host Memory Critical Alert"Symptom: "Host Memory Utilization Critical"Active: YesNotification preference: Once per state change (fire once, resolve once)

Step 3

Configure Actions:

Step 4

Test Alert (Inject Memory Pressure):

Step 5

Validation & Closed-Loop:

Validation Gate

Check: Verify lab completion

Expected: Lab exercise completed successfully

Common Errors

Alert automation executing remediation without rate limiting
Fix: A monitoring alert that triggers auto-remediation can cascade: first alert restarts service → service restart triggers second alert → second restart disrupts ongoing recovery. Implement: cooldown period (no re-trigger within 30 minutes), maximum execution count per hour, and alert deduplication.
Closed-loop remediation without rollback capability
Fix: Auto-remediation that changes configuration (restart, scale, reconfigure) without rollback capability can make the situation worse. Implement: pre-action snapshot/backup, execution logging, and automatic rollback if the remediation doesn't resolve the alert within a configured timeframe.
Not distinguishing between symptom alerts and root cause alerts
Fix: Multiple symptom alerts (high CPU, high latency, high queue depth) may share one root cause (storage controller failure). Remediating each symptom independently wastes effort. Use VCF Operations alert correlation to group related symptoms and remediate the root cause.

Final Validation

Lab completed successfully

✓ All steps completed → No errors observed

Cleanup / Restore

• Revert to snapshot if needed

Design Reflection (VCDX)

Alert automation with closed-loop remediation demonstrates AIOps capabilities in your design. VCDX panelists test whether auto-remediation is safe, controlled, and reversible.

⚠ Known Pitfalls (from Community KB)

Auto-remediation without cooldown — cascading restarts make outages worse.
Treating symptoms as independent root causes — wastes remediation effort and delays actual fix.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.