Academy/VCAP — VCF Operations (3V0-22.25)/Lab: Design Production Alert Policy with Escalation
This lab targets VCF 9.0

Lab: Design Production Alert Policy with Escalation

VCF 9.0Advancedvcap-advanced⏱ 120 min

VCF Operations alerting — symptom definitions, dynamic thresholds, alert correlation, notification chains, automated remediation

Objectives

  • Create symptom definitions with both static and dynamic thresholds
  • Build multi-symptom alert definitions with AND/OR correlation logic
  • Configure tiered notification chains with escalation timeouts
  • Implement automated remediation actions with safety controls
  • Design maintenance windows for planned change suppression
  • Understand the full alert lifecycle: fire → acknowledge → escalate → remediate → clear

Prerequisites

VCF Operations with 14+ days of metric data (required for dynamic thresholds learning), SMTP configured for email notifications, webhook endpoint available (ServiceNow, Slack, or test endpoint)

Prior labs: vcap-ops-01, vcap-ops-02

Required skills:

  • Alert design principles
  • SMTP/webhook integration concepts
  • Change management procedures

Lab Environment

VCF Operations cluster with vCenter adapter. At least one cluster with VMs that can be stress-tested to trigger alerts (stress-ng or similar). Email relay configured. Optional: ServiceNow/Slack webhook for escalation testing.

Tasks

Task 1 Production Alert Policy with Multi-Level Escalation

Production alerting follows the 'signal-to-noise' principle: every alert should be actionable. Multi-symptom correlation (CPU high AND Memory high AND Disk latency high) eliminates single-metric false positives. Dynamic thresholds adapt to each object's baseline, avoiding generic thresholds that alert on a busy-but-normal VM. Escalation ensures critical alerts don't sit unacknowledged. Automated remediation reduces MTTR for known scenarios.

Design a complete alerting framework that mirrors production operations: symptom definitions detect individual conditions, alert definitions correlate multiple symptoms to reduce false positives, notification chains escalate unacknowledged alerts, and automated remediation handles well-understood scenarios — all with safety controls to prevent runaway automation.

Step 1
Create symptom definitions — static thresholds. Configure → Alert Settings → Symptom Definitions → Add. (a) 'High-CPU-Static': Metric=cpu.usage, Condition=> 90%, Wait Cycles=3, Cancel Cycles=3. The wait cycles mean 3 consecutive violations (15 minutes at 5-min intervals) before the symptom fires — this eliminates brief spikes. (b) 'High-Memory-Static': Metric=mem.usage, Condition=> 90%, Wait Cycles=3. (c) 'High-Disk-Latency': Metric=disk.maxTotalLatency, Condition=> 30ms, Wait Cycles=2 (disk latency issues are more urgent).
Step 2
Create symptom definitions — dynamic thresholds. Add → (d) 'CPU-Anomaly-DT': Metric=cpu.usage, Threshold Type=Dynamic, Sensitivity=3 (medium). Dynamic thresholds learn each object's normal behavior over 14+ days and alert when the metric deviates significantly from its learned baseline. A VM that normally runs at 80% CPU won't alert at 82%, but a VM that normally runs at 10% will alert at 50%. (e) 'Memory-Anomaly-DT': Metric=mem.usage, Threshold Type=Dynamic, Sensitivity=3.
Step 3
Create correlated alert definition. Alert Settings → Alert Definitions → Add. Name='VM-Performance-Degradation'. Impact=Health. Criticality=Critical. Symptoms: Add all three static symptoms with AND logic: 'High-CPU-Static AND High-Memory-Static AND High-Disk-Latency'. This means the alert only fires when ALL THREE conditions are true simultaneously — dramatically reducing false positives compared to alerting on each condition independently. Object Type=VirtualMachine.
Step 4
Create anomaly alert definition. Add → Name='VM-Behavioral-Anomaly'. Impact=Risk. Criticality=Warning. Symptoms: 'CPU-Anomaly-DT OR Memory-Anomaly-DT'. Use OR logic here because any behavioral anomaly warrants investigation. This alert catches performance changes that static thresholds miss — a VM that suddenly doubles its CPU usage from 20% to 40% won't trigger the 90% static threshold but WILL trigger the dynamic threshold.
Step 5
Configure notification chain — tiered escalation. Configure → Outbound Settings → Add Plugin: (a) SMTP: server=mail.lab.local, port=25. (b) Webhook: URL=https://slack.webhook.example/alert (or ServiceNow endpoint). Then: Notification Rules → Add: Rule 1: Alert='VM-Performance-Degradation', Severity=Critical → Send email immediately to ops-team@lab.local. Rule 2: Same alert, if not acknowledged in 15 minutes → Send webhook to Slack/ServiceNow (escalation). Rule 3: Same alert, if not acknowledged in 30 minutes → Send email to manager@lab.local (management escalation).
Step 6
Configure automated remediation. Alert Settings → edit 'VM-Performance-Degradation' → Recommendations → Add Action. For CPU contention: Action='vMotion VM to host with lowest CPU usage' (requires VCF Operations + vSphere integration). Safety controls: Max executions per alert=1 (prevent vMotion loops), Cooldown=30 minutes, Mode=Automatic (or 'Recommend' for approval-required). This automates the most common response to performance degradation — move the VM to a less-loaded host.
Step 7
Configure maintenance window. Configure → Maintenance Schedules → Add. Name='Monthly-Patch-Window', Schedule=Second Saturday 22:00-06:00 UTC, Scope=select clusters being patched. During the maintenance window: all alerts for in-scope objects are suppressed (not generated), notifications are paused, automated remediation is suspended. This prevents alert storms during planned maintenance. Verify: run a stress test during a test maintenance window — no alerts should fire.
Step 8
Test the alert pipeline — trigger a real alert. SSH to a test VM and run: stress-ng --cpu 4 --vm 2 --vm-bytes 1G --timeout 1200s (stresses CPU and memory for 20 minutes). Monitor: (a) Symptom Definitions page → 'High-CPU-Static' should show 'Active' after 15 minutes (3 × 5-min cycles); (b) If memory also exceeds 90% AND disk latency spikes, the correlated alert 'VM-Performance-Degradation' fires; (c) Email notification arrives within 60 seconds of alert; (d) If unacknowledged for 15 minutes, Slack/ServiceNow webhook fires; (e) Automated remediation (vMotion) executes if configured in Automatic mode.
Step 9
Acknowledge and close the alert. In VCF Operations → Alerts → select 'VM-Performance-Degradation' alert → Acknowledge (assigns to current user, stops escalation). Add notes: 'Triggered by stress test — expected behavior.' Stop the stress-ng process on the VM. Monitor: symptoms should clear after 3 cancel cycles (15 minutes). Alert auto-closes when all symptoms clear. If symptoms don't clear: manually cancel the alert and investigate why metrics remain elevated.
Step 10
Review alert audit trail. Alerts → select the alert → Timeline tab. Shows the complete lifecycle: (a) Symptoms triggered (timestamp for each), (b) Alert fired, (c) Email notification sent, (d) Slack escalation at +15min, (e) Acknowledged by admin, (f) Remediation action executed (if applicable), (g) Symptoms cleared, (h) Alert auto-closed. This audit trail is critical for post-incident review and compliance documentation. Export the timeline for incident reports.

Validation Gate

Check: Complete alert pipeline operational: symptom → alert → notification → escalation → remediation → clear

Expected: Static and dynamic symptom definitions active, correlated alert definition firing correctly, email notification within 60 seconds, webhook escalation at 15 minutes, automated remediation executing with safety controls, maintenance window suppressing alerts during planned work

Common Errors

Dynamic threshold alerts fire constantly (too sensitive)
Fix: DT sensitivity is set too low (1-2). Increase to 3-4 for production workloads. Also verify the object has 14+ days of metric data — DT with insufficient historical data produces erratic baselines.
Correlated alert never fires despite individual symptoms being active
Fix: AND logic requires ALL symptoms active simultaneously. If CPU spikes resolve before disk latency spikes, the correlation window is missed. Consider: (a) increase wait cycles to widen the correlation window, (b) use OR logic if any single condition warrants investigation.
Email notification not received
Fix: Check: (1) SMTP plugin configured and tested (Outbound Settings → Test); (2) notification rule matches the correct alert definition and criticality; (3) email server not rejecting due to SPF/DKIM (check SMTP server logs); (4) notification rule is enabled (not disabled during testing).
Automated remediation causes vMotion loop
Fix: The target host is also at high utilization, causing the VM to be migrated back. Fix: (1) set max executions per alert to 1; (2) increase cooldown to 30+ minutes; (3) use 'Recommend' mode instead of 'Automatic' for complex remediations; (4) add a symptom that checks destination host capacity before initiating vMotion.

Final Validation

Production alerting framework operational with multi-level escalation and automated remediation

✓ Static symptoms → Fire after wait cycles met, clear after cancel cycles

✓ Dynamic thresholds → Detect behavioral anomalies relative to learned baselines

✓ Correlated alerts → Fire only when all AND symptoms are simultaneously active

✓ Email notification → Delivered within 60 seconds of alert firing

✓ Webhook escalation → Fires at +15 minutes if alert not acknowledged

✓ Automated remediation → Executes within safety controls (max 1, cooldown 30 min)

✓ Maintenance window → Suppresses alerts during scheduled period

Cleanup / Restore

• Stop stress-ng on test VM

• Cancel test alerts

• Disable automated remediation actions (switch to Recommend mode)

• Delete test maintenance window

• Optionally delete alert/symptom definitions if not needed

Design Reflection (VCDX)

Alert design is a key VCDX operational architecture topic. Demonstrate: (1) signal-to-noise ratio — multi-symptom correlation eliminates false positives; (2) dynamic thresholds adapt to workload profiles — no single static threshold fits all VMs; (3) escalation ensures nothing falls through the cracks; (4) automated remediation reduces MTTR while safety controls prevent runaway automation. In defense, discuss how you'd scale this from 100 to 10,000 VMs — what changes?

Requirements

  • Zero unacknowledged critical alerts beyond 30 minutes
  • False positive rate < 5%
  • MTTR < 15 minutes for known performance scenarios

Constraints

  • Dynamic thresholds require 14+ days of learning data
  • SMTP/webhook must be reliable — notification failure is silent
  • Automated remediation limited to safe actions (vMotion, rightsizing — not reboot, delete)

Assumptions

  • Operations team monitors email/Slack during business hours
  • After-hours alerts escalate to on-call rotation
  • vMotion is safe for all VMs (no license-locked or physical-device-dependent VMs)

Risks

  • Alert fatigue from too many low-criticality alerts — team ignores all alerts
  • DT learning phase produces false positives during first 14 days
  • Remediation action makes things worse (vMotion to overloaded host)

Self-Assessment Discussion Prompts

  1. How do you measure and improve your alert signal-to-noise ratio over time?
  2. What is the appropriate dynamic threshold sensitivity for a mixed-workload environment?
  3. How do you handle alerting for VMs that legitimately run at 95%+ CPU (e.g., HPC workloads)?
  4. What governance process controls who can enable automated remediation?

Extensions

Integrate alert notifications with PagerDuty for on-call rotation management

Build a ServiceNow ITSM integration that auto-creates incident tickets from critical alerts

Create a 'noise report' dashboard showing alert frequency per definition — identify and tune noisy alerts

Implement a phased DT rollout: enable at Sensitivity=5 (low) for 2 weeks, then tighten to 3

⚠ Known Pitfalls (from Community KB)

Creating alerts on single metrics without correlation — generates 10x more false positives than multi-symptom alerts
Setting wait cycles to 1 — every 5-minute spike triggers an alert; use 3+ for production
Enabling automated remediation in 'Automatic' mode without testing in 'Recommend' mode first
Not configuring maintenance windows — planned patches trigger alert storms that erode team trust in monitoring

References

  • VCF Operations Administration Guide — Alerts and Symptoms: techdocs.broadcom.com
  • VCF Operations Outbound Plugins Guide — Email, Webhook, SNMP: techdocs.broadcom.com
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.