Lab: Design Production Alert Policy with Escalation
Objectives
- Create symptom definitions with both static and dynamic thresholds
- Build multi-symptom alert definitions with AND/OR correlation logic
- Configure tiered notification chains with escalation timeouts
- Implement automated remediation actions with safety controls
- Design maintenance windows for planned change suppression
- Understand the full alert lifecycle: fire → acknowledge → escalate → remediate → clear
Prerequisites
VCF Operations with 14+ days of metric data (required for dynamic thresholds learning), SMTP configured for email notifications, webhook endpoint available (ServiceNow, Slack, or test endpoint)
Prior labs: vcap-ops-01, vcap-ops-02
Required skills:
- Alert design principles
- SMTP/webhook integration concepts
- Change management procedures
Lab Environment
VCF Operations cluster with vCenter adapter. At least one cluster with VMs that can be stress-tested to trigger alerts (stress-ng or similar). Email relay configured. Optional: ServiceNow/Slack webhook for escalation testing.
Tasks
Task 1 Production Alert Policy with Multi-Level Escalation
Design a complete alerting framework that mirrors production operations: symptom definitions detect individual conditions, alert definitions correlate multiple symptoms to reduce false positives, notification chains escalate unacknowledged alerts, and automated remediation handles well-understood scenarios — all with safety controls to prevent runaway automation.
Create symptom definitions — static thresholds. Configure → Alert Settings → Symptom Definitions → Add. (a) 'High-CPU-Static': Metric=cpu.usage, Condition=> 90%, Wait Cycles=3, Cancel Cycles=3. The wait cycles mean 3 consecutive violations (15 minutes at 5-min intervals) before the symptom fires — this eliminates brief spikes. (b) 'High-Memory-Static': Metric=mem.usage, Condition=> 90%, Wait Cycles=3. (c) 'High-Disk-Latency': Metric=disk.maxTotalLatency, Condition=> 30ms, Wait Cycles=2 (disk latency issues are more urgent).
Create symptom definitions — dynamic thresholds. Add → (d) 'CPU-Anomaly-DT': Metric=cpu.usage, Threshold Type=Dynamic, Sensitivity=3 (medium). Dynamic thresholds learn each object's normal behavior over 14+ days and alert when the metric deviates significantly from its learned baseline. A VM that normally runs at 80% CPU won't alert at 82%, but a VM that normally runs at 10% will alert at 50%. (e) 'Memory-Anomaly-DT': Metric=mem.usage, Threshold Type=Dynamic, Sensitivity=3.
Create correlated alert definition. Alert Settings → Alert Definitions → Add. Name='VM-Performance-Degradation'. Impact=Health. Criticality=Critical. Symptoms: Add all three static symptoms with AND logic: 'High-CPU-Static AND High-Memory-Static AND High-Disk-Latency'. This means the alert only fires when ALL THREE conditions are true simultaneously — dramatically reducing false positives compared to alerting on each condition independently. Object Type=VirtualMachine.
Create anomaly alert definition. Add → Name='VM-Behavioral-Anomaly'. Impact=Risk. Criticality=Warning. Symptoms: 'CPU-Anomaly-DT OR Memory-Anomaly-DT'. Use OR logic here because any behavioral anomaly warrants investigation. This alert catches performance changes that static thresholds miss — a VM that suddenly doubles its CPU usage from 20% to 40% won't trigger the 90% static threshold but WILL trigger the dynamic threshold.
Configure notification chain — tiered escalation. Configure → Outbound Settings → Add Plugin: (a) SMTP: server=mail.lab.local, port=25. (b) Webhook: URL=https://slack.webhook.example/alert (or ServiceNow endpoint). Then: Notification Rules → Add: Rule 1: Alert='VM-Performance-Degradation', Severity=Critical → Send email immediately to ops-team@lab.local. Rule 2: Same alert, if not acknowledged in 15 minutes → Send webhook to Slack/ServiceNow (escalation). Rule 3: Same alert, if not acknowledged in 30 minutes → Send email to manager@lab.local (management escalation).
Configure automated remediation. Alert Settings → edit 'VM-Performance-Degradation' → Recommendations → Add Action. For CPU contention: Action='vMotion VM to host with lowest CPU usage' (requires VCF Operations + vSphere integration). Safety controls: Max executions per alert=1 (prevent vMotion loops), Cooldown=30 minutes, Mode=Automatic (or 'Recommend' for approval-required). This automates the most common response to performance degradation — move the VM to a less-loaded host.
Configure maintenance window. Configure → Maintenance Schedules → Add. Name='Monthly-Patch-Window', Schedule=Second Saturday 22:00-06:00 UTC, Scope=select clusters being patched. During the maintenance window: all alerts for in-scope objects are suppressed (not generated), notifications are paused, automated remediation is suspended. This prevents alert storms during planned maintenance. Verify: run a stress test during a test maintenance window — no alerts should fire.
Test the alert pipeline — trigger a real alert. SSH to a test VM and run: stress-ng --cpu 4 --vm 2 --vm-bytes 1G --timeout 1200s (stresses CPU and memory for 20 minutes). Monitor: (a) Symptom Definitions page → 'High-CPU-Static' should show 'Active' after 15 minutes (3 × 5-min cycles); (b) If memory also exceeds 90% AND disk latency spikes, the correlated alert 'VM-Performance-Degradation' fires; (c) Email notification arrives within 60 seconds of alert; (d) If unacknowledged for 15 minutes, Slack/ServiceNow webhook fires; (e) Automated remediation (vMotion) executes if configured in Automatic mode.
Acknowledge and close the alert. In VCF Operations → Alerts → select 'VM-Performance-Degradation' alert → Acknowledge (assigns to current user, stops escalation). Add notes: 'Triggered by stress test — expected behavior.' Stop the stress-ng process on the VM. Monitor: symptoms should clear after 3 cancel cycles (15 minutes). Alert auto-closes when all symptoms clear. If symptoms don't clear: manually cancel the alert and investigate why metrics remain elevated.
Review alert audit trail. Alerts → select the alert → Timeline tab. Shows the complete lifecycle: (a) Symptoms triggered (timestamp for each), (b) Alert fired, (c) Email notification sent, (d) Slack escalation at +15min, (e) Acknowledged by admin, (f) Remediation action executed (if applicable), (g) Symptoms cleared, (h) Alert auto-closed. This audit trail is critical for post-incident review and compliance documentation. Export the timeline for incident reports.
Validation Gate
Check: Complete alert pipeline operational: symptom → alert → notification → escalation → remediation → clear
Expected: Static and dynamic symptom definitions active, correlated alert definition firing correctly, email notification within 60 seconds, webhook escalation at 15 minutes, automated remediation executing with safety controls, maintenance window suppressing alerts during planned work
Common Errors
Final Validation
Production alerting framework operational with multi-level escalation and automated remediation
✓ Static symptoms → Fire after wait cycles met, clear after cancel cycles
✓ Dynamic thresholds → Detect behavioral anomalies relative to learned baselines
✓ Correlated alerts → Fire only when all AND symptoms are simultaneously active
✓ Email notification → Delivered within 60 seconds of alert firing
✓ Webhook escalation → Fires at +15 minutes if alert not acknowledged
✓ Automated remediation → Executes within safety controls (max 1, cooldown 30 min)
✓ Maintenance window → Suppresses alerts during scheduled period
Cleanup / Restore
• Stop stress-ng on test VM
• Cancel test alerts
• Disable automated remediation actions (switch to Recommend mode)
• Delete test maintenance window
• Optionally delete alert/symptom definitions if not needed
Design Reflection (VCDX)
Alert design is a key VCDX operational architecture topic. Demonstrate: (1) signal-to-noise ratio — multi-symptom correlation eliminates false positives; (2) dynamic thresholds adapt to workload profiles — no single static threshold fits all VMs; (3) escalation ensures nothing falls through the cracks; (4) automated remediation reduces MTTR while safety controls prevent runaway automation. In defense, discuss how you'd scale this from 100 to 10,000 VMs — what changes?
Requirements
- Zero unacknowledged critical alerts beyond 30 minutes
- False positive rate < 5%
- MTTR < 15 minutes for known performance scenarios
Constraints
- Dynamic thresholds require 14+ days of learning data
- SMTP/webhook must be reliable — notification failure is silent
- Automated remediation limited to safe actions (vMotion, rightsizing — not reboot, delete)
Assumptions
- Operations team monitors email/Slack during business hours
- After-hours alerts escalate to on-call rotation
- vMotion is safe for all VMs (no license-locked or physical-device-dependent VMs)
Risks
- Alert fatigue from too many low-criticality alerts — team ignores all alerts
- DT learning phase produces false positives during first 14 days
- Remediation action makes things worse (vMotion to overloaded host)
Self-Assessment Discussion Prompts
- How do you measure and improve your alert signal-to-noise ratio over time?
- What is the appropriate dynamic threshold sensitivity for a mixed-workload environment?
- How do you handle alerting for VMs that legitimately run at 95%+ CPU (e.g., HPC workloads)?
- What governance process controls who can enable automated remediation?
Extensions
Integrate alert notifications with PagerDuty for on-call rotation management
Build a ServiceNow ITSM integration that auto-creates incident tickets from critical alerts
Create a 'noise report' dashboard showing alert frequency per definition — identify and tune noisy alerts
Implement a phased DT rollout: enable at Sensitivity=5 (low) for 2 weeks, then tighten to 3
⚠ Known Pitfalls (from Community KB)
References
- VCF Operations Administration Guide — Alerts and Symptoms: techdocs.broadcom.com
- VCF Operations Outbound Plugins Guide — Email, Webhook, SNMP: techdocs.broadcom.com