Academy/vSphere Foundation 9.0 Administrator (2V0-16.25)/Lab: Configure Custom Alarms & vSAN Health Check
This lab targets VCF 9.0

Lab: Configure Custom Alarms & vSAN Health Check

VCF 9.0Intermediatevcp-foundation⏱ 120 min

Objectives

  • Lab: Configure Custom Alarms & vSAN Health Check

Prerequisites

VCF lab environment deployed and operational

Lab Environment

Standard VCF lab environment for vSphere Foundation 9.0 Administrator

Tasks

Task 1 Lab: Configure Custom Alarms & vSAN Health Check

Custom alarms and vSAN health checks form the operational monitoring foundation. Default alarms are a starting point — production environments need customized thresholds and notification targets specific to the organization's SLAs and on-call procedures.
Step 1
vCenter UI → Alarms → Manage Alarms → Create new
Step 2

Trigger: VM CPU used > 75% for 5min

Step 3

Action: Email to ops@company.com

Step 4

Power on VM, run stress-ng cpu benchmark: stress-ng --cpu 8 --timeout 300s

Step 5

Verify alarm fires (may delay 5-10min)

Step 6
vSAN cluster → Health Check → Run now
Step 7

Review results: Any yellow/red warnings indicate misconfigurations

Step 8

Remediate: If MTU check fails, verify all vSAN port groups set to MTU 9000

Validation Gate

Check: Create a custom CPU alarm for critical VMs (warn at 75%, alert at 85%, notify via email), verify it triggers, and confirm email notification is received

Expected: Alarm triggers at 75% CPU, sends warning email. At 85%, sends critical email. When CPU drops below 75%, alarm auto-resets. Alarm actions log shows successful notification delivery.

Common Errors

Relying solely on default vCenter alarms without customization
Fix: Default alarms use generic thresholds (e.g., CPU 90% = warning, 95% = critical). For latency-sensitive workloads like databases, 75% CPU may be the appropriate warning threshold. Create custom alarm definitions for each workload tier: critical VMs at lower thresholds, dev VMs at higher thresholds.
Not configuring alarm actions (notifications)
Fix: An alarm without an action is invisible. Configure alarm actions: email notification (SMTP settings in vCenter), SNMP trap (to monitoring platform), or script execution (for auto-remediation). Most production incidents are discovered by users, not monitoring — because alarm actions were never configured.
Ignoring vSAN health check warnings
Fix: vSAN health check warnings (yellow) indicate degraded but functional state. Common warnings: 'Controller driver is VMware certified' (driver update needed), 'vSAN disk balance' (>30% imbalance), 'Network: All hosts have a vSAN vmknic configured' (misconfigured host). Address warnings proactively — they escalate to critical during host failures.
Setting alarm frequency too high causing notification flood
Fix: An alarm that triggers every minute for CPU utilization generates 1,440 alerts/day per VM. Configure: frequency (once every 30 min), tolerance (trigger only after condition persists for 5 min), and reset condition (auto-clear when condition resolves). This reduces noise and makes alerts actionable.

Final Validation

Lab completed successfully

✓ All steps completed → No errors observed

Cleanup / Restore

• Revert to snapshot if needed

Design Reflection (VCDX)

Monitoring and alerting design is part of operational architecture in VCDX. Panelists ask: how does your ops team know something is wrong? Custom alarms, notification integration, and vSAN health monitoring are expected components of a complete design.

Requirements

  • Configure custom alarm thresholds per workload tier
  • Set up notification actions (email, SNMP, webhook)
  • Monitor vSAN health checks proactively

Constraints

  • Default alarms are generic — must be customized for SLAs
  • SMTP/SNMP infrastructure must be available for notifications
  • Too many alarms cause alert fatigue

Assumptions

  • Organization has email/paging infrastructure for notifications
  • Operations team reviews and acts on alerts within SLA timeframes

Risks

  • Alert fatigue from uncustomized thresholds flooding notifications
  • Silent failures from alarms without configured actions

⚠ Known Pitfalls (from Community KB)

Presenting a VCDX design with no custom alarm definitions — this signals lack of operational planning.
Setting identical thresholds for all VM types — database VMs need different thresholds than web servers.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.