Academy/vSphere Foundation 9.0 Administrator (2V0-16.25)/Lab: Create Custom Dashboard and Configure Alerts in VCF Operations
This lab targets VCF 9.0

Lab: Create Custom Dashboard and Configure Alerts in VCF Operations

VCF 9.0Intermediateadmincloud-ops⏱ 60 min

Requires VCF Operations deployed and collecting data for at least 15 minutes (Lab 07 prerequisite)

Objectives

  • Create a custom VCF Operations dashboard with multiple widget types
  • Configure widget-to-widget interactions for drill-down analysis
  • Define a custom symptom and alert definition
  • Configure alert notification (email or webhook)
  • Understand VCF Operations policies and compliance dashboards

Prerequisites

VCF Operations deployed and collecting vCenter data for at least 15 minutes (Lab 07 completed). VCF Operations for Logs deployed (Lab 08 completed).

Prior labs: vvf-admin-07, vvf-admin-08

Required skills:

  • VCF Operations UI navigation
  • Understanding of metrics vs properties

Lab Environment

Holodeck VVF pod with VCF Operations and VCF Operations for Logs running.

Credentials

SystemUsernamePassword
VCF OperationsadminSet during Lab 07

Tasks

Task 1 Create Custom Dashboard with Interactive Widgets

Custom dashboards in VCF Operations transform raw metrics into actionable operational intelligence. The key is designing dashboards for specific personas (capacity planner, ops engineer, security analyst) rather than generic 'show everything' views.

Custom dashboards are the primary exam topic for Obj 4.3. Understanding widget types, interactions, and sharing is critical.

Step 1
VCF Operations UI → Dashboards → Create Dashboard. Name: 'VVF Cluster Health Overview'.
Step 2
Add Widget 1 — Scoreboard: Drag 'Scoreboard' widget → configure: Object Type=Cluster Compute Resource, Metric=CPU Usage (%). This shows cluster CPU utilization as a large number.
Scoreboard displays current CPU usage percentage for the cluster.
Step 3
Add Widget 2 — Heatmap: Drag 'Heatmap' widget → configure: Object Type=Host System, Color by=CPU Usage (%), Size by=Memory Usage (%). This shows hosts as colored cells.
Heatmap shows 3 host cells, colored by CPU utilization.
Step 4
Add Widget 3 — Top-N: Drag 'Top-N' widget → configure: Object Type=Virtual Machine, Metric=CPU Usage (%), Count=10. Shows top 10 VMs by CPU.
Bar chart showing top 10 VMs sorted by CPU usage.
Step 5
Add Widget 4 — Metric Chart: Drag 'Metric Chart' → configure: Object=Cluster, Metrics=[CPU Usage %, Memory Usage %, Disk Latency]. Time range=Last 1 hour.
Line chart with 3 metric series over the last hour.
Step 6
Configure Widget Interaction: Click heatmap widget settings → Widget Interactions → set 'Selected Object' output → link to Top-N widget input. Now clicking a host in heatmap filters Top-N to show only that host's VMs.
Widget interactions are a key exam concept — they enable drill-down analysis from cluster → host → VM level.
Step 7
Save dashboard. Navigate to Dashboard List → verify 'VVF Cluster Health Overview' appears.

Validation Gate

Check: Custom dashboard loads within 5 seconds, all widgets render data, and time range filter applies correctly

Expected: Dashboard renders 6-8 widgets with current data. Switching time range (4h → 24h → 7d) updates all widgets. No 'No Data' widgets.

Common Errors

Creating dashboards with too many widgets causing visual overload
Fix: Effective dashboards have 6-8 widgets maximum focused on one operational area (compute health, storage capacity, network throughput). More widgets reduce readability. Create multiple focused dashboards rather than one catch-all dashboard.
Not setting appropriate metric collection intervals
Fix: VCF Operations collects metrics at different intervals: 5 min (real-time), 5 min rolled to hourly, daily, monthly. Dashboard widgets showing 5-min intervals for a 30-day view generate excessive data points. Match collection interval to the time range: 5-min for troubleshooting (last 4 hours), hourly for daily review, daily for monthly trending.

Task 2 Define Custom Symptom and Alert

Custom symptoms and alerts in VCF Operations extend monitoring beyond default capabilities. Defining symptoms (condition detection) and alerts (notification trigger) enables proactive operations for organization-specific SLAs.

Alerting is the proactive monitoring layer — symptoms detect conditions, alerts combine them with logic, notifications ensure humans are informed. Core exam topic.

Step 1
Navigate to Alerts → Alert Settings → Symptom Definitions → Add.
Step 2

Create Symptom: Name='High Host CPU', Base Object Type='Host System', Metric='CPU Usage (%)', Condition='Above', Threshold=85, Wait Cycles=3 (15 minutes at 5-min collection).

Symptom saved with green checkmark.
Step 3
Navigate to Alert Definitions → Add. Name='Host CPU Critical', Object Type='Host System', Impact=Health, Criticality=Critical.
Step 4
Add Symptom: Select 'High Host CPU' symptom → Add. Set operator to 'All' (AND logic). Add Recommendation: 'Investigate running VMs and consider DRS migration or adding host capacity.'
Step 5

Save Alert Definition.

Alert definition 'Host CPU Critical' appears in alert definitions list.
Step 6
Configure Notification: Alerts → Alert Settings → Notification Settings → Add Outbound Setting. For lab: use 'Log File' plugin (writes alerts to local log). In production: configure SMTP for email notifications.
Production setup: SMTP (email), SNMP (monitoring), REST API (webhook to Slack/PagerDuty/ServiceNow).

Validation Gate

Check: Custom symptom detects condition correctly, alert fires with correct severity, and notification reaches the configured endpoint

Expected: Symptom triggers after configured wait cycles. Alert shows in VCF Operations with correct severity. Notification email/webhook is received at the configured endpoint.

Common Errors

Creating symptoms without appropriate 'wait cycles' causing false positives
Fix: A symptom with zero wait cycles fires on the first metric breach — CPU spikes during VM boot trigger alerts unnecessarily. Configure wait cycles: 3 cycles (15 minutes at 5-min interval) for transient conditions, 1 cycle for critical conditions (storage full, host unreachable).
Not linking alerts to notification plugins
Fix: Alerts without outbound notifications only appear in the VCF Operations UI. Configure notification plugins: email (SMTP), webhook (ServiceNow, PagerDuty), SNMP trap (legacy monitoring). Most production issues go unnoticed because alerts aren't forwarded to the paging system.

Task 3 Explore Policies and Compliance Dashboard

Compliance dashboards in VCF Operations validate that hosts and VMs adhere to security policies. Understanding how policies map to compliance checks helps maintain security posture at scale.

Policies define operational thresholds; compliance dashboards map to regulatory frameworks. Understanding both is required for Obj 4.3 sub-objectives on policies and security hardening.

Step 1
Navigate to Administration → Policies. Review the default policy: 'vSphere Solution's Default Policy'. Note the capacity, alerting, and compliance settings.
Default policy shows threshold settings for CPU, memory, storage capacity remaining.
Step 2
Clone the default policy: Select → Clone → Name='VVF Lab Custom Policy'. Edit: Change capacity remaining threshold from 30 days to 14 days (more aggressive alerting for lab).
Step 3
Apply custom policy to management cluster: Select cluster → Apply Policy → choose 'VVF Lab Custom Policy'.
Step 4
Navigate to Compliance → Overview. If CIS or DISA STIG benchmarks are enabled, review compliance status per host.
Compliance content packs may need to be installed separately. In production, CIS ESXi 8 Benchmark is the most common framework.
Step 5

Review a non-compliant host finding (if any). Note the remediation recommendation — typically an esxcli or PowerCLI command.

Compliance dashboard shows per-host green/yellow/red status. Non-compliant items list specific configuration deviations.

Validation Gate

Check: Compliance dashboard shows current posture, identifies non-compliant hosts/VMs, and provides remediation guidance

Expected: Dashboard shows compliance percentage per policy. Non-compliant items are listed with specific violations. Remediation guidance is available for each violation.

Common Errors

Not customizing compliance policies for the organization's security requirements
Fix: Default compliance policies (VMware Security Hardening Guide) are a baseline. Organizations with specific requirements (HIPAA, PCI-DSS, SOC2) need custom policies that check their specific controls. Create custom policies that map to the organization's compliance framework, not just VMware defaults.
Ignoring compliance drift over time
Fix: Initial compliance is easy — maintaining it is hard. VMs drift out of compliance as configurations change (manual edits, patching side effects, new deployments). Schedule weekly compliance scans and alert on drift. VCF Operations can auto-remediate some compliance items (e.g., disable SSH on ESXi).

Final Validation

Custom dashboard created with interactive widgets, alert definition configured with symptom, custom policy applied to cluster.

✓ Dashboard 'VVF Cluster Health Overview' loads with all 4 widgets populated → Scoreboard, heatmap, Top-N, metric chart all show data

✓ Alert definition 'Host CPU Critical' exists with symptom attached → Symptom threshold 85%, wait cycles 3, criticality Critical

✓ Custom policy applied to management cluster → Capacity threshold set to 14 days

Cleanup / Restore

Snapshot: post-vcf-ops-dashboard-alerts

• Take snapshot 'post-vcf-ops-dashboard-alerts' before Supervisor lab

Design Reflection (VCDX)

Operations monitoring design is part of VCDX defense. Panelists expect custom dashboards for capacity planning, compliance monitoring, and proactive alerting — not just default views.

Requirements

  • Design persona-specific dashboards for operational monitoring
  • Configure custom symptoms/alerts for organization-specific SLAs
  • Implement compliance monitoring with drift detection

Constraints

  • Dashboard widgets should be 6-8 per view for readability
  • Symptoms need appropriate wait cycles to avoid false positives
  • Compliance policies must map to regulatory framework

Assumptions

  • VCF Operations notification plugins are configured (SMTP, webhook)
  • Organization has defined compliance requirements

Risks

  • Alert fatigue from symptoms without wait cycles
  • Compliance drift going undetected between manual audits

⚠ Known Pitfalls (from Community KB)

Building 'show everything' dashboards instead of persona-focused views — operations engineers and capacity planners need different information.
Setting compliance policies and never reviewing them — compliance is a continuous process, not a one-time setup.

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.