Lab: Create Custom Dashboard and Configure Alerts in VCF Operations
Objectives
- Create a custom VCF Operations dashboard with multiple widget types
- Configure widget-to-widget interactions for drill-down analysis
- Define a custom symptom and alert definition
- Configure alert notification (email or webhook)
- Understand VCF Operations policies and compliance dashboards
Prerequisites
VCF Operations deployed and collecting vCenter data for at least 15 minutes (Lab 07 completed). VCF Operations for Logs deployed (Lab 08 completed).
Prior labs: vvf-admin-07, vvf-admin-08
Required skills:
- VCF Operations UI navigation
- Understanding of metrics vs properties
Lab Environment
Holodeck VVF pod with VCF Operations and VCF Operations for Logs running.
Credentials
| System | Username | Password |
|---|---|---|
| VCF Operations | admin | Set during Lab 07 |
Tasks
Task 1 Create Custom Dashboard with Interactive Widgets
Custom dashboards are the primary exam topic for Obj 4.3. Understanding widget types, interactions, and sharing is critical.
VCF Operations UI → Dashboards → Create Dashboard. Name: 'VVF Cluster Health Overview'.
Add Widget 1 — Scoreboard: Drag 'Scoreboard' widget → configure: Object Type=Cluster Compute Resource, Metric=CPU Usage (%). This shows cluster CPU utilization as a large number.
Add Widget 2 — Heatmap: Drag 'Heatmap' widget → configure: Object Type=Host System, Color by=CPU Usage (%), Size by=Memory Usage (%). This shows hosts as colored cells.
Add Widget 3 — Top-N: Drag 'Top-N' widget → configure: Object Type=Virtual Machine, Metric=CPU Usage (%), Count=10. Shows top 10 VMs by CPU.
Add Widget 4 — Metric Chart: Drag 'Metric Chart' → configure: Object=Cluster, Metrics=[CPU Usage %, Memory Usage %, Disk Latency]. Time range=Last 1 hour.
Configure Widget Interaction: Click heatmap widget settings → Widget Interactions → set 'Selected Object' output → link to Top-N widget input. Now clicking a host in heatmap filters Top-N to show only that host's VMs.
Save dashboard. Navigate to Dashboard List → verify 'VVF Cluster Health Overview' appears.
Validation Gate
Check: Custom dashboard loads within 5 seconds, all widgets render data, and time range filter applies correctly
Expected: Dashboard renders 6-8 widgets with current data. Switching time range (4h → 24h → 7d) updates all widgets. No 'No Data' widgets.
Common Errors
Task 2 Define Custom Symptom and Alert
Alerting is the proactive monitoring layer — symptoms detect conditions, alerts combine them with logic, notifications ensure humans are informed. Core exam topic.
Navigate to Alerts → Alert Settings → Symptom Definitions → Add.
Create Symptom: Name='High Host CPU', Base Object Type='Host System', Metric='CPU Usage (%)', Condition='Above', Threshold=85, Wait Cycles=3 (15 minutes at 5-min collection).
Navigate to Alert Definitions → Add. Name='Host CPU Critical', Object Type='Host System', Impact=Health, Criticality=Critical.
Add Symptom: Select 'High Host CPU' symptom → Add. Set operator to 'All' (AND logic). Add Recommendation: 'Investigate running VMs and consider DRS migration or adding host capacity.'
Save Alert Definition.
Configure Notification: Alerts → Alert Settings → Notification Settings → Add Outbound Setting. For lab: use 'Log File' plugin (writes alerts to local log). In production: configure SMTP for email notifications.
Validation Gate
Check: Custom symptom detects condition correctly, alert fires with correct severity, and notification reaches the configured endpoint
Expected: Symptom triggers after configured wait cycles. Alert shows in VCF Operations with correct severity. Notification email/webhook is received at the configured endpoint.
Common Errors
Task 3 Explore Policies and Compliance Dashboard
Policies define operational thresholds; compliance dashboards map to regulatory frameworks. Understanding both is required for Obj 4.3 sub-objectives on policies and security hardening.
Navigate to Administration → Policies. Review the default policy: 'vSphere Solution's Default Policy'. Note the capacity, alerting, and compliance settings.
Clone the default policy: Select → Clone → Name='VVF Lab Custom Policy'. Edit: Change capacity remaining threshold from 30 days to 14 days (more aggressive alerting for lab).
Apply custom policy to management cluster: Select cluster → Apply Policy → choose 'VVF Lab Custom Policy'.
Navigate to Compliance → Overview. If CIS or DISA STIG benchmarks are enabled, review compliance status per host.
Review a non-compliant host finding (if any). Note the remediation recommendation — typically an esxcli or PowerCLI command.
Validation Gate
Check: Compliance dashboard shows current posture, identifies non-compliant hosts/VMs, and provides remediation guidance
Expected: Dashboard shows compliance percentage per policy. Non-compliant items are listed with specific violations. Remediation guidance is available for each violation.
Common Errors
Final Validation
Custom dashboard created with interactive widgets, alert definition configured with symptom, custom policy applied to cluster.
✓ Dashboard 'VVF Cluster Health Overview' loads with all 4 widgets populated → Scoreboard, heatmap, Top-N, metric chart all show data
✓ Alert definition 'Host CPU Critical' exists with symptom attached → Symptom threshold 85%, wait cycles 3, criticality Critical
✓ Custom policy applied to management cluster → Capacity threshold set to 14 days
Cleanup / Restore
Snapshot: post-vcf-ops-dashboard-alerts
• Take snapshot 'post-vcf-ops-dashboard-alerts' before Supervisor lab
Design Reflection (VCDX)
Operations monitoring design is part of VCDX defense. Panelists expect custom dashboards for capacity planning, compliance monitoring, and proactive alerting — not just default views.
Requirements
- Design persona-specific dashboards for operational monitoring
- Configure custom symptoms/alerts for organization-specific SLAs
- Implement compliance monitoring with drift detection
Constraints
- Dashboard widgets should be 6-8 per view for readability
- Symptoms need appropriate wait cycles to avoid false positives
- Compliance policies must map to regulatory framework
Assumptions
- VCF Operations notification plugins are configured (SMTP, webhook)
- Organization has defined compliance requirements
Risks
- Alert fatigue from symptoms without wait cycles
- Compliance drift going undetected between manual audits
⚠ Known Pitfalls (from Community KB)
References
- VCF Operations Dashboard GuideTier 1 — Official
- VCF Operations Alerting GuideTier 1 — Official
- VMware Security Hardening GuideTier 1 — Official