Academy/VCP-VCF 9.0 Administrator (2V0-17.25)/VCF Operations Manager Dashboard Customization & Capacity Planning
This lab targets VCF 9.0

VCF Operations Manager Dashboard Customization & Capacity Planning

VCF 9.0Intermediateadminarchitectvcdx⏱ 90 min

VCF 9.0 Operations Manager (formerly vRealize Operations in VCF 5.2 / 8.x). Dashboard widgets and capacity models apply to 9.0.0 through 9.0.2.

Objectives

  • Navigate VCF Operations Manager console and understand its role in day-2 operations
  • Design and implement custom dashboards that align with VCDX design intent (capacity, compliance, availability)
  • Configure threshold-based alerting and automated remediation triggers for VCF workload domains
  • Analyze actual utilization vs. design projections to validate cluster sizing assumptions
  • Use VOM reporting to build VCDX design artifacts (capacity headroom proof, compliance posture)

Prerequisites

VCF 9.0 lab environment fully deployed with management domain + 1 workload domain (4 hosts minimum). VOM console must be accessible (https://vom-ip (TCP 443)). At least 7 days of baseline metrics available (or manually seeded test data). DNS resolution configured for vcf.local domain.

Prior labs: holodeck-02 or equivalent VCF deployment

Required skills:

  • VCF architecture fundamentals (management domain, workload domains, SDDC Manager)
  • Basic understanding of capacity planning concepts (CPU oversubscription ratios, memory overhead, storage utilization)
  • Comfort with vSphere Client and VCF management console UIs
  • Ability to interpret performance metrics and troubleshoot resource bottlenecks

Lab Environment

VCF Management Domain (SDDC Manager, vCenter, NSX Manager) + Workload Domain A (4 ESXi hosts, shared vSAN datastore). VOM is deployed as a single instance on the management domain, accessible via FQDN vom.vcf.local or direct IP. Network: 10.0.0.0/20 for management services, 10.1.1.0/24 for ESXi hosts.

Credentials

SystemUsernamePassword
SDDC Manageradministrator@vsphere.localVCF admin password from deployment spec
vCenter Serveradministrator@vsphere.localSame as SDDC Manager
VCF Operations ManageradminVOM admin password from deployment spec (default: same as SDDC Manager)
NSX ManageradminNSX admin password from deployment spec

Tasks

Task 1 Access and Navigate VCF Operations Manager Console

Manageability

Establish familiarity with VOM's role in the VCF operational lifecycle. VOM is the single-pane-of-glass for capacity, compliance, and performance visibility — critical for any VCF design that must demonstrate 'what gets monitored' and 'how it gets managed'. A VCDX panelist will ask: 'Walk me through your day-2 operational dashboards. What metrics are you tracking and why?'

Step 1

Open a web browser and navigate to https://vom.vcf.local (or https://<vom-ip> if DNS is not configured). Accept the self-signed certificate.

VCF Operations Manager login page appears with VMware by Broadcom branding
VOM is a centralized operations appliance; unlike SDDC Manager (which is single-instance but clusterable in 9.0.2+), VOM is typically deployed as a single instance in lab/small production environments.
Step 2

Log in with username 'admin' and the VOM password (use the same password as SDDC Manager administrator@vsphere.local).

VCF Operations Manager dashboard home page loads, showing empty or pre-configured dashboards
VOM requires network connectivity to SDDC Manager, vCenter, and NSX Manager. If the dashboard shows 'Data Source Unavailable', verify the VOM appliance can reach these management IPs (10.0.0.x range) via ping and https.
Step 3

Navigate to Administration > Data Sources. Verify that SDDC Manager, vCenter, and NSX Manager are registered as active data sources.

Three data sources listed: 'SDDC Manager 9.0', 'vCenter (management domain)', 'NSX Manager 9.0', each showing Connection Status = OK
If any data source shows 'Connection Failed', verify credentials in VOM > Administration > Data Sources > [source] > Edit. The password for vCenter and SDDC Manager is 'administrator@vsphere.local'; NSX Manager is 'admin'.
Step 4

Navigate to Inventory. Expand the tree to see: Workload Domains > [Workload Domain A] > Clusters > Cluster A > Hosts.

Inventory tree shows 1 workload domain with 1 cluster and 4 hosts, all marked as 'Online'
VOM's inventory is read-only and auto-synced from SDDC Manager every 5 minutes. If you see stale or missing resources, wait 5 minutes and refresh.
Step 5

Navigate to Dashboards > Built-in Dashboards. Review the pre-configured dashboards: 'Workload Domain Overview', 'Cluster Health', 'Storage Efficiency'.

Three built-in dashboards are visible, each showing current metrics or 'No Data' if metrics are still being collected
The first 24 hours after VOM deployment, dashboards show limited historical data. By 7 days, you'll have a full trend baseline.
Step 6

Click on 'Cluster Health' dashboard and inspect the widgets: Cluster Status, Host Health, vSAN Capacity, CPU/Memory Utilization. Note the time range selector (top right) — ensure it's set to 'Last 7 Days' or 'Last 24 Hours' depending on available data.

Dashboard loads with 4-5 widgets showing current and historical metrics
If widgets show 'No Data', the metrics collection interval may not have completed. VOM collects metrics every 5 minutes; allow at least 1 collection cycle (15 minutes from login) before expecting populated dashboards.

Validation Gate

Check: VOM console is accessible and showing inventory + at least one built-in dashboard with data

Expected: You can navigate to Inventory, see 1 workload domain / 1 cluster / 4 hosts, and view the 'Cluster Health' dashboard with metrics populated

Common Errors

VOM login fails with 'Unauthorized' or 'Invalid credentials'
Cause: Password mismatch between VOM admin user and SDDC Manager administrator@vsphere.local
Fix: VOM's default admin password is often set to match SDDC Manager's during deployment. If you changed the SDDC Manager password post-deployment, VOM admin password must be manually updated in VOM > Administration > Users > admin. Alternatively, check the deployment spec JSON to confirm the admin password used.
📋 KB: VCF 9.0 Operations Manager Admin Guide, Section 3.2
Data Source 'Connection Failed' for vCenter or NSX Manager
Cause: Network connectivity issue or incorrect credentials configured in VOM
Fix: From VOM admin > Data Sources, edit the failed source and re-enter credentials. Verify the vCenter/NSX IP is reachable from VOM appliance: ssh admin@vom-ip, then 'ping <vcenter-ip>'. If no network route, configure a static route on the VOM appliance to reach the management network.
Dashboards show 'No Data' or blank widgets 24+ hours after deployment
Cause: Metrics collection is not running, or data source collection is disabled
Fix: Navigate to Administration > Collection Settings and verify that collection intervals are set to 5 minutes (default). Check VOM appliance system logs: ssh admin@vom-ip, then 'tail -f /var/log/vmware/vcf-ops/collector.log' to verify metrics are being ingested.

Task 2 Design a Custom Dashboard for Capacity Planning & Design Validation

Manageability

In VCDX, you must articulate 'how will I know if my design is working as intended?' This task forces you to identify 3-4 key capacity and availability metrics that directly map to your design assumptions. For example, if you designed the cluster for 4:1 CPU oversubscription, you must monitor actual CPU contention (Ready % or %Ready) and prove it stays below design limits. This is the difference between a theoretical design and a validated deployment.

Step 1

Navigate to Dashboards > Create Dashboard. Enter the dashboard name: 'VCF Design Capacity Baseline' and description: 'Validates sizing assumptions: CPU oversubscription, memory headroom, storage utilization, and DRS/HA availability.'

New dashboard is created in edit mode with an empty canvas
Dashboard names should reflect their business intent, not just technical content. 'VCF Design Capacity Baseline' immediately tells the VCDX panelist what you're validating.
Step 2

Add Widget 1 (CPU Utilization): Click 'Add Widget' > Select 'Line Chart' > Configure as follows:

  • Data Source: vCenter (management domain) > Cluster A
  • Metrics: CPU Utilization (%), CPU Ready (%)
  • Time Range: Last 7 Days
  • Y-axis label: 'CPU %'
  • Title: 'CPU Utilization & Ready (7-day baseline)'

Thresholds: Green (0-60%), Yellow (60-75%), Red (>75%)
Add a horizontal reference line at 75% (your design's CPU threshold).

Widget displays two lines: CPU Utilization and CPU Ready over 7 days, with threshold coloring
CPU Ready % is the most critical metric for detecting oversaturation. If Ready > 5%, your cluster is CPU-contended. Monitor both Utilization (absolute capacity) and Ready (contention signal) together.
Step 3

Add Widget 2 (Memory Utilization): Click 'Add Widget' > Select 'Line Chart' > Configure as follows:

  • Metrics: Memory Utilization (%), Memory Swapped (%)
  • Time Range: Last 7 Days
  • Thresholds: Green (0-70%), Yellow (70-85%), Red (>85%)
  • Title: 'Memory Utilization & Swap Activity (7-day baseline)'

If Memory Swapped > 0%, it indicates memory pressure and possible VM performance impact.

Widget shows memory utilization trend; Memory Swapped should be 0 or near-0 for healthy clusters
Memory swap is a red flag in vSphere. If you see non-zero swap in a properly sized cluster, investigate: (a) is a VM memory balloon driver missing, (b) did a runaway VM consume memory, (c) did you undersize the physical hosts?
Step 4

Add Widget 3 (vSAN Capacity):

  • Type: Pie Chart
  • Data Source: vCenter > Cluster A > vSAN Datastore
  • Metrics: Used Capacity (GB), Free Capacity (GB)
  • Add a data label showing percentages
  • Title: 'vSAN Storage Utilization (Current Snapshot)'

If used > 80%, change pie slice color to yellow; >90%, red.

Pie chart showing vSAN capacity split between used and free, with color-coded thresholds
vSAN capacity planning is different from traditional SAN. You must account for (a) data redundancy (FTT=1 = 50% overhead), (b) rebuild capacity (reserve 1 host's worth), (c) snapshot/delta overhead. If you see 'used capacity' > 70% in a 4-node cluster with FTT=1, you are near capacity limits.
Step 5

Add Widget 4 (DRS & HA Status):

  • Type: 'Multi-stat Gauge' or 'Scorecard'
  • Metrics: (a) DRS Recommendation Count (should be ≤2 per day for balanced cluster), (b) HA Protected VMs (%), (c) Host Failures Tolerated (should be ≥1)
  • Title: 'DRS & HA Availability Status'
  • Thresholds: HA Protected <95% = Yellow, DRS recommendations >10 = Yellow.
Widget displays three KPIs: DRS Recommendations, HA Protected %, Failures Tolerated
These metrics validate your HA/DRS design. If DRS recommendations are high, your cluster may be unbalanced (e.g., VM memory/CPU imbalance). If HA Protected % < 95%, some VMs aren't covered by HA (check resource pools and admission control).
Step 6

Save the dashboard. Click 'Save' and confirm the name 'VCF Design Capacity Baseline'.

Dashboard is saved and displayed in full view (non-edit mode) with all 4 widgets populated
After saving, VOM automatically refreshes widgets every 5 minutes. The dashboard is now part of your VCDX design artifacts — include a screenshot in your design documentation.

Validation Gate

Check: Dashboard 'VCF Design Capacity Baseline' exists with 4 widgets (CPU, Memory, vSAN, DRS/HA), all displaying data

Expected: All widgets show at least 1 data point from the last 7 days; thresholds are color-coded; you can articulate what each metric means for cluster health

Common Errors

Widget shows 'No Data' even though the cluster is running VMs
Cause: Metric is not available from the selected data source, or the collection interval hasn't completed
Fix: Verify the metric exists in the data source by navigating to the data source's metric browser. For vCenter, go to Administration > Data Sources > vCenter > View Metrics to see available metrics. Some metrics require at least 1 collection cycle (5 min) to appear.
Widget displays but the threshold lines/colors don't match your configuration
Cause: Widget configuration was not saved, or the browser cache is stale
Fix: Re-save the dashboard. Refresh the browser (Ctrl+F5 hard refresh). If still not updating, edit the widget again and re-configure the thresholds, then save.
Adding a second metric to a line chart (e.g., CPU Utilization + CPU Ready) shows only one line
Cause: Metric scaling mismatch — CPU Ready may be 0-100%, while CPU Utilization is also 0-100%, but they may need separate Y-axes if ranges differ
Fix: Edit the widget and configure two Y-axes: Left Y = CPU Utilization (0-100%), Right Y = CPU Ready (0-10%) for better visual separation. VOM will then plot CPU Utilization on the left and CPU Ready on the right, making spikes more visible.

Task 3 Configure Compliance & Alerting Policies for VCF Design Governance

Manageability

VCDX requires you to define and monitor compliance with your design. This task operationalizes compliance: you identify the design rules (e.g., 'CPU Ready must stay <5%', 'vSAN capacity must not exceed 80%', 'DRS must maintain load balance'), configure automated alerts, and demonstrate how you'll be notified if the design assumptions are violated. This is the 'feedback loop' that keeps production designs healthy.

Step 1

Navigate to Administration > Alerting > Alert Policies. Click 'Create Alert Policy'.

New Alert Policy creation wizard opens
Alert policies in VOM are event-driven rules: IF [metric condition], THEN [trigger alert with severity].
Step 2

Create Alert Policy 1: CPU Oversubscription Risk

  • Policy Name: 'Cluster CPU Ready Threshold Exceeded'
  • Condition: Cluster A > CPU Ready (%) > 5 (for more than 10 minutes)
  • Severity: Warning (Yellow)
  • Action: Send email notification to 'your-email@company.com', create a ticket in SDDC Manager (if integrated)
  • Description: 'Triggers when CPU ready exceeds 5%, indicating compute contention. Design assumed 4:1 CPU oversubscription with <5% ready as acceptable threshold.'
Alert policy is created and enabled
CPU Ready > 5% is a design assumption violation. In your VCDX design documentation, you must justify why 5% is the right threshold (based on workload type, SLA, etc.). A panelist will ask: 'Why 5% and not 10%?'
Step 3

Create Alert Policy 2: Memory Pressure Detection

  • Policy Name: 'Cluster Memory Swapped Detected'
  • Condition: Cluster A > Memory Swapped (%) > 0.1 (for more than 5 minutes)
  • Severity: Critical (Red)
  • Action: Send email + create incident in SDDC Manager
  • Description: 'Memory swap should never occur in a properly sized VCF cluster. This alert indicates either a memory leak, VM memory balloon driver failure, or cluster undersizing.'
Alert policy created
Memory swap in VCP/vSphere is a critical issue. Unlike CPU contention (which can be tolerated briefly), memory swapping degrades performance immediately. This alert should be treated as 'page immediately on-call engineer'.
Step 4

Create Alert Policy 3: vSAN Capacity Threshold

  • Policy Name: 'vSAN Storage Capacity Warning'
  • Condition: Cluster A > vSAN Capacity > 80%
  • Severity: Warning (Yellow) for >80%, Critical (Red) for >90%
  • Action: Send email, create SDDC Manager ticket with recommendation: 'Provision additional vSAN capacity or implement storage tiering'
  • Description: 'vSAN clusters should maintain <80% capacity to allow for data growth and rebuilds. At >90%, cluster is in critical capacity state and one host failure may cause data loss.'
Alert policy created with two-tier thresholds
vSAN capacity is complex because used capacity includes redundancy overhead (FTT) and rebuild overhead. At 80% capacity with FTT=1, you've lost 50% of raw capacity to redundancy, and you have <20% rebuild headroom — VERY tight. Design a margin: (Raw Capacity × 2 / 3) = Safe Capacity. E.g., 100 TB raw → ~67 TB safe capacity.
Step 5

Create Alert Policy 4: DRS Load Balance Deviation

  • Policy Name: 'Cluster Load Imbalance Detected'
  • Condition: Cluster A > DRS Recommendation Count > 10 (per day)
  • Severity: Warning (Yellow)
  • Action: Send email with recommendation: 'Review VM placement. Possible causes: (1) unbalanced VM sizing, (2) affinity rules preventing DRS, (3) maintenance mode on a host.'
  • Description: 'More than 10 DRS recommendations per day indicates the cluster is not staying balanced. This may impact performance and HA redundancy.'
Alert policy created
Step 6

Navigate to Administration > Alerting > Notification Channels. Verify that Email notification channel is configured with your email address. If not, click 'Add Notification Channel', select 'Email', enter your email, and configure SMTP settings (VOM should auto-detect the SDDC Manager's SMTP server).

Email notification channel is active and can receive test emails
Test the email channel by creating a manual alert: Administration > Alerting > Test Notification. You should receive an email within 1 minute.

Validation Gate

Check: Four alert policies are enabled: CPU Ready, Memory Swap, vSAN Capacity, DRS Load Balance. Email notification channel is configured.

Expected: You can navigate to Administration > Alerting > Active Alerts and see the four policies listed as 'Enabled'. Verify that the email notification channel shows 'Status: OK'.

Common Errors

Alert policy is created but never triggers even though the metric exceeds the threshold
Cause: The metric is not being collected from the data source, or the time window (e.g., 'for more than 10 minutes') hasn't been satisfied
Fix: Verify the metric is available in the data source by checking Administration > Data Sources > vCenter > View Metrics. Ensure the time window condition is satisfied (e.g., if the threshold is 'for more than 10 minutes', wait 10+ minutes of sustained violation). Test by creating a temporary policy with a very low threshold (e.g., CPU Ready > 0%) to verify the alert fires immediately.
Alert fires immediately and repeatedly, even though the condition is briefly met
Cause: Alert is configured to fire on every metric update, instead of waiting for sustained condition. Or the threshold is too aggressive (e.g., CPU Ready > 1%).
Fix: Edit the policy and add a time window: 'Condition must be true for 10+ minutes' before alerting. This prevents alert fatigue from brief spikes.
Email notifications are not being delivered
Cause: SMTP server misconfiguration, or email address is incorrect
Fix: Navigate to Administration > Alerting > Notification Channels > Email. Verify the SMTP server (should be auto-populated from SDDC Manager's mail server). Send a test email (Administration > Alerting > Test Notification). If still failing, check VOM system logs: ssh admin@vom-ip, then 'tail -f /var/log/vmware/vcf-ops/notifications.log'.

Task 4 Create a Compliance Report: Design vs. Reality

Manageability

This is the VCDX artifact: a formal report that proves your design is working as intended. You'll compare your design assumptions (capacity projections, availability targets, cost models) against actual metrics collected over 7 days. This is the 'proof in the pudding' — design is only valid if it holds up under real workload. A panelist will ask: 'Show me your KPIs. Is the cluster performing within your design envelope?'

Step 1

Create a new dashboard called 'Compliance Report: Design Validation' (or export the 'VCF Design Capacity Baseline' dashboard as a starting point). This dashboard will serve as the basis for your compliance report.

New dashboard created with the 4 widgets from Task 2
You can export the dashboard as a PDF for your VCDX design documentation.
Step 2

Add a title widget (text box) at the top of the dashboard with the following text:
'COMPLIANCE REPORT: VCF 9.0 Design Capacity Validation
Design Period: [Enter current week]
Cluster: Workload Domain A, Cluster A
Design Intent: 4:1 CPU oversubscription, 70% memory utilization target, 80% storage capacity target, DRS load balance, HA coverage >95%'

Dashboard now displays a header with design parameters
This header ties the metrics to your design decisions. A panelist reading this dashboard will immediately understand what you were optimizing for.
Step 3

Add a fourth widget: 'Design vs. Actual Scorecard'. Create a multi-stat widget with the following KPIs:

  1. 'CPU Utilization (7-day avg)': Design projection was X%, Actual is Y%. Pass/Fail: ✓ if Y ≤ X + 10% headroom
  2. 'Memory Utilization (7-day avg)': Design projection was X%, Actual is Y%. Pass/Fail: ✓ if Y ≤ 70% and Swap = 0%
  3. 'vSAN Capacity': Design projection was X%, Actual is Y%. Pass/Fail: ✓ if Y ≤ 80%
  4. 'HA Coverage': Design requirement was >95%, Actual is Y%. Pass/Fail: ✓ if Y ≥ 95%
Scorecard widget displays design vs. actual for 4 KPIs with pass/fail indicators
For this exercise, enter estimated 'Design projection' values (you'll need to review your VCF Planning & Preparation Workbook or design spec). If you don't have actual design numbers, use typical values: CPU 60%, Memory 65%, vSAN 70%, HA 98%.
Step 4

Export the dashboard as PDF: In the dashboard view, click the menu (three dots) > Export as PDF. Name the export 'VCF-Design-Compliance-Report-<date>.pdf'.

PDF file is downloaded to your workstation
This PDF is a key VCDX artifact. Include it in your design document with the caption: 'Proof of design validation: actual cluster metrics align with design projections.'
Step 5

Review the exported PDF and write a brief analysis (in your lab notes or a separate document):

  • Are actual metrics within your design envelope?
  • If not, identify the variance and root cause (e.g., 'CPU utilization 15% higher than projected due to unexpected batch job from Finance team').
  • What design changes would you make if re-running this exercise with the actual utilization data?
  • Did any alerts fire during the 7-day period? If yes, document which alerts and what triggered them.
A written summary (3-5 paragraphs) documenting the design validation results
This analysis is the 'design maturity' signal. A panelist who sees 'we projected 60% CPU utilization, actual was 58%, within tolerance' thinks 'solid engineer.' A panelist who sees 'we projected 60%, actual was 88%, didn't bother to analyze' thinks 'risky design practices.'
Step 6

Archive the dashboard and PDF export in your VCDX design artifacts folder. Add a README.md file with the context: 'This report validates the VCF 9.0 cluster design (capacity, availability, performance) against 7 days of production metrics.'

Files are organized in your VCDX artifact repository

Validation Gate

Check: Compliance report dashboard exists with title, 4 KPI widgets, and scorecard. PDF export is available. Written analysis (3-5 paragraphs) documents design vs. actual variance.

Expected: You have a formal compliance report artifact that proves your design is working as intended, suitable for VCDX design defense.

Common Errors

PDF export is blank or missing widgets
Cause: Widgets haven't rendered data yet (metrics still being collected), or browser has cached stale widget state
Fix: Wait 5 minutes for a fresh metric collection cycle. Refresh the dashboard (Ctrl+F5). Re-export the PDF.
Design projection values are not realistic or not available
Cause: You don't have access to the original VCF Planning & Preparation Workbook or design spec
Fix: Use typical industry baseline values for a 4-node VCF cluster: CPU 55-65% utilization (assuming 4:1 oversubscription), Memory 60-70%, vSAN 65-75% (with FTT=1), HA >98%. Document your assumptions in the analysis.

Task 5 Set Up Automated Remediation & Escalation Workflow

Manageability

A mature VCF design includes automated responses to known issues. For example, if a host goes into maintenance mode, auto-trigger VM migration and DRS rebalancing. If capacity threshold is breached, auto-trigger a capacity expansion request. This task moves you from 'we monitor it' to 'the system fixes it automatically and pages us only if it can't'. This is VCDX-level operations maturity.

Step 1

Navigate to Administration > Remediation > Policies. Click 'Create Remediation Policy'.

Remediation policy creation wizard opens
Remediation policies in VOM are automated responses to alert conditions. Common remediations: restart service, rebalance VMs, trigger capacity provisioning workflow.
Step 2

Create Remediation Policy 1: Auto-Trigger DRS Load Balance

  • Policy Name: 'DRS Rebalance on High Imbalance'
  • Trigger Condition: Alert 'Cluster Load Imbalance Detected' is raised
  • Action: Execute SDDC Manager workflow: 'Trigger Manual DRS Invocation' (or run PowerCLI: Get-Cluster 'Cluster A' | Invoke-DrsRecommendation -RunAsync)
  • Escalation: If imbalance persists for 30 minutes, send email to VCF admin + create SDDC Manager ticket
  • Description: 'Automatically invoke DRS recommendations when load imbalance is detected. If manual DRS doesn't resolve within 30 minutes, escalate to on-call admin for investigation (possible affinity rule conflict or VM memory leak).'
Remediation policy created and enabled
Auto-remediation must be implemented carefully. Don't auto-remediate destructive actions (e.g., auto-restart hosts, auto-delete snapshots). Stick to safe actions: invoke existing workflows, send notifications, trigger capacity requests.
Step 3

Create Remediation Policy 2: Escalate CPU Contention

  • Policy Name: 'CPU Contention Escalation'
  • Trigger Condition: Alert 'Cluster CPU Ready Threshold Exceeded' is raised
  • First Action (immediate): Send email to 'vcf-admin-l@company.com' with subject 'CPU Contention Alert — Cluster A' and a diagnostic snippet showing current CPU metrics
  • Second Action (if persistent for 60 minutes): Escalate to on-call engineer via PagerDuty or similar (or send SMS/Slack message)
  • Third Action (if persistent for 120 minutes): Auto-open SDDC Manager incident with title 'CPU Contention — Possible workload surge' and assign to VCF team
  • Description: 'Escalating workflow for CPU contention. If it's a brief spike (VM startup, backup), it resolves within 60 minutes. If persistent, ops team needs to investigate (possible DRS misconfiguration, affinity rule blocking migration, or actual workload demand surge).'
Three-tier escalation policy created
Escalation workflows prevent alert fatigue while ensuring critical issues get human attention. The key is tuning the escalation timeline to match your SLA.
Step 4

Create Remediation Policy 3: Capacity Planning Request

  • Policy Name: 'Auto-Initiate Capacity Planning on vSAN Threshold'
  • Trigger Condition: Alert 'vSAN Storage Capacity Warning' (Yellow level, >80%)
  • Action: Auto-create a ticket in SDDC Manager (or external ticketing system like Jira, ServiceNow) with:
    • Title: 'VCF Capacity Planning: vSAN Cluster A exceeds 80% utilization'
    • Description: 'Cluster A vSAN capacity is [X]%. Recommend: (1) Add additional vSAN hosts, (2) Implement storage tiering, (3) Clean up old snapshots/clones. Estimated timeline to 90% capacity: [Y days]'
    • Priority: Medium (Yellow) for >80%, High (Red) for >90%
    • Assignee: Infrastructure/Storage Planning team
  • Description: 'Proactively requests capacity expansion when approaching limits. This prevents hitting hard capacity limits and allows for planned upgrades.'
Capacity planning remediation policy created
Step 5

Configure escalation notifications for critical alerts. Navigate to Administration > Alerting > Escalation Policies. Create an escalation policy:

  • Name: 'VCF Critical Alert Escalation'
  • Trigger: Any alert with Severity = Critical (Red)
  • Escalation Path:
    • Level 1 (0 min): Send email to 'vcf-admin-primary@company.com'
    • Level 2 (15 min): Send email to 'vcf-admin-backup@company.com' + VCF slack channel
    • Level 3 (30 min): Page on-call engineer via PagerDuty + create incident in SDDC Manager
  • Description: 'Automatic escalation for critical issues ensures on-call engineer is paged within 30 minutes of a critical alert.'
Escalation policy created
Escalation policies reduce MTTR (Mean Time To Repair) by ensuring humans are involved the moment an automated remedy can't fix the issue.
Step 6

Test the remediation workflow: Intentionally trigger one of the alert conditions to verify the remediation and escalation work. For example, (a) Create a large test VM to generate high memory utilization, (b) Monitor VOM to see the alert fire, (c) Verify the escalation email arrives, (d) Document the entire workflow in your lab notes.

Remediation policy executes and notification is received
Only test with low-impact actions (e.g., send test email). Do not trigger destructive remediations (e.g., forcibly restart a host) in a lab environment where you can't recover.

Validation Gate

Check: Three remediation policies exist and are enabled: DRS Rebalance, CPU Escalation, Capacity Planning. Escalation policy is configured. At least one policy has been tested and confirmed to send notification.

Expected: You have a fully automated ops response workflow: alerts → automatic remediation (where safe) → escalation (if needed) → human page (if critical). This is production-grade operations maturity.

Common Errors

Remediation action executes but alert is not resolved
Cause: The remediation action doesn't address the root cause of the alert. For example, DRS rebalance won't fix CPU contention caused by a memory-hungry VM.
Fix: Ensure remediation actions match the alert root cause. For CPU contention, the remediation should be: investigate workload (check top CPU consumers), then scale horizontally (add hosts) or scale workload vertically (reduce VM demand). DRS rebalance helps but isn't the primary fix.
Escalation notification never arrives even though alert fired
Cause: Notification channel (email, SMS, PagerDuty) is not configured correctly
Fix: Test the notification channel directly: Administration > Alerting > Test Notification. Verify the email address or PagerDuty service key is correct. Check spam filters.

Final Validation

VCF Operations Manager is now configured as a complete operational platform: dashboards validate design assumptions, alerts enforce SLAs, remediation workflows automate responses, and escalation policies ensure human involvement at the right time. This is the 'day-2 operations machine' that keeps your VCF design healthy.

✓ VOM console is accessible and showing 4-host workload domain with healthy metrics → Inventory shows 1 workload domain, 1 cluster, 4 hosts; Data sources all report OK

✓ Custom dashboard 'VCF Design Capacity Baseline' exists with 4 widgets: CPU, Memory, vSAN, DRS/HA → All widgets show 7-day trend data; thresholds are color-coded

✓ Four alert policies are enabled: CPU Ready, Memory Swap, vSAN Capacity, DRS Load Balance → Policies are listed in Administration > Alerting > Alert Policies with status = Enabled

✓ Compliance report dashboard and PDF export exist with design vs. actual scorecard → PDF file is available with title, metrics, and design validation scorecard

✓ Three remediation policies are enabled: DRS Rebalance, CPU Escalation, Capacity Planning → Policies are listed in Administration > Remediation > Policies with execution status logged

✓ Escalation policy configured for critical alerts with multi-tier notification (email → Slack → PagerDuty) → Escalation policy exists in Administration > Alerting > Escalation Policies

Cleanup / Restore

Snapshot: vcp-admin-01-complete

• Disable test alert policies if any were created for testing

• Take a snapshot of VOM configuration state for future labs: 'vcp-admin-01-complete'

• Archive dashboard exports and compliance reports in your VCDX design artifact folder

• Document any design insights discovered during compliance validation (e.g., 'CPU utilization was 15% higher than projected due to [reason]')

Design Reflection (VCDX)

VCDX panelists probe your operational readiness: 'How do you know your design is working? Show me your dashboards, your alerts, your playbooks.' This lab answers that question with artifacts: a compliance report proving metrics align with design intent, alert policies enforcing SLAs, and remediation workflows demonstrating architectural maturity. A panelist examining this lab will ask: (1) Why those specific metrics? (Tie to design assumptions: e.g., CPU Ready because you designed 4:1 oversubscription) (2) What's your alert tuning strategy? (Aggressive alerts = false positives and alert fatigue; soft thresholds = latent failures. Justify your choices.) (3) How do you prevent alert fatigue? (Escalation policies, deduplication, context-aware thresholds.) Be prepared to defend your threshold values — '5% CPU Ready' — by referencing VMware best practices and your own capacity modeling.

Requirements

  • VOM console must provide single-pane-of-glass visibility into cluster capacity, compliance, and performance
  • Operational dashboards must directly reflect design assumptions (oversubscription ratios, utilization targets, availability requirements)
  • Alerts must be configured to detect and report design violations (e.g., CPU contention, capacity threshold breach)
  • Alert escalation must follow a multi-tier model: notify ops → page on-call → create incident, scaled by severity and persistence
  • Remediation workflows must automate safe responses (trigger DRS, open tickets, send notifications); destructive actions require manual confirmation

Constraints

  • VOM is a single instance with no built-in HA in VCF 9.0.0/9.0.1 (clustering added in 9.0.2+)
  • Metrics collection has a 5-minute default interval; alerts cannot detect sub-5-minute events
  • Notification channels (email, SMS, PagerDuty) depend on external integrations and network connectivity
  • Dashboard export to PDF captures a static snapshot; real-time trends require the dashboard UI
  • Remediation policies can only trigger SDDC Manager workflows or external APIs (PowerCLI, REST); they cannot directly modify cluster configuration

Assumptions

  • Operations team has 24/7 on-call coverage and can respond to paged alerts within 30 minutes
  • Email, Slack, and PagerDuty are integrated and configured before alerts are deployed to production
  • Design assumptions (CPU oversubscription ratio, memory utilization target, capacity thresholds) are documented and agreed upon before deployment
  • 7-day baseline period is sufficient to establish normal operating patterns and tuning alert thresholds
  • Workload on the cluster is stable during the compliance reporting period (no major new deployments or business changes)

Risks

  • Alert threshold misconfiguration causes either alert fatigue (too aggressive) or missed failures (too lenient) — IMPACT: operational efficiency loss or late detection of design violations — MITIGATION: start with soft thresholds, monitor alert volume for 1 week, then tune to optimal false-positive rate (<5 per day)
  • VOM appliance outage prevents all operational visibility — IMPACT: blind operations until VOM recovers — MITIGATION: export dashboards regularly as PDF, configure SNMP/syslog forwarding from SDDC Manager/vCenter as backup monitoring
  • Remediation workflow fails silently (e.g., fails to create SDDC Manager ticket due to API auth failure) — IMPACT: alert condition persists, escalation doesn't occur — MITIGATION: test remediation workflows monthly, audit remediation execution logs weekly
  • Threshold tuning is never completed — operations runs with default thresholds that don't match workload characteristics — IMPACT: irrelevant alerts, missed real issues — MITIGATION: schedule monthly threshold review; adjust based on actual utilization trends

Self-Assessment Discussion Prompts

  1. Your design assumes 4:1 CPU oversubscription with <5% CPU Ready as acceptable. Walk me through how you arrived at these numbers. What workloads are running? What's the business impact if CPU Ready hits 10%?
  2. You configured an alert for CPU Ready >5%. How often do you expect this alert to fire in steady state? If it fires daily, is your threshold too aggressive?
  3. VOM's default metric collection interval is 5 minutes. This means a 2-minute CPU spike might be missed. How does this affect your alerting strategy? Should you increase collection frequency?
  4. Your remediation policy auto-triggers DRS when imbalance is detected. What happens if DRS can't rebalance (e.g., affinity rules prevent migration)? How does your escalation handle this failure case?
  5. You're comparing 7-day actual metrics against design projections. Your actual CPU utilization is 72%, but you projected 60%. Is this a design failure or a measurement artifact? How would you investigate?
  6. If vSAN capacity exceeds 80%, your policy auto-creates a capacity planning ticket. But capacity planning can take 4-8 weeks. What do you do in the meantime — tell users 'we're out of space'?
  7. Memory swap detected (critical alert). Your remediation is to send an email. But by the time the email is read, 100 VMs might already be swapping. Is 'send email' sufficient, or do you need a more aggressive auto-remediation?

Extensions

Advanced Dashboard: Workload Domain Performance Baseline

Create a second dashboard focused on workload domain performance: VM density per host (VMs/core), application response time (if APM integrated), storage latency (read/write milliseconds), network throughput (vSAN replication, NSX TEP). This dashboard proves that your workload domain can sustain the designed VM density without degradation.

moderate

VCF 5.2.x Operational Comparison

If you have access to a VCF 5.2.x environment, document the differences in VOM (formerly vRealize Operations 8.x): different UI, different metric names, different remediation workflow syntax. Compare the dashboard layout and alerting capabilities between VCF 5.2.x and 9.0. Document 5 specific operational differences.

same

Integrate Custom Metrics from vRealize Automation

If VRA is deployed in your lab, configure VOM to ingest custom metrics from VRA (e.g., provisioning success rate, cost per VM, tenant quota utilization). Add a new dashboard widget showing VRA health and tenancy compliance. This extends VOM's visibility beyond infrastructure into application/tenancy operations.

harder

Cost Allocation & Chargeback Dashboard

Create a dashboard that tracks cost per workload domain and per tenant (if multi-tenant environment). Metrics: CPU hours consumed × $/vCPU-hour, Storage GB × $/GB-month, Network egress × $/GB. Demonstrate how VOM enables cost-based capacity planning decisions. This is highly relevant for VCDX scenarios involving cost optimization.

harder

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.