VCF Operations Manager Dashboard Customization & Capacity Planning
Objectives
- Navigate VCF Operations Manager console and understand its role in day-2 operations
- Design and implement custom dashboards that align with VCDX design intent (capacity, compliance, availability)
- Configure threshold-based alerting and automated remediation triggers for VCF workload domains
- Analyze actual utilization vs. design projections to validate cluster sizing assumptions
- Use VOM reporting to build VCDX design artifacts (capacity headroom proof, compliance posture)
Prerequisites
VCF 9.0 lab environment fully deployed with management domain + 1 workload domain (4 hosts minimum). VOM console must be accessible (https://vom-ip (TCP 443)). At least 7 days of baseline metrics available (or manually seeded test data). DNS resolution configured for vcf.local domain.
Prior labs: holodeck-02 or equivalent VCF deployment
Required skills:
- VCF architecture fundamentals (management domain, workload domains, SDDC Manager)
- Basic understanding of capacity planning concepts (CPU oversubscription ratios, memory overhead, storage utilization)
- Comfort with vSphere Client and VCF management console UIs
- Ability to interpret performance metrics and troubleshoot resource bottlenecks
Lab Environment
VCF Management Domain (SDDC Manager, vCenter, NSX Manager) + Workload Domain A (4 ESXi hosts, shared vSAN datastore). VOM is deployed as a single instance on the management domain, accessible via FQDN vom.vcf.local or direct IP. Network: 10.0.0.0/20 for management services, 10.1.1.0/24 for ESXi hosts.
Credentials
| System | Username | Password |
|---|---|---|
| SDDC Manager | administrator@vsphere.local | VCF admin password from deployment spec |
| vCenter Server | administrator@vsphere.local | Same as SDDC Manager |
| VCF Operations Manager | admin | VOM admin password from deployment spec (default: same as SDDC Manager) |
| NSX Manager | admin | NSX admin password from deployment spec |
Tasks
Task 1 Access and Navigate VCF Operations Manager Console
ManageabilityEstablish familiarity with VOM's role in the VCF operational lifecycle. VOM is the single-pane-of-glass for capacity, compliance, and performance visibility — critical for any VCF design that must demonstrate 'what gets monitored' and 'how it gets managed'. A VCDX panelist will ask: 'Walk me through your day-2 operational dashboards. What metrics are you tracking and why?'
Open a web browser and navigate to https://vom.vcf.local (or https://<vom-ip> if DNS is not configured). Accept the self-signed certificate.
Log in with username 'admin' and the VOM password (use the same password as SDDC Manager administrator@vsphere.local).
Navigate to Administration > Data Sources. Verify that SDDC Manager, vCenter, and NSX Manager are registered as active data sources.
Navigate to Inventory. Expand the tree to see: Workload Domains > [Workload Domain A] > Clusters > Cluster A > Hosts.
Navigate to Dashboards > Built-in Dashboards. Review the pre-configured dashboards: 'Workload Domain Overview', 'Cluster Health', 'Storage Efficiency'.
Click on 'Cluster Health' dashboard and inspect the widgets: Cluster Status, Host Health, vSAN Capacity, CPU/Memory Utilization. Note the time range selector (top right) — ensure it's set to 'Last 7 Days' or 'Last 24 Hours' depending on available data.
Validation Gate
Check: VOM console is accessible and showing inventory + at least one built-in dashboard with data
Expected: You can navigate to Inventory, see 1 workload domain / 1 cluster / 4 hosts, and view the 'Cluster Health' dashboard with metrics populated
Common Errors
Task 2 Design a Custom Dashboard for Capacity Planning & Design Validation
ManageabilityIn VCDX, you must articulate 'how will I know if my design is working as intended?' This task forces you to identify 3-4 key capacity and availability metrics that directly map to your design assumptions. For example, if you designed the cluster for 4:1 CPU oversubscription, you must monitor actual CPU contention (Ready % or %Ready) and prove it stays below design limits. This is the difference between a theoretical design and a validated deployment.
Navigate to Dashboards > Create Dashboard. Enter the dashboard name: 'VCF Design Capacity Baseline' and description: 'Validates sizing assumptions: CPU oversubscription, memory headroom, storage utilization, and DRS/HA availability.'
Add Widget 1 (CPU Utilization): Click 'Add Widget' > Select 'Line Chart' > Configure as follows:
- Data Source: vCenter (management domain) > Cluster A
- Metrics: CPU Utilization (%), CPU Ready (%)
- Time Range: Last 7 Days
- Y-axis label: 'CPU %'
- Title: 'CPU Utilization & Ready (7-day baseline)'
Thresholds: Green (0-60%), Yellow (60-75%), Red (>75%)
Add a horizontal reference line at 75% (your design's CPU threshold).
Add Widget 2 (Memory Utilization): Click 'Add Widget' > Select 'Line Chart' > Configure as follows:
- Metrics: Memory Utilization (%), Memory Swapped (%)
- Time Range: Last 7 Days
- Thresholds: Green (0-70%), Yellow (70-85%), Red (>85%)
- Title: 'Memory Utilization & Swap Activity (7-day baseline)'
If Memory Swapped > 0%, it indicates memory pressure and possible VM performance impact.
Add Widget 3 (vSAN Capacity):
- Type: Pie Chart
- Data Source: vCenter > Cluster A > vSAN Datastore
- Metrics: Used Capacity (GB), Free Capacity (GB)
- Add a data label showing percentages
- Title: 'vSAN Storage Utilization (Current Snapshot)'
If used > 80%, change pie slice color to yellow; >90%, red.
Add Widget 4 (DRS & HA Status):
- Type: 'Multi-stat Gauge' or 'Scorecard'
- Metrics: (a) DRS Recommendation Count (should be ≤2 per day for balanced cluster), (b) HA Protected VMs (%), (c) Host Failures Tolerated (should be ≥1)
- Title: 'DRS & HA Availability Status'
- Thresholds: HA Protected <95% = Yellow, DRS recommendations >10 = Yellow.
Save the dashboard. Click 'Save' and confirm the name 'VCF Design Capacity Baseline'.
Validation Gate
Check: Dashboard 'VCF Design Capacity Baseline' exists with 4 widgets (CPU, Memory, vSAN, DRS/HA), all displaying data
Expected: All widgets show at least 1 data point from the last 7 days; thresholds are color-coded; you can articulate what each metric means for cluster health
Common Errors
Task 3 Configure Compliance & Alerting Policies for VCF Design Governance
ManageabilityVCDX requires you to define and monitor compliance with your design. This task operationalizes compliance: you identify the design rules (e.g., 'CPU Ready must stay <5%', 'vSAN capacity must not exceed 80%', 'DRS must maintain load balance'), configure automated alerts, and demonstrate how you'll be notified if the design assumptions are violated. This is the 'feedback loop' that keeps production designs healthy.
Navigate to Administration > Alerting > Alert Policies. Click 'Create Alert Policy'.
Create Alert Policy 1: CPU Oversubscription Risk
- Policy Name: 'Cluster CPU Ready Threshold Exceeded'
- Condition: Cluster A > CPU Ready (%) > 5 (for more than 10 minutes)
- Severity: Warning (Yellow)
- Action: Send email notification to 'your-email@company.com', create a ticket in SDDC Manager (if integrated)
- Description: 'Triggers when CPU ready exceeds 5%, indicating compute contention. Design assumed 4:1 CPU oversubscription with <5% ready as acceptable threshold.'
Create Alert Policy 2: Memory Pressure Detection
- Policy Name: 'Cluster Memory Swapped Detected'
- Condition: Cluster A > Memory Swapped (%) > 0.1 (for more than 5 minutes)
- Severity: Critical (Red)
- Action: Send email + create incident in SDDC Manager
- Description: 'Memory swap should never occur in a properly sized VCF cluster. This alert indicates either a memory leak, VM memory balloon driver failure, or cluster undersizing.'
Create Alert Policy 3: vSAN Capacity Threshold
- Policy Name: 'vSAN Storage Capacity Warning'
- Condition: Cluster A > vSAN Capacity > 80%
- Severity: Warning (Yellow) for >80%, Critical (Red) for >90%
- Action: Send email, create SDDC Manager ticket with recommendation: 'Provision additional vSAN capacity or implement storage tiering'
- Description: 'vSAN clusters should maintain <80% capacity to allow for data growth and rebuilds. At >90%, cluster is in critical capacity state and one host failure may cause data loss.'
Create Alert Policy 4: DRS Load Balance Deviation
- Policy Name: 'Cluster Load Imbalance Detected'
- Condition: Cluster A > DRS Recommendation Count > 10 (per day)
- Severity: Warning (Yellow)
- Action: Send email with recommendation: 'Review VM placement. Possible causes: (1) unbalanced VM sizing, (2) affinity rules preventing DRS, (3) maintenance mode on a host.'
- Description: 'More than 10 DRS recommendations per day indicates the cluster is not staying balanced. This may impact performance and HA redundancy.'
Navigate to Administration > Alerting > Notification Channels. Verify that Email notification channel is configured with your email address. If not, click 'Add Notification Channel', select 'Email', enter your email, and configure SMTP settings (VOM should auto-detect the SDDC Manager's SMTP server).
Validation Gate
Check: Four alert policies are enabled: CPU Ready, Memory Swap, vSAN Capacity, DRS Load Balance. Email notification channel is configured.
Expected: You can navigate to Administration > Alerting > Active Alerts and see the four policies listed as 'Enabled'. Verify that the email notification channel shows 'Status: OK'.
Common Errors
Task 4 Create a Compliance Report: Design vs. Reality
ManageabilityThis is the VCDX artifact: a formal report that proves your design is working as intended. You'll compare your design assumptions (capacity projections, availability targets, cost models) against actual metrics collected over 7 days. This is the 'proof in the pudding' — design is only valid if it holds up under real workload. A panelist will ask: 'Show me your KPIs. Is the cluster performing within your design envelope?'
Create a new dashboard called 'Compliance Report: Design Validation' (or export the 'VCF Design Capacity Baseline' dashboard as a starting point). This dashboard will serve as the basis for your compliance report.
Add a title widget (text box) at the top of the dashboard with the following text:
'COMPLIANCE REPORT: VCF 9.0 Design Capacity Validation
Design Period: [Enter current week]
Cluster: Workload Domain A, Cluster A
Design Intent: 4:1 CPU oversubscription, 70% memory utilization target, 80% storage capacity target, DRS load balance, HA coverage >95%'
Add a fourth widget: 'Design vs. Actual Scorecard'. Create a multi-stat widget with the following KPIs:
- 'CPU Utilization (7-day avg)': Design projection was X%, Actual is Y%. Pass/Fail: ✓ if Y ≤ X + 10% headroom
- 'Memory Utilization (7-day avg)': Design projection was X%, Actual is Y%. Pass/Fail: ✓ if Y ≤ 70% and Swap = 0%
- 'vSAN Capacity': Design projection was X%, Actual is Y%. Pass/Fail: ✓ if Y ≤ 80%
- 'HA Coverage': Design requirement was >95%, Actual is Y%. Pass/Fail: ✓ if Y ≥ 95%
Export the dashboard as PDF: In the dashboard view, click the menu (three dots) > Export as PDF. Name the export 'VCF-Design-Compliance-Report-<date>.pdf'.
Review the exported PDF and write a brief analysis (in your lab notes or a separate document):
- Are actual metrics within your design envelope?
- If not, identify the variance and root cause (e.g., 'CPU utilization 15% higher than projected due to unexpected batch job from Finance team').
- What design changes would you make if re-running this exercise with the actual utilization data?
- Did any alerts fire during the 7-day period? If yes, document which alerts and what triggered them.
Archive the dashboard and PDF export in your VCDX design artifacts folder. Add a README.md file with the context: 'This report validates the VCF 9.0 cluster design (capacity, availability, performance) against 7 days of production metrics.'
Validation Gate
Check: Compliance report dashboard exists with title, 4 KPI widgets, and scorecard. PDF export is available. Written analysis (3-5 paragraphs) documents design vs. actual variance.
Expected: You have a formal compliance report artifact that proves your design is working as intended, suitable for VCDX design defense.
Common Errors
Task 5 Set Up Automated Remediation & Escalation Workflow
ManageabilityA mature VCF design includes automated responses to known issues. For example, if a host goes into maintenance mode, auto-trigger VM migration and DRS rebalancing. If capacity threshold is breached, auto-trigger a capacity expansion request. This task moves you from 'we monitor it' to 'the system fixes it automatically and pages us only if it can't'. This is VCDX-level operations maturity.
Navigate to Administration > Remediation > Policies. Click 'Create Remediation Policy'.
Create Remediation Policy 1: Auto-Trigger DRS Load Balance
- Policy Name: 'DRS Rebalance on High Imbalance'
- Trigger Condition: Alert 'Cluster Load Imbalance Detected' is raised
- Action: Execute SDDC Manager workflow: 'Trigger Manual DRS Invocation' (or run PowerCLI: Get-Cluster 'Cluster A' | Invoke-DrsRecommendation -RunAsync)
- Escalation: If imbalance persists for 30 minutes, send email to VCF admin + create SDDC Manager ticket
- Description: 'Automatically invoke DRS recommendations when load imbalance is detected. If manual DRS doesn't resolve within 30 minutes, escalate to on-call admin for investigation (possible affinity rule conflict or VM memory leak).'
Create Remediation Policy 2: Escalate CPU Contention
- Policy Name: 'CPU Contention Escalation'
- Trigger Condition: Alert 'Cluster CPU Ready Threshold Exceeded' is raised
- First Action (immediate): Send email to 'vcf-admin-l@company.com' with subject 'CPU Contention Alert — Cluster A' and a diagnostic snippet showing current CPU metrics
- Second Action (if persistent for 60 minutes): Escalate to on-call engineer via PagerDuty or similar (or send SMS/Slack message)
- Third Action (if persistent for 120 minutes): Auto-open SDDC Manager incident with title 'CPU Contention — Possible workload surge' and assign to VCF team
- Description: 'Escalating workflow for CPU contention. If it's a brief spike (VM startup, backup), it resolves within 60 minutes. If persistent, ops team needs to investigate (possible DRS misconfiguration, affinity rule blocking migration, or actual workload demand surge).'
Create Remediation Policy 3: Capacity Planning Request
- Policy Name: 'Auto-Initiate Capacity Planning on vSAN Threshold'
- Trigger Condition: Alert 'vSAN Storage Capacity Warning' (Yellow level, >80%)
- Action: Auto-create a ticket in SDDC Manager (or external ticketing system like Jira, ServiceNow) with:
- Title: 'VCF Capacity Planning: vSAN Cluster A exceeds 80% utilization'
- Description: 'Cluster A vSAN capacity is [X]%. Recommend: (1) Add additional vSAN hosts, (2) Implement storage tiering, (3) Clean up old snapshots/clones. Estimated timeline to 90% capacity: [Y days]'
- Priority: Medium (Yellow) for >80%, High (Red) for >90%
- Assignee: Infrastructure/Storage Planning team
- Description: 'Proactively requests capacity expansion when approaching limits. This prevents hitting hard capacity limits and allows for planned upgrades.'
Configure escalation notifications for critical alerts. Navigate to Administration > Alerting > Escalation Policies. Create an escalation policy:
- Name: 'VCF Critical Alert Escalation'
- Trigger: Any alert with Severity = Critical (Red)
- Escalation Path:
- Level 1 (0 min): Send email to 'vcf-admin-primary@company.com'
- Level 2 (15 min): Send email to 'vcf-admin-backup@company.com' + VCF slack channel
- Level 3 (30 min): Page on-call engineer via PagerDuty + create incident in SDDC Manager
- Description: 'Automatic escalation for critical issues ensures on-call engineer is paged within 30 minutes of a critical alert.'
Test the remediation workflow: Intentionally trigger one of the alert conditions to verify the remediation and escalation work. For example, (a) Create a large test VM to generate high memory utilization, (b) Monitor VOM to see the alert fire, (c) Verify the escalation email arrives, (d) Document the entire workflow in your lab notes.
Validation Gate
Check: Three remediation policies exist and are enabled: DRS Rebalance, CPU Escalation, Capacity Planning. Escalation policy is configured. At least one policy has been tested and confirmed to send notification.
Expected: You have a fully automated ops response workflow: alerts → automatic remediation (where safe) → escalation (if needed) → human page (if critical). This is production-grade operations maturity.
Common Errors
Final Validation
VCF Operations Manager is now configured as a complete operational platform: dashboards validate design assumptions, alerts enforce SLAs, remediation workflows automate responses, and escalation policies ensure human involvement at the right time. This is the 'day-2 operations machine' that keeps your VCF design healthy.
✓ VOM console is accessible and showing 4-host workload domain with healthy metrics → Inventory shows 1 workload domain, 1 cluster, 4 hosts; Data sources all report OK
✓ Custom dashboard 'VCF Design Capacity Baseline' exists with 4 widgets: CPU, Memory, vSAN, DRS/HA → All widgets show 7-day trend data; thresholds are color-coded
✓ Four alert policies are enabled: CPU Ready, Memory Swap, vSAN Capacity, DRS Load Balance → Policies are listed in Administration > Alerting > Alert Policies with status = Enabled
✓ Compliance report dashboard and PDF export exist with design vs. actual scorecard → PDF file is available with title, metrics, and design validation scorecard
✓ Three remediation policies are enabled: DRS Rebalance, CPU Escalation, Capacity Planning → Policies are listed in Administration > Remediation > Policies with execution status logged
✓ Escalation policy configured for critical alerts with multi-tier notification (email → Slack → PagerDuty) → Escalation policy exists in Administration > Alerting > Escalation Policies
Cleanup / Restore
Snapshot: vcp-admin-01-complete
• Disable test alert policies if any were created for testing
• Take a snapshot of VOM configuration state for future labs: 'vcp-admin-01-complete'
• Archive dashboard exports and compliance reports in your VCDX design artifact folder
• Document any design insights discovered during compliance validation (e.g., 'CPU utilization was 15% higher than projected due to [reason]')
Design Reflection (VCDX)
VCDX panelists probe your operational readiness: 'How do you know your design is working? Show me your dashboards, your alerts, your playbooks.' This lab answers that question with artifacts: a compliance report proving metrics align with design intent, alert policies enforcing SLAs, and remediation workflows demonstrating architectural maturity. A panelist examining this lab will ask: (1) Why those specific metrics? (Tie to design assumptions: e.g., CPU Ready because you designed 4:1 oversubscription) (2) What's your alert tuning strategy? (Aggressive alerts = false positives and alert fatigue; soft thresholds = latent failures. Justify your choices.) (3) How do you prevent alert fatigue? (Escalation policies, deduplication, context-aware thresholds.) Be prepared to defend your threshold values — '5% CPU Ready' — by referencing VMware best practices and your own capacity modeling.
Requirements
- VOM console must provide single-pane-of-glass visibility into cluster capacity, compliance, and performance
- Operational dashboards must directly reflect design assumptions (oversubscription ratios, utilization targets, availability requirements)
- Alerts must be configured to detect and report design violations (e.g., CPU contention, capacity threshold breach)
- Alert escalation must follow a multi-tier model: notify ops → page on-call → create incident, scaled by severity and persistence
- Remediation workflows must automate safe responses (trigger DRS, open tickets, send notifications); destructive actions require manual confirmation
Constraints
- VOM is a single instance with no built-in HA in VCF 9.0.0/9.0.1 (clustering added in 9.0.2+)
- Metrics collection has a 5-minute default interval; alerts cannot detect sub-5-minute events
- Notification channels (email, SMS, PagerDuty) depend on external integrations and network connectivity
- Dashboard export to PDF captures a static snapshot; real-time trends require the dashboard UI
- Remediation policies can only trigger SDDC Manager workflows or external APIs (PowerCLI, REST); they cannot directly modify cluster configuration
Assumptions
- Operations team has 24/7 on-call coverage and can respond to paged alerts within 30 minutes
- Email, Slack, and PagerDuty are integrated and configured before alerts are deployed to production
- Design assumptions (CPU oversubscription ratio, memory utilization target, capacity thresholds) are documented and agreed upon before deployment
- 7-day baseline period is sufficient to establish normal operating patterns and tuning alert thresholds
- Workload on the cluster is stable during the compliance reporting period (no major new deployments or business changes)
Risks
- Alert threshold misconfiguration causes either alert fatigue (too aggressive) or missed failures (too lenient) — IMPACT: operational efficiency loss or late detection of design violations — MITIGATION: start with soft thresholds, monitor alert volume for 1 week, then tune to optimal false-positive rate (<5 per day)
- VOM appliance outage prevents all operational visibility — IMPACT: blind operations until VOM recovers — MITIGATION: export dashboards regularly as PDF, configure SNMP/syslog forwarding from SDDC Manager/vCenter as backup monitoring
- Remediation workflow fails silently (e.g., fails to create SDDC Manager ticket due to API auth failure) — IMPACT: alert condition persists, escalation doesn't occur — MITIGATION: test remediation workflows monthly, audit remediation execution logs weekly
- Threshold tuning is never completed — operations runs with default thresholds that don't match workload characteristics — IMPACT: irrelevant alerts, missed real issues — MITIGATION: schedule monthly threshold review; adjust based on actual utilization trends
Self-Assessment Discussion Prompts
- Your design assumes 4:1 CPU oversubscription with <5% CPU Ready as acceptable. Walk me through how you arrived at these numbers. What workloads are running? What's the business impact if CPU Ready hits 10%?
- You configured an alert for CPU Ready >5%. How often do you expect this alert to fire in steady state? If it fires daily, is your threshold too aggressive?
- VOM's default metric collection interval is 5 minutes. This means a 2-minute CPU spike might be missed. How does this affect your alerting strategy? Should you increase collection frequency?
- Your remediation policy auto-triggers DRS when imbalance is detected. What happens if DRS can't rebalance (e.g., affinity rules prevent migration)? How does your escalation handle this failure case?
- You're comparing 7-day actual metrics against design projections. Your actual CPU utilization is 72%, but you projected 60%. Is this a design failure or a measurement artifact? How would you investigate?
- If vSAN capacity exceeds 80%, your policy auto-creates a capacity planning ticket. But capacity planning can take 4-8 weeks. What do you do in the meantime — tell users 'we're out of space'?
- Memory swap detected (critical alert). Your remediation is to send an email. But by the time the email is read, 100 VMs might already be swapping. Is 'send email' sufficient, or do you need a more aggressive auto-remediation?
Extensions
Advanced Dashboard: Workload Domain Performance Baseline
Create a second dashboard focused on workload domain performance: VM density per host (VMs/core), application response time (if APM integrated), storage latency (read/write milliseconds), network throughput (vSAN replication, NSX TEP). This dashboard proves that your workload domain can sustain the designed VM density without degradation.
moderateVCF 5.2.x Operational Comparison
If you have access to a VCF 5.2.x environment, document the differences in VOM (formerly vRealize Operations 8.x): different UI, different metric names, different remediation workflow syntax. Compare the dashboard layout and alerting capabilities between VCF 5.2.x and 9.0. Document 5 specific operational differences.
sameIntegrate Custom Metrics from vRealize Automation
If VRA is deployed in your lab, configure VOM to ingest custom metrics from VRA (e.g., provisioning success rate, cost per VM, tenant quota utilization). Add a new dashboard widget showing VRA health and tenancy compliance. This extends VOM's visibility beyond infrastructure into application/tenancy operations.
harderCost Allocation & Chargeback Dashboard
Create a dashboard that tracks cost per workload domain and per tenant (if multi-tenant environment). Metrics: CPU hours consumed × $/vCPU-hour, Storage GB × $/GB-month, Network egress × $/GB. Demonstrate how VOM enables cost-based capacity planning decisions. This is highly relevant for VCDX scenarios involving cost optimization.
harderReferences
- VMware Cloud Foundation 9.0 Operations and Administration GuideTier 1 — Official
Official Broadcom documentation for VCF operations, VOM console, dashboard creation, alerting policies, and remediation workflows. - VCF 9.0 Planning and Preparation WorkbookTier 1 — Official
Critical reference for capacity planning assumptions: CPU oversubscription ratios, memory overhead, storage utilization targets. This is where your '4:1 CPU' and '5% Ready' assumptions come from. - vSAN 8.x Architecture and Capacity PlanningTier 1 — Official
Deep dive on vSAN capacity model: FTT overhead, rebuild capacity, compression, deduplication impact on usable capacity. Critical for understanding your vSAN 80% threshold. - VMware Best Practices: DRS and HA Configuration for VCFTier 1 — Official
Guidelines for DRS tuning (aggressiveness levels), HA admission control sizing, and load balancing strategies. Justifies your 10-recommendation/day threshold for DRS rebalance. - William Lam — VCF Operational ExcellenceTier 3 — Expert Blog
Expert blog on VCF day-2 operations, monitoring strategies, performance tuning. Often covers VOM configuration tips and gotchas.