Lab: Design Super Metrics & Custom Groups for Cost Optimization
Objectives
- Create super metrics using formula language with cross-object aggregation
- Build custom groups with dynamic membership rules for business-context organization
- Design a cost allocation model using super metrics for per-VM and per-group chargeback
- Identify underutilized VMs using efficiency super metrics and dynamic group membership
- Export super metric definitions for version control and environment migration
Prerequisites
VCF 9.0 with VCF Operations deployed and collecting metrics from vCenter for at least 7 days (required for meaningful efficiency analysis). Multiple VMs across clusters with varying utilization levels.
Required skills:
- VCF Operations UI navigation
- Basic metrics concepts (CPU usage, memory consumed)
- Cost allocation principles
Lab Environment
VCF Operations cluster connected to vCenter managing 2+ clusters with 20+ VMs of varying sizes and utilization patterns. Some VMs should be intentionally oversized or idle to demonstrate rightsizing identification.
Tasks
Task 1 Super Metrics & Custom Groups for Cost Optimization
Build a complete cost optimization workflow: create efficiency super metrics that identify waste, build dynamic custom groups that auto-populate with optimization candidates, and design a chargeback model that quantifies savings — demonstrating the full VCF Operations analytics pipeline from raw metrics to business-level cost reporting.
Navigate to super metric creation. VCF Operations → Configure → Super Metrics → Add. Name='vm_cpu_efficiency_pct'. Object Type=VirtualMachine. Formula: (${this, metric=cpu|usagemhz|average} / (${this, metric=cpu|corecount_provisioned} * ${this, metric=cpu|cpuMhz})) * 100. This calculates CPU efficiency as: actual MHz consumed / total MHz provisioned × 100. A VM with 4 vCPU @ 2600 MHz (10,400 MHz provisioned) using 1,040 MHz averages 10% efficiency — a clear rightsizing candidate.Create memory efficiency super metric. Add → Name='vm_memory_efficiency_pct'. Object Type=VirtualMachine. Formula: (${this, metric=mem|consumed|average} / ${this, metric=mem|guest|provisioned}) * 100. This shows actual memory consumption vs provisioned — capturing balloon/swap scenarios where a VM appears to use all allocated memory but actually needs far less. A value < 30% sustained over 30 days indicates oversized memory allocation.Create per-VM cost super metric. Add → Name='vm_monthly_cost_usd'. Object Type=VirtualMachine. Formula: (${this, metric=cpu|usagemhz|average} * 0.02) + (${this, metric=mem|consumed|average} / 1048576 * 5.00) + (${this, metric=virtualDisk|totalReadLatency|average} * 0.001). This assigns dollar costs based on actual resource consumption: $0.02/MHz CPU + $5.00/GB memory + $0.001/ms disk latency. Adjust rates to match your organization's infrastructure cost model.Create cluster-level contention super metric. Add → Name='cluster_cpu_contention_score'. Object Type=ClusterComputeResource. Formula: avg(${this, metric=cpu|ready|summation, depth=1}) / avg(${this, metric=cpu|usagemhz|average, depth=1}) * 100. This aggregates CPU ready time across all hosts in the cluster (depth=1 = direct children) and normalizes by utilization. Values > 5% indicate the cluster is over-committed — VMs are waiting for CPU time.Enable super metrics on object types. For each super metric, click the 'Object Types' tab and select the target: VirtualMachine for VM-level metrics, ClusterComputeResource for cluster metrics. Enable 'Monitor' to start data collection. VCF Operations will calculate and store these values at each collection interval (5 minutes). Wait 15-20 minutes for initial data population before proceeding.
Create dynamic custom group for underutilized VMs. Configure → Custom Groups → Add. Name='Rightsizing-Candidates-Oversized'. Group Type=VirtualMachine. Membership: Define Membership Criteria → Add Rule: 'Super Metric vm_cpu_efficiency_pct is less than 20' AND 'Super Metric vm_memory_efficiency_pct is less than 30' AND 'Property Summary|Runtime|Power State equals poweredOn' AND 'Property Summary|Runtime|BootTime is more than 30 days ago'. Click Preview Members to verify the dynamic query returns expected VMs. This group auto-updates as VM utilization changes.
Create custom group for idle VMs. Add → Name='Decommission-Candidates-Idle'. Membership: 'Super Metric vm_cpu_efficiency_pct is less than 2' AND 'Metric net|usage|average is less than 10 KBps' AND powered on for > 30 days. These VMs are consuming resources but doing essentially nothing — prime candidates for decommissioning. Create a third group: 'Cost-Centers-Engineering' with membership based on VM tag or folder path matching the engineering team's resources.
Build chargeback cost report. Navigate → Dashboards → Create Dashboard. Name='Cost-Optimization-Dashboard'. Add widgets: (a) Scoreboard: 'Total Monthly Cost' showing sum of vm_monthly_cost_usd across all VMs; (b) Top-N: 'Top 10 Most Expensive VMs' ranked by vm_monthly_cost_usd; (c) Scoreboard: 'Rightsizing Savings Potential' showing count of Rightsizing-Candidates group members × average cost difference; (d) List: Members of 'Decommission-Candidates-Idle' with current cost. Use widget interactions so clicking a VM in the Top-N updates a metric chart showing its efficiency trend.
Create a scheduled report for management. Configure → Reports → Add Report Template. Name='Monthly-Cost-Optimization'. Add views: (a) List view of Rightsizing-Candidates with columns: VM Name, Current CPU, Recommended CPU, Current Memory, Recommended Memory, Monthly Savings; (b) Summary view of Cost-Centers-Engineering showing total cost, VM count, average efficiency. Schedule: Monthly, first Monday, email to stakeholder distribution list. Format: PDF.
Export super metric definitions for version control. Configure → Super Metrics → select all → Export. This downloads an XML file containing all formula definitions, object type bindings, and descriptions. Store in Git for version control. To migrate to another VCF Operations instance: Import the XML file on the target. This enables dashboard-as-code and metrics-as-code practices — critical for multi-environment consistency.
Validation Gate
Check: Super metrics calculating, custom groups auto-populating, and cost dashboard operational
Expected: 4 super metrics computing values for all VMs and clusters, 3 dynamic custom groups with correct membership based on efficiency thresholds, cost dashboard showing per-VM costs and rightsizing savings potential, scheduled report configured for monthly delivery
Common Errors
Final Validation
Cost optimization workflow operational with super metrics, custom groups, and automated reporting
✓ CPU efficiency super metric → Calculating percentage values for all VMs (0-100 range)
✓ Memory efficiency super metric → Calculating percentage values for all VMs
✓ Per-VM cost super metric → Dollar values calculated for all VMs
✓ Rightsizing candidates group → Auto-populated with VMs matching efficiency criteria
✓ Cost dashboard → Displaying Top-N costs, savings potential, and drill-down interactions
✓ Scheduled report → Configured for monthly delivery with correct views and distribution
Cleanup / Restore
• Delete scheduled report
• Delete custom groups
• Delete dashboard
• Delete super metrics (in reverse order — dashboard first, then groups, then metrics)
Design Reflection (VCDX)
Cost optimization through data-driven rightsizing is a compelling VCDX story. It demonstrates: (1) Technical depth — understanding VCF Operations formula language and object model; (2) Business alignment — translating technical metrics into financial impact; (3) Operational maturity — automated reporting and dynamic grouping reduce manual effort. In defense, be prepared to justify your cost model rates and explain how dynamic thresholds avoid false positives in rightsizing recommendations.
Requirements
- Identify 20%+ cost reduction opportunity through rightsizing and decommissioning
- Automated monthly reporting to finance/management
- Dynamic candidate identification without manual VM-by-VM review
Constraints
- Super metric performance degrades beyond 50 formulas per object type
- Rightsizing requires 30+ days of historical data for reliable recommendations
- Cost model rates must be calibrated to actual infrastructure costs (hardware, power, licensing)
Assumptions
- VMs have been running with stable workloads for 30+ days
- Cost model rates approved by finance team
- Rightsizing changes are applied during maintenance windows (not live)
Risks
- Aggressive rightsizing (based on average instead of peak) causes performance degradation during burst periods
- Idle VM decommission without business validation removes disaster recovery or seasonal workloads
- Super metric formula errors produce incorrect cost data — misleading management decisions
Self-Assessment Discussion Prompts
- How do you calibrate chargeback rates to reflect actual infrastructure costs?
- When should you use conservative (peak-based) vs aggressive (average-based) rightsizing?
- How do you handle seasonal workloads that appear idle for 10 months but are critical during peak periods?
- What governance process ensures rightsizing recommendations are validated before execution?
Extensions
Build a super metric that calculates 'cost waste' = (provisioned cost - consumed cost) for each VM
Create a Slack webhook notification when the Rightsizing-Candidates group grows beyond a threshold
Implement a rightsizing automation pipeline using VCF Operations API + vSphere API
Design a multi-tenant chargeback model with per-department custom groups and separate cost rates
⚠ Known Pitfalls (from Community KB)
References
- VCF Operations Administration Guide — Super Metrics: techdocs.broadcom.com
- VCF Operations API Reference — /supermetrics endpoint: techdocs.broadcom.com