Lab: Build Capacity Planning Analysis & Rightsizing Report
Objectives
- Analyze cluster capacity using Time Remaining projections
- Run What-If scenarios for workload addition and hardware refresh
- Generate rightsizing recommendations for oversized and idle VMs
- Build a capacity planning report for CapEx decision-making
- Configure capacity alerts for proactive infrastructure procurement
Prerequisites
VCF Operations with 30+ days of metric data (required for reliable trend projections), multiple clusters with varying utilization levels, some oversized/idle VMs for rightsizing analysis
Prior labs: vcap-ops-01
Required skills:
- Capacity planning concepts
- CapEx/OpEx budgeting basics
- VCF Operations dashboards
Lab Environment
VCF Operations cluster monitoring 2+ vSphere clusters with 30+ days of historical metrics. At least one cluster near capacity (70%+ CPU or memory) for meaningful projections.
Tasks
Task 1 Capacity Planning, Rightsizing & CapEx Workflow
Build a complete capacity management workflow: analyze current capacity and Time Remaining, model future scenarios with What-If analysis, identify optimization opportunities through rightsizing, and produce a CapEx planning report that directly informs hardware procurement decisions.
Review current capacity status. VCF Operations → Optimize → Capacity → select a cluster. The Capacity Overview shows three panels: (a) Time Remaining — days until CPU, Memory, or Storage exhaustion (limiting factor determines overall Time Remaining); (b) Capacity Remaining — percentage headroom after HA and buffer reserves; (c) Reclaimable Capacity — resources recoverable through rightsizing/decommissioning. Note the limiting factor — often storage fills before CPU/Memory in vSAN environments.
Configure demand thresholds. Optimize → Capacity → Settings → Policy. Key parameters: (a) Demand Model: set to 'Allocation' (based on VM configured resources) or 'Demand' (based on actual usage — more accurate but requires historical data). (b) CPU Demand Threshold: 80% (reserves 20% for burst and HA). (c) Memory Demand Threshold: 85% (memory is less bursty than CPU). (d) HA Buffer: select 'Include High Availability buffer' and set to N+1 (one host worth of capacity reserved). These settings directly affect Time Remaining calculations.
Analyze Time Remaining projections. Click the Time Remaining chart for your cluster. The chart shows: (a) Historical trend line (actual utilization over time), (b) Projected trend line (linear regression extrapolation), (c) Demand threshold intersection (the point where projected demand hits the threshold). Example: if CPU Time Remaining = 120 days, and hardware procurement takes 90 days, you have a 30-day window to order. If < 90 days, you're already behind procurement lead time.
Run What-If scenario: Add Workloads. Optimize → Capacity → What-If → Add Scenario. Type='Add Workloads'. Configure: 50 new VMs, each with 4 vCPU / 16 GB RAM / 100 GB disk. Select target cluster. Run analysis. The results show: (a) New Time Remaining (reduced by the projected impact), (b) New utilization percentage, (c) Whether the cluster can absorb the workloads or needs additional capacity. Compare with current state side-by-side.
Run What-If scenario: Hardware Refresh. What-If → Add Scenario. Type='Add/Remove Hosts'. Configure: Remove 4 old hosts (64 GB RAM, 16-core), Add 2 new hosts (256 GB RAM, 64-core). The analysis shows whether fewer new-gen hosts can replace more old-gen hosts with net capacity gain. This directly feeds hardware refresh CapEx proposals — 'We can replace 4 hosts with 2 and still gain 30% capacity headroom.'
Generate rightsizing recommendations. Optimize → Rightsizing → filter by target cluster. VCF Operations analyzes 30+ days of usage data and recommends: (a) Oversized VMs (allocated >> consumed) — show potential savings; (b) Undersized VMs (contention indicators: CPU ready > 5%, memory balloon/swap > 0) — show recommended increase; (c) Idle VMs (< 1% utilization for 30+ days) — decommission candidates. Review each recommendation: check the metric chart to confirm the pattern is consistent, not an anomaly.
Calculate reclaimable capacity. Sum the vCPU and memory savings from all rightsizing recommendations. Example: 20 VMs each oversized by 4 vCPU and 8 GB = 80 vCPU + 160 GB reclaimable. Compare with the What-If workload addition scenario: these reclaimed resources might absorb the 50 new VMs without any hardware purchase. This is the 'optimize before you buy' argument for CapEx deferral.
Build capacity planning report. Configure → Reports → Add Report Template. Name='Quarterly-Capacity-Plan'. Sections: (a) Executive Summary: cluster Time Remaining, limiting factor, recommended action; (b) Capacity Trend: 90-day historical + projected utilization chart; (c) What-If Analysis: workload addition and hardware refresh scenarios with side-by-side comparison; (d) Rightsizing Savings: list of candidates with per-VM and total savings; (e) Recommendation: buy new hardware OR reclaim through rightsizing OR both. Schedule: Monthly or quarterly. Format: PDF. Email: capacity-planning@lab.local.
Configure capacity alerts. Alert Settings → Alert Definitions → Add. Name='Capacity-Warning-90-Days'. Symptom: Time Remaining < 90 days (matches procurement lead time). Criticality=Warning. Notification: email to infrastructure-planning@lab.local. Add second alert: 'Capacity-Critical-30-Days'. Symptom: Time Remaining < 30 days. Criticality=Critical. Notification: email + Slack webhook to management channel. These proactive alerts ensure capacity issues are identified before they become outages.
CapEx decision workflow. Combine the outputs into a procurement decision framework: (a) Time Remaining < 90 days AND no rightsizing opportunity → Immediate procurement required; (b) Time Remaining < 90 days AND rightsizing can defer by 90+ days → Execute rightsizing first, re-evaluate; (c) Time Remaining > 180 days → Monitor, no action needed; (d) What-If shows hardware refresh ROI > 30% → Include in next budget cycle. Document this framework as an operational runbook for the capacity planning team.
Validation Gate
Check: Capacity planning workflow operational from analysis to CapEx recommendation
Expected: Time Remaining projections accurate and aligned with demand thresholds, What-If scenarios modeled for workload addition and hardware refresh, rightsizing recommendations generated with savings quantified, capacity report scheduled for quarterly delivery, capacity alerts configured at 90-day and 30-day thresholds
Common Errors
Final Validation
Capacity management workflow operational with proactive alerting and CapEx reporting
✓ Time Remaining → Shows days-to-exhaustion for CPU, Memory, and Storage with correct demand thresholds
✓ What-If scenarios → Workload addition and hardware refresh scenarios produce actionable side-by-side comparison
✓ Rightsizing recommendations → Oversized, undersized, and idle VMs identified with savings quantified
✓ Capacity report → Scheduled quarterly with executive summary, trends, scenarios, and recommendations
✓ Capacity alerts → 90-day warning and 30-day critical alerts configured with correct notification channels
Cleanup / Restore
• Delete What-If scenarios if not needed
• Delete test alerts and notification rules
• Keep the capacity report template for ongoing use
Design Reflection (VCDX)
Capacity planning directly impacts CapEx decisions — one of the highest-business-impact areas in VCDX design. Demonstrate: (1) data-driven procurement (not guesswork) using Time Remaining with configurable demand thresholds; (2) 'optimize before you buy' by quantifying rightsizing savings vs new hardware cost; (3) proactive alerting aligned with procurement lead times (90-day warning = order now); (4) What-If modeling enables scenario planning for finance approval.
Requirements
- Proactive capacity alerting 90+ days before exhaustion
- Data-driven CapEx decisions with What-If scenario modeling
- 20%+ capacity reclamation through rightsizing before procurement
Constraints
- 30+ days of historical data required for reliable projections
- Linear regression assumes linear growth — may underestimate exponential expansion
- Rightsizing recommendations are averages — bursty workloads need peak-based analysis
Assumptions
- Hardware procurement lead time is 90 days
- HA buffer is N+1 (one host worth of capacity reserved)
- Demand threshold at 80% provides adequate burst headroom
Risks
- Time Remaining underestimates if workload growth is exponential (cloud migration phases)
- Rightsizing without peak analysis causes performance degradation during burst periods
- Capacity report not reviewed by management — procurement delays cause outages
Self-Assessment Discussion Prompts
- How do you set demand thresholds differently for production vs development clusters?
- What happens to Time Remaining if you change the HA buffer from N+1 to N+2?
- How do you handle capacity planning for workloads with seasonal peaks (e.g., retail)?
- When is it better to buy new hardware vs optimize existing through rightsizing?
Extensions
Build a multi-cluster capacity dashboard showing Time Remaining for all clusters on a single view
Create a what-if scenario for cloud migration — 'What if we move 30% of workloads to AWS?'
Design a chargeback model where departments 'pay' for capacity they reserve, incentivizing rightsizing
Integrate capacity alerts with ITSM to auto-create procurement tickets
⚠ Known Pitfalls (from Community KB)
References
- VCF Operations Administration Guide — Capacity Planning: techdocs.broadcom.com
- VCF Operations What-If Analysis Guide: techdocs.broadcom.com