Academy/VCAP — VCF Operations (3V0-22.25)
VCF

VCAP — VCF Operations (3V0-22.25)

VCF 9.0vcap-advanced

VCAP-level operations architecture, advanced monitoring dashboards, super metrics & custom groups, alert policies & automated remediation, capacity planning & rightsizing, API-driven extensibility, and multi-cloud observability for unified VMware Cloud Foundation fleet management.

3V0-22.25
VCAP
60
Questions
135m
Duration
300/500
Pass Score
32
Objectives

Exam Blueprint Weights

Section titles, groupings and weights below are VCDX Academy study groupings, NOT the official Broadcom blueprint structure. Broadcom publishes no section weights. Always cross-check the official exam guide. Official exam guide ↗
Section 1 — Architecture
~15%
Section 2 — Capacity Management
~20%
Section 3 — Performance Analysis
~20%
Section 4 — Automation and Remediation
~20%
Section 5 — Custom Content and Complianc
~25%
High Weight

Version Evolution

VCAP Operations covers advanced day-2 operations including performance optimization, capacity planning, and automated remediation. VCF 9.0 shifted operations tooling from Aria Operations to VCF Operations — a significant change in UI, API, and workflow. The exam tests scenario-based troubleshooting and optimization at scale.

Learning Outcomes

  • Understand VCF Operations internal architecture: node roles, Cassandra time-series database, adapter execution pipeline, and rollup policies
  • Design and implement Super Metrics for business KPIs and chargeback using formula language with cross-object references
  • Build interactive dashboards with drill-down interactions, and design executive/operations/capacity planning dashboard patterns
  • Configure multi-level alerting: symptom definitions (static + dynamic thresholds) → alert definitions (correlated symptoms) → notifications (email/webhook/SNMP) → automated remediation
  • Perform capacity planning with Time Remaining forecasting, What-If analysis, and rightsizing workflows (conservative vs aggressive)
  • Leverage the VCF Operations REST API for automation, dashboard-as-code, and integration with ITSM/SIEM systems
  • Extend monitoring to multi-cloud environments (AWS/Azure/GCP/K8s) using Management Packs

VCF Operations Internal Architecture#

VCF Operations (formerly vRealize Operations) is a distributed, horizontally scalable analytics platform. Understanding node roles, data architecture, and sharding patterns is essential for production deployments at scale.

Node Roles & Cluster Composition

VCF Operations clusters contain specialized node types:

  • Node Role
  • Responsibility
  • Typical Count
  • Storage
  • Master
  • Cluster orchestration, UI, API, decision engine (analytics)
  • 1-3
  • SSD (fast I/O for analytics)
  • Replica
  • Standby for Master failover, load balancing
  • 1-2
  • SSD (sync with Master)
  • Data (Analytics)
  • Time-series database, metric storage, rollup computation
  • 3-100+
  • SSD/NVMe (high IOPS for writes)
  • Remote Collector
  • Fetch metrics from remote vCenter (WAN links)
  • 0-10
  • Local cache (small)
  • Cloud Proxy
  • Outbound connection broker for SaaS telemetry
  • 0-2
  • Minimal

Data Architecture & Cassandra-Based Time-Series Database

VCF Operations uses Apache Cassandra for scalable, distributed metrics storage:

Write Path (Collector → Master → Data Nodes):
┌─────────────┐
│  Collector  │ (fetch metrics from vCenter)
│  (Host A)   │
└──────┬──────┘
       │
       └──→ ┌──────────────────────────────────────┐
            │ Master (Analytics Cluster)           │
            │ - Receives metric stream              │
            │ - Persists to Cassandra              │
            │ - Computes super metrics             │
            │ - Triggers alerts                     │
            └──────────┬────────────────────────────┘
                       │
           ┌───────────┼───────────┐
           │           │           │
        ┌──▼──┐    ┌──▼──┐    ┌──▼──┐
        │Data │    │Data │    │Data │  (Cassandra Nodes)
        │ N1  │    │ N2  │    │ N3  │  3x replication
        └─────┘    └─────┘    └─────┘
Read Path (Query → Rollup Aggregation):
┌──────────────┐
│  Dashboard   │ (5min resolution)
│  Query       │
└────────┬─────┘
         │
         ├─→ 5min-aggregated data (SSD cache)
         │
         ├─→ 1hr-aggregated data (compressed)
         │
         └─→ 1day-aggregated data (archive)

Rollup Policies & Data Retention

VCF Operations automatically aggregates metrics across time dimensions:

Default Rollup Configuration:

---

Raw Metrics:

  5-minute buckets → Stored 7 days

E.g., cpu.idle at [10:00-10:05], [10:05-10:10], ...

1-Hour Rollup:

  Aggregates 12 × 5-min samples → 1-hour bucket

Stored 365 days

Math: mean(cpu.idle samples in hour), max(cpu.ready samples), etc.

1-Day Rollup:

  Aggregates 24 × 1-hour samples → 1-day bucket

Stored 7+ years (per policy)

Used for trend analysis, capacity forecasting

Storage Impact:

  • Raw (7 days): 100GB
  • 1-hour (1 year): 20GB
  • 1-day (7 years): 5GB
  • Total: ~125GB for 7-year retention

Retention policies can be customized:

Development: raw=3d, 1h=90d, 1d=1y (smaller footprint)

Production: raw=30d, 1h=3y, 1d=10y (compliance retention)

Adapter Execution Model: Collect → Transform → Load

Adapters fetch data via a three-stage pipeline:

Adapter Pipeline (per collection interval, e.g., every 5 min):

---

STAGE 1: COLLECT

  • Adapter instance contacts target system (vCenter, NSX, vSAN)
  • Uses REST API, SNMP, syslog, or proprietary protocol
  • Timeout: 30s per collection (configurable)
  • Retry: 3 attempts if collection fails

STAGE 2: TRANSFORM

  - Raw data (performance counters, state info) → VCF Operations object model
  - Example: vSphere counter "cpu.usage" (MHz) → ResourceKind Host metric "cpu.usage%"
  • Applies unit conversion, normalization
  • Filters (discard if below threshold)
  • Aggregates (sum, avg across samples)

STAGE 3: LOAD

  • Transformed metrics sent to Master node
  • Master persists to Cassandra
  • Acknowledgment returned to Adapter

Adapter Concurrency:

Default: 4 parallel adapters per resource type

Configurable: 1-16 (avoid overwhelming target system)

If one adapter lags (target slow), others proceed in parallel

Example: vCenter Adapter with 4 concurrent collectors

Collector 1: Fetch host metrics from vCenter cluster A

Collector 2: Fetch host metrics from vCenter cluster B

Collector 3: Fetch VM metrics from vCenter cluster A

Collector 4: Fetch VM metrics from vCenter cluster B

All 4 run in parallel, complete independently

Solution Model & Metadata Hierarchy

VCF Operations' data model defines object types and relationships:

Solution Model - Resource Kind Hierarchy:

---
DataCenter (top-level container)

├─ Cluster (vSphere cluster)
│  ├─ Host
│  │  ├─ NUMA Node
│  │  └─ NIC
│  ├─ VM
│  │  ├─ vDisk
│  │  └─ vNIC
│  └─ vSAN Cluster (sub-object within Cluster)
│     ├─ vSAN Disk Group
│     ├─ vSAN Disk
│     └─ vSAN Component
├─ Datastore
│  └─ Datastore Disk
├─ Network
│  ├─ vSwitch
│  ├─ Portgroup
│  └─ Uplink NIC
└─ NSX Objects (if NSX is monitored)
   ├─ NSX Manager
   ├─ NSX Edge
   ├─ Logical Switch
   └─ Logical Port

Each resource kind has:

  • Metrics (numeric: cpu.ready, memory.usage, etc.)
  • Properties (text: hostname, version, config, etc.)
  • Relationships (parent/

Key Takeaways

  • Node roles: Master (UI, API, analytics), Replica (standby), Data (time-series storage), Remote Collector (WAN-fed metrics). Cassandra backend handles distributed metric writes with 3x replication.
  • Adapter execution: Collect → Transform → Load pipeline; default 4 parallel adapters per resource type; adapter timeouts configurable (30s typical).
  • Rollup policies: raw metrics 7 days, 1-hour rollup 365 days, 1-day rollup 7+ years; storage efficient for long-term trend analysis and capacity forecasting.
  • Solution model hierarchy: Datacenter → Cluster → Host → VM; each resource kind has metrics (numeric), properties (text), and relationships (parent/child). Custom objects definable via custom groups.

Super Metrics, Custom Groups & Multi-Tenancy#

Super Metrics enable computed KPIs that aggregate raw metrics across objects. Custom Groups organize infrastructure by business context rather than physical topology. Together they power multi-tenant chargeback, executive reporting, and cost optimization.

Super Metric Architecture

Super Metrics are user-defined formulas that combine metrics from one or more objects into a single derived value. They execute at collection time (every 5 minutes by default) and are stored as first-class metrics — available in dashboards, alerts, views, and reports.

Formula Language & Operators:

---

Basic Operators: + - * / ()

Aggregation Functions:

  avg(${this, metric=cpu.usagemhz.average})    # Average across children
  sum(${this, metric=mem.usage.average})         # Sum across children
  max(${this, metric=disk.maxTotalLatency.latest})  # Peak across children
  count(${adaptertype=VMWARE, resourcekind=VirtualMachine, attribute=cpu|usage_average, depth=1, where="State = Powered On"})

Cross-Object References:

  ${adaptertype=VMWARE, resourcekind=ClusterComputeResource, attribute=cpu|usagemhz|average}
  ${this, metric=...}  # Relative to assigned object
  ${...depth=1}         # Direct children only
  ${...depth=-1}        # All descendants (recursive)

Conditional Logic:

if(${this, metric=cpu.ready.summation} > 1000, 1, 0) # Binary flag

Key Super Metric Design Patterns:

---

  1. Cluster CPU Contention Score:

Formula: avg(${this, metric=cpu.ready.summation, depth=1}) / avg(${this, metric=cpu.usagemhz.average, depth=1}) * 100

Assigns to: ClusterComputeResource

Purpose: Ratio of CPU ready time to active CPU — values > 5% indicate contention

  1. Per-VM Cost Allocation:

Formula: (${this, metric=cpu.usagemhz.average} 0.02) + (${this, metric=mem.consumed.average} / 1024 / 1024 0.005) + (${this, metric=virtualDisk.totalReadLatency.average} * 0.001)

Assigns to: VirtualMachine

Purpose: Dollar-cost per VM based on resource consumption (CPU rate + Memory rate + Storage rate)

  1. vSAN Health Score:

Formula: if(${this, metric=vsan.health.overall} == 0, 100, if(${this, metric=vsan.health.overall} == 1, 75, if(${this, metric=vsan.health.overall} == 2, 50, 25)))

Assigns to: ClusterComputeResource

Purpose: Normalize vSAN health enum to 0-100 score for trending

Custom Groups & Dynamic Membership:

---
Custom Groups organize objects by business logic rather than infrastructure topology:

Group Types:

Static: Manually selected objects (useful for fixed sets)

Dynamic: Membership rules evaluated at collection time

Rule criteria: Object type, metric threshold, property match, tag match, relationship
Example: "All VMs where cpu.usage > 80% AND tag contains 'production'"

Use Cases:

  • Application mapping: Group all VMs/hosts/datastores for "ERP System" across clusters
  • Cost center allocation: Group by department tag for chargeback
  • Compliance scope: Group PCI-relevant workloads for audit reporting
  • Tier classification: Gold/Silver/Bronze based on SLA tags

Multi-Tenancy with Custom Groups:

Each tenant gets a custom group scoping their infrastructure view

Dashboards, alerts, and reports are filtered to the group scope

RBAC: assign group-level read access to tenant admins

Chargeback: super metrics applied at group level calculate per-tenant costs

Best Practices for Super Metric Design:

---

  1. Always test formulas on a small group before deploying to production objects
  2. Avoid deep recursion (depth=-1) on large hierarchies — performance degrades with 10,000+ objects
  3. Use meaningful names: "cluster_cpu_contention_pct" not "sm_001"
  4. Document the formula purpose and expected range in the description field
  5. Set appropriate alert thresholds on super metrics — they behave like any other metric
  6. Export super metric definitions as XML for version control and migration between environments

Key Takeaways

  • Super Metrics combine raw metrics into business KPIs using formula language with aggregation functions (avg/sum/max/count) and cross-object references (depth, where clauses).
  • Custom Groups enable business-context organization (application, cost center, compliance scope) with static or dynamic membership rules evaluated at collection time.
  • Multi-tenancy pattern: Custom Group per tenant + RBAC scoping + Super Metrics for chargeback = tenant-specific dashboards and cost allocation.
  • Super Metric design: test on small groups first, avoid deep recursion on large hierarchies, use meaningful names, document formula purpose.

Dashboards, Views & Executive Reporting#

VCF Operations dashboards provide real-time and historical visualization of infrastructure health, performance, and capacity. Views and reports enable structured data export for compliance, capacity planning, and executive communication.

Dashboard Architecture:

---

Dashboard Components:

Widgets: Individual visualization elements (chart, heatmap, scoreboard, topology map, etc.)
Interactions: Cross-widget linking (click a host in a list, other widgets filter to that host)

Self-Provider: Widget has its own data source (static metric selection)

External-Provider: Widget receives data from another widget's selection (interactive drill-down)

Widget Types & Use Cases:

  • Metric Chart: Time-series line/bar chart for trending (cpu.usage over 7 days)
  • Scoreboard: Single KPI value with threshold coloring (green/yellow/red)
  • Heat Map: Grid showing relative metric intensity across objects (VM density per cluster)
  • Top-N: Ranked list of objects by a metric (Top 10 VMs by CPU usage)
  Topology Map: Visual relationship map (Cluster → Host → VM → Datastore)

Distribution: Histogram showing metric distribution (VM CPU allocation spread)
Object Relationship: Parent-child traversal for root cause analysis

Dashboard Design Patterns for VCDX:

  • ---
  • Pattern 1: Executive NOC Dashboard (Non-Technical Audience)
  • Row 1: Scoreboard widgets — Overall Health (green/yellow/red), Total VMs, Active Alerts
  • Row 2: Heat map — Cluster utilization (CPU + Memory + Storage)
  • Row 3: Trend chart — Week-over-week workload growth
  • Key principle: No technical jargon, traffic-light colors, big numbers

Pattern 2: Operations Drill-Down Dashboard

Left panel: Object list (hosts, filtered by alert status)

Right panel: Metric chart (auto-updates based on left panel selection)

Bottom panel: Alert list (filtered to selected object)

Key principle: Click-to-drill from summary to detail

Pattern 3: Capacity Planning Dashboard

  • Widget 1: Time Remaining (days until CPU/Memory/Storage exhaustion)
  • Widget 2: Demand vs Capacity trend lines
  • Widget 3: What-If scenarios (add 50 VMs — when does capacity run out?)
  • Widget 4: Rightsizing recommendations (oversized VMs wasting capacity)
  • Key principle: Forward-looking, actionable recommendations

Views & Reports:

---

Views are structured data queries that return tabular results:

  • List View: All VMs with columns (Name, CPU Usage, Memory Usage, Disk Latency, Power State)
  • Trend View: Selected metrics over time for selected objects
  • Distribution View: Statistical spread of a metric across objects

Reports combine views into formatted documents:

Templates: Pre-built (Cluster Utilization, VM Rightsizing, Chargeback)

Custom: User-defined combining multiple views with headers/footers

Scheduling: Daily/weekly/monthly auto-generation and email delivery
Formats: PDF, CSV, Excel

Report Design for Compliance:

  • PCI DSS: Monthly report on security-related metrics (failed logins, config changes)
  • SOX: Quarterly capacity and change audit report
  • Internal SLA: Weekly uptime and performance report per service tier

Dashboard Best Practices:

---

1. Use interactions (self-provider → external-provider) for drill-down — don't create separate dashboards
  1. Limit widgets per dashboard to 8-12 — too many causes performance issues
  2. Use consistent color schemes: green=good, yellow=warning, red=critical (match alert definitions)
  3. Include time context: always show the time range in the dashboard title or header
  4. Export dashboard definitions as JSON for version control and migration

Key Takeaways

  • Dashboard architecture: Self-Provider widgets (static data source) vs External-Provider (dynamic, fed by other widget selections) enable interactive drill-down without separate dashboards.
  • Three dashboard design patterns: Executive NOC (traffic-light, big numbers), Operations Drill-Down (click-to-detail), Capacity Planning (forward-looking, what-if scenarios).
  • Views query structured data (list, trend, distribution); Reports combine views into scheduled documents (PDF/CSV/Excel) for compliance and executive reporting.
  • Dashboard limits: 8-12 widgets max per dashboard for performance; use interactions for drill-down; export definitions as JSON for versioning.

Alert Policies, Symptom Definitions & Automated Remediation#

VCF Operations uses a multi-level alerting framework: Symptom Definitions (individual metric conditions) combine into Alert Definitions (correlated conditions), which trigger Notifications and optionally Automated Remediation actions. Understanding this hierarchy is critical for designing production-grade monitoring.

Alerting Architecture — Symptom → Alert → Notification Pipeline:
---
LEVEL 1: SYMPTOM DEFINITION
  What: A single condition on a metric or property
  Example: "CPU usage > 90% for 3 consecutive collection cycles"
  Types:
    Metric/Property: threshold on a numeric metric (static or dynamic)
    Metric Event: rapid change detection (spike >50% in 5 minutes)
    Log Event: pattern match in log messages
    Health: object health state change (green → yellow)

Wait Cycles: # of consecutive violations before symptom fires (reduces flapping)
Cancel Cycles: # of consecutive clear readings before symptom clears

Dynamic Thresholds (DT):

VCF Operations learns normal behavior per object over 7+ days

  • DT adjusts thresholds based on historical patterns (business hours vs night)
  • Example: CPU usage normal at 70% during business hours but only 20% at night
  • A spike to 50% at 3 AM triggers DT alert, but 70% at 2 PM does not
  • DT sensitivity: 1 (very sensitive, more alerts) to 5 (less sensitive)

LEVEL 2: ALERT DEFINITION

What: One or more symptoms combined with logic (AND/OR)

  • Example: "CPU usage > 90% AND Memory usage > 85% AND Disk latency > 20ms"
  • Criticality: Info, Warning, Immediate, Critical
  • Impact: Health, Risk, Efficiency (determines badge type on objects)
  • Type: Application, Virtualization, Hardware, Storage, Network
  • Correlation: Multiple symptoms must fire simultaneously to trigger the alert
  • This reduces alert noise by ensuring multiple conditions confirm a real issue

LEVEL 3: NOTIFICATION

What: Action taken when alert fires/clears

Channels:

Email: SMTP integration with customizable templates

  • Webhook: REST callback to external systems (PagerDuty, ServiceNow, Slack)
  • SNMP Trap: v1/v2c/v3 to network management systems
  • Log File: Write to local log for external collection
  Escalation: Chain notifications (email immediately → webhook after 15 min → SNMP after 30 min)

Suppression: Maintenance windows to silence alerts during planned work

LEVEL 4: AUTOMATED REMEDIATION

What: Action executed when alert fires (requires explicit enable)

Supported actions via VCF Operations + vSphere integration:

  • Power on/off VM
  • Increase VM CPU/Memory (hot-add if enabled)
  • vMotion VM to less-loaded host
  • Snapshot VM before risky change
  • Run custom script via SSH or REST API

Safety controls:

  • Maximum actions per alert per hour (default: 1)
  • Require approval before execution (optional)
  • Dry-run mode: log what WOULD happen without executing
  • Cooldown period between consecutive actions

Alert Design Anti-Patterns to Avoid:

---

1. Alert storms: Too many low-threshold alerts → operations team ignores all alerts
   Fix: Use dynamic thresholds, set appropriate wait cycles (3+), suppress during maintenance
2. Symptomless alerts: Alert on single metric without correlation → high false positive rate
   Fix: Always use multi-symptom alerts (CPU + Memory + Disk = real problem, not noise)
3. Missing escalation: All alerts go to same channel → critical mixed with informational
   Fix: Design tiered notification chains: Info → log only; Warning → email; Critical → webhook + email
4. No remediation policy: All alerts require manual intervention → slow MTTR

Fix: Automate safe actions (vMotion, rightsizing) for well-understood scenarios; reserve manual for complex issues

Key Takeaways

  • Alerting pipeline: Symptom Definition (single condition) → Alert Definition (correlated symptoms with AND/OR logic) → Notification (email/webhook/SNMP) → Optional Automated Remediation (vMotion, resize, script).
  • Dynamic Thresholds learn per-object baselines over 7+ days, adapting to business-hour vs off-hour patterns — reduces false positives from static thresholds.
  • Multi-symptom correlation is key: alert on 'CPU > 90% AND Memory > 85%' is far more meaningful than alerting on CPU alone.
  • Alert design: use wait cycles (3+) to reduce flapping, tiered notifications for escalation, maintenance windows for suppression, and dry-run remediation before production automation.

Capacity Planning, Rightsizing & What-If Analysis#

Capacity planning in VCF Operations uses predictive analytics based on historical trends to forecast resource exhaustion, identify rightsizing opportunities, and model what-if scenarios for infrastructure changes. This is one of the highest-value VCAP Ops topics — it directly impacts CapEx decisions and is a frequent VCDX design defense topic.

Capacity Planning Model:

---

VCF Operations calculates three capacity dimensions for each cluster:

  1. Time Remaining: Days until a resource (CPU, Memory, Storage) reaches the demand threshold

Calculation: Linear regression on historical usage trend, projected forward to intersection with capacity limit

Demand threshold: Configurable (default 100%, typical production: 80-85% to maintain headroom)

Buffer: Additional headroom for HA, maintenance, burst (e.g., N+1 = one host worth of capacity reserved)

  1. Recommended Size: Optimal resource allocation based on actual usage patterns

Rightsizing for VMs:

     Oversized: VM allocated 8 vCPU but avg usage 15% → recommend 2 vCPU (saves 6 vCPU)
     Undersized: VM allocated 2 vCPU but avg usage 95% + high ready time → recommend 4 vCPU
     Idle: VM powered on but <1% CPU/Memory/Network for 30+ days → candidate for decommission
   Rightsizing for clusters:
     Over-provisioned: cluster 30% utilized → candidate for consolidation
     Under-provisioned: cluster 85%+ utilized → add hosts or redistribute workloads
  1. Capacity Remaining: Headroom in each resource after accounting for HA, buffer, and committed capacity

Formula: Total Capacity - HA Reserve - Buffer - Current Demand = Remaining
HA Reserve: typically 1 host worth of CPU/Memory (N+1) or 2 hosts (N+2 for critical workloads)

What-If Analysis:

---
What-If scenarios model the impact of hypothetical changes before committing resources:

Scenario Types:

Add Workloads: "What if we deploy 50 new VMs with 4 vCPU / 16 GB RAM each?"

    → Shows: new Time Remaining, resource utilization after addition
  Remove Workloads: "What if we decommission the legacy ERP cluster?"
    → Shows: capacity freed, potential consolidation opportunity
  Add Hosts: "What if we add 4 new hosts to the cluster?"
    → Shows: extended Time Remaining, new utilization percentage
  Hardware Refresh: "What if we replace 10 old hosts with 5 new (higher-spec) hosts?"
    → Shows: capacity comparison (old vs new), consolidation ratio

What-If Execution:

  1. Select target cluster
  2. Choose scenario type
  3. Input parameters (VM count, resource profile, host spec)
  4. Run analysis → VCF Operations recalculates capacity projections
  1. Compare results: current state vs. proposed state side-by-side

Rightsizing Workflow:

---
Step 1: Identify candidates

  VCF Operations → Optimize → Rightsizing → filter by oversized/undersized/idle

Default: analyzes last 30 days of metric data

Step 2: Review recommendations

  • Each VM shows: Current allocation, Recommended allocation, Expected savings
  • Conservative mode: recommends based on peak usage (safer)
  • Aggressive mode: recommends based on 95th percentile (more savings, higher risk)

Step 3: Apply changes

Manual: Admin reviews and applies in vSphere Client

Automated: VCF Operations + Action Adapter can auto-resize (hot-add if supported)
Scheduled: Apply rightsizing during maintenance window

Step 4: Monitor post-change

After rightsizing, monitor for 7-14 days: CPU ready time increase? Memory ballooning? Disk latency spike?
If degradation detected: revert and adjust recommendation sensitivity

Cost Optimization with Chargeback/Showback:

  • ---
  • Chargeback: actual billing to business units based on resource consumption
  • Showback: informational report without actual billing (transparency without enforcement)

Cost Model Configuration:

  • CPU Rate: $0.02 per MHz per month
  • Memory Rate: $0.005 per GB per month
  • Storage Rate: $0.10 per GB per month
  • Network Rate: $0.001 per Mbps per month

Cost Report Output:

Per-VM: Monthly cost based on actual consumption × rates

Per-Group: Aggregated cost for custom groups (departments, applications)

Per-Cluster: Infrastructure cost allocated proportionally

VCDX Design Consideration: Capacity planning parameters significantly impact procurement cycles. Setting demand threshold too high (95%) means late warning — procurement lead time for new hardware is typically 8-16 weeks. Setting too low (60%) means premature investment. The sweet spot is 80% demand threshold with 12-week Time Remaining alarm.

Key Takeaways

  • Three capacity dimensions: Time Remaining (days to exhaustion via linear regression), Recommended Size (rightsizing based on actual usage), Capacity Remaining (headroom after HA/buffer reserves).
  • What-If analysis models hypothetical changes (add/remove workloads, add hosts, hardware refresh) before committing resources — key for CapEx planning.
  • Rightsizing workflow: Identify candidates (oversized/undersized/idle) → Review recommendations (conservative vs aggressive) → Apply changes → Monitor post-change for 7-14 days.
  • Chargeback/showback: cost model with per-resource rates (CPU/Memory/Storage/Network) → per-VM/group/cluster cost reports. Set demand threshold at 80% with 12-week Time Remaining alarm to align with hardware procurement cycles.

VCF Operations API, Extensibility & Multi-Cloud Observability#

VCF Operations exposes a comprehensive REST API for automation, integration, and extensibility. The API enables programmatic access to all operations functions — metrics, alerts, dashboards, capacity, and configuration — enabling infrastructure-as-code patterns for operations management.

REST API Architecture:

  • ---
  • Base URL: https://<vcf-ops-fqdn>/suite-api/api/
  • Authentication: Token-based (acquire via POST /auth/token/acquire with username/password)
  • Response Format: JSON (default), XML (via Accept header)

API Categories:

  • /resources: Query objects (hosts, VMs, clusters) with metric data
  • /alerts: Read/update/cancel alerts, manage alert definitions
  • /supermetrics: CRUD operations on super metric definitions
  • /dashboards: Export/import dashboard definitions
  • /capacity: Access capacity planning projections and what-if results
  • /reports: Generate and download reports programmatically
  • /custom-groups: Manage custom group membership and rules
  • /recommendations: Access rightsizing and optimization recommendations

Common API Workflows:

---

  1. Bulk Metric Export (for external analytics):

GET /resources?resourceKind=VirtualMachine&pageSize=1000

   → Returns list of VM resource IDs
   POST /resources/stats/query (with resource IDs + metric keys + time range)
   → Returns time-series metric data for all VMs

Use case: Feed VCF Operations data into Splunk, Elasticsearch, or a data lake

  1. Alert-to-Ticket Integration:

GET /alerts?status=ACTIVE&criticality=CRITICAL

   → Returns active critical alerts with resource details

POST to ServiceNow API to create incident ticket

PUT /alerts/{alertId} (update owner and notes)

Use case: Bi-directional sync between VCF Operations and ITSM

  1. Dashboard-as-Code:
   GET /dashboards/{id}/export → JSON definition of dashboard
   Store in Git repository for version control
   POST /dashboards/import → Deploy dashboard to new environment
   Use case: Promote dashboards from dev → staging → production
  1. Automated Rightsizing Pipeline:

GET /recommendations?type=RIGHTSIZING&resourceKind=VirtualMachine

   → Returns list of rightsizing recommendations with current/proposed values

Filter by confidence score > 80%

Apply via vSphere API (reconfigure VM with new CPU/Memory values)

POST back acknowledgment to VCF Operations

Management Packs (Extensibility):

---

Management Packs extend VCF Operations to monitor non-VMware systems:

Built-in Packs:

  vCenter: vSphere infrastructure (hosts, VMs, clusters, datastores)
  vSAN: vSAN health, performance, capacity
  NSX: Network virtualization components

VCF: Cloud Foundation lifecycle and health

Marketplace Packs (VMware-authored):

AWS: EC2, EBS, RDS, S3, Lambda monitoring

Azure: Virtual Machines, Storage, SQL, AKS monitoring

Kubernetes: Cluster, node, pod, container metrics

Dell/HP/Lenovo: Hardware health via iDRAC/iLO/XCC SNMP

NetApp/Pure/Dell Storage: SAN/NAS performance and capacity

Custom Packs (User-authored):

Built using Management Pack SDK

  • Define: adapter (data collection), resource kinds, dashboards, alerts
  • Package as PAK file for distribution
  • Use case: Monitor custom applications or proprietary systems

Multi-Cloud Observability:

---

VCF Operations provides a single pane of glass across:

On-premises: vSphere/vSAN/NSX via native adapters

AWS: via AWS management pack (CloudWatch metrics + Config)

Azure: via Azure management pack (Monitor metrics + Resource Graph)

GCP: via GCP management pack (Stackdriver metrics)

Kubernetes: via Kubernetes management pack (kube-state-metrics + cAdvisor)

Cross-Cloud Correlation:

Application spans on-prem VMs + AWS EC2 + K8s pods

Custom group contains all objects regardless of cloud

Super metrics aggregate CPU/Memory/Cost across clouds

Single dashboard shows application health across all environments

API Security Best Practices:

---

  1. Use service accounts for API integrations (not personal admin accounts)
  2. Token expiry: default 6 hours; for long-running integrations, implement token refresh
  3. Rate limiting: API enforces 100 requests/second per session; batch queries where possible
  4. TLS: All API calls must use HTTPS; verify certificates in production integrations
  5. Audit: All API calls are logged in VCF Operations audit trail — used for compliance reporting

Key Takeaways

  • REST API categories: /resources (objects+metrics), /alerts (CRUD), /supermetrics, /dashboards (export/import), /capacity, /reports, /recommendations (rightsizing). Token-based auth with 6-hour default expiry.
  • Key API workflows: bulk metric export to SIEM/data lake, alert-to-ticket integration with ITSM, dashboard-as-code via Git, automated rightsizing pipeline.
  • Management Packs extend monitoring to AWS/Azure/GCP/K8s/hardware/storage — enabling multi-cloud observability from a single platform.
  • Multi-cloud correlation: custom groups span clouds, super metrics aggregate cross-cloud KPIs, single dashboards show application health across all environments.

Exam Mapping: 3V0-22.25 — Advanced VCF 9.0 Operations

  • See Advanced VCF 9.0 Operations exam blueprint for detailed objectives

Labs in This Section

Lab: Design Super Metrics & Custom Groups for Cost Optimization

VCF 9.0Advanced⏱ 120 min

Lab: Build Executive KPI Dashboard with Drill-Down

VCF 9.0Advanced⏱ 90 min

Lab: Design Production Alert Policy with Escalation

VCF 9.0Advanced⏱ 120 min

Lab: Build Capacity Planning Analysis & Rightsizing Report

VCF 9.0Advanced⏱ 90 min

Lab: Plan & Execute VCF Operations Cluster Upgrade

VCF 9.0Advanced⏱ 120 min
📝 Quiz (65)
🃏 Flashcards (79)

📝 Quiz — VCAP Advanced Operations

0/65 correct

Architecture

Q1
A remote datacenter has limited WAN and must forward metrics locally before sending to the analytics cluster. Which component do you deploy?
  • Second analytics cluster
  • Remote Collector or Cloud Proxy
  • Additional vCenter
  • SDDC Manager clone
Remote Collector/Cloud Proxy collects data at remote sites with limited WAN and forwards to the analytics cluster, minimizing bandwidth usage. A second analytics cluster is overkill. Additional vCenter doesn't help with monitoring. SDDC Manager clones aren't a thing.
Q2
Which deployment model provides horizontal scale for metric ingest and analytics?
  • Standalone single node
  • HA pair
  • Multi-node analytics cluster
  • Witness-only node
Multi-node analytics clusters provide horizontal scale for metric ingest and analytics, distributing load across nodes. Standalone single node doesn't scale. HA pair provides redundancy but limited scale. Witness nodes support quorum, not analytics.
Q3
An 'object' in VCF Operations is best defined as:
  • A dashboard widget
  • A monitored entity exposing metrics, properties, and relationships
  • A syslog source only
  • A disk group
An 'object' in VCF Operations is a monitored entity exposing metrics, properties, and relationships — VMs, hosts, datastores, etc. form the object model. It's not a dashboard widget, syslog source, or disk group.
Q4
VCF Operations for Logs stores long-term archived logs using:
  • Only local disk
  • NFS or object-storage archive targets
  • vSAN File Services exclusively
  • Email archive
VCF Operations for Logs uses NFS or object-storage archive targets for long-term log retention. Local disk alone lacks scalability. vSAN File Services isn't exclusively used. Email archive is not a log storage mechanism.
Q5
VCF Identity Broker centralizes authentication across:
  • VCF Operations, Automation, and related services via AD/LDAP/SAML/OIDC
  • Only vCenter
  • Only NSX Manager
  • Only SDDC Manager
VCF Identity Broker centralizes authentication across VCF Operations, Automation, and related services via AD/LDAP/SAML/OIDC integration. It's not limited to only vCenter, NSX Manager, or SDDC Manager.

Capacity Management

Performance Analysis

Automation and Remediation

Custom Content and Compliance

🃏 Flashcards — VCAP Advanced Operations

79 cards
Card 1 of 79
Analytics Node
A VCF Operations cluster role running the metric, property, and event processing pipeline, backing the analytics UI. Multiple analytics nodes partition data for scale. They host adapters, collectors-of-last-resort, and alerting.

Labs in this section

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.