VCAP-level operations architecture, advanced monitoring dashboards, super metrics & custom groups, alert policies & automated remediation, capacity planning & rightsizing, API-driven extensibility, and multi-cloud observability for unified VMware Cloud Foundation fleet management.
3V0-22.25
VCAP
60
Questions
135m
Duration
300/500
Pass Score
32
Objectives
Exam Blueprint Weights
Section titles, groupings and weights below are VCDX Academy study groupings, NOT the official Broadcom blueprint structure. Broadcom publishes no section weights. Always cross-check the official exam guide. Official exam guide ↗
Section 1 — Architecture
~15%
Section 2 — Capacity Management
~20%
Section 3 — Performance Analysis
~20%
Section 4 — Automation and Remediation
~20%
Section 5 — Custom Content and Complianc
~25%
High Weight
Version Evolution
VCAP Operations covers advanced day-2 operations including performance optimization, capacity planning, and automated remediation. VCF 9.0 shifted operations tooling from Aria Operations to VCF Operations — a significant change in UI, API, and workflow. The exam tests scenario-based troubleshooting and optimization at scale.
VCF Operations (formerly vRealize Operations) is a distributed, horizontally scalable analytics platform. Understanding node roles, data architecture, and sharding patterns is essential for production deployments at scale.
Super Metrics enable computed KPIs that aggregate raw metrics across objects. Custom Groups organize infrastructure by business context rather than physical topology. Together they power multi-tenant chargeback, executive reporting, and cost optimization.
Super Metric Architecture
Super Metrics are user-defined formulas that combine metrics from one or more objects into a single derived value. They execute at collection time (every 5 minutes by default) and are stored as first-class metrics — available in dashboards, alerts, views, and reports.
Formula Language & Operators:
---
Basic Operators: + - * / ()
Aggregation Functions:
avg(${this, metric=cpu.usagemhz.average}) # Average across children
sum(${this, metric=mem.usage.average}) # Sum across children
max(${this, metric=disk.maxTotalLatency.latest}) # Peak across children
count(${adaptertype=VMWARE, resourcekind=VirtualMachine, attribute=cpu|usage_average, depth=1, where="State = Powered On"})
Cross-Object References:
${adaptertype=VMWARE, resourcekind=ClusterComputeResource, attribute=cpu|usagemhz|average}
${this, metric=...} # Relative to assigned object
${...depth=1} # Direct children only
${...depth=-1} # All descendants (recursive)
Conditional Logic:
if(${this, metric=cpu.ready.summation} > 1000, 1, 0) # Binary flag
Purpose: Normalize vSAN health enum to 0-100 score for trending
Custom Groups & Dynamic Membership:
--- Custom Groups organize objects by business logic rather than infrastructure topology:
Group Types:
Static: Manually selected objects (useful for fixed sets)
Dynamic: Membership rules evaluated at collection time
Rule criteria: Object type, metric threshold, property match, tag match, relationship Example: "All VMs where cpu.usage > 80% AND tag contains 'production'"
Use Cases:
Application mapping: Group all VMs/hosts/datastores for "ERP System" across clusters
Cost center allocation: Group by department tag for chargeback
Compliance scope: Group PCI-relevant workloads for audit reporting
Tier classification: Gold/Silver/Bronze based on SLA tags
Multi-Tenancy with Custom Groups:
Each tenant gets a custom group scoping their infrastructure view
Dashboards, alerts, and reports are filtered to the group scope
RBAC: assign group-level read access to tenant admins
Chargeback: super metrics applied at group level calculate per-tenant costs
Best Practices for Super Metric Design:
---
Always test formulas on a small group before deploying to production objects
Avoid deep recursion (depth=-1) on large hierarchies — performance degrades with 10,000+ objects
Use meaningful names: "cluster_cpu_contention_pct" not "sm_001"
Document the formula purpose and expected range in the description field
Set appropriate alert thresholds on super metrics — they behave like any other metric
Export super metric definitions as XML for version control and migration between environments
Key Takeaways
Super Metrics combine raw metrics into business KPIs using formula language with aggregation functions (avg/sum/max/count) and cross-object references (depth, where clauses).
Custom Groups enable business-context organization (application, cost center, compliance scope) with static or dynamic membership rules evaluated at collection time.
Multi-tenancy pattern: Custom Group per tenant + RBAC scoping + Super Metrics for chargeback = tenant-specific dashboards and cost allocation.
Super Metric design: test on small groups first, avoid deep recursion on large hierarchies, use meaningful names, document formula purpose.
VCF Operations dashboards provide real-time and historical visualization of infrastructure health, performance, and capacity. Views and reports enable structured data export for compliance, capacity planning, and executive communication.
Dashboard Architecture:
---
Dashboard Components:
Widgets: Individual visualization elements (chart, heatmap, scoreboard, topology map, etc.) Interactions: Cross-widget linking (click a host in a list, other widgets filter to that host)
Self-Provider: Widget has its own data source (static metric selection)
External-Provider: Widget receives data from another widget's selection (interactive drill-down)
Widget Types & Use Cases:
Metric Chart: Time-series line/bar chart for trending (cpu.usage over 7 days)
Scoreboard: Single KPI value with threshold coloring (green/yellow/red)
Heat Map: Grid showing relative metric intensity across objects (VM density per cluster)
Top-N: Ranked list of objects by a metric (Top 10 VMs by CPU usage)
Internal SLA: Weekly uptime and performance report per service tier
Dashboard Best Practices:
---
1. Use interactions (self-provider → external-provider) for drill-down — don't create separate dashboards
Limit widgets per dashboard to 8-12 — too many causes performance issues
Use consistent color schemes: green=good, yellow=warning, red=critical (match alert definitions)
Include time context: always show the time range in the dashboard title or header
Export dashboard definitions as JSON for version control and migration
Key Takeaways
Dashboard architecture: Self-Provider widgets (static data source) vs External-Provider (dynamic, fed by other widget selections) enable interactive drill-down without separate dashboards.
Three dashboard design patterns: Executive NOC (traffic-light, big numbers), Operations Drill-Down (click-to-detail), Capacity Planning (forward-looking, what-if scenarios).
Views query structured data (list, trend, distribution); Reports combine views into scheduled documents (PDF/CSV/Excel) for compliance and executive reporting.
Dashboard limits: 8-12 widgets max per dashboard for performance; use interactions for drill-down; export definitions as JSON for versioning.
VCF Operations uses a multi-level alerting framework: Symptom Definitions (individual metric conditions) combine into Alert Definitions (correlated conditions), which trigger Notifications and optionally Automated Remediation actions. Understanding this hierarchy is critical for designing production-grade monitoring.
Alerting Architecture — Symptom → Alert → Notification Pipeline:
---
LEVEL 1: SYMPTOM DEFINITION
What: A single condition on a metric or property
Example: "CPU usage > 90% for 3 consecutive collection cycles"
Types:
Metric/Property: threshold on a numeric metric (static or dynamic)
Metric Event: rapid change detection (spike >50% in 5 minutes)
Log Event: pattern match in log messages
Health: object health state change (green → yellow)
Wait Cycles: # of consecutive violations before symptom fires (reduces flapping) Cancel Cycles: # of consecutive clear readings before symptom clears
Dynamic Thresholds (DT):
VCF Operations learns normal behavior per object over 7+ days
DT adjusts thresholds based on historical patterns (business hours vs night)
Example: CPU usage normal at 70% during business hours but only 20% at night
A spike to 50% at 3 AM triggers DT alert, but 70% at 2 PM does not
DT sensitivity: 1 (very sensitive, more alerts) to 5 (less sensitive)
LEVEL 2: ALERT DEFINITION
What: One or more symptoms combined with logic (AND/OR)
Example: "CPU usage > 90% AND Memory usage > 85% AND Disk latency > 20ms"
Criticality: Info, Warning, Immediate, Critical
Impact: Health, Risk, Efficiency (determines badge type on objects)
Correlation: Multiple symptoms must fire simultaneously to trigger the alert
This reduces alert noise by ensuring multiple conditions confirm a real issue
LEVEL 3: NOTIFICATION
What: Action taken when alert fires/clears
Channels:
Email: SMTP integration with customizable templates
Webhook: REST callback to external systems (PagerDuty, ServiceNow, Slack)
SNMP Trap: v1/v2c/v3 to network management systems
Log File: Write to local log for external collection
Escalation: Chain notifications (email immediately → webhook after 15 min → SNMP after 30 min)
Suppression: Maintenance windows to silence alerts during planned work
LEVEL 4: AUTOMATED REMEDIATION
What: Action executed when alert fires (requires explicit enable)
Supported actions via VCF Operations + vSphere integration:
Power on/off VM
Increase VM CPU/Memory (hot-add if enabled)
vMotion VM to less-loaded host
Snapshot VM before risky change
Run custom script via SSH or REST API
Safety controls:
Maximum actions per alert per hour (default: 1)
Require approval before execution (optional)
Dry-run mode: log what WOULD happen without executing
Cooldown period between consecutive actions
Alert Design Anti-Patterns to Avoid:
---
1. Alert storms: Too many low-threshold alerts → operations team ignores all alerts
Fix: Use dynamic thresholds, set appropriate wait cycles (3+), suppress during maintenance
2. Symptomless alerts: Alert on single metric without correlation → high false positive rate
Fix: Always use multi-symptom alerts (CPU + Memory + Disk = real problem, not noise)
3. Missing escalation: All alerts go to same channel → critical mixed with informational
Fix: Design tiered notification chains: Info → log only; Warning → email; Critical → webhook + email
4. No remediation policy: All alerts require manual intervention → slow MTTR
Fix: Automate safe actions (vMotion, rightsizing) for well-understood scenarios; reserve manual for complex issues
Dynamic Thresholds learn per-object baselines over 7+ days, adapting to business-hour vs off-hour patterns — reduces false positives from static thresholds.
Multi-symptom correlation is key: alert on 'CPU > 90% AND Memory > 85%' is far more meaningful than alerting on CPU alone.
Alert design: use wait cycles (3+) to reduce flapping, tiered notifications for escalation, maintenance windows for suppression, and dry-run remediation before production automation.
Capacity planning in VCF Operations uses predictive analytics based on historical trends to forecast resource exhaustion, identify rightsizing opportunities, and model what-if scenarios for infrastructure changes. This is one of the highest-value VCAP Ops topics — it directly impacts CapEx decisions and is a frequent VCDX design defense topic.
Capacity Planning Model:
---
VCF Operations calculates three capacity dimensions for each cluster:
Time Remaining: Days until a resource (CPU, Memory, Storage) reaches the demand threshold
Calculation: Linear regression on historical usage trend, projected forward to intersection with capacity limit
Buffer: Additional headroom for HA, maintenance, burst (e.g., N+1 = one host worth of capacity reserved)
Recommended Size: Optimal resource allocation based on actual usage patterns
Rightsizing for VMs:
Oversized: VM allocated 8 vCPU but avg usage 15% → recommend 2 vCPU (saves 6 vCPU)
Undersized: VM allocated 2 vCPU but avg usage 95% + high ready time → recommend 4 vCPU
Idle: VM powered on but <1% CPU/Memory/Network for 30+ days → candidate for decommission
Rightsizing for clusters:
Over-provisioned: cluster 30% utilized → candidate for consolidation
Under-provisioned: cluster 85%+ utilized → add hosts or redistribute workloads
Capacity Remaining: Headroom in each resource after accounting for HA, buffer, and committed capacity
Formula: Total Capacity - HA Reserve - Buffer - Current Demand = Remaining HA Reserve: typically 1 host worth of CPU/Memory (N+1) or 2 hosts (N+2 for critical workloads)
What-If Analysis:
--- What-If scenarios model the impact of hypothetical changes before committing resources:
Scenario Types:
Add Workloads: "What if we deploy 50 new VMs with 4 vCPU / 16 GB RAM each?"
→ Shows: new Time Remaining, resource utilization after addition
Remove Workloads: "What if we decommission the legacy ERP cluster?"
→ Shows: capacity freed, potential consolidation opportunity
Add Hosts: "What if we add 4 new hosts to the cluster?"
→ Shows: extended Time Remaining, new utilization percentage
Hardware Refresh: "What if we replace 10 old hosts with 5 new (higher-spec) hosts?"
→ Shows: capacity comparison (old vs new), consolidation ratio
4. Run analysis → VCF Operations recalculates capacity projections
Compare results: current state vs. proposed state side-by-side
Rightsizing Workflow:
--- Step 1: Identify candidates
VCF Operations → Optimize → Rightsizing → filter by oversized/undersized/idle
Default: analyzes last 30 days of metric data
Step 2: Review recommendations
Each VM shows: Current allocation, Recommended allocation, Expected savings
Conservative mode: recommends based on peak usage (safer)
Aggressive mode: recommends based on 95th percentile (more savings, higher risk)
Step 3: Apply changes
Manual: Admin reviews and applies in vSphere Client
Automated: VCF Operations + Action Adapter can auto-resize (hot-add if supported) Scheduled: Apply rightsizing during maintenance window
Step 4: Monitor post-change
After rightsizing, monitor for 7-14 days: CPU ready time increase? Memory ballooning? Disk latency spike? If degradation detected: revert and adjust recommendation sensitivity
Cost Optimization with Chargeback/Showback:
---
Chargeback: actual billing to business units based on resource consumption
Showback: informational report without actual billing (transparency without enforcement)
Cost Model Configuration:
CPU Rate: $0.02 per MHz per month
Memory Rate: $0.005 per GB per month
Storage Rate: $0.10 per GB per month
Network Rate: $0.001 per Mbps per month
Cost Report Output:
Per-VM: Monthly cost based on actual consumption × rates
Per-Group: Aggregated cost for custom groups (departments, applications)
VCDX Design Consideration: Capacity planning parameters significantly impact procurement cycles. Setting demand threshold too high (95%) means late warning — procurement lead time for new hardware is typically 8-16 weeks. Setting too low (60%) means premature investment. The sweet spot is 80% demand threshold with 12-week Time Remaining alarm.
Key Takeaways
Three capacity dimensions: Time Remaining (days to exhaustion via linear regression), Recommended Size (rightsizing based on actual usage), Capacity Remaining (headroom after HA/buffer reserves).
What-If analysis models hypothetical changes (add/remove workloads, add hosts, hardware refresh) before committing resources — key for CapEx planning.
Rightsizing workflow: Identify candidates (oversized/undersized/idle) → Review recommendations (conservative vs aggressive) → Apply changes → Monitor post-change for 7-14 days.
Chargeback/showback: cost model with per-resource rates (CPU/Memory/Storage/Network) → per-VM/group/cluster cost reports. Set demand threshold at 80% with 12-week Time Remaining alarm to align with hardware procurement cycles.
VCF Operations exposes a comprehensive REST API for automation, integration, and extensibility. The API enables programmatic access to all operations functions — metrics, alerts, dashboards, capacity, and configuration — enabling infrastructure-as-code patterns for operations management.
REST API Architecture:
---
Base URL: https://<vcf-ops-fqdn>/suite-api/api/
Authentication: Token-based (acquire via POST /auth/token/acquire with username/password)
Response Format: JSON (default), XML (via Accept header)
API Categories:
/resources: Query objects (hosts, VMs, clusters) with metric data
/supermetrics: CRUD operations on super metric definitions
/dashboards: Export/import dashboard definitions
/capacity: Access capacity planning projections and what-if results
/reports: Generate and download reports programmatically
/custom-groups: Manage custom group membership and rules
/recommendations: Access rightsizing and optimization recommendations
Common API Workflows:
---
Bulk Metric Export (for external analytics):
GET /resources?resourceKind=VirtualMachine&pageSize=1000
→ Returns list of VM resource IDs
POST /resources/stats/query (with resource IDs + metric keys + time range)
→ Returns time-series metric data for all VMs
Use case: Feed VCF Operations data into Splunk, Elasticsearch, or a data lake
Alert-to-Ticket Integration:
GET /alerts?status=ACTIVE&criticality=CRITICAL
→ Returns active critical alerts with resource details
POST to ServiceNow API to create incident ticket
PUT /alerts/{alertId} (update owner and notes)
Use case: Bi-directional sync between VCF Operations and ITSM
Dashboard-as-Code:
GET /dashboards/{id}/export → JSON definition of dashboard
Store in Git repository for version control
POST /dashboards/import → Deploy dashboard to new environment
Use case: Promote dashboards from dev → staging → production
Automated Rightsizing Pipeline:
GET /recommendations?type=RIGHTSIZING&resourceKind=VirtualMachine
→ Returns list of rightsizing recommendations with current/proposed values
Filter by confidence score > 80%
Apply via vSphere API (reconfigure VM with new CPU/Memory values)
POST back acknowledgment to VCF Operations
Management Packs (Extensibility):
---
Management Packs extend VCF Operations to monitor non-VMware systems:
Rate limiting: API enforces 100 requests/second per session; batch queries where possible
TLS: All API calls must use HTTPS; verify certificates in production integrations
Audit: All API calls are logged in VCF Operations audit trail — used for compliance reporting
Key Takeaways
REST API categories: /resources (objects+metrics), /alerts (CRUD), /supermetrics, /dashboards (export/import), /capacity, /reports, /recommendations (rightsizing). Token-based auth with 6-hour default expiry.
Key API workflows: bulk metric export to SIEM/data lake, alert-to-ticket integration with ITSM, dashboard-as-code via Git, automated rightsizing pipeline.
Management Packs extend monitoring to AWS/Azure/GCP/K8s/hardware/storage — enabling multi-cloud observability from a single platform.
Multi-cloud correlation: custom groups span clouds, super metrics aggregate cross-cloud KPIs, single dashboards show application health across all environments.
Remote Collector/Cloud Proxy collects data at remote sites with limited WAN and forwards to the analytics cluster, minimizing bandwidth usage. A second analytics cluster is overkill. Additional vCenter doesn't help with monitoring. SDDC Manager clones aren't a thing.
Multi-node analytics clusters provide horizontal scale for metric ingest and analytics, distributing load across nodes. Standalone single node doesn't scale. HA pair provides redundancy but limited scale. Witness nodes support quorum, not analytics.
An 'object' in VCF Operations is a monitored entity exposing metrics, properties, and relationships — VMs, hosts, datastores, etc. form the object model. It's not a dashboard widget, syslog source, or disk group.
VCF Operations for Logs uses NFS or object-storage archive targets for long-term log retention. Local disk alone lacks scalability. vSAN File Services isn't exclusively used. Email archive is not a log storage mechanism.
VCF Identity Broker centralizes authentication across VCF Operations, Automation, and related services via AD/LDAP/SAML/OIDC integration. It's not limited to only vCenter, NSX Manager, or SDDC Manager.
Cloud Proxy / Remote Collector is the data collection component that gathers metrics from endpoints. Analytics nodes process data. Replica nodes provide HA. Web nodes serve the UI.
The Replica node (plus data nodes) provides search capabilities and resilience for the primary analytics node. Cloud Proxy collects data. Collector nodes gather metrics. Witness appliances are for vSAN quorum.
The NSX-T Adapter / NSX Management Pack provides visibility into NSX-specific objects like transport nodes, segments, and edge clusters beyond what vCenter provides. Storage and Service Discovery adapters serve different purposes. Ping adapters test reachability only.
The vSAN adapter collects vSAN-specific performance, capacity, and health metrics for dashboards and alerts. It doesn't provision datastores, configure HCL compatibility, or replace the built-in vSAN Health service.
VCF Operations computes resource costs using pricing cards and provides cost data consumed by Automation catalog for chargeback/showback reports. Automation doesn't calculate independently. Cost isn't only from vCenter or CSV uploads.
VCF 5.2 supports cross-vCenter vMotion between sites with up to 150 ms RTT latency. Note this differs from the 5 ms requirement for stretched vSAN clusters.
The VCF 5.2 Design Guide defines 5 topology design blueprints: (1) Single Instance - Single AZ, (2) Consolidated Single Instance - Single AZ, (3) Single Instance - Multiple AZ, (4) Multiple Instance - Single AZ, and (5) Multiple Instance - Multiple AZ.
VCF 9.0 publishes five blueprints; Blueprint 1 (Single-Site, Minimal Footprint) is the entry-level design. VCF 9.0 official documentation (Design, Deployment, Administration, and NSX 9.0 guides).
Pattern 4 features dual-level management where each region owns a VCF Instance; Central IT owns LCM, certs, global policy; Regional IT handles Day-N ops. VCF 9.0 official documentation (Design, Deployment, Administration, and NSX 9.0 guides).
The NSX Overlay Stretched Segment fleet-components networking model is for DR/IP mobility between VCF Instances and requires NSX Federation with Active/Standby global managers. VCF 9.0 official documentation (Design, Deployment, Administration, and NSX 9.0 guides).
VPC with Full Services uses an Active/Standby T0 and is the only VPC model that offers Default Auto SNAT and VPNs. VCF 9.0 official documentation (Design, Deployment, Administration, and NSX 9.0 guides).
What-if scenario analysis projects adding 50 VMs against current capacity trends over 90 days. Real-time metrics show current state only. Historical alerts list past issues. vSAN Skyline is for vSAN health, not capacity modeling.
Reclaimable waste includes idle VMs, oversized VMs, powered-off VMs, and orphaned VMDKs — all consuming resources without delivering value. Active databases, vCenter, and Edge nodes are operational resources.
Chargeback/showback cost drivers can be scoped to datacenters, clusters, or custom groups, enabling granular cost attribution. Entire fleet only is too coarse. Individual vMotion events aren't cost units. Single host only is too narrow.
Right-sizing recommendations need several weeks of stable utilization data to establish reliable baselines. One hour or a single peak event is insufficient. Host reboots don't generate utilization trends.
Time-to-full forecasting projects average growth trends against available capacity to predict exhaustion dates. Current instantaneous usage doesn't predict trends. Static input isn't dynamic forecasting. Backup retention is unrelated.
Demand-based capacity with configurable buffers accounts for HA admission control, failover overhead, and operational buffers. Allocation model only counts provisioned resources. Raw hardware total ignores overhead. A flat 10% deduction is arbitrary.
What-If Scenarios in Capacity Planning model the impact of adding workloads of specific profiles. Troubleshoot → Events is for incident analysis. Log Explorer searches logs. Inventory Explorer browses objects.
Time Remaining (Capacity) for CPU forecasts when CPU capacity will be exhausted. CPU Ready % measures contention. Demand vs Usage shows current load. Co-Stop is a multi-vCPU scheduling metric.
Reclaimable capacity highlights idle VMs, powered-off VMs, snapshots, and oversized VMs — all representing waste. Only powered-on VMs is too narrow. vSAN witness and NSX Edge are operational infrastructure, not waste.
VCF 5.2 supports four vSAN witness host sizes: Tiny (up to 10 hosts per site), Medium (up to 15), Large (up to 25), and Extra Large (up to 32). The witness must use ESA-compatible storage for ESA clusters.
VCF-EXT-REQD-NET-001: DNS with forward and reverse records. VCF-EXT-REQD-NET-002: NTP time synchronization. VCF-EXT-REQD-NET-003: Active Directory or LDAP for identity. VCF-EXT-REQD-NET-004: SFTP/FTP server for backup storage.
A vSAN Storage Cluster (ESA-required, ≥4 nodes) serves vSAN Compute Clusters in the same workload domain via HCI Mesh. VCF 9.0 official documentation (Design, Deployment, Administration, and NSX 9.0 guides).
High CPU Ready with low host CPU use indicates per-VM CPU scheduling contention, typically from oversized vCPU counts or anti-affinity rules creating scheduling constraints. Memory ballooning affects memory. vSAN cache is storage. DNS is networking.
Heat maps provide quick cluster-wide visualization of hotspots by color-coding resource utilization across objects. Static text widgets don't visualize data. Pie charts show composition. Topology lists show relationships.
vSAN latency troubleshooting should focus on p90/p99 to capture tail latency that affects user experience. p50 misses outliers. Averages hide spikes. Minimums are irrelevant for troubleshooting.
Correlation analysis finds related symptoms across CPU, memory, disk, and network, revealing root causes behind cascading issues. It doesn't change disk config, upgrade SDDC Manager, or set lease policies.
DRS migration recommendations surface in VCF Operations dashboards for workload placement optimization. Storage policy recompliance is vSAN-specific. VM power reset and snapshot consolidation are separate operations.
CPU Ready time means the VM is waiting for physical CPU scheduling — the vCPUs are ready to run but no physical core is available. Storage latency causes I/O delays. Network packet loss is networking. Guest swap is memory pressure.
Operations Overview / Environment Overview dashboard with Health, Risk, Efficiency badges provides the holistic view as the first landing page. Inventory browses objects. Log and Cost dashboards are specific focus areas.
Alert Definitions combine multiple Symptom Definitions with impact and recommendations into a single actionable alert through automatic correlation. Super Metrics create custom KPIs. Custom Groups organize objects. Report Schedules automate delivery.
Datastore → Top-N VMs by IOPS/latency/throughput identifies the noisiest-neighbor VM consuming the most storage resources. vCenter Events show operations. Host CPU view is compute-focused. NSX flow view is network-focused.
For a 4-host management domain cluster in VCF 5.2, the recommended HA admission control is 25% (one host failure tolerance). For 3-host clusters it is 33%, and for 2-host clusters it is 50%.
Three Management Zones with Isolated Workload Zones splits both control-plane and workload clusters across 3 zones — the most resilient Supervisor model. VCF 9.0 official documentation (Design, Deployment, Administration, and NSX 9.0 guides).
An alert definition combines symptom definition(s) with criticality and recommended actions. Just a metric threshold is a symptom, not a complete alert. Custom dashboards visualize data. Syslog patterns are log-based, not metric-based.
IPX/SPX is a legacy Novell protocol not supported as a notification channel. VCF Operations natively supports SMTP email, REST webhooks, and SNMP traps for alert notifications.
Alert symptom + recommended action can invoke an Orchestrator workflow for automated remediation when specific conditions are met. Generic cron jobs lack context. Manual scripts don't auto-trigger. vSphere HA events alone don't drive Operations remediation.
Anomaly/predictive analytics with dynamic baselines detect behavioral deviations by learning normal patterns over time. Static thresholds miss anomalies within normal ranges. Manual review and ping checks are reactive, not proactive.
Automated Action on an Alert Definition (Policy-driven) leveraging the Actions Adapter can execute actions like Power Off when alerts trigger. Log Insight forwarding is for log delivery. vCenter alarms are separate. Super Metrics calculate KPIs.
Power Off or Set CPU/Memory (right-size) reclaims resources from idle VMs without deletion. Optionally delete snapshots to reclaim storage. Deleting the VM is destructive. Moving to another vCenter doesn't reclaim resources. Guest OS reset doesn't reduce allocation.
Outbound notification / webhook/REST plug-in or the ServiceNow management pack enables ITSM integration for automatic ticket creation. Direct SQL writes are unsupported. Manual email is not automated. NSX API doesn't manage ITSM tickets.
A policy is configured for a Custom Group of production VMs to tighten thresholds and enable automated remediation. What is the advantage over modifying the default policy?
It allows differentiated monitoring/remediation behavior per object scope without affecting other environments
Custom Group-scoped policies allow differentiated monitoring and remediation per scope without affecting other environments — production VMs get tighter thresholds while dev VMs keep defaults. It doesn't disable alerts globally or remove Super Metrics.
Workload Optimization can recommend or automatically move workloads across clusters within a Custom Datacenter based on business intent, going beyond single-cluster DRS. It doesn't replace DRS entirely. It works on multi-host clusters. It doesn't require stretched clusters.
The Supervisor Management Zone topology (Single-Zone vs Three-Zone) must be chosen at activation and is not changeable afterward. VCF 9.0 official documentation (Design, Deployment, Administration, and NSX 9.0 guides).
Which VCF 9.0 Automation tenancy deployment model gives each organization its own dedicated VCF Instance while a centralized VCF Automation manages them all?
Model 1 — Consolidated
Model 2 — Shared WLD via Namespaces + NSX
Model 3 — Centralized with Dedicated WLDs per Org
Model 4 — Centralized VCF Automation Managing Dedicated VCF Instances per Org
Model 4 is centralized VCF Automation managing dedicated per-org VCF Instances; Model 5 is fully dedicated fleets per org. VCF 9.0 official documentation (Design, Deployment, Administration, and NSX 9.0 guides).
A super metric formula aggregating over the cluster object hierarchy computes composite KPIs like average cluster CPU demand. Static properties don't compute. Dashboard filters display data. Log patterns extract text, not metrics.
Drift detection alerts when an object's configuration deviates from its baseline template, catching unauthorized or accidental changes. VM power-off, new objects, and log retention reaching limits don't trigger drift alerts.
Management packs from the VMware Marketplace extend monitoring to third-party systems (Dell, NetApp, Cisco, etc.) with additional adapters and content. They don't upgrade vCenter, replace Identity Broker, or rotate certificates.
Log masking/filter rules on the ingestion pipeline redact sensitive data like credit-card numbers before storage. Custom dashboards visualize data. Super metrics compute KPIs. Management packs add monitoring capabilities.
A Super Metric using aggregation (e.g., avg) over members of a Custom Group computes composite KPIs. Alert and Symptom Definitions define thresholds. View widgets display data but don't compute cross-object KPIs.
A Custom Group with dynamic membership based on vSphere tag criteria groups VMs across multiple vCenters for dashboards and policies. DRS groups are cluster-scoped. vCenter folders and Resource Pools are per-vCenter.
VCF 9.0 CA model limits AZ-to-AZ RTT to 10 ms (peaks to 20 ms during 20-s intervals); witness-to-cluster up to 30 ms. VCF 9.0 official documentation (Design, Deployment, Administration, and NSX 9.0 guides).
Fleet-Wide SSO uses one Identity Broker for the entire fleet. Cross-VCF-Instance SSO covers multiple instances but not necessarily all; Single-Instance SSO scopes per instance. VCF 9.0 official documentation (Design, Deployment, Administration, and NSX 9.0 guides).
A VCF Operations cluster role running the metric, property, and event processing pipeline, backing the analytics UI. Multiple analytics nodes partition data for scale. They host adapters, collectors-of-last-resort, and alerting.
Remote Collector
A lightweight VCF Operations node deployed at a remote site to aggregate adapter data and forward it to the main cluster, reducing WAN bandwidth. It has no analytics or UI role. It is the standard remote-site deployment unit.
Adapter Instance
A configured connection to a data source (vCenter, NSX, vSAN, third-party) owned by a specific collector. Adapters collect inventory, metrics, properties, and events. They are the ingestion boundary of VCF Operations.
Object Relationship
A parent/child or peer link between VCF Operations objects (for example, a VM is a child of a host, which is a child of a cluster) enabling cross-object analysis. Relationships power many dashboards and 'find the ancestor' workflows. They are discovered automatically by adapters.
VCF Identity Broker (Operations)
The centralized identity service in VCF 9 that authenticates VCF Operations users via AD/LDAP/SAML/OIDC and maps groups to Operations roles. It replaces per-product local IdP configuration. It enables SSO across the fleet.
VCF Operations Node Roles
Primary (leader), Replica (HA peer for primary), Data (analytics scale-out), Remote Collector (proxy, no data store), Witness (HA tiebreaker in stretched deployments). Sizing mixes these to meet object and metric scale.
HA Mode (VCF Operations)
Activates a Replica node that mirrors the Primary. On failover the cluster promotes the Replica, preserving UI/API. Witness node recommended across sites to prevent split-brain.
Cluster Continuous Availability (CA)
A stretched deployment across two fault domains with a Witness in a third site. Tolerates full-site loss while preserving analytics. Latency between sites must stay under 5 ms RTT.
Content Pack Management Pack
A downloadable bundle that installs dashboards, views, alerts, and adapters for a specific solution (SRM, NSX, HCX, Horizon). Installed via Admin → Repository or MP Marketplace.
Operations Identity Sources
VCF Ops integrates with VCF Identity Broker (OIDC), LDAP/AD, and SAML. Roles and object-scope mappings determine what each user sees; fine-grained with Custom Groups.
Five VCF 5.2 topology design blueprints?
1) Single Instance-Single AZ, 2) Consolidated Single Instance-Single AZ, 3) Single Instance-Multiple AZ, 4) Multiple Instance-Single AZ, 5) Multiple Instance-Multiple AZ.
Cross-vCenter vMotion max latency in VCF 5.2?
150 ms RTT between sites. NSX Federation: <500 ms between instances.
Blueprint 1 — Single Site Minimal Footprint
Entry-level VCF 9.0 design; smallest viable fleet with one Instance in one site.
Blueprint 2 — Single Site
Production single-site VCF 9.0 fleet; typical starting blueprint for small-medium DCs.
Blueprint 3 — Multiple Sites, Single Region
Multiple VCF Instances in one region; uses stretched clusters for intra-region HA.
Blueprint 4 — Multiple Sites across Multiple Regions
Largest design — DR across regions plus intra-region HA; supports Fault Domains + DR model.
Blueprint 5 — Single Region + Additional Region(s)
Hybrid blueprint adding disaster-recovery regions to a single-region primary.
A VCF Operations projection of how many days remain until a container (cluster, datastore) exhausts CPU, memory, or storage at current growth. It drives procurement planning. Anomalous growth makes the forecast unreliable, which the UI calls out.
Reclaimable Waste
VCF Operations' quantification of resources that could be returned to the pool by right-sizing oversized VMs, powering off idle VMs, or removing orphaned disks. It is the single biggest knob for efficient operations. It can be acted on from Automation Central.
Right-Sizing Recommendation
A specific VCF Operations suggestion to increase or decrease a VM's CPU/memory based on observed utilization over a window. Each recommendation includes predicted impact. Bulk acceptance turns recommendations into reclaimable capacity.
Cost Driver
A configurable line item (hardware depreciation, licensing, labor, facilities) that contributes to per-object cost in VCF Operations. Drivers feed chargeback and showback reports. They are the bridge between infrastructure and finance.
Capacity Planning Report
A scheduled VCF Operations artifact combining current utilization, reclaimable waste, time-to-full, and procurement recommendations. It is typically produced monthly for infrastructure leadership. It is the executive deliverable of capacity management.
Capacity Buffer
A percentage subtracted from total usable capacity before forecasting to leave room for HA/bursts. Set per cluster or resource type; default is commonly 10–20% and is a key tunable in accurate Time-to-Full.
Demand vs Allocation Model
Capacity can be calculated against Demand (actual consumption) or Allocation (provisioned reservations). Allocation model is stricter and used for overcommit-sensitive environments; Demand is for typical virtualization workloads.
Stress Score
A derived score (0–100) combining CPU/memory demand, contention, and peak behavior over a rolling window. High stress signals that a cluster is capacity-constrained despite nominal average utilization.
Reclaim — Idle VMs
Definition of 'idle' in VCF Ops: CPU < 100 MHz and network < 1 Kbps over a configurable window (default 90 days). Reclaim reports suggest power-off/snapshot/delete actions and quantify reclaimable resources.
Cost Driver Categories
Categories include Server Hardware, Storage Hardware, Licenses, Maintenance & Support, Labor, Network, Facilities, and Additional Costs. Rates feed per-VM showback and cluster cost-per-GB/GHz.
vSAN witness host sizes in VCF 5.2?
Tiny (≤10 hosts/site), Medium (≤15), Large (≤25), Extra Large (≤32).
Four external service requirements in VCF 5.2?
VCF-EXT-REQD-NET-001: DNS. VCF-EXT-REQD-NET-002: NTP. VCF-EXT-REQD-NET-003: AD/LDAP. VCF-EXT-REQD-NET-004: SFTP/FTP for backups.
vSAN ESA Storage Cluster
Disaggregated ESA-only vSAN cluster; ≥4 nodes; consumed by vSAN Compute Clusters via HCI Mesh.
A VCF Operations technique that ranks metrics most related to an object's anomaly during a time window — for example, CPU ready correlated with memory ballooning. It turns 'where do I look?' into a short prioritized list. It is built into the Troubleshooting Workbench.
Performance Dashboard
A curated VCF Operations view showing latency, IOPS, CPU, and memory across a chosen scope (cluster, application, custom group). Out-of-the-box dashboards ship for common scenarios. Custom dashboards are the typical VCAP-level deliverable.
Latency Percentile (p99)
A statistical measure reporting the latency below which 99% of samples fall, exposing tail behavior hidden by averages. vSAN and VCF Operations expose p50/p90/p99 for IO latency. Regressions in p99 often predict user-visible incidents.
Cache Hit Rate
A vSAN OSA hybrid performance metric showing the fraction of reads served from the cache tier (70% read cache / 30% write buffer) rather than capacity HDDs. Low hit rate on random-read workloads predicts high latency. On OSA all-flash, the cache is write-only; on ESA there is no cache tier, so this metric does not apply.
DRS Recommendation
A vCenter suggestion (or automated move) generated by the Distributed Resource Scheduler to rebalance VMs across hosts. VCF Operations surfaces DRS data alongside its own recommendations to explain observed rebalances. It bridges platform automation and observability.
Troubleshooting Workbench
Guided troubleshooting UI that correlates alerts, property changes, events, and metrics for a selected object across a chosen time range. Starting point for root-cause analysis in VCF Ops 8/9.
Metric Correlation Engine
ML-based engine that surfaces metrics strongly correlated with a target metric anomaly on the same or related objects. Helps pinpoint symptom vs cause (e.g., latency correlates with queue depth on a specific datastore).
Anomaly Score
Dynamic threshold score indicating deviation from learned baseline for each metric. Anomaly badge colors are derived from object rollups of anomaly scores.
Workload Badge
Composite badge combining demand and capacity; Green/Yellow/Orange/Red signals whether a host/cluster is being stressed by current load independent of raw utilization.
KPI Alert
Alert definition bound to a specific KPI metric with symptoms using static, dynamic, or HT threshold. Drives notifications and optional automated actions.
HA admission control for 4-host management cluster?
Type A: single pNIC (lab/PoC). Type B: two pNICs same speed. Type C: two pNICs different speeds.
Supervisor models — most resilient
Three Management Zones with Isolated Workload Zones: mgmt + workloads both spread across 3 zones.
Simplified Supervisor limits
1 CP VM, 1 vNIC, no LB; VM Service only — no vSphere Pods, no VKS.
Alert Definition
A VCF Operations composite construct combining one or more symptoms, a criticality level, and a set of recommended actions. Alerts are raised on objects and flow through notifications. They are the operational unit of monitoring.
Notification Rule
A VCF Operations mapping from alert criteria to an outbound channel — SMTP, REST webhook, SNMP, syslog. Rules are scoped by object, criticality, and tag. They route the right alerts to the right teams.
Automated Remediation Workflow
A VCF Operations action (sometimes via vRO) that runs when an alert fires — for example, snapshotting before remediating, or powering off a runaway VM. Workflows can require approval. They are the 'self-healing' backbone of mature operations.
Maintenance Schedule
A window during which VCF Operations suppresses alerts for specified objects, preventing noise during planned changes. Schedules can be one-time or recurring and scoped by tag or group. They are essential for stable notification posture.
Predictive Anomaly Detection
A VCF Operations ML capability that models metric behavior against historical norms and flags deviations before thresholds fire. It complements static symptoms. It is tuned per metric through policies.
Automation via Aria Orchestrator
Alerts can trigger vRO workflows (for example, migrate a VM off a hot host, extend a datastore). Integration uses a registered vRO endpoint; inputs are populated from the alert object context.
Action Framework
Built-in catalog of safe actions (Power Off VM, Set CPU/Memory, Delete Idle VM, Move VM, Shut Down Guest). Actions require proper adapter credentials and can be run ad-hoc, from alerts, or on schedule.
Maintenance Schedule Effects
When an object enters a scheduled maintenance window, alerts are suppressed, metrics are flagged as 'in maintenance' for forecasting, and reports can optionally exclude that window.
Symptom Definition
The atomic trigger used inside Alert Definitions — metric symptom, property symptom, event symptom, fault symptom, or message symptom. Multiple symptoms combined with AND/OR logic form the alert.
Outbound Notification
Configurable via SMTP, SNMP, REST, Webhook, ServiceNow, or Slack. Template-driven payload using alert properties; filters decide which alerts go to which channel.
Tenancy Model 4
Centralized VCF Automation managing dedicated VCF Instances per organization.
Tenancy Model 5
Dedicated VCF Fleet per organization with independent VCF Automation.
Interactive Dashboard
A VCF Operations advanced dashboard where widgets pass context — the object selected in one widget drives the data displayed in others. It is the recommended format for multi-object troubleshooting. It requires configuring widget interactions explicitly.
Super Metric (Operations)
A derived metric computed from a formula across one or more objects and their relationships, for example 'average host CPU ready by cluster'. Super metrics feed dashboards, alerts, and reports. They are the primary extensibility mechanism for custom analytics.
Content Pack (VCF Operations)
A bundle of dashboards, views, reports, symptoms, and super metrics for a specific workload type installed into VCF Operations for Logs or VCF Operations. Packs are shipped via Marketplace or built in-house. They accelerate new-environment onboarding.
Compliance Benchmark (CIS)
A VCF Operations rule set that checks ESXi, vCenter, NSX, and VKS against the Center for Internet Security's hardening guides. Violations surface as risk-badge contributors. It is the default baseline when customers have no custom compliance framework.
Configuration Drift
A VCF Operations detection capability that watches properties against a baseline template and alerts when they change outside approved channels. It catches out-of-band edits and compliance regressions. It pairs naturally with Configuration Profiles in vSphere 9.
Super Metric Syntax
Super Metrics use a formula language combining metric references, functions (avg, sum, max, percentile), and literals, e.g., sum($This, cpu|demand_without_overhead_kb). Scope is a resource kind; result is a new virtual metric.
Super Metric Policy Binding
A Super Metric only shows up when added to a Policy applied to the target object type. Without policy binding, the metric is defined but not collected.
Compliance Benchmark Types
Built-in benchmarks include CIS, DISA STIG, NIST, PCI, HIPAA, and VMware Security Configuration Guide (SCG). Each benchmark is a content pack of rules compared against configuration properties.
Configuration Drift Scoring
Drift is the delta between a baseline configuration (captured or templated) and current state. Reports list per-property changes and map them to compliance rules.
Custom Dashboard Widgets
Widgets include Object List, Scoreboard, Metric Chart, Heatmap, Top-N, Object Relationship, Topology, and Text. Widgets can be linked via self-provider or interaction so selecting an object in one updates others.
CA model latency
≤10 ms RTT (peaks to 20 ms during 20-s windows) between AZs; ≤30 ms to witness.
Fleet-Wide SSO
A single Identity Broker serves every Instance in a Fleet — lowest operational overhead, largest blast radius.