Academy/VCAP — VCF Operations (3V0-22.25)/Lab: Plan & Execute VCF Operations Cluster Upgrade
This lab targets VCF 9.0

Lab: Plan & Execute VCF Operations Cluster Upgrade

VCF 9.0Advancedvcap-advanced⏱ 120 min

VCF Operations lifecycle — cluster upgrade procedures, pre-checks, PAK file management, rollback planning, multi-node upgrade orchestration

Objectives

  • Plan a VCF Operations cluster upgrade with pre-flight validation
  • Execute a staged upgrade across Master, Replica, and Data nodes
  • Verify upgrade completion and data integrity post-upgrade
  • Understand PAK file management and offline upgrade procedures
  • Design a rollback plan for failed upgrade scenarios
  • Monitor upgrade progress and handle common failure modes

Prerequisites

VCF Operations cluster (multi-node preferred: Master + 2 Data nodes minimum) running a version that can be upgraded, upgrade PAK file downloaded, snapshot/backup of VCF Operations nodes taken

Required skills:

  • VCF Operations administration
  • vSphere snapshot management
  • Change management procedures

Lab Environment

VCF Operations cluster with minimum 3 nodes (Master, Data Node 1, Data Node 2) or single-node for smaller lab. PAK file for the target version available. vCenter backup taken. SSH access to VCF Operations nodes for troubleshooting.

Tasks

Task 1 VCF Operations Cluster Upgrade with Rollback Planning

VCF Operations upgrades follow a staged process: Master node first (orchestrates the upgrade), then Replica, then Data nodes (one at a time). During the upgrade, the cluster operates in degraded mode — metrics collection continues but analytics may be delayed. The upgrade is non-reversible once all nodes are upgraded; rollback requires restoring from snapshot. In VCF 9.0, VCF Operations upgrades may also be orchestrated through VCF Lifecycle Management (LCM), which handles dependency ordering across the entire VCF stack.

Execute a controlled upgrade of a VCF Operations cluster following production best practices: pre-flight validation, staged node-by-node upgrade, post-upgrade verification, and rollback planning. Understanding the upgrade lifecycle is critical for VCAP-level operations and ensures zero data loss during version transitions.

Step 1
Pre-flight: Check current version and compatibility. VCF Operations → Administration → Cluster Status. Note: current version (e.g., 8.14.x), node count, node health. Verify the target version is compatible with your current version (check the upgrade path matrix on techdocs.broadcom.com — some versions require intermediate upgrades). Verify: vCenter version compatibility, NSX version compatibility, browser compatibility.
Step 2
Pre-flight: Verify cluster health. Administration → Cluster Status → all nodes should show 'Online' with green health. Check: (a) Master Availability=Available, (b) All adapters collecting data (no 'Down' adapters), (c) No active critical alerts on the VCF Operations cluster itself, (d) Cassandra cluster health: SSH to Master → $VCOPS_BASE/cassandra/apache-cassandra-*/bin/nodetool status — all nodes should show 'UN' (Up Normal).
Step 3

Pre-flight: Take snapshots. In vSphere Client, take memory snapshots of ALL VCF Operations nodes (Master, Replica, Data nodes) with the same timestamp. Name: 'Pre-Upgrade-<version>-<date>'. These snapshots are the rollback safety net — if the upgrade fails, revert all nodes to these snapshots simultaneously. Also verify: vCenter backup is current (VCF Operations upgrade can affect vCenter integration).

Step 4

Pre-flight: Check disk space. SSH to Master node: df -h /storage. VCF Operations upgrade requires at least 5 GB free space for PAK extraction and upgrade staging. If space is low: (a) reduce metric retention temporarily, (b) clean log files: $VCOPS_BASE/user/log/*.gz, (c) extend the data disk in vSphere if needed (VCF Operations supports online disk expansion).

Step 5
Upload PAK file. Administration → Software Update → Upload. Select the PAK file (e.g., vRealize_Operations-8.16.x-xxxxx.pak or vcf-operations-xxx.pak). The upload process validates the file integrity (checksum) and compatibility with the current version. If validation fails: re-download the PAK file (corrupt download), or check the upgrade path matrix. For air-gapped environments, transfer the PAK via SCP to /tmp on the Master node.
Step 6
Execute upgrade. Software Update → Install. The upgrade proceeds in stages: (a) Master node enters maintenance mode (API remains available, UI may be intermittent), (b) Master services stop, upgrade files are applied, services restart, (c) Master validates upgrade success, (d) If multi-node: Master signals each node to upgrade in sequence — Replica first, then Data nodes one at a time. During the upgrade: metrics collection pauses on each node as it upgrades (other nodes continue collecting). Total expected duration: 30-60 minutes for Master, 15-30 minutes per additional node.
Step 7
Monitor upgrade progress. The UI shows a progress bar per node. For CLI monitoring (if UI is unavailable during Master upgrade): SSH to Master → tail -f $VCOPS_BASE/user/log/upgradeStatus.log. Key log entries to watch: 'Upgrade started', 'Database migration in progress' (longest phase), 'Services restarting', 'Upgrade completed'. If a node is stuck > 60 minutes: check upgradeStatus.log for errors, do NOT reboot (may corrupt the database migration).
Step 8
Post-upgrade verification. After all nodes show 'Upgrade Complete': (a) Administration → Cluster Status — all nodes Online with new version number; (b) Check adapter status — all adapters should resume collection within 15 minutes; (c) Verify dashboards display data (some may need refresh); (d) Check super metrics — custom formulas should survive the upgrade, but verify values are computing; (e) Run a Cassandra health check: nodetool status — all nodes 'UN'; (f) Verify API access: curl -s -k https://<vcf-ops>/suite-api/api/versions — should return the new version.
Step 9

Post-upgrade: Validate data continuity. Compare key metrics before and after upgrade: (a) Check a VM's CPU usage chart — historical data should show a continuous line (no gap during upgrade window); (b) Verify super metric calculations match pre-upgrade values; (c) Check custom group membership — dynamic groups should have the same member count; (d) Verify scheduled reports are still configured and scheduled. Any discrepancy indicates a data migration issue — escalate to VMware support before deleting pre-upgrade snapshots.

Step 10

Cleanup and documentation. After 48 hours of successful operation post-upgrade: (a) Delete pre-upgrade snapshots (they consume significant disk space and degrade storage performance); (b) Update CMDB with new VCF Operations version; (c) Document the upgrade in the change management system: start time, duration per node, any issues encountered, rollback snapshots deleted; (d) If this was a VCF LCM-managed upgrade, verify the LCM inventory shows the correct version. Rollback option: if critical issues are found post-upgrade, power off ALL nodes simultaneously, revert ALL to pre-upgrade snapshots, power on Master first, then other nodes. This restores the exact pre-upgrade state.

Validation Gate

Check: VCF Operations cluster upgraded with data integrity verified

Expected: All nodes upgraded to target version, cluster health green, all adapters collecting, metrics data continuous through upgrade window, super metrics and dashboards functional, scheduled reports intact

Common Errors

PAK upload fails with 'Incompatible version'
Fix: The target version may require an intermediate upgrade. Check the VCF Operations upgrade path matrix: some major versions require stepping through intermediate releases. Example: 8.10 → 8.14 → 8.16, not 8.10 → 8.16 directly.
Node stuck at 'Database migration in progress' for > 60 minutes
Fix: Large Cassandra databases take longer to migrate. Do NOT reboot. Check disk I/O with iostat — migration is I/O intensive. If disk I/O is saturated, wait. If disk I/O is idle but status is stuck, check upgradeStatus.log for specific migration step errors. Common cause: insufficient disk space during migration compaction.
Adapters not collecting after upgrade
Fix: Some adapters require re-registration post-upgrade, especially third-party management packs. Check: Administration → Solutions → Adapters → verify each adapter shows 'Collecting'. For vCenter adapter: re-enter credentials if prompted (certificate may have changed during upgrade).
Super metrics showing incorrect values post-upgrade
Fix: Formula syntax may have changed between versions. Review: Configure → Super Metrics → check for 'Invalid' status. Edit and re-validate each formula. If formulas reference deprecated metric keys (renamed in new version), update the metric path references.

Final Validation

VCF Operations cluster successfully upgraded with verified data integrity and operational continuity

✓ Cluster version → All nodes show target version in Administration → Cluster Status

✓ Node health → All nodes Online with green health status

✓ Adapter collection → All adapters showing 'Collecting' status with current timestamps

✓ Data continuity → Historical metric charts show no gaps during upgrade window

✓ Super metrics → All custom super metrics computing values correctly

✓ Dashboards and reports → All dashboards display data, scheduled reports intact

Cleanup / Restore

• Delete pre-upgrade snapshots after 48 hours of stable operation

• Update change management records with upgrade details

• Archive upgrade PAK file for rollback reference

Design Reflection (VCDX)

Upgrade planning demonstrates operational maturity — a VCDX differentiator. Key points: (1) pre-flight validation prevents failed upgrades (disk space, cluster health, snapshot safety net); (2) staged upgrade minimizes monitoring gap (only one node offline at a time); (3) data continuity verification ensures no metric loss; (4) 48-hour bake period before snapshot cleanup provides rollback window. In VCF 9.0, discuss how VCF LCM orchestrates upgrades across the entire stack (vCenter, NSX, VCF Operations) with dependency ordering.

Requirements

  • Zero data loss during upgrade
  • Maximum 30-minute monitoring gap per node during upgrade
  • Rollback capability for 48 hours post-upgrade

Constraints

  • Upgrade is non-reversible once all nodes complete — rollback requires snapshot revert
  • Some adapters may require re-registration post-upgrade
  • Air-gapped environments require manual PAK file transfer

Assumptions

  • Pre-upgrade snapshots are taken on all nodes with consistent timestamps
  • Sufficient disk space (5+ GB) for PAK extraction and migration
  • vCenter and NSX versions are compatible with target VCF Operations version

Risks

  • Database migration failure during upgrade — long outage if rollback is needed
  • Snapshot revert after partial upgrade leaves cluster in inconsistent state — must revert ALL nodes simultaneously
  • Third-party management pack incompatibility with new version — breaks monitoring for non-VMware systems

Self-Assessment Discussion Prompts

  1. How do you plan VCF Operations upgrades when the cluster monitors the infrastructure being upgraded?
  2. What is the difference between upgrading via VCF LCM vs direct PAK upgrade?
  3. How do you handle upgrade testing in environments without a non-production VCF Operations instance?
  4. What monitoring gaps are acceptable during upgrade, and how do you mitigate them?

Extensions

Practice the upgrade using VCF LCM instead of direct PAK upgrade

Design an upgrade runbook with pre/post checks as a reusable template

Test rollback by intentionally reverting snapshots after a successful upgrade to validate the procedure

Create an automated pre-flight check script using the VCF Operations API

⚠ Known Pitfalls (from Community KB)

Upgrading without taking snapshots — no rollback option if upgrade fails
Taking snapshots at different times — inconsistent state if rollback is needed; snapshot ALL nodes within the same 5-minute window
Rebooting a node during database migration — corrupts Cassandra data, requiring full rebuild from backup
Keeping pre-upgrade snapshots longer than 48-72 hours — snapshot delta files degrade storage performance significantly

References

  • VCF Operations Upgrade Guide: techdocs.broadcom.com
  • VCF 9.0 Lifecycle Management Guide — Operations Upgrade: techdocs.broadcom.com
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.