Lab: Plan & Execute VCF Operations Cluster Upgrade
Objectives
- Plan a VCF Operations cluster upgrade with pre-flight validation
- Execute a staged upgrade across Master, Replica, and Data nodes
- Verify upgrade completion and data integrity post-upgrade
- Understand PAK file management and offline upgrade procedures
- Design a rollback plan for failed upgrade scenarios
- Monitor upgrade progress and handle common failure modes
Prerequisites
VCF Operations cluster (multi-node preferred: Master + 2 Data nodes minimum) running a version that can be upgraded, upgrade PAK file downloaded, snapshot/backup of VCF Operations nodes taken
Required skills:
- VCF Operations administration
- vSphere snapshot management
- Change management procedures
Lab Environment
VCF Operations cluster with minimum 3 nodes (Master, Data Node 1, Data Node 2) or single-node for smaller lab. PAK file for the target version available. vCenter backup taken. SSH access to VCF Operations nodes for troubleshooting.
Tasks
Task 1 VCF Operations Cluster Upgrade with Rollback Planning
Execute a controlled upgrade of a VCF Operations cluster following production best practices: pre-flight validation, staged node-by-node upgrade, post-upgrade verification, and rollback planning. Understanding the upgrade lifecycle is critical for VCAP-level operations and ensures zero data loss during version transitions.
Pre-flight: Check current version and compatibility. VCF Operations → Administration → Cluster Status. Note: current version (e.g., 8.14.x), node count, node health. Verify the target version is compatible with your current version (check the upgrade path matrix on techdocs.broadcom.com — some versions require intermediate upgrades). Verify: vCenter version compatibility, NSX version compatibility, browser compatibility.
Pre-flight: Verify cluster health. Administration → Cluster Status → all nodes should show 'Online' with green health. Check: (a) Master Availability=Available, (b) All adapters collecting data (no 'Down' adapters), (c) No active critical alerts on the VCF Operations cluster itself, (d) Cassandra cluster health: SSH to Master → $VCOPS_BASE/cassandra/apache-cassandra-*/bin/nodetool status — all nodes should show 'UN' (Up Normal).
Pre-flight: Take snapshots. In vSphere Client, take memory snapshots of ALL VCF Operations nodes (Master, Replica, Data nodes) with the same timestamp. Name: 'Pre-Upgrade-<version>-<date>'. These snapshots are the rollback safety net — if the upgrade fails, revert all nodes to these snapshots simultaneously. Also verify: vCenter backup is current (VCF Operations upgrade can affect vCenter integration).
Pre-flight: Check disk space. SSH to Master node: df -h /storage. VCF Operations upgrade requires at least 5 GB free space for PAK extraction and upgrade staging. If space is low: (a) reduce metric retention temporarily, (b) clean log files: $VCOPS_BASE/user/log/*.gz, (c) extend the data disk in vSphere if needed (VCF Operations supports online disk expansion).
Upload PAK file. Administration → Software Update → Upload. Select the PAK file (e.g., vRealize_Operations-8.16.x-xxxxx.pak or vcf-operations-xxx.pak). The upload process validates the file integrity (checksum) and compatibility with the current version. If validation fails: re-download the PAK file (corrupt download), or check the upgrade path matrix. For air-gapped environments, transfer the PAK via SCP to /tmp on the Master node.
Execute upgrade. Software Update → Install. The upgrade proceeds in stages: (a) Master node enters maintenance mode (API remains available, UI may be intermittent), (b) Master services stop, upgrade files are applied, services restart, (c) Master validates upgrade success, (d) If multi-node: Master signals each node to upgrade in sequence — Replica first, then Data nodes one at a time. During the upgrade: metrics collection pauses on each node as it upgrades (other nodes continue collecting). Total expected duration: 30-60 minutes for Master, 15-30 minutes per additional node.
Monitor upgrade progress. The UI shows a progress bar per node. For CLI monitoring (if UI is unavailable during Master upgrade): SSH to Master → tail -f $VCOPS_BASE/user/log/upgradeStatus.log. Key log entries to watch: 'Upgrade started', 'Database migration in progress' (longest phase), 'Services restarting', 'Upgrade completed'. If a node is stuck > 60 minutes: check upgradeStatus.log for errors, do NOT reboot (may corrupt the database migration).
Post-upgrade verification. After all nodes show 'Upgrade Complete': (a) Administration → Cluster Status — all nodes Online with new version number; (b) Check adapter status — all adapters should resume collection within 15 minutes; (c) Verify dashboards display data (some may need refresh); (d) Check super metrics — custom formulas should survive the upgrade, but verify values are computing; (e) Run a Cassandra health check: nodetool status — all nodes 'UN'; (f) Verify API access: curl -s -k https://<vcf-ops>/suite-api/api/versions — should return the new version.
Post-upgrade: Validate data continuity. Compare key metrics before and after upgrade: (a) Check a VM's CPU usage chart — historical data should show a continuous line (no gap during upgrade window); (b) Verify super metric calculations match pre-upgrade values; (c) Check custom group membership — dynamic groups should have the same member count; (d) Verify scheduled reports are still configured and scheduled. Any discrepancy indicates a data migration issue — escalate to VMware support before deleting pre-upgrade snapshots.
Cleanup and documentation. After 48 hours of successful operation post-upgrade: (a) Delete pre-upgrade snapshots (they consume significant disk space and degrade storage performance); (b) Update CMDB with new VCF Operations version; (c) Document the upgrade in the change management system: start time, duration per node, any issues encountered, rollback snapshots deleted; (d) If this was a VCF LCM-managed upgrade, verify the LCM inventory shows the correct version. Rollback option: if critical issues are found post-upgrade, power off ALL nodes simultaneously, revert ALL to pre-upgrade snapshots, power on Master first, then other nodes. This restores the exact pre-upgrade state.
Validation Gate
Check: VCF Operations cluster upgraded with data integrity verified
Expected: All nodes upgraded to target version, cluster health green, all adapters collecting, metrics data continuous through upgrade window, super metrics and dashboards functional, scheduled reports intact
Common Errors
Final Validation
VCF Operations cluster successfully upgraded with verified data integrity and operational continuity
✓ Cluster version → All nodes show target version in Administration → Cluster Status
✓ Node health → All nodes Online with green health status
✓ Adapter collection → All adapters showing 'Collecting' status with current timestamps
✓ Data continuity → Historical metric charts show no gaps during upgrade window
✓ Super metrics → All custom super metrics computing values correctly
✓ Dashboards and reports → All dashboards display data, scheduled reports intact
Cleanup / Restore
• Delete pre-upgrade snapshots after 48 hours of stable operation
• Update change management records with upgrade details
• Archive upgrade PAK file for rollback reference
Design Reflection (VCDX)
Upgrade planning demonstrates operational maturity — a VCDX differentiator. Key points: (1) pre-flight validation prevents failed upgrades (disk space, cluster health, snapshot safety net); (2) staged upgrade minimizes monitoring gap (only one node offline at a time); (3) data continuity verification ensures no metric loss; (4) 48-hour bake period before snapshot cleanup provides rollback window. In VCF 9.0, discuss how VCF LCM orchestrates upgrades across the entire stack (vCenter, NSX, VCF Operations) with dependency ordering.
Requirements
- Zero data loss during upgrade
- Maximum 30-minute monitoring gap per node during upgrade
- Rollback capability for 48 hours post-upgrade
Constraints
- Upgrade is non-reversible once all nodes complete — rollback requires snapshot revert
- Some adapters may require re-registration post-upgrade
- Air-gapped environments require manual PAK file transfer
Assumptions
- Pre-upgrade snapshots are taken on all nodes with consistent timestamps
- Sufficient disk space (5+ GB) for PAK extraction and migration
- vCenter and NSX versions are compatible with target VCF Operations version
Risks
- Database migration failure during upgrade — long outage if rollback is needed
- Snapshot revert after partial upgrade leaves cluster in inconsistent state — must revert ALL nodes simultaneously
- Third-party management pack incompatibility with new version — breaks monitoring for non-VMware systems
Self-Assessment Discussion Prompts
- How do you plan VCF Operations upgrades when the cluster monitors the infrastructure being upgraded?
- What is the difference between upgrading via VCF LCM vs direct PAK upgrade?
- How do you handle upgrade testing in environments without a non-production VCF Operations instance?
- What monitoring gaps are acceptable during upgrade, and how do you mitigate them?
Extensions
Practice the upgrade using VCF LCM instead of direct PAK upgrade
Design an upgrade runbook with pre/post checks as a reusable template
Test rollback by intentionally reverting snapshots after a successful upgrade to validate the procedure
Create an automated pre-flight check script using the VCF Operations API
⚠ Known Pitfalls (from Community KB)
References
- VCF Operations Upgrade Guide: techdocs.broadcom.com
- VCF 9.0 Lifecycle Management Guide — Operations Upgrade: techdocs.broadcom.com