Academy/vSphere Foundation 9.0 Support (2V0-18.25)/Lab: Collect Support Bundle & Analyze Logs
This lab targets VCF 9.0

Lab: Collect Support Bundle & Analyze Logs

VCF 9.0Intermediatevcp-foundation⏱ 90 min

Objectives

  • Lab: Collect Support Bundle & Analyze Logs

Prerequisites

VCF lab environment deployed and operational

Lab Environment

Standard VCF lab environment for vSphere Foundation 9.0 Support (Support Specialist)

Tasks

Task 1 Lab: Collect Support Bundle & Analyze Logs

Support bundle collection is the first step in any troubleshooting engagement. Knowing which bundle to collect, when to collect it, and how to analyze the contents separates reactive firefighting from systematic root cause analysis.
Step 1
vCenter UI → Administration → Bundles → Request Support Bundle (initiates background collection)
Step 2

Wait 5-10 minutes; download from UI when complete

Step 3

Extract: tar -xzf vc_support_bundle-*.tgz

Step 4

Examine vCenter log: grep -i "error\|warning" vsan/vsanmgmtd.log | head -20

Step 5

Check ESXi host logs: grep "vmkernel" hostd.log | tail -10

Step 6

Export summary: List 5 most critical errors for support case

Validation Gate

Check: Collect support bundles from ESXi, vCenter, and NSX Manager. Verify each bundle is complete (non-zero size) and contains logs from the expected time window.

Expected: Three bundles collected: vm-support.tgz (ESXi), vc-support.tgz (vCenter), nsx-support-bundle.tgz (NSX). Each contains logs covering the incident window. Total size is reasonable (not truncated).

Common Errors

Collecting support bundle AFTER restarting the failed service
Fix: Restarting a service clears in-memory state and may rotate logs. Always collect the support bundle BEFORE any remediation action. The bundle captures the current state — restart erases the evidence. This is the #1 mistake that makes root cause analysis impossible.
Collecting only the ESXi bundle when the issue involves vCenter
Fix: Match the bundle to the component: ESXi issue → vm-support on the host. vCenter issue → vc-support from VCSA. NSX issue → support bundle from NSX Manager. SDDC Manager issue → sddc-support bundle. VCF-wide issue → collect ALL bundles. Incomplete log collection leads to 'need more data' back-and-forth with support.
Not filtering log analysis by time window
Fix: Support bundles contain weeks of logs. Searching the entire bundle is time-consuming and produces noise. First identify the incident timestamp (from vCenter events, user report, or monitoring alert), then filter logs to a window of +/- 30 minutes around the incident. Key log locations: /var/log/vmkernel.log (ESXi), /var/log/vmware/vpxd/vpxd.log (vCenter).
Sending uncompressed bundles to support
Fix: ESXi vm-support bundles can be 500MB+. vCenter bundles can exceed 1GB. Always compress before transfer. The support bundle tools produce .tgz format by default — do not extract and re-upload individual files unless support specifically requests them.

Final Validation

Lab completed successfully

✓ All steps completed → No errors observed

Cleanup / Restore

• Revert to snapshot if needed

Design Reflection (VCDX)

Troubleshooting methodology is part of operational design. VCDX panelists may ask about your log collection strategy, incident response procedures, and how you ensure root cause analysis is possible after remediation.

Requirements

  • Collect appropriate support bundles for each component
  • Filter log analysis to incident time window
  • Preserve evidence before remediation

Constraints

  • Bundle must be collected before service restart
  • Bundle size can be large — compress before transfer
  • Log retention varies by component — collect promptly

Assumptions

  • SSH access is available for bundle collection
  • Storage space available for bundle collection on each component

Risks

  • Evidence loss from restarting services before bundle collection
  • Incomplete analysis from collecting wrong component's bundle

⚠ Known Pitfalls (from Community KB)

Collecting support bundle AFTER restarting the service — this is the single most common mistake that prevents root cause analysis.
Sending unfiltered 1GB+ bundles to support without identifying the incident time window — causes delays.
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.