Academy/VCAP — VCF Administrator (3V0-11.26)/Lab: SoS Diagnostics & VCF API Troubleshooting
This lab targets VCF 9.0

Lab: SoS Diagnostics & VCF API Troubleshooting

VCF 9.0Advancedvcap-advanced⏱ 90 min

SoS utility, log bundle collection, REST API diagnostics, failed task analysis

Objectives

  • Execute SoS health checks and interpret results
  • Collect targeted log bundles for support cases
  • Use SDDC Manager REST API for diagnostic queries
  • Analyze failed lifecycle tasks via API subtask inspection
  • Build a troubleshooting decision tree for common VCF issues

Prerequisites

VCF 9.0 Instance operational, SSH access to SDDC Manager, REST API client (curl, Postman, or similar)

Required skills:

  • Linux CLI basics
  • REST API concepts
  • Log analysis

Lab Environment

VCF 9.0 with SDDC Manager accessible via SSH and HTTPS. At least one completed lifecycle task (domain creation, host addition) in task history for API analysis.

Tasks

Task 1 SoS Diagnostics & API-Driven Troubleshooting

Effective VCF troubleshooting follows a top-down approach: (1) SoS health-check for broad assessment, (2) API task/subtask inspection for specific operation failures, (3) targeted log collection for root cause analysis. Never start with logs — start with health checks to narrow the scope.

Master VCF's built-in diagnostic tools: SoS for health checks and log collection, SDDC Manager REST API for programmatic diagnostics, and log correlation for root cause analysis — the primary troubleshooting toolkit for VCAP-level VCF administration.

Step 1
Run SoS health check. SSH to SDDC Manager VM: sudo /opt/vmware/sddc-support/sos --health-check --domain <management-domain-name>. The health check examines: DNS resolution (forward + reverse for all components), NTP synchronization (skew < 1 second), certificate validity (all certs checked), service status (all VCF services running), component connectivity (SDDC Mgr ↔ vCenter ↔ NSX ↔ ESXi), password validity (credential store vs actual), disk space (all components), vSAN health. Output: HTML report at /tmp/sos-<timestamp>/health-check-report.html.
Step 2

Interpret SoS results. Open the HTML report in a browser (SCP to local machine or browse via vCenter console). Each check shows: GREEN (pass), YELLOW (warning), RED (fail). Common findings: (a) NTP skew > 1s on specific hosts — causes certificate validation failures; (b) DNS reverse lookup failure — blocks future lifecycle operations; (c) certificate expiry within 30 days — immediate rotation needed; (d) disk space low on SDDC Manager — log rotation needed. Document all RED and YELLOW findings as an action list.

Step 3

Collect targeted log bundle. For a support case, collect logs for a specific time window and component: sudo /opt/vmware/sddc-support/sos --log-collection --domain <domain> --component vCenter --start-date 2026-05-13 --end-date 2026-05-14. This produces a smaller, targeted bundle vs. collecting everything. Available component filters: vCenter, NSX, ESXi, SDDC-Manager. Bundle location: /tmp/sos-<timestamp>/.

Step 4

Authenticate to SDDC Manager API. Using curl: TOKEN=$(curl -s -k -X POST 'https://<sddc-mgr>/v1/tokens' -H 'Content-Type: application/json' -d '{"username":"admin@local","password":"<password>"}' | python3 -c 'import sys,json; print(json.load(sys.stdin)["accessToken"])'). Verify: curl -s -k -H "Authorization: Bearer $TOKEN" 'https://<sddc-mgr>/v1/system/health' | python3 -m json.tool. This returns the overall VCF health summary.

Step 5

Query lifecycle task history. curl -s -k -H "Authorization: Bearer $TOKEN" 'https://<sddc-mgr>/v1/tasks?status=FAILED' | python3 -m json.tool. This returns all failed tasks with: task ID, type (HOST_COMMISSION, DOMAIN_CREATE, UPGRADE, etc.), timestamp, status. For the most recent failure, note the task ID for detailed inspection.

Step 6
Deep-dive into failed task subtasks. curl -s -k -H "Authorization: Bearer $TOKEN" 'https://<sddc-mgr>/v1/tasks/<task-id>/subtasks' | python3 -m json.tool. Each subtask represents a step in the lifecycle operation. Find the subtask with status='FAILED' — its 'description' and 'errors' fields identify the exact failure point. Example: 'NSX Transport Node Preparation' failed → look at NSX Manager logs for the specific host.
Step 7

Correlate API findings with component logs. Based on the failed subtask, SSH to the relevant component and examine logs: (a) If vCenter-related: /var/log/vmware/vpxd/vpxd.log on vCenter; (b) If NSX-related: /var/log/proton/nsxapi.log on NSX Manager; (c) If ESXi-related: /var/log/hostd.log on the specific host; (d) If SDDC Manager: /var/log/vmware/vcf/domainmanager/domainmanager.log. Use the subtask timestamp to find relevant log entries: grep '<timestamp>' <logfile>.

Step 8
Password validation via API. curl -s -k -H "Authorization: Bearer $TOKEN" 'https://<sddc-mgr>/v1/credentials' → lists all managed credentials with: component, username, credential type, last rotation date. For password-related issues: compare SoS password-check results with API credential inventory. If a password was changed outside VCF, the API shows the old password while the component has the new one — causing authentication failures in lifecycle operations.
Step 9

Retry a failed task via API. After resolving the root cause: curl -s -k -X PATCH -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' 'https://<sddc-mgr>/v1/tasks/<task-id>' -d '{"status":"RETRY"}'. This restarts the failed task from the failed subtask (not from the beginning). Monitor: poll GET /v1/tasks/<task-id> until status changes to SUCCESSFUL or FAILED again.

Step 10
Build troubleshooting decision tree. Document the diagnostic workflow as a decision tree: (1) Symptom observed → (2) Run SoS health-check → (3) Any RED findings? → Yes: remediate RED items first → (4) Check API task history for recent failures → (5) Inspect failed subtasks → (6) Correlate with component logs → (7) Resolve root cause → (8) Retry via API → (9) Verify with SoS health-check post-resolution. This decision tree is a VCAP exam deliverable and a production operations artifact.

Validation Gate

Check: SoS and API diagnostic workflow mastered

Expected: SoS health-check executed and interpreted, log bundle collected with targeted filters, API authentication and queries working, failed task subtasks inspected and correlated with logs, task retry procedure validated

Common Errors

SoS health-check fails to run
Fix: Check: (a) running as root (sudo required); (b) Python dependencies installed; (c) sufficient disk space for output (/tmp needs 5+ GB for log collection). If SoS script itself fails: check /opt/vmware/sddc-support/sos --version for compatibility.
API token request returns 401
Fix: Credentials incorrect or account locked. Default API user is admin@local (not administrator@vsphere.local). Check: password expiry, account lockout policy. If locked: wait for lockout timeout or reset via SDDC Manager console.
Task retry fails with same error
Fix: Root cause was not fully resolved. Re-examine subtask errors — the failure may have a prerequisite that was missed. Common: DNS entry was added but NTP skew wasn't fixed — the retried task fails on the next subtask that depends on time sync.
Log bundle is too large to upload (>10 GB)
Fix: Use targeted collection: --component and --start-date/--end-date filters reduce bundle size significantly. For support cases: collect only the relevant component and a 24-hour window around the failure.

Final Validation

VCF troubleshooting toolkit mastered — SoS, API, and log correlation

✓ SoS health-check → HTML report generated with all checks evaluated

✓ Targeted log bundle → Compressed bundle created with component and time filters

✓ API diagnostics → Token auth, task query, subtask inspection all functional

✓ Log correlation → Failed subtask matched to specific component log entries

✓ Task retry → Failed task retried via API after root cause resolution

Cleanup / Restore

• Clean up SoS output files from /tmp

• Close API session (tokens expire after 6 hours)

Design Reflection (VCDX)

Troubleshooting methodology demonstrates operational maturity. VCDX defense: show you have a systematic approach (SoS → API → logs) rather than random log searching. Discuss: how do you scale troubleshooting across a fleet with 50+ VCF Instances?

Requirements

  • Sub-30-minute root cause identification for lifecycle failures
  • Targeted log collection for efficient support case resolution
  • Programmatic task retry without re-running entire workflows

Constraints

  • SoS health-check takes 10-30 minutes depending on environment size
  • Log bundles can be 10+ GB without targeted filtering
  • API token expires after 6 hours — long-running scripts need token refresh

Assumptions

  • SSH access to SDDC Manager is available
  • REST API client (curl) is available on the admin workstation
  • Component logs have not been rotated or purged since the failure occurred

Risks

  • SoS health-check during active lifecycle operation may show false positives
  • Retrying a task without understanding root cause makes the situation worse
  • Unfiltered log collection fills /tmp and causes SDDC Manager disk space issues

Self-Assessment Discussion Prompts

  1. What information should be collected before opening a VMware support case?
  2. How do you troubleshoot a VCF lifecycle failure when SoS itself fails to run?
  3. When would you use SoS vs direct component log analysis?

Extensions

Script the full diagnostic workflow: SoS → API query → log extraction → email report

Build a monitoring dashboard that runs SoS health-check daily and trends results

Create a PowerShell module for common SDDC Manager API diagnostic queries

⚠ Known Pitfalls (from Community KB)

Running SoS log-collection without filters — generates massive bundles that fill /tmp and are too large to upload
Starting troubleshooting with logs instead of SoS health-check — wastes time looking at the wrong component
Retrying failed tasks before running SoS to verify overall health — the root cause may be a systemic issue (DNS, NTP) not a task-specific failure
Forgetting to correlate SoS findings with the failed subtask — SoS may reveal the root cause before you even look at logs

References

  • VCF 9.0 Troubleshooting Guide — SoS Utility: techdocs.broadcom.com
  • SDDC Manager API Reference: developer.broadcom.com
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.