Lab: SoS Diagnostics & VCF API Troubleshooting
Objectives
- Execute SoS health checks and interpret results
- Collect targeted log bundles for support cases
- Use SDDC Manager REST API for diagnostic queries
- Analyze failed lifecycle tasks via API subtask inspection
- Build a troubleshooting decision tree for common VCF issues
Prerequisites
VCF 9.0 Instance operational, SSH access to SDDC Manager, REST API client (curl, Postman, or similar)
Required skills:
- Linux CLI basics
- REST API concepts
- Log analysis
Lab Environment
VCF 9.0 with SDDC Manager accessible via SSH and HTTPS. At least one completed lifecycle task (domain creation, host addition) in task history for API analysis.
Tasks
Task 1 SoS Diagnostics & API-Driven Troubleshooting
Master VCF's built-in diagnostic tools: SoS for health checks and log collection, SDDC Manager REST API for programmatic diagnostics, and log correlation for root cause analysis — the primary troubleshooting toolkit for VCAP-level VCF administration.
Run SoS health check. SSH to SDDC Manager VM: sudo /opt/vmware/sddc-support/sos --health-check --domain <management-domain-name>. The health check examines: DNS resolution (forward + reverse for all components), NTP synchronization (skew < 1 second), certificate validity (all certs checked), service status (all VCF services running), component connectivity (SDDC Mgr ↔ vCenter ↔ NSX ↔ ESXi), password validity (credential store vs actual), disk space (all components), vSAN health. Output: HTML report at /tmp/sos-<timestamp>/health-check-report.html.
Interpret SoS results. Open the HTML report in a browser (SCP to local machine or browse via vCenter console). Each check shows: GREEN (pass), YELLOW (warning), RED (fail). Common findings: (a) NTP skew > 1s on specific hosts — causes certificate validation failures; (b) DNS reverse lookup failure — blocks future lifecycle operations; (c) certificate expiry within 30 days — immediate rotation needed; (d) disk space low on SDDC Manager — log rotation needed. Document all RED and YELLOW findings as an action list.
Collect targeted log bundle. For a support case, collect logs for a specific time window and component: sudo /opt/vmware/sddc-support/sos --log-collection --domain <domain> --component vCenter --start-date 2026-05-13 --end-date 2026-05-14. This produces a smaller, targeted bundle vs. collecting everything. Available component filters: vCenter, NSX, ESXi, SDDC-Manager. Bundle location: /tmp/sos-<timestamp>/.
Authenticate to SDDC Manager API. Using curl: TOKEN=$(curl -s -k -X POST 'https://<sddc-mgr>/v1/tokens' -H 'Content-Type: application/json' -d '{"username":"admin@local","password":"<password>"}' | python3 -c 'import sys,json; print(json.load(sys.stdin)["accessToken"])'). Verify: curl -s -k -H "Authorization: Bearer $TOKEN" 'https://<sddc-mgr>/v1/system/health' | python3 -m json.tool. This returns the overall VCF health summary.
Query lifecycle task history. curl -s -k -H "Authorization: Bearer $TOKEN" 'https://<sddc-mgr>/v1/tasks?status=FAILED' | python3 -m json.tool. This returns all failed tasks with: task ID, type (HOST_COMMISSION, DOMAIN_CREATE, UPGRADE, etc.), timestamp, status. For the most recent failure, note the task ID for detailed inspection.
Deep-dive into failed task subtasks. curl -s -k -H "Authorization: Bearer $TOKEN" 'https://<sddc-mgr>/v1/tasks/<task-id>/subtasks' | python3 -m json.tool. Each subtask represents a step in the lifecycle operation. Find the subtask with status='FAILED' — its 'description' and 'errors' fields identify the exact failure point. Example: 'NSX Transport Node Preparation' failed → look at NSX Manager logs for the specific host.
Correlate API findings with component logs. Based on the failed subtask, SSH to the relevant component and examine logs: (a) If vCenter-related: /var/log/vmware/vpxd/vpxd.log on vCenter; (b) If NSX-related: /var/log/proton/nsxapi.log on NSX Manager; (c) If ESXi-related: /var/log/hostd.log on the specific host; (d) If SDDC Manager: /var/log/vmware/vcf/domainmanager/domainmanager.log. Use the subtask timestamp to find relevant log entries: grep '<timestamp>' <logfile>.
Password validation via API. curl -s -k -H "Authorization: Bearer $TOKEN" 'https://<sddc-mgr>/v1/credentials' → lists all managed credentials with: component, username, credential type, last rotation date. For password-related issues: compare SoS password-check results with API credential inventory. If a password was changed outside VCF, the API shows the old password while the component has the new one — causing authentication failures in lifecycle operations.
Retry a failed task via API. After resolving the root cause: curl -s -k -X PATCH -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' 'https://<sddc-mgr>/v1/tasks/<task-id>' -d '{"status":"RETRY"}'. This restarts the failed task from the failed subtask (not from the beginning). Monitor: poll GET /v1/tasks/<task-id> until status changes to SUCCESSFUL or FAILED again.
Build troubleshooting decision tree. Document the diagnostic workflow as a decision tree: (1) Symptom observed → (2) Run SoS health-check → (3) Any RED findings? → Yes: remediate RED items first → (4) Check API task history for recent failures → (5) Inspect failed subtasks → (6) Correlate with component logs → (7) Resolve root cause → (8) Retry via API → (9) Verify with SoS health-check post-resolution. This decision tree is a VCAP exam deliverable and a production operations artifact.
Validation Gate
Check: SoS and API diagnostic workflow mastered
Expected: SoS health-check executed and interpreted, log bundle collected with targeted filters, API authentication and queries working, failed task subtasks inspected and correlated with logs, task retry procedure validated
Common Errors
Final Validation
VCF troubleshooting toolkit mastered — SoS, API, and log correlation
✓ SoS health-check → HTML report generated with all checks evaluated
✓ Targeted log bundle → Compressed bundle created with component and time filters
✓ API diagnostics → Token auth, task query, subtask inspection all functional
✓ Log correlation → Failed subtask matched to specific component log entries
✓ Task retry → Failed task retried via API after root cause resolution
Cleanup / Restore
• Clean up SoS output files from /tmp
• Close API session (tokens expire after 6 hours)
Design Reflection (VCDX)
Troubleshooting methodology demonstrates operational maturity. VCDX defense: show you have a systematic approach (SoS → API → logs) rather than random log searching. Discuss: how do you scale troubleshooting across a fleet with 50+ VCF Instances?
Requirements
- Sub-30-minute root cause identification for lifecycle failures
- Targeted log collection for efficient support case resolution
- Programmatic task retry without re-running entire workflows
Constraints
- SoS health-check takes 10-30 minutes depending on environment size
- Log bundles can be 10+ GB without targeted filtering
- API token expires after 6 hours — long-running scripts need token refresh
Assumptions
- SSH access to SDDC Manager is available
- REST API client (curl) is available on the admin workstation
- Component logs have not been rotated or purged since the failure occurred
Risks
- SoS health-check during active lifecycle operation may show false positives
- Retrying a task without understanding root cause makes the situation worse
- Unfiltered log collection fills /tmp and causes SDDC Manager disk space issues
Self-Assessment Discussion Prompts
- What information should be collected before opening a VMware support case?
- How do you troubleshoot a VCF lifecycle failure when SoS itself fails to run?
- When would you use SoS vs direct component log analysis?
Extensions
Script the full diagnostic workflow: SoS → API query → log extraction → email report
Build a monitoring dashboard that runs SoS health-check daily and trends results
Create a PowerShell module for common SDDC Manager API diagnostic queries
⚠ Known Pitfalls (from Community KB)
References
- VCF 9.0 Troubleshooting Guide — SoS Utility: techdocs.broadcom.com
- SDDC Manager API Reference: developer.broadcom.com