Troubleshooting Holodeck Deployment Failures
Objectives
- Navigate the Holodeck log directory structure and distinguish between Cloud Builder, SDDC Manager, and component-specific logs
- Diagnose HoloRouter failures (FRR, DHCP, DNS, SSH) using low-level commands and kernel logs
- Identify and resolve nested ESXi provisioning failures related to CPU compatibility, boot.cfg, and datastore issues
- Troubleshoot VCF bring-up failures including Cloud Builder initialization, NSX version mismatches, and SDDC Manager deployment stalls
- Construct a personal troubleshooting runbook by mapping symptoms to log locations, root causes, and fixes using the Holodeck KB
Prerequisites
Holodeck 9.0.2.19 with a FAILED deployment attempt (or snapshot from holodeck-02 that can be rolled back mid-deployment). Knowledge of both successful and failed deployment logs from holodeck-02.
Prior labs: holodeck-02
Required skills:
- Basic Linux command-line navigation (find, grep, tail, less)
- PowerShell log filtering and string matching
- SSH access to nested appliances
- JSON editing and log file parsing
- Understanding of VCF bring-up sequence from holodeck-02
Lab Environment
Same as holodeck-02: Holodeck pod with HoloRouter, HoloConsole, 4x nested ESXi hosts, and optionally a partially-failed deployment to analyze. Candidate can optionally trigger failures intentionally to practice troubleshooting (see extension 'Deliberately Introduce and Recover').
graph TB PHY[Physical ESXi Host] --> HR[HoloRouter 10.1.1.1] PHY --> HC[HoloConsole/Webtop] HR --> ESXi1[ESXi-01 10.1.1.101] HR --> CB[Cloud Builder 10.1.1.200] CB -->|logs monitored| LOGS[/holodeck-runtime/logs/] CB -.->|Prepare Phase| FRR[FRR/BGP] CB -.->|Bring-up Phase| VC[vCenter] CB -.->|Bring-up Phase| NSX[NSX Manager]
IP Addressing
| Network | Purpose | VLAN |
|---|---|---|
10.1.1.0/20 | Management network (Cloud Builder, SDDC Manager, vCenter, NSX) | VLAN 1644 |
10.1.1.0/24 | ESXi host management VMkernel | VLAN 1644 |
10.1.2.0/24 | vMotion | VLAN 1645 |
10.1.3.0/24 | vSAN | VLAN 1646 |
10.1.4.0/24 | NSX Host TEP | VLAN 1647 |
10.1.5.0/24 | NSX Edge TEP | VLAN 1648 |
Credentials
| System | Username | Password |
|---|---|---|
| HoloConsole (Webtop) | holodeck | Default Holodeck password from Toolkit docs |
| HoloRouter (SSH) | admin | Default admin password for FRR appliance |
| Cloud Builder | admin@local | Set by Holodeck config.json cbPassword field |
| Nested ESXi Hosts (SSH) | root | Set by Holodeck config.json esxiPassword field |
| SDDC Manager (if deployed) | administrator@vsphere.local | Set in VCF bring-up spec JSON |
| API Authentication | administrator@vsphere.local | For API calls, use doubled password: VMware123!VMware123! (not VMware123!) |
Tasks
Task 1 Understand Holodeck log structure and reading strategy
manageabilityIn a production VCF environment, operators spend 70% of troubleshooting time in logs. Mastering log navigation is a foundational VCDX skill — panelists will expect you to articulate where to look first when something fails. This task builds that muscle memory in the Holodeck context.
From HoloConsole PowerShell, navigate to the Holodeck runtime directory: cd C:\Holodeck\holodeck-runtime
List the log directory structure: Get-ChildItem .\logs -Recurse | Select-Object -Property FullName | Format-List
Review the Prepare phase log: Get-Content .\logs\prepare-*.log -Tail 100 | Out-String
Review the Start phase log: Get-Content .\logs\start-*.log -Tail 100 | Out-String
Review Cloud Builder logs via PowerShell. Execute: Get-ChildItem .\logs\cloud-builder-*.log | ForEach-Object { Write-Host "File: $($_.Name)"; Get-Content $_.FullName -Tail 50 }
Examine SDDC Manager and vCenter logs inside the nested environment via SSH. Run: ssh root@10.1.1.101 -c 'tail -100 /var/log/vmware/vcf/sddc-manager/domainmanager.log' (replace IP with any management host)
Build a personal 'Log Map' document. Create a table with columns: [Component, Log File, Expected Success Indicators, Common Failure Keywords]. Example row: HoloRouter | /var/log/frr.log | 'bgpd: BGP Initialized' | 'Connection refused', 'FRR crashed'.
Validation Gate
Check: You can identify and access the following logs: (1) Prepare phase log, (2) Start phase log, (3) Cloud Builder logs, (4) Nested SDDC Manager logs via SSH. You have created a personal Log Map document.
Expected: All four log sources accessible and documented. Log Map has at least 5 component entries.
Common Errors
Task 2 Diagnose HoloRouter failures (FRR, DHCP, DNS, SSH)
recoverabilityHoloRouter is the network heart of the Holodeck pod. When it fails, the entire nested environment is isolated. VCDX panelists expect you to understand networking fundamentals: BGP, DHCP, DNS, and what happens when the gateway goes down. This task exercises that knowledge in a remediation context.
Verify HoloRouter connectivity: ping 10.1.1.1 from HoloConsole PowerShell
SSH into the HoloRouter and check FRR status: ssh admin@10.1.1.1 -c 'systemctl status frr'
If FRR is not running, restart it: ssh admin@10.1.1.1 -c 'systemctl restart frr && sleep 5 && systemctl status frr'
Check FRR logs for errors: ssh admin@10.1.1.1 -c 'tail -50 /var/log/frr/frr.log'
Verify DHCP service on HoloRouter is operational (it assigns IPs to nested ESXi hosts): ssh admin@10.1.1.1 -c 'systemctl status dnsmasq'
Check DHCP lease assignments: ssh admin@10.1.1.1 -c 'cat /var/lib/misc/dnsmasq.leases'
Verify DNS resolution from a nested ESXi host: ssh root@10.1.1.101 -c 'nslookup sddc-manager.site-a.vcf.lab 10.1.1.1'
Create a HoloRouter health check script. In PowerShell on HoloConsole, save the following as check-holorouter.ps1:
$checks = @(
@{ Name='HoloRouter Ping'; Cmd='ping -c 1 10.1.1.1' },
@{ Name='FRR Status'; Cmd='ssh admin@10.1.1.1 systemctl status frr' },
@{ Name='DHCP Status'; Cmd='ssh admin@10.1.1.1 systemctl status dnsmasq' },
@{ Name='BGP Neighbors'; Cmd='ssh admin@10.1.1.1 vtysh -c "show ip bgp summary"' }
)
foreach ($check in $checks) {
Write-Host "\n[$($check.Name)]" -ForegroundColor Cyan
Invoke-Expression $check.Cmd
}
Validation Gate
Check: HoloRouter ping succeeds. FRR status is 'active (running)'. DHCP shows at least 4 leases. DNS resolution works from a nested ESXi host. You have created a health check script.
Expected: HoloRouter fully operational with all connectivity paths verified. Health check script ready for future deployments.
Common Errors
Task 3 Diagnose ESXi host provisioning failures (CPU compat, boot.cfg, datastore)
availabilityNested ESXi provisioning is where hardware constraints bite hardest. Legacy CPU support, boot.cfg settings, and datastore capacity are common culprits. VCDX panelists probe your understanding of ESXi boot sequences and hardware compatibility — this task ensures you can diagnose these issues methodically.
Check the physical ESXi host CPU capabilities: On physical host PowerShell, run: Get-VMHost | Select-Object -Property Name, @{N='CPUModel'; E={$_.ProcessorType}}, CpuCores
Check if legacy CPU support is needed: Review the Holodeck deployment logs for any mentions of 'legacy CPU', 'CPUID', or 'unknown processor'. Run: Select-String -Path '.\logs\*' -Pattern 'legacy|CPUID|unknown processor'
If legacy CPU support is required, check the nested ESXi boot.cfg. From a nested ESXi host, run: ssh root@10.1.1.101 -c 'cat /bootbank/boot.cfg | grep -i allowLegacyCPU'
If allowLegacyCPU is missing, add it. SSH to the nested ESXi host and edit boot.cfg: ssh root@10.1.1.101 -c 'echo "allowLegacyCPU=TRUE" >> /bootbank/boot.cfg && cat /bootbank/boot.cfg | grep allowLegacy'
Verify physical datastore capacity and accessibility. On physical ESXi, run: Get-Datastore | Select-Object -Property Name, @{N='FreeGB'; E={[Math]::Round($_.FreeSpaceGB, 1)}}, @{N='CapacityGB'; E={[Math]::Round($_.CapacityGB, 1)}}
Verify the datastore specified in config.json exists and is accessible. From HoloConsole, run: Get-Content .\templates\config.json | ConvertFrom-Json | Select-Object -Property targetDatastore | Format-List
Check nested ESXi VM disk provisioning. List the nested ESXi VM files on the datastore: ssh root@10.1.1.101 -c 'df -h /'
Create a pre-flight checklist script. In PowerShell on HoloConsole, save the following:
$checks = @(
@{ Check='Physical ESXi CPU'; Cmd='Get-VMHost | Select-Object ProcessorType' },
@{ Check='Datastore Free Space'; Cmd='Get-Datastore | Where-Object Name -eq (Get-Content .\templates\config.json | ConvertFrom-Json).targetDatastore | Select-Object Name, FreeSpaceGB' },
@{ Check='Nested ESXi allowLegacyCPU'; Cmd='ssh root@10.1.1.101 grep allowLegacyCPU /bootbank/boot.cfg' }
)
foreach ($check in $checks) {
Write-Host "[$($check.Check)]" -ForegroundColor Cyan
Invoke-Expression $check.Cmd
}
Validation Gate
Check: You can identify physical CPU type, verify datastore capacity is >1.5 TB, confirm boot.cfg has (or needs) allowLegacyCPU setting, and check nested ESXi disk usage. Pre-flight checklist script created.
Expected: All four environmental checks completed. Pre-flight script ready for future deployments.
Common Errors
Task 4 Diagnose VCF bring-up failures (Cloud Builder, NSX, SDDC Manager)
availabilityThe bring-up phase is where most failures occur. Cloud Builder initialization, NSX bundle version mismatches, and SDDC Manager startup issues are the big three. VCDX panelists expect you to methodically narrow down which component failed and why — this task builds that diagnostic skill.
Monitor Cloud Builder UI during bring-up (or check logs post-failure). From HoloConsole Firefox, navigate to https://10.1.1.200 and log in with admin@local credentials. Watch the 'Bring-up Progress' dashboard.
Check Cloud Builder initialization logs: ssh admin@10.1.1.200 -c 'tail -100 /opt/vmware/vcf/domainmanager/logs/domainmanager.log | grep -i "error\|fail\|warn"'
Check NSX bundle version in the Cloud Builder logs. Search for 'NSX' and 'version': ssh admin@10.1.1.200 -c 'grep -i "nsx" /opt/vmware/vcf/domainmanager/logs/domainmanager.log | grep -i "version\|bundle"'
If NSX version mismatch is detected, verify the content manifest and NSX bundle. From HoloConsole, run: Get-Content .\templates\config.json | ConvertFrom-Json | Select-Object -Property contentManifest | Format-List
Check SDDC Manager service status and logs on the Cloud Builder or deployed SDDC Manager. Run: ssh root@10.1.1.4 -c 'systemctl status domainmanager && tail -50 /opt/vmware/vcf/sddc-manager/logs/domainmanager.log'
Verify DNS and NTP on Cloud Builder (critical for all VCF services). Run: ssh admin@10.1.1.200 -c 'echo "=== DNS ===" && cat /etc/resolv.conf && echo "=== NTP ===" && timedatectl status'
Create a VCF bring-up health check. In PowerShell, save the following:
Write-Host "VCF Bring-up Diagnostics" -ForegroundColor Green
Write-Host "\n[1] Cloud Builder Accessibility"
ping -c 1 10.1.1.200 | Select-Object -Last 1
Write-Host "\n[2] Cloud Builder Service Status"
ssh admin@10.1.1.200 'systemctl status vcf-bringup | grep Active'
Write-Host "\n[3] NSX Bundle Version"
ssh admin@10.1.1.200 'grep -i nsx /opt/vmware/vcf/domainmanager/logs/domainmanager.log | grep version | tail -1'
Write-Host "\n[4] SDDC Manager Status"
ssh root@10.1.1.4 'systemctl status domainmanager | grep Active'
Write-Host "\n[5] DNS and NTP"
ssh admin@10.1.1.200 'timedatectl status | grep -i synchronized'
Validation Gate
Check: You can access Cloud Builder UI, read domainmanager.log for errors, identify NSX version, verify DNS/NTP, and check SDDC Manager service status. Health check script created.
Expected: All five bring-up diagnostic checks working. Health check script ready for deployment monitoring.
Common Errors
Task 5 Build a personal troubleshooting runbook from symptoms to recovery
manageabilityA troubleshooting runbook is a decision tree that maps symptoms (what you observe) to root causes and fixes. In production, this is the single most valuable document an operator can create. VCDX panelists expect you to demonstrate this capability — not just reactive troubleshooting, but proactive documentation of the troubleshooting process itself.
Create a runbook template in JSON or YAML format (or a markdown table). The runbook should have columns: [Symptom, Where to Look First, Root Cause Check Commands, Likely Causes, Remediation Steps, KB Thread Reference]. Start by documenting 3 failures you've already analyzed (HoloRouter down, ESXi provisioning failure, NSX version mismatch).
Add three more failures to your runbook by reading the KB threads directly. For each of the following KB threads, extract: symptom, first diagnostic command, root cause, fix, and any KB reference:
- Thread 8: ESXi CPU compatibility (allowLegacyCPU)
- Thread 51: Datastore selection failure
- Thread 42: HoloRouter DHCP configuration
Format each as a runbook entry.
Create a decision tree flowchart (ASCII art or mermaid diagram) that shows the diagnostic flow: [Deployment fails] -> [Check phase] -> [Check component] -> [Read logs] -> [Identify root cause] -> [Apply fix]. Map each branch to a runbook entry.
Example pseudocode:
if phase == 'Prepare' then
check HoloRouter ping
if fails: HoloRouter down (see Runbook #1)
else: check FRR status
else if phase == 'Start' then
check Cloud Builder logs
if NSX mismatch: update manifest (see Runbook #5)
else if DNS fails: verify HoloRouter dnsmasq (see Runbook #3)
Test your runbook against a real failure (or a simulated one). Take the snapshot 'holodeck-02-complete' and intentionally introduce a failure:
Option A: Kill the HoloRouter FRR service: ssh admin@10.1.1.1 -c 'systemctl stop frr'
Option B: Delete the NSX bundle from the content directory (to trigger manifest mismatch)
Option C: Misconfigure the datastore name in config.json (to trigger 'Cannot select datastore' error)
Then follow your runbook to diagnose the failure end-to-end.
Document lessons learned. For each failure you diagnosed (real or simulated), record: (a) Time spent diagnosing, (b) Key insight that helped you find the root cause, (c) What you would do differently next time. Add this as a 'Lessons' section to your runbook.
Share your runbook with your study group or save it in your VCDX study materials folder. A well-written runbook is a VCDX panel gold mine — panelists will probe how you would troubleshoot, and your runbook is proof you've done it systematically.
Validation Gate
Check: Runbook has at least 6 failure scenarios, each with: symptom, diagnostic commands, root cause, fix, KB reference. Decision tree flowchart completed. Runbook tested against at least one real/simulated failure. Lessons documented.
Expected: A comprehensive, tested troubleshooting runbook ready for production use and VCDX panel discussion.
Common Errors
Final Validation
You have built operational maturity in Holodeck troubleshooting. You understand the log architecture, can diagnose the most common failure points (HoloRouter, ESXi, Cloud Builder, NSX), and have documented a personal troubleshooting runbook grounded in real experience. This is exactly what VCDX panelists probe during the design defense.
✓ Task 1: Personal Log Map created with at least 5 component entries → Log Map document accessible and complete
✓ Task 2: HoloRouter health check script created and tested → Script runs and confirms FRR, DHCP, DNS operational
✓ Task 3: ESXi pre-flight checklist script created → Script checks CPU type, datastore capacity, allowLegacyCPU setting
✓ Task 4: VCF bring-up health check script created → Script monitors Cloud Builder, NSX version, SDDC Manager, DNS/NTP
✓ Task 5: Troubleshooting runbook with 6+ scenarios, decision tree, lessons learned → Runbook tested against at least one real/simulated failure
Cleanup / Restore
Snapshot: holodeck-03-complete
• Save all scripts and runbook to C:\Holodeck\holodeck-runtime\scripts\ and your VCDX study folder
• If you intentionally triggered failures in Task 5, revert to a clean snapshot: Restore-VMSnapshot -Snapshot (Get-VM 'Holo-*' | Get-Snapshot -Name 'holodeck-02-complete' | Select-Object -First 1) -Confirm:$false
• Document which KB threads you found most helpful and add them to your study notes
Design Reflection (VCDX)
A VCDX panelist examining your troubleshooting approach will ask: 'Walk me through how you would diagnose a failed VCF deployment.' They're listening for: (1) systematic log navigation (not random guessing), (2) understanding of component dependencies (HoloRouter -> nested ESXi -> Cloud Builder -> SDDC Manager -> NSX), (3) knowledge of common pitfalls from the community, (4) evidence that you've documented the process. Your runbook IS the evidence. Be prepared to explain how you built it, which scenarios you've tested, and what you learned. VCDX is about recoverability and manageability — demonstrate that you can recover from failure methodically.
Requirements
- Deployment failures must be diagnosed systematically using logs, not guesswork
- Troubleshooting process must be repeatable and documented for future operators
- Solutions must address root cause, not symptoms (e.g., fix boot.cfg, not just restart)
- Lessons learned from each failure must feed back into operational procedures
Constraints
- Log volume in Holodeck can be overwhelming — requires filtering strategy (grep, Select-String)
- SSH access to nested VMs assumes HoloRouter is at least partially functional
- Some failures (e.g., corrupted datastore) cannot be fully recovered without a rebuild
- Runbook relies on assumption that KB threads are accurate — test entries empirically before trusting them
- Time-sensitive troubleshooting during bring-up (e.g., NTP drift) requires real-time monitoring, not post-failure log analysis
Assumptions
- Operator has SSH access to HoloRouter, nested ESXi, and Cloud Builder
- HoloRouter is reachable (at least for pings) to establish baseline connectivity
- Holodeck logs are preserved (not deleted mid-deployment)
- Runbook is kept current — adding new entries as new failure modes are discovered
- Operator has read access to Holodeck KB and can extract lessons from forum discussions
Risks
- Over-reliance on runbook without understanding root causes — IMPACT: runbook entries become cargo-cult procedures, MITIGATION: understand the 'why' behind each fix, not just the 'how'
- Runbook outdated after VCF or Holodeck version upgrade — IMPACT: procedures fail on new versions, MITIGATION: version-tag runbook entries and review them after each upgrade
- False confidence from testing on 'clean' snapshots — IMPACT: real-world failures are messier than lab scenarios, MITIGATION: intentionally vary failure injection (kill services, misconfigure DNS, etc.)
- Documentation burden overwhelms troubleshooting velocity — IMPACT: you skip writing runbook to focus on recovery, MITIGATION: allocate 30 minutes post-incident for runbook update
Self-Assessment Discussion Prompts
- Walk me through your runbook structure. How did you organize the failure scenarios? What would change if you had to troubleshoot a VCF 5.2.x deployment?
- You encounter a failure not in your runbook. How do you approach it? (Answer: decision tree + log reading + hypothesis testing, then add to runbook.)
- Describe the dependency chain in a Holodeck deployment. If HoloRouter goes down, what cascades? How would you recover? (Answer: all nested VMs lose network access, nested ESXi can't PXE boot, deployment hangs indefinitely.)
- NSX version mismatch is common. Why does VCF 9.0.2 specifically require NSX 9.0.2.0.25150386? (Answer: SDDC Manager/vCenter APIs expect specific NSX REST endpoint versions; mismatch causes validation failures.)
- Boot.cfg and legacy CPU support. Why is this a gotcha? When do you need it? (Answer: older hardware doesn't support modern VMX instructions; nested ESXi fails to boot without allowLegacyCPU=TRUE.
- Your runbook helped solve a failure in 10 minutes. How do you know it's actually correct, not just lucky? (Answer: test it on a fresh deployment with the same failure condition; if it works twice, it's probably correct.)
Extensions
Build a Holodeck monitoring dashboard with real-time alerts
Create a PowerShell script that monitors HoloRouter, Cloud Builder, SDDC Manager, and vCenter health every 60 seconds and logs alerts (email or Teams webhook) when services go down or logs show ERROR keywords. This simulates production monitoring practices and is directly applicable to VCDX Day 2 operational excellence discussions.
harderReproduce and document 5 different Holodeck failure modes
Intentionally trigger each of these failures in sequence, document the symptoms, run through your runbook to diagnose, fix, and recover. Failures: (1) HoloRouter FRR crash, (2) Cloud Builder DNS misconfiguration, (3) Datastore full, (4) NSX version mismatch, (5) ESXi boot.cfg corruption. For each, measure time-to-diagnose before and after runbook.
much harderTranslate Holodeck troubleshooting to production VCF
Compare your Holodeck runbook to production VCF troubleshooting practices. Which principles carry over? What changes in a non-nested environment? (Answer: HoloRouter gone, physical network replaces it; NSX version mismatches work the same; datastore issues same logic but at physical scale.) Document the translation as a 'Holodeck to Production' mapping.
same⚠ Known Pitfalls (from Community KB)
References
- VMware Cloud Foundation 9.0 Troubleshooting GuideTier 1 — Official
Official Broadcom documentation on VCF troubleshooting — covers bring-up failures, component health checks, and recovery procedures - Holodeck Toolkit GitHub IssuesTier 1 — Official
Community-reported Holodeck issues and official Broadcom responses. Tracks known bugs and workarounds. - VCF Holodeck Broadcom Community ForumTier 1 — Official
63 community threads on Holodeck troubleshooting. 'Holodeck KB normalized' used in this lab is derived from these threads. - PowerShell Select-String and log filtering best practicesTier 2 — VMware Press
Microsoft documentation on efficient log searching in PowerShell. Key for rapid troubleshooting. - FRR (Free Range Routing) BGP TroubleshootingTier 2 — VMware Press
FRR documentation covering BGP diagnostics. Applies to HoloRouter BGP troubleshooting (systemctl restart frr, vtysh diagnostics). - William Lam — VCF Deployment and TroubleshootingTier 3 — Expert Blog
VMware Staff Engineer's blog with deep dives on VCF bring-up failures, log analysis, and recovery procedures.