Academy/Holodeck Lab Setup & Operations/Troubleshooting Holodeck Deployment Failures
This lab targets VCF 9.0.2

Troubleshooting Holodeck Deployment Failures

VCF 9.0.2Intermediateadminsupportvcdx⏱ 120 min

Troubleshooting methodology applies across VCF 9.0.0–9.0.2. Log paths and component versions are specific to 9.0.2 + Holodeck 9.0.2.19 + ESXi 8.0 U3.

Objectives

  • Navigate the Holodeck log directory structure and distinguish between Cloud Builder, SDDC Manager, and component-specific logs
  • Diagnose HoloRouter failures (FRR, DHCP, DNS, SSH) using low-level commands and kernel logs
  • Identify and resolve nested ESXi provisioning failures related to CPU compatibility, boot.cfg, and datastore issues
  • Troubleshoot VCF bring-up failures including Cloud Builder initialization, NSX version mismatches, and SDDC Manager deployment stalls
  • Construct a personal troubleshooting runbook by mapping symptoms to log locations, root causes, and fixes using the Holodeck KB

Prerequisites

Holodeck 9.0.2.19 with a FAILED deployment attempt (or snapshot from holodeck-02 that can be rolled back mid-deployment). Knowledge of both successful and failed deployment logs from holodeck-02.

Prior labs: holodeck-02

Required skills:

  • Basic Linux command-line navigation (find, grep, tail, less)
  • PowerShell log filtering and string matching
  • SSH access to nested appliances
  • JSON editing and log file parsing
  • Understanding of VCF bring-up sequence from holodeck-02
📸 Starting State: S3-52 — Management Domain Deployed (VCF 5.2)

Lab Environment

Same as holodeck-02: Holodeck pod with HoloRouter, HoloConsole, 4x nested ESXi hosts, and optionally a partially-failed deployment to analyze. Candidate can optionally trigger failures intentionally to practice troubleshooting (see extension 'Deliberately Introduce and Recover').

graph TB
  PHY[Physical ESXi Host] --> HR[HoloRouter 10.1.1.1]
  PHY --> HC[HoloConsole/Webtop]
  HR --> ESXi1[ESXi-01 10.1.1.101]
  HR --> CB[Cloud Builder 10.1.1.200]
  CB -->|logs monitored| LOGS[/holodeck-runtime/logs/]
  CB -.->|Prepare Phase| FRR[FRR/BGP]
  CB -.->|Bring-up Phase| VC[vCenter]
  CB -.->|Bring-up Phase| NSX[NSX Manager]

IP Addressing

NetworkPurposeVLAN
10.1.1.0/20Management network (Cloud Builder, SDDC Manager, vCenter, NSX)VLAN 1644
10.1.1.0/24ESXi host management VMkernelVLAN 1644
10.1.2.0/24vMotionVLAN 1645
10.1.3.0/24vSANVLAN 1646
10.1.4.0/24NSX Host TEPVLAN 1647
10.1.5.0/24NSX Edge TEPVLAN 1648

Credentials

SystemUsernamePassword
HoloConsole (Webtop)holodeckDefault Holodeck password from Toolkit docs
HoloRouter (SSH)adminDefault admin password for FRR appliance
Cloud Builderadmin@localSet by Holodeck config.json cbPassword field
Nested ESXi Hosts (SSH)rootSet by Holodeck config.json esxiPassword field
SDDC Manager (if deployed)administrator@vsphere.localSet in VCF bring-up spec JSON
API Authenticationadministrator@vsphere.localFor API calls, use doubled password: VMware123!VMware123! (not VMware123!)

Tasks

Task 1 Understand Holodeck log structure and reading strategy

manageability

In a production VCF environment, operators spend 70% of troubleshooting time in logs. Mastering log navigation is a foundational VCDX skill — panelists will expect you to articulate where to look first when something fails. This task builds that muscle memory in the Holodeck context.

Step 1

From HoloConsole PowerShell, navigate to the Holodeck runtime directory: cd C:\Holodeck\holodeck-runtime

PowerShell prompt at the holodeck-runtime directory
Step 2

List the log directory structure: Get-ChildItem .\logs -Recurse | Select-Object -Property FullName | Format-List

A tree of log files organized by phase and component. Typical structure: logs/prepare-*.log, logs/start-*.log, logs/cloud-builder/*.log, logs/sddc-manager/*.log
Step 3

Review the Prepare phase log: Get-Content .\logs\prepare-*.log -Tail 100 | Out-String

Last 100 lines of Prepare log. Look for milestone lines like 'Prepare phase completed successfully' (success) or ERROR/FAIL keywords (failure).
Prepare phase logs are your first stop. They tell you if the HoloRouter and nested ESXi staging succeeded. If Prepare fails, the deployment never reaches the bring-up phase.
Step 4

Review the Start phase log: Get-Content .\logs\start-*.log -Tail 100 | Out-String

Last 100 lines of Start log. Track the phases: 'Nested ESXi deployment' -> 'Cloud Builder initialization' -> 'VCF bring-up' -> 'Host commissioning' -> 'Completion'.
The Start log is sequential. If the last lines are 'Step X of Y', the deployment got stuck on that step. Cross-reference with Cloud Builder logs for details.
Step 5

Review Cloud Builder logs via PowerShell. Execute: Get-ChildItem .\logs\cloud-builder-*.log | ForEach-Object { Write-Host "File: $($_.Name)"; Get-Content $_.FullName -Tail 50 }

Cloud Builder log files showing SDDC Manager deployment progress, component bring-up sequence, and validation results
Cloud Builder logs are voluminous and span multiple files. Use grep/Select-String to search for keywords: 'WARN', 'ERROR', 'FAILED', 'Validation', 'NSX'.
Step 6

Examine SDDC Manager and vCenter logs inside the nested environment via SSH. Run: ssh root@10.1.1.101 -c 'tail -100 /var/log/vmware/vcf/sddc-manager/domainmanager.log' (replace IP with any management host)

Domain Manager logs showing component initialization status, licensing checks, and cluster formation steps
SSH access to nested hosts requires the nested ESXi password from config.json. If SSH fails, the nested environment may not be healthy enough for in-depth troubleshooting.
Step 7

Build a personal 'Log Map' document. Create a table with columns: [Component, Log File, Expected Success Indicators, Common Failure Keywords]. Example row: HoloRouter | /var/log/frr.log | 'bgpd: BGP Initialized' | 'Connection refused', 'FRR crashed'.

A reference document (text or spreadsheet) you can consult during future deployments
This log map is your troubleshooting North Star. Invest 10 minutes now; it will save 1 hour in a real failure scenario.

Validation Gate

Check: You can identify and access the following logs: (1) Prepare phase log, (2) Start phase log, (3) Cloud Builder logs, (4) Nested SDDC Manager logs via SSH. You have created a personal Log Map document.

Expected: All four log sources accessible and documented. Log Map has at least 5 component entries.

Common Errors

Log files not found or directory is empty
Cause: Logs are not being generated because Holodeck Toolkit was run without -Verbose flag, or logs were cleaned up
Fix: Re-run deployment with -Verbose: New-HoloDeckInstance -ConfigFile .\templates\config.json -Verbose. Logs are generated only during active operations.
📋 KB: General practice — always use -Verbose for troubleshooting
SSH to nested ESXi fails: 'Connection refused' or timeout
Cause: Nested ESXi host not fully booted, or SSH service not running, or network isolation issue
Fix: Wait 5 minutes for the host to fully initialize. Verify the HoloRouter is reachable: ping 10.1.1.1. If HoloRouter is down, the entire nested network is isolated.
📋 KB: Thread 37: new 9.0.1 instance will not create custom iso

Task 2 Diagnose HoloRouter failures (FRR, DHCP, DNS, SSH)

recoverability

HoloRouter is the network heart of the Holodeck pod. When it fails, the entire nested environment is isolated. VCDX panelists expect you to understand networking fundamentals: BGP, DHCP, DNS, and what happens when the gateway goes down. This task exercises that knowledge in a remediation context.

Step 1

Verify HoloRouter connectivity: ping 10.1.1.1 from HoloConsole PowerShell

Reply from 10.1.1.1: bytes=32 time<1ms
If ping fails, the HoloRouter VM is down or its network is misconfigured. Check the physical ESXi host to see if the HoloRouter VM is still running: Get-VM -Name 'Holo-Router*' | Select-Object -Property Name, PowerState
Step 2

SSH into the HoloRouter and check FRR status: ssh admin@10.1.1.1 -c 'systemctl status frr'

Output showing 'active (running)' and no errors in the status message
If FRR status shows 'inactive (dead)' or 'failed', proceed to Step 3 to restart it.
FRR (Free Range Routing) provides the BGP daemon that advertises routes to the physical network. If it's dead, routing between the physical host and nested environment breaks.
Step 3

If FRR is not running, restart it: ssh admin@10.1.1.1 -c 'systemctl restart frr && sleep 5 && systemctl status frr'

Confirmation that FRR restarted and is now 'active (running)'
After restart, allow 30 seconds for BGP peering to stabilize. Then verify ping from HoloConsole again.
Step 4

Check FRR logs for errors: ssh admin@10.1.1.1 -c 'tail -50 /var/log/frr/frr.log'

Log entries showing BGP initialization, neighbor peering status, and no critical errors
Look for lines like 'bgpd: BGP Initialized' (good) or 'Connection refused' (bad). If you see connection refused, the BGP neighbor (physical host) may not be reachable.
Step 5

Verify DHCP service on HoloRouter is operational (it assigns IPs to nested ESXi hosts): ssh admin@10.1.1.1 -c 'systemctl status dnsmasq'

Output showing 'active (running)'
dnsmasq provides both DNS forwarding and DHCP service. If it's down, nested hosts cannot get IPs and the deployment stalls.
Step 6

Check DHCP lease assignments: ssh admin@10.1.1.1 -c 'cat /var/lib/misc/dnsmasq.leases'

A list of IP leases (one per line) showing nested ESXi hosts, Cloud Builder, and management components with their assigned IPs and MAC addresses
If you see fewer leases than expected, some VMs haven't obtained IPs — a sign of DHCP starvation or VM boot failure.
Step 7

Verify DNS resolution from a nested ESXi host: ssh root@10.1.1.101 -c 'nslookup sddc-manager.site-a.vcf.lab 10.1.1.1'

Name server response showing the resolved IP (typically 10.1.1.4 for SDDC Manager)
If DNS fails, the bring-up sequence hangs because services cannot reach each other by FQDN. The community thread 'Holodeck 9.0.2 DNS troubleshooting' (Thread 24 in KB) covers this.
Step 8

Create a HoloRouter health check script. In PowerShell on HoloConsole, save the following as check-holorouter.ps1:

$checks = @(
@{ Name='HoloRouter Ping'; Cmd='ping -c 1 10.1.1.1' },
@{ Name='FRR Status'; Cmd='ssh admin@10.1.1.1 systemctl status frr' },
@{ Name='DHCP Status'; Cmd='ssh admin@10.1.1.1 systemctl status dnsmasq' },
@{ Name='BGP Neighbors'; Cmd='ssh admin@10.1.1.1 vtysh -c "show ip bgp summary"' }
)

foreach ($check in $checks) {
Write-Host "\n[$($check.Name)]" -ForegroundColor Cyan
Invoke-Expression $check.Cmd
}

A reusable script you can run daily to confirm HoloRouter health
Save this script as C:\Holodeck\holodeck-runtime\scripts\check-holorouter.ps1 for future troubleshooting sessions.

Validation Gate

Check: HoloRouter ping succeeds. FRR status is 'active (running)'. DHCP shows at least 4 leases. DNS resolution works from a nested ESXi host. You have created a health check script.

Expected: HoloRouter fully operational with all connectivity paths verified. Health check script ready for future deployments.

Common Errors

Ping to 10.1.1.1 fails: 'Request timed out' or 'Destination host unreachable'
Cause: HoloRouter VM is down, powered off, or network is misconfigured on the physical host
Fix: On physical ESXi, verify the HoloRouter VM is powered on: Get-VM 'Holo-Router*' | Where-Object { $_.PowerState -eq 'PoweredOff' } | Start-VM. Wait 30 seconds for services to initialize.
📋 KB: General VCF operations practice
FRR restart fails: 'Job for frr.service failed because the control process exited with error code'
Cause: FRR configuration is corrupted, or a dependency (network interface) is missing
Fix: SSH to HoloRouter and reload the config: vtysh -c 'reload'. If that fails, check interface status: ip link show. If interfaces are down, restart networking: systemctl restart networking && systemctl restart frr
DHCP leases show only 1-2 entries, not 5+ for nested pods
Cause: DHCP is starving because the HoloRouter doesn't have a large enough lease pool defined, or nested VMs are failing to boot
Fix: Check HoloRouter dnsmasq config: cat /etc/dnsmasq.conf | grep dhcp-range. Typical range is 10.1.1.100-10.1.1.200 (100+ leases). If the range is too small, redeploy the HoloRouter with corrected config.
📋 KB: Thread 42: Configuration of HoloRouter Failed. Encountered error

Task 3 Diagnose ESXi host provisioning failures (CPU compat, boot.cfg, datastore)

availability

Nested ESXi provisioning is where hardware constraints bite hardest. Legacy CPU support, boot.cfg settings, and datastore capacity are common culprits. VCDX panelists probe your understanding of ESXi boot sequences and hardware compatibility — this task ensures you can diagnose these issues methodically.

Step 1

Check the physical ESXi host CPU capabilities: On physical host PowerShell, run: Get-VMHost | Select-Object -Property Name, @{N='CPUModel'; E={$_.ProcessorType}}, CpuCores

Output showing CPU model and core count. Example: 'Intel(R) Xeon(R) CPU E5-2697 v4' with 36 cores
Note the exact CPU model. If it's older than 2012, nested ESXi may fail to boot with legacy CPU support errors.
Step 2

Check if legacy CPU support is needed: Review the Holodeck deployment logs for any mentions of 'legacy CPU', 'CPUID', or 'unknown processor'. Run: Select-String -Path '.\logs\*' -Pattern 'legacy|CPUID|unknown processor'

Either (a) no matches (CPU is modern and supported), or (b) matches indicating legacy CPU mode is required
If legacy CPU errors appear in logs, the nested ESXi hosts will fail to boot. See Step 3 to fix this.
Step 3

If legacy CPU support is required, check the nested ESXi boot.cfg. From a nested ESXi host, run: ssh root@10.1.1.101 -c 'cat /bootbank/boot.cfg | grep -i allowLegacyCPU'

Either (a) a line showing 'allowLegacyCPU=TRUE' (correct), or (b) no output (missing — must be added)
According to KB Thread 8, if this line is missing and the host has a legacy CPU, the deployment will fail post-deployment during workload domain brings.
Step 4

If allowLegacyCPU is missing, add it. SSH to the nested ESXi host and edit boot.cfg: ssh root@10.1.1.101 -c 'echo "allowLegacyCPU=TRUE" >> /bootbank/boot.cfg && cat /bootbank/boot.cfg | grep allowLegacy'

Confirmation that the line was appended: 'allowLegacyCPU=TRUE'
This change persists across reboots. However, it must be applied BEFORE the Cloud Builder bring-up completes. If applied after, re-run the bring-up.
Step 5

Verify physical datastore capacity and accessibility. On physical ESXi, run: Get-Datastore | Select-Object -Property Name, @{N='FreeGB'; E={[Math]::Round($_.FreeSpaceGB, 1)}}, @{N='CapacityGB'; E={[Math]::Round($_.CapacityGB, 1)}}

List of datastores with their free capacity. Holodeck requires >1.5 TB free on the target datastore.
If free space < 500 GB, the ESXi provisioning step will fail. The community thread 'Holodeck vcf 9 - Cannot select datastore' (Thread 51) documents this as the #1 datastore issue.
Step 6

Verify the datastore specified in config.json exists and is accessible. From HoloConsole, run: Get-Content .\templates\config.json | ConvertFrom-Json | Select-Object -Property targetDatastore | Format-List

The datastore name from config.json. Cross-reference it against the Get-Datastore output from Step 5 to confirm it exists.
If the datastore name is misspelled or doesn't exist, Holodeck will fail with 'Cannot select datastore' error during nested VM provisioning.
Step 7

Check nested ESXi VM disk provisioning. List the nested ESXi VM files on the datastore: ssh root@10.1.1.101 -c 'df -h /'

Nested ESXi root filesystem usage (typically very low, ~5% used). If >90%, the host is running out of disk space.
Nested ESXi hosts have a small 32 GB root filesystem. If they're running high on disk, logs are being spammed or there are orphaned VM files.
Step 8

Create a pre-flight checklist script. In PowerShell on HoloConsole, save the following:

$checks = @(
@{ Check='Physical ESXi CPU'; Cmd='Get-VMHost | Select-Object ProcessorType' },
@{ Check='Datastore Free Space'; Cmd='Get-Datastore | Where-Object Name -eq (Get-Content .\templates\config.json | ConvertFrom-Json).targetDatastore | Select-Object Name, FreeSpaceGB' },
@{ Check='Nested ESXi allowLegacyCPU'; Cmd='ssh root@10.1.1.101 grep allowLegacyCPU /bootbank/boot.cfg' }
)

foreach ($check in $checks) {
Write-Host "[$($check.Check)]" -ForegroundColor Cyan
Invoke-Expression $check.Cmd
}

A pre-flight check script for future deployments
Run this before every deployment. It takes <1 minute and catches 90% of CPU/datastore issues before they waste 2 hours.

Validation Gate

Check: You can identify physical CPU type, verify datastore capacity is >1.5 TB, confirm boot.cfg has (or needs) allowLegacyCPU setting, and check nested ESXi disk usage. Pre-flight checklist script created.

Expected: All four environmental checks completed. Pre-flight script ready for future deployments.

Common Errors

Nested ESXi VMs created but fail to boot: 'This CPU is not supported'
Cause: Physical host has a legacy CPU (pre-2012), and allowLegacyCPU=TRUE was not set in boot.cfg
Fix: Destroy the failed instance, add allowLegacyCPU=TRUE to config.json (Holodeck will inject it on re-provision), and re-run Prepare. Reference: KB Thread 8.
📋 KB: Thread 8: Issue with legacy CPU support (boot.cfg)
'Cannot select datastore' error during nested VM provisioning
Cause: Target datastore from config.json doesn't exist, is full, or is inaccessible due to network/permissions issue
Fix: Verify datastore name in config.json against Get-Datastore output. Ensure >1.5 TB free. If shared with other VMs, shut them down to free space. Reference: KB Thread 51.
📋 KB: Thread 51: Holodeck vcf 9 - Cannot select datastore
Nested ESXi host shows 100% disk usage after deployment
Cause: VM logs are being spammed (e.g., by continuous errors), filling the 32 GB root filesystem
Fix: SSH to the host and check log size: du -sh /var/log/*. Clear old logs: find /var/log -name '*.log' -mtime +7 -delete. Restart the logging service: systemctl restart hostd.

Task 4 Diagnose VCF bring-up failures (Cloud Builder, NSX, SDDC Manager)

availability

The bring-up phase is where most failures occur. Cloud Builder initialization, NSX bundle version mismatches, and SDDC Manager startup issues are the big three. VCDX panelists expect you to methodically narrow down which component failed and why — this task builds that diagnostic skill.

Step 1

Monitor Cloud Builder UI during bring-up (or check logs post-failure). From HoloConsole Firefox, navigate to https://10.1.1.200 and log in with admin@local credentials. Watch the 'Bring-up Progress' dashboard.

A multi-stage progress tracker showing phases: Pre-Checks, SDDC Manager Deployment, vCenter Deployment, NSX Manager Deployment, Host Commissioning, Cluster Creation, Validation.
If a step shows 'Failed' or is stuck on 'In Progress' for >10 minutes, click the step to see detailed error messages. This is often more informative than PowerShell logs.
Step 2

Check Cloud Builder initialization logs: ssh admin@10.1.1.200 -c 'tail -100 /opt/vmware/vcf/domainmanager/logs/domainmanager.log | grep -i "error\|fail\|warn"'

Either (a) no matches (initialization succeeded), or (b) matches showing specific error: e.g., 'DNS resolution failed', 'NTP not synced', 'License validation failed'
The most common Cloud Builder issue is DNS or NTP not being configured on the HoloRouter. If you see DNS/NTP errors, jump to Step 6 (DNS verification).
Step 3

Check NSX bundle version in the Cloud Builder logs. Search for 'NSX' and 'version': ssh admin@10.1.1.200 -c 'grep -i "nsx" /opt/vmware/vcf/domainmanager/logs/domainmanager.log | grep -i "version\|bundle"'

Lines showing NSX bundle version being validated. Example: 'NSX install image version: 9.0.2.0.25150386' (correct for VCF 9.0.2) or 'NSX version mismatch' (error).
According to KB Thread 23, VCF 9.0.1/9.0.2 expects NSX 9.0.2.0.25150386 specifically. If the bundled NSX is 9.0.1.x, the validation will fail.
Step 4

If NSX version mismatch is detected, verify the content manifest and NSX bundle. From HoloConsole, run: Get-Content .\templates\config.json | ConvertFrom-Json | Select-Object -Property contentManifest | Format-List

Path to the manifest file. Example: '.\content\manifests\vcf-9.0.2-manifest.json'
Open the manifest file and search for 'nsx' to see the bundled NSX version and file checksum. If it's not 9.0.2.0.25150386, download the correct version and update the manifest.
Step 5

Check SDDC Manager service status and logs on the Cloud Builder or deployed SDDC Manager. Run: ssh root@10.1.1.4 -c 'systemctl status domainmanager && tail -50 /opt/vmware/vcf/sddc-manager/logs/domainmanager.log'

Service status (should be 'active (running)') and log tail showing successful initialization or specific error if deployment got stuck
If SDDC Manager logs show 'Waiting for cluster to stabilize' for >30 minutes, vCenter or NSX initialization is stalled. Check those components separately.
Step 6

Verify DNS and NTP on Cloud Builder (critical for all VCF services). Run: ssh admin@10.1.1.200 -c 'echo "=== DNS ===" && cat /etc/resolv.conf && echo "=== NTP ===" && timedatectl status'

resolv.conf showing nameserver entries (10.1.1.1 for HoloRouter, or external DNS). timedatectl showing 'synchronized: yes'.
If NTP is not synchronized (shows 'synchronized: no'), the SDDC Manager bring-up will hang waiting for clock skew checks. Restart NTP: systemctl restart systemd-timesyncd
Step 7

Create a VCF bring-up health check. In PowerShell, save the following:

Write-Host "VCF Bring-up Diagnostics" -ForegroundColor Green
Write-Host "\n[1] Cloud Builder Accessibility"
ping -c 1 10.1.1.200 | Select-Object -Last 1
Write-Host "\n[2] Cloud Builder Service Status"
ssh admin@10.1.1.200 'systemctl status vcf-bringup | grep Active'
Write-Host "\n[3] NSX Bundle Version"
ssh admin@10.1.1.200 'grep -i nsx /opt/vmware/vcf/domainmanager/logs/domainmanager.log | grep version | tail -1'
Write-Host "\n[4] SDDC Manager Status"
ssh root@10.1.1.4 'systemctl status domainmanager | grep Active'
Write-Host "\n[5] DNS and NTP"
ssh admin@10.1.1.200 'timedatectl status | grep -i synchronized'

A reusable diagnostic script showing all critical bring-up services in one view
Run this script every 30 seconds during bring-up to catch failures early. Saves hours of post-failure log analysis.

Validation Gate

Check: You can access Cloud Builder UI, read domainmanager.log for errors, identify NSX version, verify DNS/NTP, and check SDDC Manager service status. Health check script created.

Expected: All five bring-up diagnostic checks working. Health check script ready for deployment monitoring.

Common Errors

Cloud Builder stuck on 'VCF Installer is not ready yet' for >30 minutes
Cause: Cloud Builder VM is not fully initialized — DNS, NTP, or firewall blocking initialization services
Fix: SSH to Cloud Builder: ssh admin@10.1.1.200. Verify DNS: nslookup google.com. Verify NTP: timedatectl status. Verify services: systemctl status vcf-bringup. If services are dead, restart: systemctl restart vcf-bringup. Reference: KB Thread 5.
📋 KB: Thread 5: VCF Installer is not ready yet
NSX validation fails: 'Validate NSX install image fails'
Cause: NSX bundle version in content pack doesn't match what VCF 9.0.x bring-up expects (should be 9.0.2.0.25150386 for VCF 9.0.2)
Fix: Download NSX 9.0.2.0.25150386 from Broadcom support portal. Update contentManifest to reference the correct NSX bundle. Destroy instance and re-prepare. Reference: KB Thread 23.
📋 KB: Thread 23: Holodeck 9.0.1 VCF Deployment fails
SDDC Manager deployment starts but never completes: 'Waiting for cluster to stabilize'
Cause: vCenter or NSX Manager initialization is slow/failed, blocking SDDC Manager from fully deploying
Fix: Check vCenter status: ssh root@10.1.1.6 'systemctl status vpostgres && systemctl status vpxd'. Check NSX status: ssh admin@10.1.1.10 'systemctl status -l nsx-manager'. If either is dead, restart the service or check logs for root cause.

Task 5 Build a personal troubleshooting runbook from symptoms to recovery

manageability

A troubleshooting runbook is a decision tree that maps symptoms (what you observe) to root causes and fixes. In production, this is the single most valuable document an operator can create. VCDX panelists expect you to demonstrate this capability — not just reactive troubleshooting, but proactive documentation of the troubleshooting process itself.

Step 1

Create a runbook template in JSON or YAML format (or a markdown table). The runbook should have columns: [Symptom, Where to Look First, Root Cause Check Commands, Likely Causes, Remediation Steps, KB Thread Reference]. Start by documenting 3 failures you've already analyzed (HoloRouter down, ESXi provisioning failure, NSX version mismatch).

A runbook document (JSON, YAML, Markdown, or spreadsheet) with at least 3 failure scenarios fully documented
Use the exact symptoms, commands, and log paths from Tasks 1-4. This ensures the runbook is grounded in real troubleshooting.
Step 2

Add three more failures to your runbook by reading the KB threads directly. For each of the following KB threads, extract: symptom, first diagnostic command, root cause, fix, and any KB reference:

  • Thread 8: ESXi CPU compatibility (allowLegacyCPU)
  • Thread 51: Datastore selection failure
  • Thread 42: HoloRouter DHCP configuration

Format each as a runbook entry.

Runbook updated with 6 total scenarios (3 from hands-on troubleshooting + 3 from KB)
Step 3

Create a decision tree flowchart (ASCII art or mermaid diagram) that shows the diagnostic flow: [Deployment fails] -> [Check phase] -> [Check component] -> [Read logs] -> [Identify root cause] -> [Apply fix]. Map each branch to a runbook entry.

Example pseudocode:
if phase == 'Prepare' then
check HoloRouter ping
if fails: HoloRouter down (see Runbook #1)
else: check FRR status
else if phase == 'Start' then
check Cloud Builder logs
if NSX mismatch: update manifest (see Runbook #5)
else if DNS fails: verify HoloRouter dnsmasq (see Runbook #3)

A flowchart showing the decision logic for diagnosing failures
Step 4

Test your runbook against a real failure (or a simulated one). Take the snapshot 'holodeck-02-complete' and intentionally introduce a failure:

Option A: Kill the HoloRouter FRR service: ssh admin@10.1.1.1 -c 'systemctl stop frr'
Option B: Delete the NSX bundle from the content directory (to trigger manifest mismatch)
Option C: Misconfigure the datastore name in config.json (to trigger 'Cannot select datastore' error)

Then follow your runbook to diagnose the failure end-to-end.

Runbook successfully used to diagnose and fix the intentional failure (without consulting Task 1-4 instructions)
Step 5

Document lessons learned. For each failure you diagnosed (real or simulated), record: (a) Time spent diagnosing, (b) Key insight that helped you find the root cause, (c) What you would do differently next time. Add this as a 'Lessons' section to your runbook.

Lessons documented with timestamps and insights for each scenario
Step 6

Share your runbook with your study group or save it in your VCDX study materials folder. A well-written runbook is a VCDX panel gold mine — panelists will probe how you would troubleshoot, and your runbook is proof you've done it systematically.

Runbook finalized and stored for future reference (and VCDX preparation)

Validation Gate

Check: Runbook has at least 6 failure scenarios, each with: symptom, diagnostic commands, root cause, fix, KB reference. Decision tree flowchart completed. Runbook tested against at least one real/simulated failure. Lessons documented.

Expected: A comprehensive, tested troubleshooting runbook ready for production use and VCDX panel discussion.

Common Errors

Runbook is too generic or doesn't include actual command output
Cause: Copying KB descriptions without hands-on testing and validation
Fix: Re-test each runbook entry by running the actual commands in your environment. Include real command output (sanitized of IPs/passwords). This grounds the runbook in reality.
Runbook doesn't cover the failure you're currently experiencing
Cause: Runbook is incomplete — it only covers previously-seen failures, not the new one
Fix: Diagnose the new failure using the decision tree and log reading skills from Tasks 1-4. Document it as a new runbook entry so future you (or someone using your runbook) benefits.

Final Validation

You have built operational maturity in Holodeck troubleshooting. You understand the log architecture, can diagnose the most common failure points (HoloRouter, ESXi, Cloud Builder, NSX), and have documented a personal troubleshooting runbook grounded in real experience. This is exactly what VCDX panelists probe during the design defense.

✓ Task 1: Personal Log Map created with at least 5 component entries → Log Map document accessible and complete

✓ Task 2: HoloRouter health check script created and tested → Script runs and confirms FRR, DHCP, DNS operational

✓ Task 3: ESXi pre-flight checklist script created → Script checks CPU type, datastore capacity, allowLegacyCPU setting

✓ Task 4: VCF bring-up health check script created → Script monitors Cloud Builder, NSX version, SDDC Manager, DNS/NTP

✓ Task 5: Troubleshooting runbook with 6+ scenarios, decision tree, lessons learned → Runbook tested against at least one real/simulated failure

Cleanup / Restore

Snapshot: holodeck-03-complete

• Save all scripts and runbook to C:\Holodeck\holodeck-runtime\scripts\ and your VCDX study folder

• If you intentionally triggered failures in Task 5, revert to a clean snapshot: Restore-VMSnapshot -Snapshot (Get-VM 'Holo-*' | Get-Snapshot -Name 'holodeck-02-complete' | Select-Object -First 1) -Confirm:$false

• Document which KB threads you found most helpful and add them to your study notes

Design Reflection (VCDX)

A VCDX panelist examining your troubleshooting approach will ask: 'Walk me through how you would diagnose a failed VCF deployment.' They're listening for: (1) systematic log navigation (not random guessing), (2) understanding of component dependencies (HoloRouter -> nested ESXi -> Cloud Builder -> SDDC Manager -> NSX), (3) knowledge of common pitfalls from the community, (4) evidence that you've documented the process. Your runbook IS the evidence. Be prepared to explain how you built it, which scenarios you've tested, and what you learned. VCDX is about recoverability and manageability — demonstrate that you can recover from failure methodically.

Requirements

  • Deployment failures must be diagnosed systematically using logs, not guesswork
  • Troubleshooting process must be repeatable and documented for future operators
  • Solutions must address root cause, not symptoms (e.g., fix boot.cfg, not just restart)
  • Lessons learned from each failure must feed back into operational procedures

Constraints

  • Log volume in Holodeck can be overwhelming — requires filtering strategy (grep, Select-String)
  • SSH access to nested VMs assumes HoloRouter is at least partially functional
  • Some failures (e.g., corrupted datastore) cannot be fully recovered without a rebuild
  • Runbook relies on assumption that KB threads are accurate — test entries empirically before trusting them
  • Time-sensitive troubleshooting during bring-up (e.g., NTP drift) requires real-time monitoring, not post-failure log analysis

Assumptions

  • Operator has SSH access to HoloRouter, nested ESXi, and Cloud Builder
  • HoloRouter is reachable (at least for pings) to establish baseline connectivity
  • Holodeck logs are preserved (not deleted mid-deployment)
  • Runbook is kept current — adding new entries as new failure modes are discovered
  • Operator has read access to Holodeck KB and can extract lessons from forum discussions

Risks

  • Over-reliance on runbook without understanding root causes — IMPACT: runbook entries become cargo-cult procedures, MITIGATION: understand the 'why' behind each fix, not just the 'how'
  • Runbook outdated after VCF or Holodeck version upgrade — IMPACT: procedures fail on new versions, MITIGATION: version-tag runbook entries and review them after each upgrade
  • False confidence from testing on 'clean' snapshots — IMPACT: real-world failures are messier than lab scenarios, MITIGATION: intentionally vary failure injection (kill services, misconfigure DNS, etc.)
  • Documentation burden overwhelms troubleshooting velocity — IMPACT: you skip writing runbook to focus on recovery, MITIGATION: allocate 30 minutes post-incident for runbook update

Self-Assessment Discussion Prompts

  1. Walk me through your runbook structure. How did you organize the failure scenarios? What would change if you had to troubleshoot a VCF 5.2.x deployment?
  2. You encounter a failure not in your runbook. How do you approach it? (Answer: decision tree + log reading + hypothesis testing, then add to runbook.)
  3. Describe the dependency chain in a Holodeck deployment. If HoloRouter goes down, what cascades? How would you recover? (Answer: all nested VMs lose network access, nested ESXi can't PXE boot, deployment hangs indefinitely.)
  4. NSX version mismatch is common. Why does VCF 9.0.2 specifically require NSX 9.0.2.0.25150386? (Answer: SDDC Manager/vCenter APIs expect specific NSX REST endpoint versions; mismatch causes validation failures.)
  5. Boot.cfg and legacy CPU support. Why is this a gotcha? When do you need it? (Answer: older hardware doesn't support modern VMX instructions; nested ESXi fails to boot without allowLegacyCPU=TRUE.
  6. Your runbook helped solve a failure in 10 minutes. How do you know it's actually correct, not just lucky? (Answer: test it on a fresh deployment with the same failure condition; if it works twice, it's probably correct.)

Extensions

Build a Holodeck monitoring dashboard with real-time alerts

Create a PowerShell script that monitors HoloRouter, Cloud Builder, SDDC Manager, and vCenter health every 60 seconds and logs alerts (email or Teams webhook) when services go down or logs show ERROR keywords. This simulates production monitoring practices and is directly applicable to VCDX Day 2 operational excellence discussions.

harder

Reproduce and document 5 different Holodeck failure modes

Intentionally trigger each of these failures in sequence, document the symptoms, run through your runbook to diagnose, fix, and recover. Failures: (1) HoloRouter FRR crash, (2) Cloud Builder DNS misconfiguration, (3) Datastore full, (4) NSX version mismatch, (5) ESXi boot.cfg corruption. For each, measure time-to-diagnose before and after runbook.

much harder

Translate Holodeck troubleshooting to production VCF

Compare your Holodeck runbook to production VCF troubleshooting practices. Which principles carry over? What changes in a non-nested environment? (Answer: HoloRouter gone, physical network replaces it; NSX version mismatches work the same; datastore issues same logic but at physical scale.) Document the translation as a 'Holodeck to Production' mapping.

same

⚠ Known Pitfalls (from Community KB)

VCF Installer is not ready yet LIKELY_RESOLVED
Problem: Cloud Builder initialization hangs waiting for services to start, blocking bring-up from progressing
Resolution: Verify DNS, NTP, and firewall on Cloud Builder. Restart vcf-bringup service if stalled >30 minutes.
Issue with legacy CPU support (boot.cfg) RESOLVED
Problem: Deployment fails on older hardware with legacy CPUs because allowLegacyCPU=TRUE not set in boot.cfg
Resolution: Add allowLegacyCPU=TRUE to nested ESXi boot.cfg before or during Cloud Builder bring-up
Holodeck 9.0.1 VCF Deployment fails at NSX validation RESOLVED
Problem: NSX bundle version in content pack doesn't match VCF 9.0.x bring-up expectations
Resolution: Download NSX 9.0.2.0.25150386 and update content manifest to reference correct bundle
Waiting for FRR service LIKELY_RESOLVED
Problem: Prepare phase hangs waiting for FRR (BGP) service to initialize on HoloRouter
Resolution: SSH to HoloRouter and restart FRR: systemctl restart frr. Check BGP neighbor status: vtysh -c 'show ip bgp summary'
Configuration of HoloRouter Failed. Encountered error RESOLVED
Problem: HoloRouter management IP assignment fails during Prepare phase due to DHCP or network misconfiguration
Resolution: Verify port group has VLAN trunking enabled. Ensure DHCP is available on management network. Check HoloRouter logs: journalctl -u dnsmasq
Offline depot validation and licensing issues RESOLVED
Problem: Deployment fails during pre-flight checks when trial licenses or offline depots are misconfigured
Resolution: Ensure offline depot is valid. Use evaluation licenses bundled with Holodeck. Disable HTTPS validation if needed for offline mirror.
Holodeck vcf 9 - Cannot select datastore RESOLVED
Problem: Nested ESXi VM provisioning fails because target datastore is full, inaccessible, or misnamed
Resolution: Verify datastore exists, has >1.5 TB free, and is accessible from physical host. Update config.json targetDatastore if name was wrong.

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.