Lab: VCF Host Lifecycle — Commission, Add, Maintain, Decommission
Objectives
- Execute full host lifecycle from commission to decommission
- Understand vSAN data evacuation modes and their impact
- Troubleshoot host commission failures using SoS and API diagnostics
- Manage host maintenance operations with zero workload downtime
Prerequisites
VCF 9.0 with a spare ESXi host meeting HCL requirements, management domain operational, at least one workload domain with a cluster that can accept an additional host
Required skills:
- VCF SDDC Manager navigation
- ESXi host management
- vSAN operations
Lab Environment
VCF 9.0 Instance with management domain and one workload domain. Spare ESXi host with network connectivity to all required VLANs (management, vMotion, vSAN, TEP). Host must meet VCF HCL.
Tasks
Task 1 Full Host Lifecycle Operations
Walk through the complete host lifecycle in VCF — from initial commission through operational use to decommission — demonstrating the operational procedures, troubleshooting techniques, and design considerations at each stage.
Pre-flight: Validate host readiness. SSH to the spare host and verify: (a) ESXi version matches VCF BOM (esxcli system version get); (b) NTP configured and synced (esxcli system ntp get); (c) DNS resolution works for SDDC Manager FQDN and vCenter FQDN; (d) Network connectivity: vmkping from management vmkernel to SDDC Manager IP, vCenter IP. Run SoS connectivity check from SDDC Manager to verify all required ports are open.
Commission host. SDDC Manager UI → Inventory → Hosts → Commission Host. Provide: FQDN, IP, credentials (root password), network pool assignment. The commission process: (a) validates connectivity, (b) checks hardware against HCL, (c) stores credentials in VCF credential store, (d) applies baseline ESXi configuration (NTP, DNS, syslog). Monitor: GET /v1/tasks/{id} → wait for 'SUCCESSFUL'. If failed: check /var/log/vmware/vcf/domainmanager/domainmanager.log for specific error.Add host to workload domain. Inventory → Domains → select workload domain → Hosts → Add Host. Select the commissioned host. VCF orchestrates: (a) ESXi added to vCenter as cluster member, (b) vSAN disk claiming (NVMe drives added to ESA pool), (c) NSX transport node preparation (VIBs installed, TEP configured), (d) DRS/HA integration. This multi-step process takes 15-30 minutes. Monitor each subtask via API: GET /v1/tasks/{id}/subtasks.Verify host integration. Post-add checks: (a) vSphere Client → host shows as Connected in cluster; (b) vSAN → host disks appear in storage pool, health shows no issues; (c) NSX Manager → host appears as transport node with status 'Success'; (d) DRS → host participating in load balancing (VMs should migrate within minutes if cluster is imbalanced). Run SoS health-check to confirm clean integration.
Enter maintenance mode. vSphere Client → right-click host → Maintenance Mode → Enter. Select vSAN data evacuation mode: 'Ensure Accessibility' (for a quick reboot) or 'Full Data Migration' (for extended maintenance). Monitor: vSAN resyncing tab shows data movement progress. VMs automatically vMotion to other hosts (DRS-managed). Verify: no VMs remain on the host, vSAN evacuation complete, host status = 'Maintenance Mode'.
Simulate maintenance tasks. While host is in maintenance mode: (a) update firmware via iLO/iDRAC if needed; (b) verify ESXi patch level; (c) check hardware health (esxcli hardware platform get). Note: in VCF, patching should be done through LCM (not manual ESXi updates) to maintain BOM compliance. Exit maintenance mode: right-click → Exit Maintenance Mode. vSAN begins rebalancing, DRS migrates VMs back.
Remove host from domain. SDDC Manager → Domains → select domain → Hosts → Remove Host. VCF orchestrates reverse of add: (a) VMs evacuated via DRS, (b) vSAN full data evacuation (mandatory — all data rebuilt on remaining hosts), (c) NSX transport node deconfigured, (d) host removed from vCenter cluster. This is the longest operation — vSAN evacuation can take hours for TB-scale data. Monitor via API.
Decommission host. SDDC Manager → Inventory → Hosts → select host → Decommission. This removes the host from VCF inventory, clears stored credentials, and returns the host to standalone ESXi. The host can now be repurposed or physically removed. Verify: host no longer appears in VCF inventory, credentials removed from credential store.
Troubleshooting exercise: Simulate commission failure. Attempt to commission a host with incorrect DNS (no reverse DNS entry). Observe the failure: API shows task FAILED, domainmanager.log shows 'DNS validation failed'. Fix: add reverse DNS entry on DNS server, retry commission. This is the most common commission failure in production environments.
Document operational runbook. Create a host lifecycle runbook documenting: (a) pre-flight checklist (HCL, network, DNS, NTP, firmware); (b) commission procedure with expected duration; (c) domain add procedure with subtask monitoring; (d) maintenance mode decision matrix (which evacuation mode for which scenario); (e) removal/decommission procedure; (f) common failure modes and resolution steps. This runbook is a deliverable for the VCAP exam scenario.
Validation Gate
Check: Full host lifecycle executed successfully
Expected: Host commissioned, added to domain with vSAN/NSX integration, maintenance mode entered/exited, host removed and decommissioned, commission failure troubleshooted
Common Errors
Final Validation
Host lifecycle operations mastered with troubleshooting capability
✓ Commission → Host added to VCF inventory with credentials stored
✓ Domain add → Host integrated with vSAN, NSX, and DRS in workload domain cluster
✓ Maintenance mode → vSAN evacuation completed, VMs migrated, host in maintenance
✓ Removal → Host removed from domain with full data evacuation
✓ Decommission → Host removed from VCF inventory
✓ Troubleshooting → DNS commission failure diagnosed and resolved
Cleanup / Restore
• Ensure decommissioned host is in a known state
• Verify cluster health after host removal (vSAN rebalanced, DRS stable)
Design Reflection (VCDX)
Host lifecycle demonstrates operational maturity. VCDX defense: justify vSAN evacuation mode selection based on maintenance duration and cluster resilience level. Discuss blast radius: removing a host from a 4-host cluster with FTT=1 leaves zero fault tolerance during evacuation — plan accordingly.
Requirements
- Zero workload downtime during host lifecycle operations
- Full data protection during vSAN evacuation
- Sub-30-minute host commission for rapid scale-out
Constraints
- vSAN evacuation time proportional to data volume — can take hours
- Minimum cluster size after removal must maintain FTT policy
- Host must meet VCF HCL for commission
Assumptions
- DNS/NTP infrastructure is stable and pre-configured
- Sufficient cluster capacity to absorb evacuated workloads
- Network connectivity to all VCF VLANs validated before commission
Risks
- Commission with incorrect credentials locks VCF from managing the host
- vSAN evacuation during peak hours impacts cluster performance
- Removing host from undersized cluster violates FTT policy
Self-Assessment Discussion Prompts
- How do you handle host replacement when the cluster is at minimum size for FTT?
- What is the impact of choosing 'No Data Migration' if another host fails during maintenance?
- How do you automate host lifecycle for scale-out events?
Extensions
Automate host commission via SDDC Manager API (POST /v1/hosts)
Script a pre-flight validation tool that checks all commission prerequisites
Test parallel host addition (2 hosts simultaneously) and observe vSAN behavior
⚠ Known Pitfalls (from Community KB)
References
- VCF 9.0 Administration Guide — Host Management: techdocs.broadcom.com