vCenter Service Recovery
Objectives
- Map vCenter Server service architecture and startup dependency chains
- Diagnose service failures using systemctl status, log analysis, and vpxd health checks
- Execute correct service restart sequences (PSC/SSO then vpxd then vsphere-ui then remaining)
- Troubleshoot vPostgres database connectivity and disk space exhaustion
- Perform vCenter file-based backup and restore operations
- Document service recovery procedures using RCAR methodology for VCDX defense
Prerequisites
VCF lab environment with operational vCenter Server Appliance (VCSA)
Required skills:
- SSH access to VCSA and ESXi hosts
- Basic Linux service management (systemctl)
- Understanding of vCenter role in VCF management domain
Lab Environment
VCF management domain with VCSA, SDDC Manager, and minimum 4 ESXi hosts
Tasks
Task 1 vCenter Service Architecture & Dependency Mapping
Build a complete mental model of VCSA service architecture — which services depend on which, the correct startup order, and how service health cascades through the VCF stack.
SSH into the VCSA appliance. Run 'service-control --status' to list all vCenter services and their current state. Document every service name and status (RUNNING/STOPPED).
Map the service dependency chain. Start from the bottom: (1) vmware-vpostgres (database), (2) vmware-stsd (SSO/STS token service), (3) vmware-vpxd (core vCenter daemon), (4) vsphere-ui (web client), (5) remaining services. Draw a dependency diagram showing which services block which.
Check the VCSA VAMI (https://<vcsa>:5480) health dashboard. Document the health status of each subsystem: Database, Overall, SSO, Content Library, etc. Compare VAMI health with service-control output.
Document the VCF management domain impact chain: vCenter down means SDDC Manager loses inventory visibility, NSX Manager cannot sync with vCenter, no new workload provisioning possible, but existing VMs continue running unmanaged. Verify by checking SDDC Manager UI while vCenter is healthy.
Validation Gate
Check: Complete service dependency diagram and document VCF impact chain
Expected: All 20+ VCSA services identified and mapped in dependency order. VCF impact chain documented showing management plane vs. data plane separation.
Common Errors
Task 2 Service Failure Diagnosis & Log Analysis
Develop systematic log analysis skills to identify root cause of vCenter service failures — from vpxd crashes to SSO token expiry to database corruption.
Map the critical log locations for vCenter troubleshooting. SSH to VCSA and verify each path exists: /var/log/vmware/vpxd/vpxd.log (core vCenter), /var/log/vmware/sso/ssoAdminServer.log (SSO), /var/log/vmware/vpostgres/postgresql-*.log (database), /var/log/vmware/vsphere-ui/logs/vsphere_client_virgo.log (UI), /var/log/vmware/rhttpproxy/rhttpproxy.log (reverse proxy). Document file sizes and last-modified timestamps.
Simulate a common failure scenario: check disk space on the VCSA. Run 'df -h' and examine /storage/db (database), /storage/log (logs), /storage/seat (stats), and root filesystem. Document current usage percentages. Note: VCSA generates alerts at 75% disk usage.
Practice log analysis for three common failure patterns. Search vpxd.log for: (a) 'ODBC connection failed' indicating database connectivity loss, (b) 'SSO initialization failed' indicating SSO/STS service down, (c) 'Failed to initialize VMware VirtualCenter' indicating critical startup failure. Use grep -i to search for these patterns.
Generate a vCenter support bundle for practice. Navigate to VAMI > Support > Create Support Bundle. Alternatively, use the CLI: /usr/lib/vmware-vmon/vmon-cli --get-state-all. Document what is included in the bundle and the typical file size.
Check for STS certificate expiry — a common vCenter 7.x/8.x issue that carries into VCF 9.0 upgrades. Run: /usr/lib/vmware-vmca/bin/certool --status. Also check: for VCF managed environments, verify the SDDC Manager certificate status via API: GET /v1/certificates. Document expiry dates for all certificates.
Validation Gate
Check: Document all critical log locations, disk space status, and certificate expiry dates
Expected: Complete log location reference created, disk space verified healthy, all certificates checked with expiry dates documented.
Common Errors
Task 3 Service Recovery Procedures & Database Repair
Execute correct service recovery sequences for different failure scenarios — from simple service restart to database repair to full VCSA restore from backup.
Practice the correct service restart sequence. First, stop all services in reverse dependency order: 'service-control --stop --all'. Wait for confirmation. Then start services: 'service-control --start --all'. Monitor startup progress in /var/log/vmware/vmon/vmon-startup.log. Document the startup time.
Practice individual service recovery for the most common failure: vpxd crash. Stop only vpxd: 'service-control --stop vmware-vpxd'. Check vpxd.log for the last error before crash. Start vpxd: 'service-control --start vmware-vpxd'. Monitor startup in vpxd.log — look for 'VirtualCenter started successfully' message.
Practice vPostgres database health check and repair. Connect to the embedded database: /opt/vmware/vpostgres/current/bin/psql -U postgres -d VCDB. Run basic health queries: SELECT count(*) FROM vpx_vm; (VM count), SELECT pg_database_size('VCDB'); (DB size). Check for long-running queries: SELECT pid, now() - pg_stat_activity.query_start AS duration, query FROM pg_stat_activity WHERE state = 'active';
Practice disk space recovery procedures for /storage/db exhaustion. Steps: (1) Reduce stats retention via vSphere Client Administration System Configuration Stats Level; (2) Clean old log files: find /storage/log -name '*.log' -mtime +30 and review before deleting; (3) Purge old events/tasks using vCenter MOB or PowerCLI to remove events older than 30 days; (4) Verify space recovered: df -h /storage/db.
Practice VCSA file-based backup and restore. Configure backup schedule in VAMI Backup Configure Backup Schedule. Set target: SFTP server or NFS share. Run an ad-hoc backup. Verify backup completion and document the backup size and duration. Review restore procedure: VCSA ISO Restore point to backup location.
Validation Gate
Check: Successfully execute service restart, database health check, and backup/restore procedures
Expected: Service restart completed in correct dependency order. Database health verified. Backup completed and stored on external target. Restore procedure documented.
Common Errors
Task 4 vCenter HA Architecture & VCDX Defense
Understand vCenter HA (active/passive/witness) architecture, its role in VCF availability design, and prepare to defend vCenter recovery strategies in VCDX.
Document vCenter HA architecture. Three nodes: Active (runs all services), Passive (receives continuous replication), Witness (quorum arbitrator, lightweight). Network requirements: management network for client access, HA network (dedicated, isolated, low-latency) for replication. Replication mechanism: vPostgres native streaming replication for database, file-level replication for configuration.
Analyze the design decision: vCenter HA deployment in VCF management domain. Document: Requirement (R-001: vCenter RPO=0, RTO<10 min), Constraint (VCF management domain must maintain operational continuity), Assumption (dedicated HA network available between management domain hosts), Risk (HA failover during VCF lifecycle operation may leave SDDC Manager in inconsistent state).
Build a vCenter recovery decision tree: (1) Single service failure: restart that service (RTO: 2-5 min); (2) vpxd crash loop: analyze logs, restart all services (RTO: 5-15 min); (3) Database corruption: restore from file-based backup (RTO: 30-60 min); (4) VCSA VM failure: vCenter HA automatic failover (RTO: 5-10 min); (5) Full site failure: DR recovery at secondary site (RTO: per SRM/DR design). Document this as a runbook.
Prepare four VCDX defense responses for vCenter recovery design: (1) 'Why not use SRM for vCenter protection?' — SRM requires vCenter running at protected and recovery sites, creating circular dependency; vCenter HA provides independent protection. (2) 'What happens to running VMs during vCenter failover?' — zero impact, VMs continue on ESXi hosts, only management operations are paused. (3) 'How do you handle vCenter HA in a stretched cluster?' — place active and passive on different sites, witness on third site, same as vSAN witness placement strategy. (4) 'What is your vCenter backup strategy for VCF?' — daily file-based backup to external SFTP, coordinated with SDDC Manager and NSX Manager backups, tested quarterly.
Validation Gate
Check: Complete vCenter HA architecture documentation and VCDX defense preparation
Expected: vCenter HA architecture documented with network requirements. Recovery decision tree created as runbook. Four VCDX defense responses prepared and practiced.
Common Errors
Final Validation
Complete vCenter service recovery with graduated response procedures and VCDX defense preparation
✓ Service dependency chain documented → All 20+ services mapped with startup order and dependencies
✓ Log analysis capability demonstrated → Critical log locations known, failure patterns identified
✓ Recovery procedures tested → Service restart, database health check, and backup/restore completed
✓ vCenter HA architecture understood → Active/passive/witness documented with VCF management domain context
✓ VCDX defense responses prepared → Four defense scenarios documented with concise technical responses
Cleanup / Restore
• Verify all vCenter services are running: service-control --status
• Confirm SDDC Manager shows healthy vCenter connectivity
• Remove any temporary support bundles to reclaim disk space
• Revert to snapshot if any configuration changes were made during practice
Design Reflection (VCDX)
vCenter recovery design demonstrates understanding of management plane vs. data plane separation, graduated response methodology, and the critical role of vCenter in VCF lifecycle management. The design decision to deploy vCenter HA with coordinated backup strategy shows operational maturity expected at VCDX level.
Requirements
- R-001: vCenter availability >=99.99% (<=52 min/year downtime)
- R-002: vCenter RPO = 0 for automated failover, RPO <= 24h for backup-based recovery
- R-003: VCF management domain operational continuity during component failures
Constraints
- VCSA must run on management domain — cannot be hosted externally
- VCF lifecycle operations (upgrades, expansion) require vCenter availability
- SDDC Manager, NSX Manager, and Aria Suite all depend on vCenter for inventory
Assumptions
- Dedicated HA network available between management domain hosts
- External SFTP/NFS target available for file-based backups
- Operations team trained on graduated recovery procedures
Risks
- vCenter HA failover during active SDDC Manager workflow may leave workflow in failed state — requires manual retry
- STS certificate expiry causes cascading authentication failure across all VCF components — monitor proactively
- Database corruption from disk exhaustion requires full restore — implement proactive disk monitoring
Self-Assessment Discussion Prompts
- How would you design vCenter recovery differently for a VCF environment with 500+ hosts across 5 workload domains?
- What metrics would you monitor to predict vCenter failures before they cause outages?
- How does vCenter recovery priority change during a VCF upgrade window?
- If you could only implement one preventive measure for vCenter availability, what would it be and why?
Extensions
vCenter HA Failover Testing
Deploy vCenter HA in your Holodeck lab. Simulate active node failure by powering off the active VCSA VM. Measure actual failover time and verify all VCF management services reconnect. Document the failover timeline and any service recovery steps required post-failover.
Automated vCenter Health Monitoring
Build a PowerCLI or Python script that checks vCenter service status, disk space, certificate expiry, and database size on a schedule. Output a health report and send alerts when thresholds are breached. Integrate with Aria Operations for centralized monitoring.
Multi-vCenter Recovery Coordination
Design a recovery runbook for a VCF environment with management domain vCenter and 3 workload domain vCenters. Document the recovery sequence, inter-vCenter dependencies, and SDDC Manager coordination required. Practice with Holodeck multi-domain deployment.
⚠ Known Pitfalls (from Community KB)
References
- vCenter Server 9.0 Administration Guide — Service Management and Troubleshooting chapter
- VMware KB 2147144 — How to stop, start, or restart vCenter Server services
- VMware VCF 9.0 Administration Guide — Management Domain Backup and Recovery
- VMware KB 2149253 — vCenter Server Appliance file-based backup and restore
- vCenter HA Architecture and Deployment Guide