VCAP — VCF Administrator (3V0-11.26)
VCAP-level VCF administration covering fleet lifecycle management, Private Cloud/Fleet/Instance hierarchy, identity & certificate management, configuration drift & compliance, advanced troubleshooting with SoS & log bundles, vSAN ESA operations, and VCF 9.0 day-2 operational excellence.
Version Evolution
VCAP-Admin exam covers advanced VCF day-2 operations. VCF 9.0 shifted management from SDDC Manager UI to VCF Operations and vSphere Client. vSAN ESA replaces OSA as the default storage architecture. Identity management gains OIDC federation support. The exam tests scenario-based troubleshooting and operational decision-making at scale.
Learning Outcomes
- Master VCF 9.0 fleet lifecycle management: bring-up, expand, shrink, upgrade, and decommission operations across Private Cloud, Fleet, Instance, and Domain hierarchy
- Configure and manage identity federation (vCenter SSO, AD/LDAP, OIDC), certificate lifecycle (VMCA, custom CA, certificate rotation), and password policies at scale
- Detect, analyze, and remediate configuration drift using VCF Operations compliance dashboards and desired-state configuration baselines
- Execute advanced troubleshooting using SoS utility, log bundles, REST API diagnostics, and component health checks
- Manage vSAN ESA operations: disk lifecycle, maintenance mode, data evacuation, and fault domain management
- Implement VCF security hardening: STIG compliance, ESXi lockdown mode, TPM attestation, and audit logging
VCF 9.0 Fleet Lifecycle Management#
VCF 9.0 introduces a hierarchical management model: Private Cloud → Fleet → Instance → Domain. Understanding this hierarchy and the lifecycle operations at each level is critical for VCAP-level administration.
VCF 9.0 Taxonomy & Hierarchy:
---
Private Cloud (highest level):
- Represents the entire VMware Cloud Foundation deployment
- Contains one or more Fleets
- Managed by VCF Operations (central management plane)
- Policy and compliance scope: organization-wide baselines
Fleet:
- A logical grouping of VCF Instances
- Used for grouping by geography, business unit, or operational boundary
- Fleet-level policies cascade to all Instances within
- Example: "EMEA Fleet", "APAC Fleet", "Production Fleet"
Instance:
- A single VCF deployment with its own management components
- Contains: 1 management domain + 0-N workload domains
- Each Instance has its own: vCenter, NSX Manager, SDDC Manager (deprecated UI)
- In VCF 9.0: Instance management shifts from SDDC Manager to VCF Operations and vSphere Client
Domain:
- Management Domain: infrastructure services (vCenter, NSX, VCF Operations, etc.)
- Workload Domain: tenant/application workloads
- Each domain has: vSphere cluster(s), vSAN storage, NSX networking
Lifecycle Operations by Level:
---
Instance-Level Operations:
Bring-Up: Initial VCF deployment (Cloud Builder → management domain)
- Commission/Decommission Hosts: Add/remove ESXi hosts from the Instance pool
- Create Workload Domain: Deploy new domain with cluster, vSAN, NSX
- Expand Domain: Add cluster or hosts to existing domain
- Shrink Domain: Remove hosts from domain (with vSAN data evacuation)
- Delete Workload Domain: Remove domain entirely (destructive)
Fleet-Level Operations:
Add/Remove Instance: Manage Instance membership in Fleet
Fleet Policy Push: Apply configuration baselines to all Instances
Fleet Health Monitoring: Aggregate health view across Instances
Upgrade Lifecycle:
VCF Lifecycle Management (LCM) orchestrates upgrades across the entire stack:
- Pre-check: compatibility validation, health check, disk space
- Download: fetch update bundles from depot (online) or import (offline)
- Stage: prepare update on target components
4. Upgrade sequence: NSX → vCenter → ESXi → vSAN → VCF Operations
Upgrade Modes:
- Rolling: One host at a time, VMs migrate via vMotion (zero workload downtime)
- Parallel: Multiple hosts simultaneously (faster but requires more spare capacity)
- Maintenance Window: All hosts at once (requires VM downtime)
Host Lifecycle in VCF:
---
- Commission: Add host to SDDC Manager inventory
- Host must meet VCF HCL requirements
- Network connectivity validated (management, vMotion, vSAN, TEP VLANs)
- Credentials stored in VCF credential store
- Add to Domain: Assign commissioned host to a workload domain cluster
- ESXi configuration applied automatically (NTP, DNS, syslog, lockdown mode)
- vSAN disk groups claimed
- NSX transport node configuration applied
- Host enters cluster and participates in HA/DRS
- Maintenance: Enter maintenance mode for patching/hardware
- vSAN data evacuation mode selection:
a. Full data evacuation (safest, longest, ensures all data replicated to other hosts)
b. Ensure accessibility (faster, maintains access but may reduce resilience)
c. No data evacuation (fastest, risk of data loss if another host fails during maintenance)
- Remove from Domain: Reverse of add — vSAN evacuation, VM migration, NSX cleanup
- Decommission: Remove from SDDC Manager inventory (host returns to standalone ESXi)
Key Takeaways
- VCF 9.0 hierarchy: Private Cloud (org-wide) → Fleet (logical grouping) → Instance (single VCF deployment) → Domain (management + workload). Each level has distinct lifecycle operations.
- Host lifecycle: Commission → Add to Domain → Operational → Maintenance → Remove → Decommission. vSAN data evacuation mode selection is critical during maintenance operations.
- LCM orchestrates upgrades across the full stack in dependency order: NSX → vCenter → ESXi → vSAN → VCF Operations. Rolling upgrades enable zero workload downtime.
- SDDC Manager UI is deprecated in VCF 9.0 — functions migrated to VCF Operations and vSphere Client, but SDDC Manager API remains available during transition.
Identity, Certificate & Password Management#
VCF identity management spans SSO federation, certificate lifecycle, and password rotation across all components. At VCAP level, you must understand the interactions between vCenter SSO, Active Directory, OIDC providers, and the certificate chain of trust across the VCF stack.
Identity Architecture:
---
vCenter SSO (Platform Services Controller):
- Central identity provider for all vSphere/VCF components
- Embedded PSC in vCenter 7.0+ (no external PSC deployment)
- Default domain: vsphere.local
- Supports: local users, AD/LDAP, OIDC identity providers
AD/LDAP Integration:
vCenter → Administration → Single Sign-On → Identity Sources → Add
Types:
- Active Directory (Integrated Windows Authentication): Kerberos-based, machine account join
- AD over LDAP: LDAP bind with service account (simpler but less secure)
- Open LDAP: Non-Microsoft LDAP directories
Best Practices:
- Use AD/LDAP for human administrators (centralized password policy)
- Keep a local vsphere.local admin account for break-glass access
- Map AD groups to vCenter roles (not individual users)
- Service accounts for API integrations should be local SSO accounts
OIDC Federation (VCF 9.0):
VCF 9.0 supports OIDC identity providers (Okta, Azure AD, etc.)
Configuration: vCenter → Identity Provider → Add OIDC Provider
Benefits: MFA support, centralized identity governance, SSO across VCF + cloud services
Limitations: OIDC does not replace SSO — it federates INTO SSO as an identity source
Certificate Architecture:
---
VCF uses a hierarchical certificate model:
VMCA (VMware Certificate Authority):
- Embedded CA in each vCenter (PSC component)
- Issues certificates to: ESXi hosts, vCenter services, NSX components, Solution Users
- Default: VMCA is the root CA (self-signed root)
- Production option: VMCA as subordinate CA under enterprise PKI
Certificate Hierarchy:
Enterprise Root CA (optional)
└── VMCA (subordinate or self-signed root)
├── ESXi host certificates
├── vCenter Machine SSL certificate
├── vCenter Solution User certificates
├── NSX Manager certificate
└── VCF component certificatesCertificate Operations:
Replace VMCA root certificate:
- Generate CSR from VMCA
- Sign CSR with enterprise CA
- Import signed certificate + CA chain into VMCA
- VMCA re-issues all leaf certificates automatically
Replace individual certificates (custom mode):
- Generate CSR for specific service
- Sign with enterprise CA
- Import signed certificate via Certificate Manager
- Repeat for each service (labor-intensive, not recommended)
Certificate rotation in VCF:
- VCF tracks certificate expiry across all components
- LCM can orchestrate certificate rotation during upgrade cycles
- VCF Operations monitors certificate health and alerts on expiry (30-day default warning)
Password Management:
---
VCF manages passwords for all deployed components:
Password Lifecycle:
- VCF credential store holds passwords for: vCenter, NSX, ESXi, SDDC Manager, VCF Operations
- Password rotation: VCF supports auto-rotation with configurable schedule
- Rotation scope: per-component or fleet-wide
- Emergency rotation: forced rotation after security incident
Password Policies:
- Minimum length, complexity, history, lockout threshold
- ESXi: /etc/pam.d/passwd (local) or vCenter-managed
- vCenter: SSO password policy (Administration → Configuration → Local Accounts)
- NSX: CLI and API password policies (independent from vCenter SSO)
Best Practices:
- Use AD/LDAP for human access — single password policy, centralized audit
- Rotate service account passwords quarterly via VCF auto-rotation
- Maintain a break-glass procedure for the vsphere.local administrator account
- Monitor VCF Operations for password expiry alerts across all components
- Document password dependencies — some passwords (ESXi root) affect VCF operations if changed outside VCF
Key Takeaways
- Identity stack: vCenter SSO (central IdP) ← AD/LDAP (human admins) + OIDC (MFA/federation in VCF 9.0) + local accounts (break-glass). Map AD groups to vCenter roles, never individual users.
- Certificate hierarchy: Enterprise Root CA → VMCA (subordinate) → leaf certs (ESXi, vCenter, NSX). VMCA subordinate mode is the production best practice — enables enterprise PKI integration while VMCA handles automated leaf cert issuance.
- Password management: VCF credential store holds all component passwords. Auto-rotation is configurable. Never change ESXi root password outside VCF — it breaks lifecycle operations.
- Certificate rotation: VCF Operations monitors expiry across all components. LCM can orchestrate rotation during upgrade cycles. 30-day default warning alert.
Configuration Drift Detection & Compliance#
Configuration drift occurs when infrastructure components deviate from their desired-state baseline. VCF Operations provides drift detection, compliance dashboards, and automated remediation for maintaining consistent configurations across the fleet.
Drift Detection Architecture:
---
VCF Operations continuously compares actual configurations against defined baselines:
Configuration Sources (what is collected):
ESXi hosts: NTP, DNS, syslog, lockdown mode, SSH status, firewall rules, power policy, BIOS settings
vCenter: SSO configuration, permission assignments, HA/DRS settings, alarm definitions
NSX: transport zone membership, DFW rules, Edge configuration, certificate status
vSAN: storage policy compliance, disk group health, fault domain alignment
Baseline Types:
- VCF Default Baseline: Built-in best practices from VMware (applied automatically by VCF during host provisioning)
- Custom Baseline: Organization-specific standards (e.g., "All ESXi hosts must have NTP pointing to 10.0.0.1 and 10.0.0.2")
- Regulatory Baseline: STIG, CIS Benchmark, PCI DSS profiles (pre-built or imported)
- Golden Image Baseline: Snapshot of a known-good host configuration used as reference
Drift Detection Process:
- Collection: VCF Operations adapters fetch current configuration (every 5-30 minutes)
- Comparison: Current config compared against active baseline
- Classification: Deviations classified as Critical/Warning/Informational
- Alert: Drift alerts generated for Critical deviations
- Report: Compliance dashboard shows drift percentage per cluster/host
Compliance Dashboard Design:
---
Dashboard Components:
- Compliance Score: Percentage of hosts matching baseline (100% = no drift)
- Drift Distribution: Heat map showing which configuration parameters drift most frequently
- Top Drifters: Hosts with the most deviations from baseline
- Trend: Compliance score over time (is drift getting better or worse?)
Remediation Options:
Manual: Admin reviews drift report and manually corrects each deviation
Semi-Automated: VCF Operations generates remediation runbook with exact commands
Automated: VCF Operations + Host Profile/desired state applies corrections automatically
- Safety consideration: Not all drift is bad — sometimes deliberate exceptions exist
- (e.g., specific host needs different NTP for regulatory reasons)
- Tag exceptions explicitly in the baseline to avoid false positive drift alerts
STIG Compliance (Security Technical Implementation Guide):
---
DoD STIG profiles for vSphere/VCF are available as pre-built compliance baselines:
Categories:
- CAT I: High severity — exploitable vulnerabilities (must fix immediately)
- CAT II: Medium severity — potential vulnerabilities (fix within 30 days)
- CAT III: Low severity — best practice deviations (fix within 90 days)
Common STIG checks for ESXi:
- SSH disabled (CAT II)
- Lockdown mode enabled (CAT II)
- DCUI timeout configured (CAT III)
- Persistent logging to remote syslog (CAT II)
- TLS 1.2 minimum for all management interfaces (CAT I)
- SNMP v3 only (v1/v2c disabled) (CAT II)
- vMotion traffic encrypted (CAT II)
CIS Benchmark compliance follows similar structure with Level 1 (essential) and Level 2 (defense-in-depth) profiles.
Key Takeaways
- Drift detection: VCF Operations compares actual config against baselines (VCF default, custom, regulatory/STIG, golden image) at each collection interval. Deviations classified by severity.
- Compliance dashboard: compliance score percentage, drift distribution heatmap, top drifters, trend over time. Target: 100% compliance on critical parameters.
- Remediation: manual, semi-automated (generated runbook), or automated (desired state enforcement). Tag deliberate exceptions to avoid false positive drift alerts.
- STIG/CIS compliance: pre-built profiles with CAT I/II/III severity. Common checks: SSH disabled, lockdown mode, TLS 1.2, encrypted vMotion, persistent syslog.
Advanced Troubleshooting: SoS, Log Bundles & API Diagnostics#
VCAP-level troubleshooting requires mastery of VCF's built-in diagnostic tools: the SoS (Service-on-Service) utility for health checks, log bundle generation for support cases, REST API diagnostics for programmatic analysis, and component-specific troubleshooting procedures.
SoS Utility (Support & Operations Script):
---
Location: SDDC Manager VM → /opt/vmware/sddc-support/sos
Purpose: Comprehensive health check and log collection for all VCF components
Key SoS Commands:
Health Check (non-disruptive diagnostic):
- sudo /opt/vmware/sddc-support/sos --health-check
- Checks: DNS resolution, NTP sync, certificate validity, service status,
- connectivity between all components, password expiry, disk space,
- vSAN health, NSX status
- Output: HTML report with pass/fail/warning for each check
Log Bundle Collection:
- sudo /opt/vmware/sddc-support/sos --log-collection
- Collects logs from: SDDC Manager, vCenter, ESXi hosts, NSX, VCF Operations
- Output: Compressed bundle for VMware support case attachment
Options:
--domain <name>: Limit to specific domain (reduces bundle size)
--component <name>: Limit to specific component (vCenter, NSX, ESXi)
--start-date --end-date: Time-bounded collectionPassword Check:
- sudo /opt/vmware/sddc-support/sos --password-check
- Validates all stored passwords against actual component passwords
- Identifies mismatches (password changed outside VCF)
Connectivity Check:
- sudo /opt/vmware/sddc-support/sos --connectivity-check
- Tests network connectivity between all VCF components
- Validates: management network, vMotion network, vSAN network, TEP network
Log Architecture & Key Log Locations:
---
SDDC Manager:
- /var/log/vmware/vcf/sddc-manager-ui-app/: UI application logs
- /var/log/vmware/vcf/commonsvcs/: Common services (LCM, inventory)
- /var/log/vmware/vcf/domainmanager/: Domain lifecycle operations
- /var/log/vmware/vcf/operationsmanager/: Operations and health
- vCenter:
- /var/log/vmware/vpxd/vpxd.log: Core vCenter service
- /var/log/vmware/vsan-health/: vSAN health service
- /var/log/vmware/sso/: SSO authentication
- /var/log/vmware/content-library/: Content library operations
ESXi:
- /var/log/vmkernel.log: Kernel messages (PSOD, hardware, storage)
- /var/log/hostd.log: Host daemon (VM operations, API)
- /var/log/vobd.log: VMware Observation Broker (events)
- /var/log/esxupdate.log: Update/patch operations
- /var/log/fdm.log: HA fault domain manager
NSX Manager:
- /var/log/proton/nsxapi.log: API requests
- /var/log/corfu/corfu.log: Distributed database (CCP)
- /var/log/policy/policy.log: Policy engine operations
REST API Diagnostics:
---
VCF SDDC Manager API:
Base: https://<sddc-mgr>/v1/
Authentication: Bearer token (POST /v1/tokens with credentials)
Key diagnostic endpoints:
GET /v1/system/health: Overall VCF health summary
GET /v1/hosts: All commissioned hosts with status
GET /v1/domains: Domain inventory and health
GET /v1/tasks: Lifecycle operation history and status
GET /v1/credentials: Password inventory and expiry dates
GET /v1/bundles: Available update bundles
GET /v1/upgrades: Upgrade history and current status
Troubleshooting workflow:
1. GET /v1/tasks?status=FAILED → find failed operations
2. GET /v1/tasks/{id} → detailed task info with subtasks
3. GET /v1/tasks/{id}/subtasks → each subtask status with error messages- Correlate subtask errors with component logs for root cause
Common Troubleshooting Scenarios:
---
- Host commission fails:
Check: SoS connectivity check → network path to host
Check: DNS resolution (forward and reverse) for host FQDN
Check: Host firmware/driver versions against VCF HCL
Logs: /var/log/vmware/vcf/domainmanager/domainmanager.log
- Workload domain creation stalls:
Check: GET /v1/tasks/{id}/subtasks → identify stuck subtaskCommon causes: vSAN disk claim failure, NSX TEP connectivity, vCenter deployment timeout
Logs: domainmanager.log + component-specific logs based on stuck subtask
- LCM upgrade fails midway:
Check: GET /v1/upgrades/{id} → which component failed
Action: Do NOT retry without understanding the failure
Logs: /var/log/vmware/vcf/lcm/ → upgrade orchestration logsRecovery: Retry the failed task via API (PUT /v1/upgrades/{id}/retry) after resolving root cause
- Certificate expiry warnings:
Check: SoS health-check → certificate section
Action: VCF certificate rotation via LCM or manual vCenter Certificate Manager
Prevention: Set VCF Operations alert for certificate expiry at 60-day threshold
Key Takeaways
- SoS utility: health-check (HTML report for all components), log-collection (support bundle with domain/component/time filters), password-check, connectivity-check. Run health-check first in any troubleshooting scenario.
- Key log paths: SDDC Manager (/var/log/vmware/vcf/), vCenter (/var/log/vmware/vpxd/), ESXi (/var/log/vmkernel.log, hostd.log), NSX (/var/log/proton/, /var/log/corfu/).
- REST API diagnostics: GET /v1/tasks?status=FAILED → GET /v1/tasks/{id}/subtasks → correlate subtask errors with component logs. This is the primary troubleshooting workflow for failed lifecycle operations.
- Recovery pattern: identify failed subtask → check component logs → resolve root cause → retry via API (PUT /v1/upgrades/{id}/retry). Never blindly retry without understanding the failure.
vSAN ESA Operations & Day-2 Management#
vSAN Express Storage Architecture (ESA) is the default storage architecture in VCF 9.0, replacing the original storage architecture (OSA). VCAP Admins must understand ESA's operational model, disk lifecycle, maintenance procedures, and fault domain management.
vSAN ESA vs OSA — Operational Differences:
---
Architecture Comparison:
OSA (Original Storage Architecture):
- Two-tier: Cache (SSD) + Capacity (SSD or HDD)
- Disk Groups: 1 cache + 1-7 capacity disks per group
- RAID-1 mirroring for data protection (FTT=1 uses 2x storage)
- Maximum 5 disk groups per host
ESA (Express Storage Architecture):
- Single-tier: All NVMe drives in a flat pool (no cache/capacity distinction)
- No disk groups: All disks contribute to a single storage pool per host
- RAID-5/6 erasure coding for data protection (OSA RAID-5 FTT=1 = 1.33x; ESA adaptive RAID-5 FTT=1 = 1.25x at 6+ hosts or 1.5x below; RAID-6 FTT=2 = 1.5x)
- Minimum 128 GB RAM per host
- Significant storage efficiency improvement (RAID-5 vs RAID-1)
ESA Host Requirements (VCF 9.0):
- CPU: 2 sockets minimum, 32+ cores recommended
- Memory: 128 GB minimum (ESA metadata caching is memory-intensive)
- Storage: 2+ NVMe drives per host (no SATA/SAS SSD)
- Network: 25 GbE minimum for vSAN traffic (10 GbE insufficient for ESA performance)
vSAN ESA Day-2 Operations:
---
Disk Lifecycle:
Add Disk:
- Insert new NVMe drive into host
- ESXi detects drive automatically
- vSAN claims disk into the storage pool (no disk group management needed)
- Data automatically rebalances across all disks (background operation)
Remove/Replace Disk:
- Mark disk for decommission in vSAN UI
- vSAN evacuates data from the disk (rebuilds on remaining disks)
- Wait for evacuation to complete (monitor vSAN health)
- Physically remove disk
5. Insert replacement if needed → auto-claimed
Disk Failure Handling:
- vSAN detects failure immediately via NVMe health monitoring
- Failed disk marked as 'Degraded' or 'Absent'
- vSAN begins automatic rebuild on remaining disks
- Rebuild time depends on data volume and cluster load
- FTT policy determines if data is still accessible during rebuild
Host Maintenance Mode in vSAN ESA:
---
Three evacuation modes (same as OSA, but ESA behavior differs):
- Full Data Migration:
- All data on the host's disks is rebuilt on other hosts
- Safest option: cluster can tolerate additional failures during maintenance
- Longest duration: proportional to data on the host (can take hours for TB-scale)
- Use for: extended maintenance (hardware replacement, firmware update)
- Ensure Accessibility:
- Only data components that would become inaccessible are migrated
- Faster than full migration
- Cluster resilience is reduced during maintenance (cannot tolerate additional failure)
- Use for: short maintenance (reboot, quick patch)
- No Data Migration:
- No data is moved (immediate entry to maintenance mode)
- Risk: if another host fails, data may be inaccessible
- Use for: emergency maintenance only or when cluster has sufficient redundancy (FTT=2)
Fault Domain Management:
---
Fault domains define failure boundaries for data placement:
- Without fault domains: vSAN distributes replicas across hosts (host = failure unit)
- With fault domains: vSAN distributes replicas across domains (rack/enclosure = failure unit)
Configuration:
vSphere Client → Cluster → Configure → vSAN → Fault Domains → Add Domain
Assign hosts to domains based on physical topology (rack, power circuit, network switch)
Rules:
- Minimum 3 fault domains for FTT=1 (data on 2 domains, witness on 3rd)
- Data replicas/stripes are placed in DIFFERENT fault domains
- If a domain has fewer hosts, it limits the fault domain's capacity contribution
Stretched Cluster Configuration:
Two data sites + 1 witness site
Latency requirements: < 5ms RTT between data sites (no geographic distance restriction)
Site affinity: VMs prefer local storage reads but replicate writes to remote site
Storage Policy Management:
---
vSAN storage policies define data protection and performance characteristics:
Key policy parameters:
- FTT (Failures to Tolerate): 0-3 (ESA supports higher efficiency via RAID-5/6)
- RAID: RAID-1 (mirroring), RAID-5/6 (erasure coding — ESA preferred)
- Object Space Reservation: Thin (0%) or Thick (100%)
- IOPS Limit: Per-object IOPS cap for noisy neighbor prevention
ESA-specific policies:
Compression: Always on (ESA compression is near-zero CPU overhead)
Deduplication: Enabled by default (inline, not post-process)
Encryption: At-rest encryption with per-host key management
Key Takeaways
- ESA vs OSA: Single-tier NVMe pool (no disk groups), RAID-5/6 erasure coding (ESA adaptive RAID-5: 1.25x at 6+ hosts, 1.5x below; RAID-6 4+2 = 1.5x - versus 2x for RAID-1 mirroring), 128 GB RAM minimum, 25 GbE minimum. ESA is the VCF 9.0 default.
- Disk lifecycle: ESA disks are auto-claimed into a flat pool. Add = auto-claim + rebalance. Remove = evacuate data + physical removal. Failure = auto-detect + auto-rebuild.
- Maintenance mode: Full Migration (safest, longest), Ensure Accessibility (faster, reduced resilience), No Migration (emergency only). Choice depends on maintenance duration and cluster FTT level.
- Fault domains: map to physical topology (rack, power, switch). Minimum 3 domains for FTT=1. Stretched cluster: 2 data sites + witness, < 5ms RTT requirement.
VCF Security Hardening & Audit#
VCF security hardening covers ESXi lockdown mode, TPM attestation, STIG compliance, audit logging, and secure lifecycle operations. VCAP Admins must implement defense-in-depth security while maintaining operational access for management tools.
ESXi Lockdown Mode:
---
Lockdown mode restricts direct access to ESXi hosts, forcing all management through vCenter:
Modes:
Normal Lockdown: DCUI access allowed, direct SSH/API blocked
- Only vCenter can manage the host
- DCUI available for emergency access (root login at console)
- Exception users: configured in vCenter for specific access needs
Strict Lockdown: DCUI access also blocked
- Only vCenter management path available
- No local console access (even DCUI)
- WARNING: If vCenter is unavailable, the host is unmanageable
- Use with caution — requires reliable vCenter HA
Exception Users:
Added via vCenter → Host → Configure → Security Profile → Lockdown Mode
Purpose: Allow specific service accounts (backup agents, monitoring) direct access
Best practice: Minimize exception users; route through vCenter API where possible
VCF Impact:
VCF requires certain service accounts to have direct ESXi access
SDDC Manager/VCF Operations use exception user accounts for lifecycle operations
Do NOT enable strict lockdown on management domain hosts without testing VCF operations
TPM Attestation:
---
Trusted Platform Module (TPM) 2.0 enables host attestation:
- Verifies ESXi boot integrity (Secure Boot + TPM)
- vCenter displays attestation status per host
- VCF can enforce TPM requirement during host commission
Key Manager Integration:
- vSAN encryption and VM encryption require key management:
- Native Key Provider: vCenter built-in (uses TPM for key persistence)
- External KMS: KMIP-compliant server (Thales, Hytrust, etc.)
- VCF integrates KMS at the Instance level — all domains share the KMS
Audit Logging Architecture:
---
Comprehensive audit trail across VCF components:
ESXi: /var/log/audit/ — records all authenticated actions
vCenter: Events database — all configuration changes, logins, permissions
- Export via syslog: vCenter → Administration → System Configuration → Syslog → configure remote target
NSX: Audit logs via API or syslog — all policy changes, login events
SDDC Manager: /var/log/vmware/vcf/sddc-manager-audit.log — lifecycle operations
Centralized Logging Best Practice:
- Configure all components to forward syslog to centralized collector
(VCF Operations for Logs / Aria Operations for Logs / Splunk)
- Set syslog protocol to TCP+TLS (not UDP — UDP drops logs under load)
- Retain audit logs per compliance requirement (PCI: 1 year, SOX: 7 years, HIPAA: 6 years)
- Configure alerts on critical audit events:
- Root login to ESXi
- Permission changes on vCenter
- DFW policy changes in NSX
- VCF credential rotation events
Security Hardening Checklist (VCF 9.0):
---
- ESXi Hosts:
- [ ] Lockdown mode enabled (Normal minimum)
- [ ] SSH disabled (enable only for troubleshooting, with timeout)
- [ ] NTP configured (time sync critical for certificate validation and log correlation)
- [ ] Syslog forwarding to centralized collector
- [ ] TLS 1.2 minimum for all management interfaces
- [ ] vMotion encryption enabled
- [ ] vSAN encryption enabled (at-rest + in-transit)
- [ ] Firewall rules restricted to VCF-required ports only
- vCenter:
- [ ] AD/LDAP integration for human admin accounts
- [ ] MFA enabled via OIDC federation (VCF 9.0)
- [ ] Password policy: 15+ characters, complexity, 90-day rotation
- [ ] Least-privilege roles (no blanket Administrator access)
- [ ] Session timeout: 30 minutes (configurable)
- [ ] Audit logging to remote syslog
- NSX:
- [ ] Management plane certificates: CA-signed (not self-signed)
- [ ] API rate limiting enabled
- [ ] Audit logging for all policy changes
- [ ] DFW default deny for east-west traffic
- VCF:
- [ ] Credential auto-rotation enabled (90-day cycle)
- [ ] LCM depot secured (HTTPS only, authenticated access)
- [ ] Break-glass procedure documented and tested quarterly
Key Takeaways
- Lockdown mode: Normal (DCUI allowed, SSH/API blocked) vs Strict (all local access blocked). VCF requires exception users for lifecycle operations — never enable Strict without testing.
- TPM 2.0 attestation verifies host boot integrity. Native Key Provider uses TPM for key persistence. External KMS (KMIP) required for enterprise key management with vSAN/VM encryption.
- Audit logging: ESXi (/var/log/audit/), vCenter (events DB + syslog), NSX (API/syslog), SDDC Manager (audit.log). Centralize via TCP+TLS syslog to a SIEM with compliance-aligned retention.
- Hardening checklist: lockdown mode, SSH disabled, TLS 1.2, encrypted vMotion/vSAN, AD/LDAP + MFA, least-privilege roles, credential auto-rotation, break-glass procedure documented.
Exam Mapping: 3V0-11.26 — VCAP — VMware Cloud Foundation Administrator
- Section 1: Fleet Lifecycle Management — bring-up, expand, shrink, upgrade, decommission
- Section 2: Identity, Certificate & Password Management — SSO, AD/LDAP, OIDC, VMCA, credential rotation
- Section 3: Configuration Drift & Compliance — baselines, drift detection, STIG/CIS, remediation
- Section 4: Advanced Troubleshooting — SoS, log bundles, API diagnostics, component health
- Section 5: vSAN ESA Operations — disk lifecycle, maintenance mode, fault domains, storage policies
- Section 6: Security Hardening — lockdown mode, TPM, audit logging, encryption
Labs in This Section
Lab: VCF Host Lifecycle — Commission, Add, Maintain, Decommission
VCF 9.0AdvancedLab: Certificate Lifecycle — VMCA Subordinate CA & Rotation
VCF 9.0AdvancedLab: Configuration Drift Detection & STIG Compliance
VCF 9.0AdvancedLab: SoS Diagnostics & VCF API Troubleshooting
VCF 9.0AdvancedLab: VCF Security Hardening — Lockdown, Encryption & Audit
VCF 9.0AdvancedReferences
- VCAP-Admin Exam Guide (3V0-11.26)Tier 1 — Official
- VCF 9.0 Administration GuideTier 1 — Official
- vSAN ESA Operations GuideTier 1 — Official