Lab: VCF Security Hardening — Lockdown, Encryption & Audit
Objectives
- Configure ESXi lockdown mode with VCF-compatible exception users
- Enable vSAN and vMotion encryption with key management
- Configure centralized audit logging with syslog forwarding
- Implement STIG CAT I controls on ESXi hosts
- Build a security posture dashboard for continuous monitoring
Prerequisites
VCF 9.0 Instance operational, syslog collector available (VCF Operations for Logs or Splunk), TPM 2.0 enabled on hosts (optional for attestation lab)
Prior labs: vcap-admin-01, vcap-admin-03
Required skills:
- ESXi security configuration
- Encryption concepts
- Syslog/SIEM basics
Lab Environment
VCF 9.0 with 2+ ESXi hosts in a workload domain. Syslog collector (Aria Operations for Logs, Splunk, or rsyslog VM) accessible from all hosts. Optional: TPM 2.0 hardware for attestation.
Tasks
Task 1 VCF Security Hardening & Audit Configuration
Implement defense-in-depth security across a VCF Instance: ESXi lockdown mode (restrict management paths), encryption at rest and in transit (vSAN + vMotion), centralized audit logging (compliance trail), and continuous security monitoring — building a hardened environment suitable for regulated industries.
Configure ESXi lockdown mode. vSphere Client → Host → Configure → Security Profile → Lockdown Mode → Edit. Set to 'Normal' (not Strict — Strict blocks DCUI which is needed for emergency access). Add exception users: (a) the VCF service account (used by SDDC Manager for lifecycle operations — check VCF documentation for the exact account name); (b) monitoring service accounts (VCF Operations collector). Test: attempt SSH to the host as root — should be denied. Access via vCenter should still work.
Disable SSH and configure shell timeout. Host → Configure → Security Profile → Services → SSH → Stop and set to 'Start and stop manually' (not with host). Configure shell timeout: Host → Configure → System → Advanced System Settings → UserVars.ESXiShellInteractiveTimeOut = 600 (10 minutes). UserVars.ESXiShellTimeOut = 600. These settings ensure that if SSH is temporarily enabled for troubleshooting, it times out automatically.
Enable vSAN encryption. vSphere Client → Cluster → Configure → vSAN → Services → Data-At-Rest Encryption → Enable. Key provider: select Native Key Provider (uses TPM for key persistence) or configure an external KMS (KMIP server). vSAN encrypts all data on disk — both cache and capacity tiers. Encryption is per-cluster. Performance impact: < 5% with AES-NI hardware acceleration (standard on modern CPUs). Also enable: Data-In-Transit Encryption → all vSAN network traffic encrypted.
Enable vMotion encryption. Cluster → Configure → VM Overrides → Edit. Select encryption policy: 'Required' (all vMotion traffic encrypted, VM migration fails if encryption not possible), 'Opportunistic' (encrypt if both hosts support it, fall back to unencrypted), 'Disabled'. Set to 'Required' for regulated environments. vMotion encryption uses AES-256 — performance impact is minimal with AES-NI support. Verify: perform a test vMotion and check the task log for 'Encrypted vMotion' confirmation.
Configure centralized syslog. For each ESXi host: Host → Configure → System → Advanced System Settings → Syslog.global.logHost = tcp://10.0.0.50:514 (use TCP, not UDP — TCP guarantees delivery). Also set: Syslog.global.logDir = [] (ensure local logging continues), Syslog.global.logLevel = info. For vCenter: Administration → System Configuration → select vCenter → Syslog → configure target. For NSX: System → Fabric → Settings → Syslog → add collector. All components should send to the same centralized collector for correlated analysis.
Configure audit-specific logging. On ESXi: ensure /var/log/audit/ is being forwarded via syslog. On vCenter: configure audit events forwarding — Administration → System Configuration → Events → Syslog → enable. Key audit events to monitor: (a) Root logins, (b) Permission changes, (c) VM power state changes, (d) Configuration changes, (e) Certificate operations. On the syslog collector, create alerts for these high-severity events.
Implement STIG CAT I controls. Apply the highest-severity STIG requirements: (a) TLS 1.2 minimum: esxcli system security tls set -p TLSv1.2; (b) Disable TLS 1.0/1.1: verify with esxcli system security tls get; (c) Verify Secure Boot enabled: /usr/lib/vmware/secureboot/bin/secureBoot.py -s (should return 'Enabled'); (d) Verify no unauthorized users in local ESXi accounts: esxcli system account list; (e) Verify firewall restricts management to specific subnets: esxcli network firewall ruleset list. Document compliance status for each control.
Build security posture dashboard. VCF Operations → create dashboard 'Security Posture'. Widgets: (a) Scoreboard: Lockdown Mode compliance (% hosts with lockdown enabled); (b) Scoreboard: Encryption status (vSAN + vMotion encryption enabled across fleet); (c) Alert List: Security-related alerts (certificate expiry, SSH enabled, lockdown disabled); (d) Trend: Security compliance score over 30 days. Configure alerts: 'SSH Enabled on Host' → Criticality=Warning, immediate notification. 'Lockdown Mode Disabled' → Criticality=Critical.
Verify end-to-end security. Perform validation tests: (a) Attempt SSH to ESXi host — should fail (lockdown mode); (b) Attempt unauthenticated API call to vCenter — should fail (no anonymous access); (c) Check vSAN data encryption: esxcli vsan encryption info get → should show 'Enabled'; (d) Perform vMotion and verify encryption in task log; (e) Check syslog collector for log entries from all components; (f) Verify TLS version: openssl s_client -connect <host>:443 → should show TLS 1.2 or 1.3.
Create security operations runbook. Document: (a) Lockdown mode configuration with VCF exception users; (b) Encryption key management procedures (key rotation, backup, recovery); (c) Syslog architecture diagram (which components send to which collectors); (d) STIG compliance status with exception justifications; (e) Security incident response: how to enable SSH temporarily for troubleshooting (with audit trail), emergency lockdown mode bypass via DCUI; (f) Quarterly security review checklist.
Validation Gate
Check: Defense-in-depth security hardening operational with continuous monitoring
Expected: Lockdown mode enabled with VCF-compatible exceptions, vSAN and vMotion encryption active, centralized syslog collecting from all components, STIG CAT I controls implemented, security posture dashboard operational with alerting
Common Errors
Final Validation
VCF security hardening implemented with defense-in-depth and continuous monitoring
✓ Lockdown mode → Normal lockdown enabled on all hosts with VCF exception users configured
✓ SSH disabled → SSH service stopped, shell timeout configured at 600s
✓ vSAN encryption → Data-at-rest and data-in-transit encryption enabled
✓ vMotion encryption → 'Required' policy enforced, verified in vMotion task logs
✓ Centralized syslog → All components (ESXi, vCenter, NSX) sending to syslog collector via TCP
✓ STIG CAT I → TLS 1.2 enforced, Secure Boot enabled, unauthorized accounts removed
✓ Security dashboard → Posture dashboard operational with real-time alerting
Cleanup / Restore
• Keep security hardening in place (this is the desired production state)
• Document any temporary relaxations made during testing
• Verify all VCF lifecycle operations still work with hardening in place
Design Reflection (VCDX)
Security hardening is a VCDX defense essential. Demonstrate: (1) defense-in-depth: lockdown (access control) + encryption (data protection) + audit (accountability) + monitoring (detection); (2) operability balance: hardening that breaks management is counterproductive; (3) compliance alignment: STIG/CIS controls mapped to regulatory requirements; (4) key management: the most critical operational procedure for encrypted environments.
Requirements
- STIG CAT I compliance on all ESXi hosts
- All data encrypted at rest (vSAN) and in transit (vMotion)
- Centralized audit trail with 1-year retention for compliance
Constraints
- Lockdown mode requires VCF exception users — cannot fully lock down without breaking lifecycle
- Encryption requires key management infrastructure — key loss = data loss
- TLS 1.2 enforcement breaks legacy management tools
Assumptions
- AES-NI hardware acceleration available on all hosts (modern CPUs)
- Key management infrastructure (Native KP or external KMS) is reliable and backed up
- Syslog collector has sufficient storage for 1+ year retention
Risks
- Lockdown mode without exception users breaks VCF operations — always test before production enablement
- Key provider failure makes encrypted vSAN data inaccessible — back up keys to secure offline location
- Syslog collector failure creates audit gap — implement collector HA or dual-forward to two collectors
Self-Assessment Discussion Prompts
- How do you balance maximum security with VCF operational requirements?
- What is your key management backup and recovery procedure?
- How do you handle security hardening across a fleet of 50+ ESXi hosts consistently?
Extensions
Configure TPM 2.0 attestation and verify host boot integrity in vCenter
Implement dual syslog forwarding for audit trail redundancy
Build a security incident response playbook using VCF diagnostic tools
Configure vSphere Trust Authority for workload attestation
⚠ Known Pitfalls (from Community KB)
References
- vSphere 8.0 Security Configuration Guide: techdocs.broadcom.com
- VCF 9.0 Security Hardening Guide: techdocs.broadcom.com
- DISA vSphere STIG: public.cyber.mil