Lab: Configuration Drift Detection & STIG Compliance
Objectives
- Create and apply configuration baselines for ESXi hosts
- Detect configuration drift using VCF Operations compliance features
- Implement STIG compliance checks for DoD/regulated environments
- Build a compliance dashboard with drift trending
- Execute semi-automated drift remediation
Prerequisites
VCF 9.0 with VCF Operations monitoring ESXi hosts, at least 2 hosts in a cluster for drift comparison
Prior labs: vcap-admin-01
Required skills:
- ESXi advanced settings
- VCF Operations dashboards
- Security compliance concepts
Lab Environment
VCF 9.0 management domain with VCF Operations collecting host configuration data. 2+ ESXi hosts in a cluster. One host intentionally misconfigured to demonstrate drift detection.
Tasks
Task 1 Drift Detection, STIG Compliance & Remediation
Build a complete compliance management workflow: define configuration baselines, detect drift across the fleet, assess STIG compliance posture, and remediate deviations — demonstrating the operational process for maintaining consistent, secure configurations at scale.
Create a configuration baseline. Identify a 'golden' host in the cluster that represents the desired configuration state. Document its key settings: (a) NTP servers (esxcli system ntp get); (b) DNS servers (esxcli network ip dns server list); (c) Syslog target (esxcli system syslog config get); (d) Lockdown mode (vim-cmd hostsvc/hosthardware | grep lockdownMode); (e) SSH status (chkconfig --list | grep SSH); (f) Power policy (esxcli hardware power policy get); (g) Advanced settings: Security.AccountLockFailures, UserVars.ESXiShellInteractiveTimeOut, etc. Export this as a baseline document or Host Profile.
Introduce intentional drift. On a second host (not the golden host), make changes: (a) Change NTP server: esxcli system ntp set --server=10.0.0.99; (b) Enable SSH: vim-cmd hostsvc/enable_ssh; (c) Disable lockdown mode: vim-cmd hostsvc/hosthardware lockdownmode --disable; (d) Change syslog target: esxcli system syslog config set --loghost=udp://10.0.0.99:514. These represent typical drift scenarios in production environments.
Detect drift with VCF Operations. In VCF Operations, navigate to the compliance view (Environment → Compliance or use a custom dashboard). If using vSphere Host Profiles: vCenter → Policies and Profiles → Host Profiles → select golden host → Extract Profile → Attach to cluster → Check Compliance. The compliance check compares all hosts against the baseline and lists deviations per host per setting.
Build compliance dashboard. VCF Operations → Dashboards → Create. Add widgets: (a) Scoreboard: 'Fleet Compliance Score' = percentage of hosts matching baseline; (b) Heat Map: hosts colored by deviation count (green=0, red=>5); (c) List: 'Top Drift Items' — which settings drift most frequently across the fleet; (d) Trend: Compliance score over 30 days (is drift being remediated or accumulating?).
STIG compliance assessment. Apply a STIG compliance profile: either import a VMware-published STIG baseline (available from public.cyber.mil for vSphere) or manually check key STIG controls: CAT I: (a) TLS 1.2 minimum (esxcli system tls server get # ESXi 8.0 U3+); (b) No unauthorized root access. CAT II: (c) SSH disabled; (d) Lockdown mode enabled; (e) Persistent logging configured; (f) SNMP v3 only. CAT III: (g) DCUI timeout < 600s; (h) Shell timeout < 600s. Document findings as a compliance report.
Remediate drift. For the intentionally drifted host: (a) NTP: esxcli system ntp set --server=<correct-ntp>; (b) SSH: vim-cmd hostsvc/disable_ssh; (c) Lockdown: vim-cmd hostsvc/hosthardware lockdownmode --enable; (d) Syslog: esxcli system syslog config set --loghost=tcp://<correct-syslog>:514. After remediation, re-run compliance check — all deviations should clear. For production: use Host Profiles to remediate entire clusters in one operation.
Automate drift detection alerting. VCF Operations → create alert definition: 'Configuration-Drift-Detected'. Symptom: compliance score < 100% (or specific configuration properties changed). Notification: email to security-ops@lab.local. This ensures drift is detected and reported immediately, not discovered during quarterly audits.
Handle legitimate exceptions. Some hosts may have intentional configuration differences (e.g., a host running GPU workloads needs specific power policy). Document exceptions: in VCF Operations, tag the host with 'compliance-exception=power-policy' and exclude it from the power policy drift check. This prevents false positive alerts while maintaining visibility into the exception.
Generate compliance report. VCF Operations → Reports → create template: 'Monthly-Compliance-Report'. Sections: (a) Executive summary: overall compliance score, trend, critical deviations; (b) STIG compliance: CAT I/II/III pass rates; (c) Drift details: per-host deviation list with remediation status; (d) Exceptions: documented legitimate deviations with business justification. Schedule: Monthly. Format: PDF. Distribution: security-team + audit-committee.
Continuous compliance process. Design the ongoing workflow: (1) Weekly: automated drift scan via VCF Operations; (2) Real-time: drift alert notifications to security ops; (3) Monthly: compliance report to management; (4) Quarterly: baseline review and update (new security patches may require baseline changes); (5) Annual: full STIG/CIS audit with external assessor. Document this process as a compliance operations runbook.
Validation Gate
Check: Compliance management workflow operational from baseline to reporting
Expected: Configuration baseline defined, drift detected and quantified, STIG compliance assessed, drift remediated, compliance dashboard and reporting operational, exception management in place
Common Errors
Final Validation
Configuration compliance management operational with drift detection, STIG assessment, and automated reporting
✓ Baseline defined → Golden host configuration documented with key settings
✓ Drift detection → Intentional drift detected and quantified per setting per host
✓ STIG compliance → CAT I/II/III findings documented with pass/fail status
✓ Remediation → Drifted settings corrected, compliance score restored to 100%
✓ Dashboard and reporting → Compliance dashboard operational, monthly report scheduled
Cleanup / Restore
• Remediate all intentional drift on test host
• Verify cluster compliance at 100%
• Keep compliance dashboard and reporting for ongoing use
Design Reflection (VCDX)
Compliance management is a VCDX differentiator for regulated environments. Key points: (1) proactive drift detection vs reactive audit findings; (2) exception management with documented business justification; (3) STIG/CIS alignment for government/financial sector deployments; (4) automated reporting reduces audit preparation burden.
Requirements
- 100% compliance on CAT I STIG controls
- Weekly drift detection with real-time alerting
- Monthly compliance reports for audit committee
Constraints
- VCF-managed settings cannot be remediated manually
- Host Profile remediation requires maintenance mode (brief host impact)
- STIG controls may conflict with VCF operational requirements
Assumptions
- Golden host represents the desired production state
- Legitimate exceptions are documented before drift detection is enabled
- Security team reviews and acts on drift alerts within 24 hours
Risks
- Undocumented exceptions generate false positive alerts — erodes trust in compliance system
- Remediation of VCF-managed settings breaks lifecycle operations
- Compliance score regression after VCF upgrades if baseline is not updated
Self-Assessment Discussion Prompts
- How do you handle drift that is introduced by VCF lifecycle operations?
- What compensating controls do you implement for STIG exceptions required by VCF?
- How would you scale compliance management from 10 to 1,000 ESXi hosts?
Extensions
Import the DoD vSphere STIG baseline from public.cyber.mil and run a full assessment
Automate drift remediation using PowerCLI scripts triggered by VCF Operations alerts
Build a multi-cluster compliance comparison dashboard
⚠ Known Pitfalls (from Community KB)
References
- DISA vSphere STIG: public.cyber.mil
- CIS ESXi Benchmark: cisecurity.org
- VCF 9.0 Security Hardening Guide: techdocs.broadcom.com