Academy/VCAP — VCF Administrator (3V0-11.26)/Lab: Configuration Drift Detection & STIG Compliance
This lab targets VCF 9.0

Lab: Configuration Drift Detection & STIG Compliance

VCF 9.0Advancedvcap-advanced⏱ 90 min

Drift detection, compliance baselines, STIG/CIS profiles, remediation workflows

Objectives

  • Create and apply configuration baselines for ESXi hosts
  • Detect configuration drift using VCF Operations compliance features
  • Implement STIG compliance checks for DoD/regulated environments
  • Build a compliance dashboard with drift trending
  • Execute semi-automated drift remediation

Prerequisites

VCF 9.0 with VCF Operations monitoring ESXi hosts, at least 2 hosts in a cluster for drift comparison

Prior labs: vcap-admin-01

Required skills:

  • ESXi advanced settings
  • VCF Operations dashboards
  • Security compliance concepts

Lab Environment

VCF 9.0 management domain with VCF Operations collecting host configuration data. 2+ ESXi hosts in a cluster. One host intentionally misconfigured to demonstrate drift detection.

Tasks

Task 1 Drift Detection, STIG Compliance & Remediation

Configuration drift is inevitable — manual changes, failed patches, firmware updates all introduce deviations. The goal is not zero drift (impossible) but rapid detection and remediation. VCF Operations provides the detection layer; Host Profiles and automation provide the remediation layer. Always tag deliberate exceptions to avoid false positive alerts.

Build a complete compliance management workflow: define configuration baselines, detect drift across the fleet, assess STIG compliance posture, and remediate deviations — demonstrating the operational process for maintaining consistent, secure configurations at scale.

Step 1

Create a configuration baseline. Identify a 'golden' host in the cluster that represents the desired configuration state. Document its key settings: (a) NTP servers (esxcli system ntp get); (b) DNS servers (esxcli network ip dns server list); (c) Syslog target (esxcli system syslog config get); (d) Lockdown mode (vim-cmd hostsvc/hosthardware | grep lockdownMode); (e) SSH status (chkconfig --list | grep SSH); (f) Power policy (esxcli hardware power policy get); (g) Advanced settings: Security.AccountLockFailures, UserVars.ESXiShellInteractiveTimeOut, etc. Export this as a baseline document or Host Profile.

Step 2

Introduce intentional drift. On a second host (not the golden host), make changes: (a) Change NTP server: esxcli system ntp set --server=10.0.0.99; (b) Enable SSH: vim-cmd hostsvc/enable_ssh; (c) Disable lockdown mode: vim-cmd hostsvc/hosthardware lockdownmode --disable; (d) Change syslog target: esxcli system syslog config set --loghost=udp://10.0.0.99:514. These represent typical drift scenarios in production environments.

Step 3
Detect drift with VCF Operations. In VCF Operations, navigate to the compliance view (Environment → Compliance or use a custom dashboard). If using vSphere Host Profiles: vCenter → Policies and Profiles → Host Profiles → select golden host → Extract Profile → Attach to cluster → Check Compliance. The compliance check compares all hosts against the baseline and lists deviations per host per setting.
Step 4
Build compliance dashboard. VCF Operations → Dashboards → Create. Add widgets: (a) Scoreboard: 'Fleet Compliance Score' = percentage of hosts matching baseline; (b) Heat Map: hosts colored by deviation count (green=0, red=>5); (c) List: 'Top Drift Items' — which settings drift most frequently across the fleet; (d) Trend: Compliance score over 30 days (is drift being remediated or accumulating?).
Step 5

STIG compliance assessment. Apply a STIG compliance profile: either import a VMware-published STIG baseline (available from public.cyber.mil for vSphere) or manually check key STIG controls: CAT I: (a) TLS 1.2 minimum (esxcli system tls server get # ESXi 8.0 U3+); (b) No unauthorized root access. CAT II: (c) SSH disabled; (d) Lockdown mode enabled; (e) Persistent logging configured; (f) SNMP v3 only. CAT III: (g) DCUI timeout < 600s; (h) Shell timeout < 600s. Document findings as a compliance report.

Step 6

Remediate drift. For the intentionally drifted host: (a) NTP: esxcli system ntp set --server=<correct-ntp>; (b) SSH: vim-cmd hostsvc/disable_ssh; (c) Lockdown: vim-cmd hostsvc/hosthardware lockdownmode --enable; (d) Syslog: esxcli system syslog config set --loghost=tcp://<correct-syslog>:514. After remediation, re-run compliance check — all deviations should clear. For production: use Host Profiles to remediate entire clusters in one operation.

Step 7
Automate drift detection alerting. VCF Operations → create alert definition: 'Configuration-Drift-Detected'. Symptom: compliance score < 100% (or specific configuration properties changed). Notification: email to security-ops@lab.local. This ensures drift is detected and reported immediately, not discovered during quarterly audits.
Step 8

Handle legitimate exceptions. Some hosts may have intentional configuration differences (e.g., a host running GPU workloads needs specific power policy). Document exceptions: in VCF Operations, tag the host with 'compliance-exception=power-policy' and exclude it from the power policy drift check. This prevents false positive alerts while maintaining visibility into the exception.

Step 9
Generate compliance report. VCF Operations → Reports → create template: 'Monthly-Compliance-Report'. Sections: (a) Executive summary: overall compliance score, trend, critical deviations; (b) STIG compliance: CAT I/II/III pass rates; (c) Drift details: per-host deviation list with remediation status; (d) Exceptions: documented legitimate deviations with business justification. Schedule: Monthly. Format: PDF. Distribution: security-team + audit-committee.
Step 10

Continuous compliance process. Design the ongoing workflow: (1) Weekly: automated drift scan via VCF Operations; (2) Real-time: drift alert notifications to security ops; (3) Monthly: compliance report to management; (4) Quarterly: baseline review and update (new security patches may require baseline changes); (5) Annual: full STIG/CIS audit with external assessor. Document this process as a compliance operations runbook.

Validation Gate

Check: Compliance management workflow operational from baseline to reporting

Expected: Configuration baseline defined, drift detected and quantified, STIG compliance assessed, drift remediated, compliance dashboard and reporting operational, exception management in place

Common Errors

Host Profile compliance check shows false positives
Fix: Host Profiles capture ALL settings — many are host-specific (MAC addresses, serial numbers). Edit the profile to exclude hardware-specific settings. Only include settings that should be consistent across all hosts.
STIG baseline shows failures for VCF-required settings
Fix: Some VCF operational requirements conflict with STIG controls (e.g., VCF needs certain service accounts with direct ESXi access that STIG lockdown would block). Document these as 'operational exceptions' with compensating controls.
Drift remediation breaks VCF operations
Fix: Some settings that appear non-compliant are VCF-managed (e.g., specific firewall rules for VCF inter-component communication). Never remediate VCF-managed settings manually — use VCF lifecycle operations or consult VCF documentation for expected values.
Compliance score drops after every VCF upgrade
Fix: VCF upgrades may change ESXi settings as part of the upgrade process. Update the baseline after each upgrade cycle to reflect the new expected state.

Final Validation

Configuration compliance management operational with drift detection, STIG assessment, and automated reporting

✓ Baseline defined → Golden host configuration documented with key settings

✓ Drift detection → Intentional drift detected and quantified per setting per host

✓ STIG compliance → CAT I/II/III findings documented with pass/fail status

✓ Remediation → Drifted settings corrected, compliance score restored to 100%

✓ Dashboard and reporting → Compliance dashboard operational, monthly report scheduled

Cleanup / Restore

• Remediate all intentional drift on test host

• Verify cluster compliance at 100%

• Keep compliance dashboard and reporting for ongoing use

Design Reflection (VCDX)

Compliance management is a VCDX differentiator for regulated environments. Key points: (1) proactive drift detection vs reactive audit findings; (2) exception management with documented business justification; (3) STIG/CIS alignment for government/financial sector deployments; (4) automated reporting reduces audit preparation burden.

Requirements

  • 100% compliance on CAT I STIG controls
  • Weekly drift detection with real-time alerting
  • Monthly compliance reports for audit committee

Constraints

  • VCF-managed settings cannot be remediated manually
  • Host Profile remediation requires maintenance mode (brief host impact)
  • STIG controls may conflict with VCF operational requirements

Assumptions

  • Golden host represents the desired production state
  • Legitimate exceptions are documented before drift detection is enabled
  • Security team reviews and acts on drift alerts within 24 hours

Risks

  • Undocumented exceptions generate false positive alerts — erodes trust in compliance system
  • Remediation of VCF-managed settings breaks lifecycle operations
  • Compliance score regression after VCF upgrades if baseline is not updated

Self-Assessment Discussion Prompts

  1. How do you handle drift that is introduced by VCF lifecycle operations?
  2. What compensating controls do you implement for STIG exceptions required by VCF?
  3. How would you scale compliance management from 10 to 1,000 ESXi hosts?

Extensions

Import the DoD vSphere STIG baseline from public.cyber.mil and run a full assessment

Automate drift remediation using PowerCLI scripts triggered by VCF Operations alerts

Build a multi-cluster compliance comparison dashboard

⚠ Known Pitfalls (from Community KB)

Using Host Profiles without excluding hardware-specific settings — generates hundreds of false positive deviations
Remediating VCF-managed ESXi settings manually — breaks lifecycle operations and credential management
Not updating baselines after VCF upgrades — causes compliance score to drop, eroding trust in the compliance system
Applying STIG controls without testing VCF operations — some controls break VCF inter-component communication

References

  • DISA vSphere STIG: public.cyber.mil
  • CIS ESXi Benchmark: cisecurity.org
  • VCF 9.0 Security Hardening Guide: techdocs.broadcom.com
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.