Advanced storage design and deployment covering vSAN architecture, storage policies, capacity management, encryption, stretched clusters, and HCI integration within VCF environments.
3V0-23.25
VCAP
60
Questions
135m
Duration
300/500
Pass Score
33
Objectives
Exam Blueprint Weights
Section titles, groupings and weights below are VCDX Academy study groupings, NOT the official Broadcom blueprint structure. Broadcom publishes no section weights. Always cross-check the official exam guide. Official exam guide ↗
Section 1 — vSAN Architecture
~20%
Section 2 — Storage Policies
~20%
Section 3 — Cluster Operations
~20%
Section 4 — Performance and Capacity
~20%
Section 5 — Troubleshooting
~20%
vSAN iSCSI Target Service and vSAN Transport (RDT vs RDMA)#
vSAN iSCSI TARGET SERVICE
A cluster service (since vSAN 6.5) presenting vSAN-backed LUNs over iSCSI to initiators OUTSIDE the cluster. SUPPORTED CONSUMERS
├── Remote physical servers and VMs outside the vSAN cluster
├── WSFC (Windows Server Failover Cluster) nodes — vSAN 6.7+
└── Physical Oracle RAC — the docs call it the "preferred method for mapping VMDKs to physical Oracle RAC deployments"
NOT SUPPORTED: other vSphere/ESXi hosts or initiators, third-party hypervisors, RDM migrations, MCS.
MAXIMA: 1024 LUNs per cluster, 128 targets per cluster, 256 LUNs per target, 62 TB per LUN, 64 initiators per LUN. Stretched-cluster support added in 7.0 U1.
ESA in VCF 9.x: supported — the current usage guide is written for VCF 9.1 and its iSCSI LUN policies are managed by Auto-RAID, which is ESA-only.
DESIGN STEER: for VIRTUALIZED WSFC and Oracle RAC, Broadcom recommends native shared VMDKs (SCSI3-PR for WSFC, multiwriter for Oracle RAC) over iSCSI. Reserve the iSCSI Target Service for PHYSICAL clustered hosts and legacy block consumers.
vSAN TRANSPORT — RDT, AND WHERE RDMA FITS
By default vSAN carries I/O as UNICAST TCP over Ethernet, using its purpose-built Reliable Datagram Transport (RDT) protocol on the vSAN-tagged VMkernel port. RDT is lightweight and built to set up and tear down sessions quickly.
RDMA (RoCE v2) is an OPTIONAL alternative introduced in vSAN 7 U2, recommended mainly for ESA. Critical operational property: if any single host loses RDMA, the ENTIRE cluster falls back to TCP.
RDMA constraints: certified NICs only, DCB/PFC-capable switches, 32 hosts maximum, no LACP or IP-hash teaming, and incompatible with 2-node, stretched clusters, datastore sharing and vSAN storage clusters.
COMMON EXAM ERROR: RDT did not replace RDMA. RDT long PREDATES RDMA support in vSAN — RDMA was added later as an accelerator. Nothing about this changed in VCF 9.0.
Key Takeaways
vSAN iSCSI Target Service serves initiators outside the cluster — physical servers, VMs, WSFC nodes (6.7+) and physical Oracle RAC. NOT other ESXi hosts, third-party hypervisors, RDM migrations or MCS. Supported on ESA in VCF 9.x.
For virtualized WSFC/Oracle RAC use native shared VMDKs (SCSI3-PR / multiwriter), not iSCSI — reserve iSCSI for physical clustered hosts.
vSAN's default transport is unicast TCP over Ethernet using RDT. RDMA (RoCE v2) is optional, added in vSAN 7 U2; a single host losing RDMA drops the whole cluster back to TCP.
Exam trap: RDT did NOT replace RDMA — RDT is long-standing, RDMA came later as an optional accelerator.
vSAN Capacity Reserves and RAID Overhead (verified reference)#
RAID OVERHEAD AND HOST MINIMUMS
├── RAID-1 FTT=1 = 2.0x overhead, 3 hosts minimum
├── RAID-1 FTT=2 = 3.0x, 5 hosts
├── RAID-1 FTT=3 = 4.0x, 7 hosts
├── OSA RAID-5 (3+1) FTT=1 = 1.33x, 4 hosts — OSA only, always 3+1
├── ESA RAID-5 ADAPTIVE FTT=1: 2+1 = 1.5x when the cluster has fewer than 6 hosts (3 host min); 4+1 = 1.25x at 6+ hosts (5 host min). vSAN re-evaluates 24 HOURS after a host-count change
└── RAID-6 (4+2) FTT=2 = 1.5x, 6 hosts minimum (7 recommended) — same in ESA and OSA
VCF 9.1 AUTO-RAID retires the 4+1 scheme: RAID-5 is always 2+1 and used only at 3-5 hosts; RAID-6 4+2 takes over at 6+.
SLACK SPACE IS OBSOLETE TERMINOLOGY
The flat 25-30% "slack space" rule applied only BEFORE vSAN 7 U1. From 7 U1 it is replaced by Reserved Capacity:
├── Operations Reserve (OR): NOT a fixed percentage — varies with capacity-device size/count, disk-group count and dedup/compression. Documented worked examples span 17% / 12% / 10% / 8% / 7% / 6%
└── Host Rebuild Reserve (HRR): approximately 1/N of the cluster — 25% at 4 hosts, ~8% at 12 hosts
Combined OR+HRR under 10% is achievable. EXCEPTIONS: with fault domains configured the OR/HRR toggles are unavailable, so fall back to "about 25% of free capacity"; in VCF 9.1 both reserves are removed in favour of the Auto-RAID Effective Capacity view. Size with the vSAN ReadyNode Sizer, not a flat multiplier. Exam-style arithmetic using a 0.85 or 0.75 factor is a teaching shorthand, not Broadcom guidance.
WITNESS APPLIANCE
Four sizes — Tiny, Medium, Large, Extra Large. ESA DOES NOT SUPPORT THE TINY WITNESS; ESA starts at Medium. ESA and OSA use DIFFERENT witness OVAs with different vCPU/RAM:
├── ESA: XL 8 vCPU/64 GB · L 4/32 · M 4/16 · Tiny not supported
└── OSA: XL 6 vCPU/32 GB · L 2/32 · Normal 2/16 · Tiny 2/8
Capacity tiers: Tiny 750 components / <=10 VMs; Medium 21,833 / 500; Large 45,000 / >500; XL 64,000 / >500.
ESA MEDIA REQUIREMENTS
ESA = EXPRESS Storage Architecture (not "Elastic"). NVMe TLC flash only — at least one NVMe TLC device per storage pool. SAS and SATA devices are NOT supported, nor RAID/tri-mode controllers or HBAs. There is no HDD support of any kind. Minimum 1.6 TB production drives, 4+ devices per host recommended, host RAM >=128 GB. Only OSA (disk groups) can use magnetic media, and only in hybrid mode — which is deprecated as of the vSAN 9.0 announcement.
Key Takeaways
ESA RAID-5 is ADAPTIVE: 2+1 (1.5x) below 6 hosts, 4+1 (1.25x) at 6+, re-evaluated 24h after a host-count change. OSA RAID-5 is always 3+1 (1.33x, 4 hosts). RAID-6 4+2 FTT=2 = 1.5x, never 1.33x.
'Slack space' 25-30% is pre-vSAN 7 U1 terminology. It is now Reserved Capacity = Operations Reserve (hardware-dependent, 6-17% in documented examples) + Host Rebuild Reserve (~1/N: 25% at 4 hosts, ~8% at 12). Size with the ReadyNode Sizer.
ESA does NOT support the Tiny witness — it starts at Medium — and ESA/OSA use different witness OVAs with different vCPU/RAM.
ESA = Express Storage Architecture, NVMe TLC only. No SAS, no SATA, no HDD, no RAID/tri-mode controllers. Only OSA hybrid uses magnetic media, and hybrid OSA is deprecated.
📝 Quiz (50)
🃏 Flashcards (53)
📝 Quiz — VCAP Advanced Storage
0/50 correct
vSAN Architecture
Q1
An ESA all-NVMe cluster must adapt RAID protection dynamically for performance vs. space efficiency. Which statement is correct?
ESA uses fixed RAID-1 only
ESA uses a log-structured engine with adaptive RAID-5/6 that can shift based on cluster size
ESA uses a log-structured engine with adaptive RAID-5 erasure coding that adjusts its stripe width based on cluster size — using 2+1 for smaller clusters and 4+1 for 6+ hosts. RAID-5 and RAID-6 remain separate policy choices (not dynamically swapped). ESA is NOT fixed RAID-1 only, does NOT require dedicated cache disks (single-tier NVMe), and is architecturally different from OSA.
An OSA disk group consists of one cache tier device (SSD/NVMe for read cache + write buffer) and up to seven capacity devices. ESA eliminated this two-tier model. Witness devices are for stretched clusters. SCM is used in ESA for metadata.
Objects are logical storage units (VMDKs, VM home namespaces, swap files). Components are the replica/witness/parity pieces of those objects distributed across hosts/disk groups. Objects aren't physical disks, RAID groups, or clusters.
In VCF, the management cluster uses vSAN (OSA or ESA) as principal storage, deployed during SDDC bring-up. NFSv3, FC-only, and local VMFS are supplemental storage options, not the principal storage for VCF management domains.
ESA's single-tier architecture eliminates the separate cache tier used in OSA disk groups. All NVMe drives participate as a single pool. ESA still requires NVMe drives, supports RAID-5/6, and uses erasure coding.
vSphere CSI driver with Cloud Native Storage (CNS) uses vSAN storage policies to dynamically provision persistent volumes for Kubernetes. Manual iSCSI, direct NFS per pod, and local hostPath volumes don't provide policy-driven integration.
In a vSAN ESA cluster, an architect is designing for maximum write performance on mixed read/write OLTP workloads. Which statement accurately describes ESA's write path?
All writes are mirrored to a dedicated SSD cache tier before destaging to capacity
ESA uses a log-structured file system (LFS) and a 'durability component' at the performance leg so RAID-5/6 writes avoid the read-modify-write penalty
Writes bypass NVMe and go directly to NVDIMM on each host
ESA requires a minimum of two cache devices per host to function
ESA uses a log-structured file system (LFS) with a 'durability component' at the performance leg, eliminating the read-modify-write penalty of RAID-5/6 that OSA suffered. This dramatically improves write performance. ESA doesn't use a separate SSD cache tier, NVDIMM, or require two cache devices.
Principal storage is deployed during WLD creation by SDDC Manager and is mandatory. Supplemental storage (FC, NFS, vVols) is added post-deployment and can host VMs. Supplemental is not limited to ISOs, not restricted to one datastore, and doesn't require vVols.
CNS leverages vSAN File Services to provision an NFS v4.1 file share for RWX PVCs. Standard block VMDKs are RWO only. vSAN does support RWX through File Services. vSphere Pods don't fall back to hostPath.
RAID-6 (FTT=2) provides dual-failure tolerance with best space efficiency — 4+2 erasure coding uses ~67% of raw capacity vs RAID-1 FTT=2 which uses only 33%. RAID-0 has no protection. FTT=0 can't tolerate any failures.
The question specifies a 4+1 RAID-5 scheme, which requires 5 hosts (4 data + 1 parity across 5 fault domains). Note: vSAN OSA also supports a 3+1 RAID-5 scheme requiring only 4 hosts minimum. The 4+1 scheme provides better space efficiency (80% usable vs 75% for 3+1). In practice, VMware recommends having at least one additional host beyond the minimum for maintenance operations without degrading protection.
Object Space Reservation (OSR) controls thick vs thin provisioning on vSAN — 100% OSR reserves the full object space upfront. It doesn't control cache sizing, compression ratios, or network bandwidth.
PFTT=1 mirrors across sites for site-level protection. SFTT=1 adds local host-level protection within each site using RAID-1 or RAID-5/6. PFTT=0 provides no site protection. Only FTT=1 without site awareness doesn't protect against site failure.
IOPS limit per object protects against noisy neighbors by capping I/O on individual objects. Encryption, thin provisioning, and checksum settings don't address I/O contention between tenants.
vSAN data-at-rest encryption requires a KMIP-compliant KMS server trusted by vCenter for key management. No external components means no key management. Self-signed certs don't provide KMS. Deduplication is independent of encryption.
Data-in-transit encryption encrypts intra-cluster vSAN traffic between hosts, protecting data moving between components during replication and rebuild. It doesn't replace vCenter TLS, isn't ESA-only, and doesn't require dedicated NICs.
An architect must maximize usable capacity on a 12-host ESA cluster storing archive data. Which policy choice provides the best space efficiency while tolerating dual failures?
RAID-6 FTT=2 (4+2) with compression on a 12-host ESA cluster maximizes usable capacity (~67% of raw) while tolerating dual failures. RAID-1 FTT=2 uses only ~33% usable. RAID-5 FTT=1 only tolerates single failures.
Given a 1 TB VMDK with RAID-1 FTT=1 on OSA (no compression), and the same VMDK using RAID-5 FTT=1 on ESA, what are the approximate raw capacity consumptions?
On a vSAN stretched cluster, a tier-1 VM must survive the loss of an entire site AND tolerate one host failure at the surviving site. Which policy settings are required?
PFTT=0, SFTT=1, site affinity to preferred
PFTT=1 (mirroring across sites), SFTT=1 (RAID-1 or RAID-5/6 within site)
PFTT=1 provides site mirroring (survives site loss). SFTT=1 adds local host protection within each site. Together they ensure VM survival through both site and host failures. PFTT=0 doesn't protect sites. FTT=1 alone without PFTT/SFTT distinction doesn't specify site behavior.
The witness must reside at a third site with independent connectivity to both data sites, ensuring quorum decisions aren't affected by inter-site link failures. Placing it on a data site defeats the quorum purpose. Nested clusters and physical ToRs are inappropriate locations.
'Ensure Accessibility' moves only the minimum data needed to maintain availability for the remaining replicas. It does NOT fully evacuate data (that's 'Full Data Migration'). The host isn't forcibly removed. vSAN isn't disabled.
HCI Mesh allows one vSAN cluster (client) to mount a remote vSAN datastore from another cluster (server), sharing capacity across clusters. It doesn't share VMFS with NFS, migrate pods to AWS, or replace NFS with ESA.
vSAN File Services provides NFS v3/v4.1 and SMB file shares on top of vSAN storage. It doesn't provide iSCSI targets (that's vSAN iSCSI Target Service), FC LUNs, or NVMe passthrough.
New disks are added to a new disk group, and data is rebalanced over time through proactive rebalancing. Disks don't immediately receive all data. They don't need VMFS formatting. They can host capacity objects after rebalancing.
vSAN iSCSI Target Service supports external physical servers and clustered applications (like MSCS) that need shared storage access. It's not limited to VMs only, ESA clusters only, or Kubernetes pods only.
Site affinity rules pin VMs preferentially to a given site's hosts, controlling where VMs restart after a site failure. They don't prevent encryption, disable HA, or require manual failover.
When the witness loses connectivity to both sites, the preferred site continues servicing I/O for objects with site affinity. Cross-site objects without preference may become inaccessible until quorum is restored. VMs don't all power off, the secondary doesn't become witness, and vCenter doesn't cast tiebreaker votes.
An admin places a host in maintenance mode with 'Ensure Accessibility'. Two days later a second host fails. What is the most accurate outcome for FTT=1 objects whose components resided on the first host?
Data is fully protected because Ensure Accessibility rebuilds replicas
Objects may become inaccessible because Ensure Accessibility does not rebuild the missing replica; a second failure removes the last copy
vSAN automatically escalates to Full Data Migration
'Ensure Accessibility' does NOT rebuild the missing replica — it only ensures the remaining copies stay accessible. If a second host then fails, objects lose their last copy and become inaccessible. vSAN doesn't auto-escalate to Full Data Migration, and HA can't restart on the witness.
Which statement about HCI Mesh (cross-cluster capacity sharing) is correct?
The client cluster must run ESA; server cluster must run OSA
A vSAN cluster (client) mounts a remote vSAN datastore from another cluster (server) in the same vCenter; storage policies and I/O are serviced by the server cluster
HCI Mesh only works with vSAN File Services enabled
In HCI Mesh, a vSAN cluster (client) mounts a remote vSAN datastore from another cluster (server) in the same vCenter; storage policies and I/O are serviced by the server cluster. Both clusters don't need identical hardware. It's not limited to ESA client/OSA server or File Services only.
vSAN IOInsight provides detailed I/O profiling including block size distribution, read/write ratio, and I/O patterns. This data informs policy and sizing decisions. It doesn't handle procurement quoting, key rotation, or HA replacement.
OSA slack space of 25-30% is recommended for rebuilds, resyncs, and normal operations. 0% leaves no room for recovery. 50% is excessive. 5% is insufficient for rebuild headroom.
Moving from RAID-1 (50% usable) to RAID-6 4+2 (67% usable) reduces overhead while maintaining dual-failure protection. Usable capacity does not double or remain unchanged.
High congestion on OSA indicates cache tier destage back-pressure — the cache can't flush to capacity fast enough, throttling incoming I/O. Network congestion, NSX Edge, and vCenter DB locks are unrelated to vSAN congestion metrics.
VCF Operations vSAN capacity dashboards provide what-if analysis for vSAN capacity planning. ESXi shell has limited visibility. NSX Traceflow is for network paths. vRO workflows automate tasks but don't provide capacity analytics.
vSAN snapshot replication provides asynchronous, policy-driven RPO replication, enabling DR with configurable recovery point objectives. It's not synchronous-only, not NFS mirroring, and not array-based stretched LUNs.
Low cache hit rate means the working set exceeds cache capacity, increasing destage pressure as more reads miss cache and hit capacity tier. The workload is not well-sized. Disabling encryption or removing RAID-6 won't fix cache sizing issues.
25-30% operational reserve is recommended for rebuilds, resyncs, policy changes, and to avoid performance degradation as clusters fill. 10% for logs, 5% for core dumps, and 50% always are all incorrect recommendations.
vSAN IOInsight reports an average I/O size of 4 KB and 95% reads for a database workload, yet latency p99 is elevated. Which policy adjustment most directly addresses the issue on ESA?
Switch from RAID-6 to RAID-1 to eliminate parity reads on the hot path
For small random read-heavy workloads with elevated p99 latency, switching from RAID-6 to RAID-1 eliminates the parity read overhead on the hot path. Encryption, OSR, and checksum changes don't address read amplification from parity calculations.
vSAN cluster partition means vSAN VMkernel port communication has failed between subsets of hosts, creating isolated partitions. It's not caused by vCenter restarts, disk group rebuilds, or policy misconfiguration.
When a capacity device fails, vSAN automatically rebuilds components on remaining healthy disks/hosts to restore policy compliance. The host isn't removed. VMs don't power off. The disk group isn't preserved with a failed drive indefinitely (after timeout, rebuild starts).
Non-compliant objects: first check component placement and resync status in vSAN health to understand which components are missing or degraded. vCenter backups, NSX routing, and DRS rules don't address vSAN object compliance.
'esxcli vsan' provides low-level vSAN diagnostics on ESXi including storage, network, health, and cluster commands. 'esxcli software profile' manages VIBs. 'esxcli nfs41' is for NFS. 'esxcli system account' manages local accounts.
Persistent high latency across stretched cluster sites indicates inter-site link RTT exceeding supported thresholds (5ms recommended, max varies by version) or ISL bandwidth saturation. Witness CPU, vCenter certs, and NTP drift on vCenter alone don't cause cross-site storage latency.
Ruby vSphere Console (RVC) provides advanced vSAN diagnostics and cluster-level commands not easily accessible through the GUI. It doesn't edit policies via GUI, replace the vCenter UI, or install ESXi.
VM CPU Ready % is a compute metric and is LEAST likely to cause NFS APD (All Paths Down). Array controller failover, network path/MTU mismatch, and NFS export settings are all storage/network causes of APD.
A vSAN health check reports 'vSAN cluster partition' with two subclusters of 3 and 2 hosts. The VMs on the 2-host side lose storage access. What is the underlying quorum rule?
Whichever subcluster has the highest host numbering wins
A component requires more than 50% of its votes to be accessible; the 2-host subcluster lacks quorum for most objects
vCenter decides the winner based on uptime
The witness always sides with the smaller partition
vSAN quorum requires >50% of component votes for accessibility. In a 3+2 partition, the 3-host subcluster has majority and retains quorum. The 2-host side lacks quorum for most objects. vCenter doesn't decide winners, host numbering is irrelevant, and witnesses are for stretched clusters.
After a capacity-tier NVMe drive fails in an ESA host, the admin sees 'resync' objects in Skyline Health. Which statement is correct about ESA rebuild behavior?
ESA requires manual re-creation of the storage pool before rebuild starts
ESA rebuilds using remaining healthy NVMe drives across hosts based on policy; because there is no cache tier, rebuilds leverage all drives in parallel for higher throughput
Rebuild is blocked until the failed drive is physically replaced
Only the dedicated cache drive can participate in rebuilds
ESA rebuilds using remaining healthy NVMe drives across hosts in parallel. Since there's no cache tier, all drives participate in the rebuild, providing higher rebuild throughput. ESA doesn't require manual pool re-creation, doesn't wait for physical replacement, and has no dedicated cache drive.
The logical storage unit in vSAN (a VMDK, namespace, or swap) composed of multiple components distributed across hosts by policy. Objects are opaque to the guest but visible in vSAN management. They are the atomic unit of placement and repair.
vSAN Component
A piece of a vSAN object (data mirror, RAID stripe, or witness) placed on a specific disk or host. Components are the unit vSAN rebuilds on host/disk failure. Per-host limit is 9,000 components (both OSA and ESA); component counts drive some design decisions, especially with high stripe-width or high FTT policies on OSA.
Witness Component
A metadata-only object piece used as tiebreaker between full data components, consuming negligible capacity but crucial for quorum after a failure. Witnesses live on hosts or on dedicated witness appliances in stretched clusters. They are invisible to storage capacity math but counted toward component limits.
ESA Log-Structured Engine
The vSAN ESA write path that collects incoming writes into a compact log before committing to NVMe capacity, enabling efficient RAID-5/6 with small random writes. It eliminates the OSA cache-tier destage step. It is why ESA delivers mirror-like performance with erasure-coded efficiency.
Adaptive RAID
A vSAN ESA behavior that automatically chooses RAID-5 or RAID-6 layouts based on cluster size, optimizing capacity efficiency at the storage level. The admin's policy simply specifies FTT; the engine selects the width. It is a key simplification over OSA erasure coding.
Disaggregated Storage Cluster
A vSAN deployment mode where storage hosts serve vSAN capacity to compute clusters without running those VMs themselves, enabling independent scaling of compute and storage. It is closely related to but distinct from HCI Mesh. It is a core use case of vSAN Max.
Supervisor vSAN Context
The integration where vSAN provides persistent volumes to Kubernetes workloads on a Supervisor via the CNS/CSI stack. Each PVC maps to a vSAN object on the vSAN datastore. It is the storage foundation of VKS stateful workloads.
ESA Storage Pool
In vSAN ESA every claimed NVMe device joins a single per-host storage pool (no cache/capacity tiers). Data services run across the pool natively, simplifying hardware layout and removing cache-tier sizing.
Stretched Cluster Requirements
Two data fault domains plus one Witness host/appliance on a third site. ≤5 ms RTT between data sites, ≤200 ms RTT to Witness, minimum 10 Gbps between sites (ESA requires ≥25 Gbps for best performance).
vSAN Max (Disaggregated)
A dedicated vSAN ESA storage cluster that exports datastores over RDMA to compute clusters via HCI Mesh. Supports very large storage footprints (up to petabytes) decoupled from compute sizing.
Object Space Reservation (OSR)
A storage-policy setting that pre-reserves a percentage of a thin-provisioned object's logical size, trading flexibility for guaranteed provisioning. A value of 100% is effectively thick-provisioned. It interacts with slack-space planning.
IOPS Limit Policy
A storage-policy rule capping IOPS for a VM or VMDK to prevent noisy-neighbor contention in shared clusters. Limits are enforced at the vSAN layer. It is a common tool in multi-tenant environments.
PFTT
Primary Failures to Tolerate in a stretched cluster, controlling how many site-level failures a policy can survive — typically 1 for dual-site clusters. It is separate from SFTT, which controls intra-site failures. Together they form the stretched-cluster protection level.
SFTT
Secondary Failures to Tolerate in a stretched cluster, controlling how many host-level failures within a site a policy can survive. Policies can specify both PFTT and SFTT independently. This two-level protection is unique to stretched clusters.
vSAN Encryption Policy
vSAN data-at-rest encryption is enabled at the cluster level on both OSA and ESA and encrypts the entire datastore. Keys can be supplied by an external KMIP-compliant KMS or by the vSphere Native Key Provider (NKP). Per-VM encryption policy-based control is provided by vSphere VM Encryption (separate from vSAN encryption). KMS/NKP connectivity is a hard prerequisite.
Data-in-Transit Encryption
A vSAN cluster-level setting that encrypts traffic between hosts for writes and replication, independent of at-rest encryption. It does not require external KMS because keys are generated internally. It is an important control for in-rack compromise scenarios.
Storage Rule Compliance
The match-state between a VM's assigned storage policy and the actual layout of its vSAN components. Objects can be Compliant, Non-Compliant, or Reconfiguring during policy change. Monitoring compliance is a continuous operational task in advanced storage environments.
Checksum Policy
Enable/Disable end-to-end checksum per storage policy. ESA always uses checksum on the ingress log path; OSA exposes it as a policy toggle. Disabling trades integrity for marginal CPU reduction (not recommended).
Force Provisioning
Allows an object to be created even if the cluster cannot currently satisfy full policy (e.g., insufficient hosts for RAID-6). Object remains in non-compliant state and reflows when resources become available.
Host Affinity / Storage Tag Rule
Storage policy rule using host/fault-domain tags to steer placement (e.g., 'data on site A only'). Used for data sovereignty and stretched-cluster site-local objects.
Auto-Policy Management
vSAN can auto-assign storage policy based on cluster type (stretched, ESA) and FTT availability, ensuring objects use an optimal erasure-coding scheme as cluster size changes.
Witness Appliance
A small ESXi virtual appliance hosting only witness components for a vSAN stretched cluster, deployed at a third site. It does not run workloads. It must remain reachable from both data sites for quorum decisions.
Site Affinity Rule
A vSAN stretched-cluster storage-policy option that pins a VM's data to a specific site, reducing cross-site latency at the cost of site-failure protection. Use for workloads sized to fit one site. It is configured per-policy.
Decommission Mode
The host-maintenance option controlling how vSAN handles components on an exiting host: Ensure Accessibility (fastest, weaker protection), Full Data Migration (slowest, strongest), or No Action. Choice is dictated by cluster size and urgency. It is the first prompt when entering maintenance mode.
vSAN File Services
A per-host NFS/SMB file-server stack running as VMs on each vSAN cluster node, presenting shares from the vSAN datastore. It supports access control, quotas, and snapshots. It is enabled per cluster and integrates with AD for SMB.
vSAN iSCSI Target
A service that exposes vSAN objects as iSCSI LUNs to external initiators, typically for physical servers or specialized appliances. It extends vSAN beyond virtualized consumers. It is configured per target group with CHAP security.
HCI Mesh Capacity Sharing
A vSAN feature where a client cluster mounts a server cluster's vSAN datastore remotely, unifying scattered capacity. It improves utilization without physically moving disks. It is a primary use case for disaggregated design.
KMS Integration
The configuration that registers a Key Management Server with vCenter so vSAN can request encryption keys. VCF 9 supports KMIP-compliant vendors and vSphere Native Key Provider. Failures surface as encryption-operation errors and policy non-compliance.
Maintenance Mode Data Options
Ensure accessibility (move only components that would otherwise go inaccessible), Full data migration (evacuate all), No data migration (fastest, may reduce redundancy). Choice drives rebuild traffic and maintenance-window length.
Fault Domain Sizing
Fault domains group hosts (typically by rack). A cluster needs at least 2·FTT+1 fault domains for the target FTT level; recommended to keep equal host counts per FD for balanced rebuild.
Cluster Shutdown Wizard
Guided workflow that shuts down VMs, disables HA/DRS, and places all hosts in maintenance in the correct order for planned cluster shutdown (DC migration, power work). Reverse wizard brings cluster back up.
vSAN IOInsight
A diagnostic tool that samples live I/O on selected VMs or hosts, reporting block-size histograms, queue depths, and read/write ratios. It is invaluable when workload characteristics are disputed. It is non-disruptive and time-bounded.
Raw vs Usable Capacity
The distinction between total drive bytes (raw) and the capacity actually available for VM data after FTT/FTM overhead, slack space, metadata, and compression estimates (usable). VCAP candidates must calculate this for typical OSA and ESA policies. It drives sizing proposals.
Slack Space Requirement
The 25-30% of vSAN capacity that must remain free for rebalancing, resync, and transient overhead. Exceeding 70-75% utilization triggers warnings and degrades availability. It is non-negotiable in any competent design.
Compression Factor
The data-reduction ratio achieved by vSAN ESA's compression, applied at the top of the stack (before writes traverse the network). Compression is enabled by default on every ESA cluster; it cannot be disabled cluster-wide via a single toggle but can be disabled per-VM via a custom storage policy. New writes are compressed; pre-existing uncompressed blocks are not retroactively compressed.
vSAN-to-vSAN Snapshot Replication
A DR capability where vSAN snapshots are asynchronously replicated to another vSAN cluster on a configurable RPO. It is a VMware-native alternative to third-party replication. It is configured per VM or per policy.
Capacity What-If Scenario (vSAN)
A VCF Operations or vSAN-specific projection that models adding workloads, changing policies, or adding hosts and reports resulting usable capacity and protection impact. It is the tool for sizing conversations with app teams. It replaces guesswork with numeric tradeoffs.
Congestion Threshold
The vSAN DOM back-pressure point above which writes are slowed to protect the cluster from overload. Persistent congestion above ~75 indicates an undersized cluster or a saturated component. It is visible in Performance Service charts.
Dedup Scope (OSA vs ESA)
In OSA, dedup/compression operates per disk group; in ESA, compression is per-object with global dedup in the log path. ESA typically achieves better ratios with lower overhead.
vSAN Capacity Reporting Layers
Reports distinguish Raw, Used (after FTT), Free, Used by VMs (logical), and Efficiency savings. Slack/transient capacity is held back for rebuild, snapshots, and internal operations.
Performance Service
Per-cluster stats DB (stored as a vSAN object) collecting IOPS/latency/bandwidth per object, host, disk group, and network. Required for historical perf views; sized automatically by cluster capacity.
Adaptive Resync
Dynamic scheduler that allocates bandwidth between VM I/O, resync, and namespace ops based on current load. Replaces manual resync throttling in newer releases.
Skyline Health (vSAN)
The primary vSAN health dashboard covering hardware compatibility, network, cluster, data, and performance checks with direct links to KBs. Green/yellow/red status drives most vSAN triage. Ignoring Skyline findings is a common cause of avoidable outages.
Object Non-Compliance
A state where an object's current placement does not match its storage policy — typically during rebuild or insufficient fault domains. Non-compliance does not always mean data loss risk, but it always warrants action. vCenter VM > Monitor > vSAN > Physical Disk Placement exposes the details.
Rebuild Behavior
How vSAN regenerates missing or degraded components after a failure, driven by object policy and available capacity. ESA rebuilds at disk granularity; OSA cache failures force disk-group rebuild. Understanding rebuild behavior is essential for predicting resync duration.
Split-Brain (vSAN)
A stretched-cluster condition where both sites temporarily believe they have quorum, typically due to a witness partition and simultaneous link issues. vSAN preferred-site logic prevents data divergence but can pause writes. Post-incident, cmmds-tool confirms component state across partitions.
RVC (Ruby vSphere Console)
A command-line toolkit that exposes deep vSAN and vSphere inspection commands useful for advanced diagnostics. Historically the primary vSAN CLI, it is now complemented by esxcli vsan and UI-based tools. It is still invaluable for niche cases.
vSAN Observer
A legacy real-time diagnostic front-end (launched from RVC) that renders vSAN internal metrics for deep performance analysis. Most of its capabilities are now in vSAN Performance Service. It remains useful when capturing point-in-time traces during support engagements.
Resync Throttling
The operational practice of limiting vSAN's rebuild bandwidth to protect production I/O during large rebuild events. Throttling is configured cluster-level and is a trade-off: slower repair vs. less latency impact. It is rarely used but essential to know.
CMMDS
Cluster Monitoring, Membership, and Directory Service — the in-memory distributed DB that tracks object/component state across hosts. Key tool: 'cmmds-tool find' to inspect object records during troubleshooting.
esxcli vsan commands
Core CLI: 'esxcli vsan cluster get' (status), 'esxcli vsan storage list' (disks), 'esxcli vsan network list' (vmk), 'esxcli vsan debug object list' (object health). Used when UI is unavailable or for scripted checks.
vmkping with vSAN MTU
Verification of vSAN vmknic MTU end-to-end with 'vmkping -I vmkN -d -s 8972 <peer>' (for 9000 MTU). Silent drops caused by MTU mismatch are a common cause of degraded clusters.
vSAN Health Findings
Skyline Health groups findings into Network, Physical Disk, Data, Cluster, Capacity, Performance, Online Health, Hardware Compatibility. Each finding links to a KB and remediation steps.
Was this page useful?Thanks — noted.
Type to search. ↑↓ to move,
Enter to open, Esc to close.