vSAN Storage Design with Erasure Coding
Objectives
- Design vSAN ESA storage architecture with appropriate failure tolerance
- Calculate usable capacity for mirroring vs erasure coding configurations
- Create storage policies aligned to workload tier requirements
- Plan vSAN encryption with native key management
- Evaluate deduplication and compression ratios for mixed workloads
Prerequisites
Access to VCF 9.0 and vSAN documentation; capacity planning spreadsheets
Prior labs: vcp-architect-01, vcp-architect-02
Required skills:
- Understanding of vSAN concepts (OSA vs ESA)
- Storage capacity planning
- Familiarity with RAID levels and failure domain concepts
Lab Environment
Design exercise — validate storage policy concepts on Holodeck vSAN cluster if available
Tasks
Task 1 vSAN ESA Architecture & Capacity Planning
Design the vSAN ESA storage foundation with accurate capacity math showing raw-to-usable conversion under different protection schemes.
Document vSAN ESA vs OSA architecture differences (VCF 9.0):
vSAN ESA (Express Storage Architecture) — Default for VCF 9.0:
- Single storage tier: NVMe devices only (no cache/capacity split)
- Minimum: 1 NVMe TLC device per storage pool per host
- Native snapshots with improved performance (no CoW chain degradation)
- Improved compression/dedup efficiency (operates on original data, not cached data)
- Support for RAID-5/6 erasure coding from 4-host clusters (vs 5+ for OSA)
- Object write path: direct to storage pool, no caching tier
vSAN OSA (Original Storage Architecture):
- Two tiers: Cache (SSD/NVMe) + Capacity (SSD/HDD)
- Cache tier: 10% of capacity tier recommended
- Disk groups: max 5 per host, 1 cache + max 7 capacity per group
- More complex failure domain: disk group failure affects all capacity disks in group
Design decision: Select ESA for new VCF 9.0 deployments. OSA only for brownfield with existing HDD investments.
Calculate raw capacity for the cluster:
Assume: 8-host workload cluster, each host has 4 × 3.84 TB NVMe SSDs
Raw capacity: 8 hosts × 4 disks × 3.84 TB = 122.88 TB raw
Capacity deductions (in order):
1. vSAN metadata overhead: ~1-2% → 122.88 × 0.98 = 120.42 TB 2. Slack space (for rebalancing/maintenance): 25-30% reserve → 120.42 × 0.70 = 84.30 TB Note: 25% is minimum for maintenance mode evacuation; 30% recommended for production 3. Filesystem overhead: ~2% → 84.30 × 0.98 = 82.61 TB after overhead
Available for FTT protection: 82.61 TB
Compare protection schemes and usable capacity:
| Protection | Method | Min Hosts | Overhead | Usable from 82.61 TB | When to Use |
|---|---|---|---|---|---|
| FTT=1 Mirror | RAID-1 | 3 | 2× | 41.30 TB | Small clusters, high IOPS |
| FTT=1 RAID-5 | Erasure Coding | 4 | 1.33× | 62.03 TB | ≥4 hosts, read-heavy |
| FTT=2 Mirror | RAID-1 (triple) | 5 | 3× | 27.54 TB | Mission-critical, small |
| FTT=2 RAID-6 | Erasure Coding | 6 | 1.5× | 55.07 TB | ≥6 hosts, critical data |
Capacity gain from erasure coding vs mirroring:
- FTT=1: RAID-5 provides 50% more usable capacity than mirroring (62 vs 41 TB)
- FTT=2: RAID-6 provides 100% more than triple-mirror (55 vs 28 TB)
Performance trade-off:
- Erasure coding has higher CPU overhead for parity calculation
- Write amplification: ~1.33× for RAID-5 vs 2× for mirror, but parity is compute-intensive
- ESA mitigates this with improved code paths — erasure coding is production-viable on ESA
Map storage requirements to capacity:
Workload storage needs (from capacity model):
- Tier-1 VMs (80): avg 200 GB × 80 = 16 TB
- Tier-2 VMs (300): avg 150 GB × 300 = 45 TB
- Tier-3 VMs (220): avg 100 GB × 220 = 22 TB
- Total required: 83 TB usable
Fit analysis:
- FTT=1 Mirror: 41 TB usable → INSUFFICIENT (need 83 TB) - FTT=1 RAID-5: 62 TB usable → INSUFFICIENT
- Options: Add more disks, add hosts, or split protection by tier
Recommended: Tiered protection strategy:
- Tier-1 (16 TB): FTT=2 RAID-6 on Tier-1 cluster (separate domain, own capacity)
- Tier-2/3 (67 TB): FTT=1 RAID-5 on consolidated cluster
- This requires separate vSAN datastores per workload domain (VCF native architecture)
Validation Gate
Check: Capacity calculations show full raw-to-usable chain with FTT overhead for each protection scheme
Expected: At minimum: raw → metadata → slack → filesystem → FTT usable capacity, with comparison table of mirroring vs erasure coding
Common Errors
Task 2 Storage Policy Design & Datastore Layout
Create vSAN storage policies that align protection levels and performance characteristics with workload tiers.
Design storage policies for each workload tier:
Policy: Tier-1-Critical
- Failures to tolerate (FTT): 2
- Failure tolerance method: RAID-6 (erasure coding)
- Object space reservation (thick provisioning): 25% (for database workloads)
- IOPS limit: None (unlimited)
- Force provisioning: No
- Deduplication and compression: Enabled (ESA — compression-only for database if dedup ratio < 1.5:1)
- Encryption: Enabled (vSAN native encryption, AES-256)
- Checksum: Enabled (object-level CRC)
Policy: Tier-2-Business
- FTT: 1
- Failure tolerance method: RAID-5
- Object space reservation: 0% (thin provisioned)
- IOPS limit: 2,000 per VM (prevent noisy neighbor)
- Force provisioning: No
- Deduplication and compression: Enabled
- Encryption: Enabled
Policy: Tier-3-DevTest
- FTT: 1
- Failure tolerance method: RAID-5
- Object space reservation: 0%
- IOPS limit: 1,000 per VM
- Force provisioning: Yes (allow provisioning even if policy cannot be fully satisfied — acceptable for dev/test)
- Deduplication and compression: Enabled
- Encryption: Optional (depends on data classification)
Design IOPS and throughput controls:
vSAN QoS (IOPS limits):
- Purpose: Prevent noisy-neighbor effect where one VM starves others
- Tier-1: No limit (critical workloads get priority)
- Tier-2: 2,000 IOPS per VMDK (sufficient for most business apps)
- Tier-3: 1,000 IOPS per VMDK (dev/test can tolerate throttling)
Storage I/O Control (SIOC) — not available on vSAN (uses native QoS instead)
Document expected IOPS per host:
- ESA with NVMe: ~200K-400K IOPS per host (depending on device and workload)
- 8-host cluster: ~1.6M-3.2M IOPS total (theoretical)
- Overhead from RAID-5: ~15-20% reduction in write IOPS
- Target: maintain <5ms latency at P99 for Tier-1 workloads
Plan deduplication and compression:
ESA dedup/compression behavior:
- Operates inline (during write) — no post-process cycle needed
- ESA dedup is block-level (4K granularity)
- Compression: LZ4 (fast, moderate ratio) applied after dedup
- Cannot selectively enable/disable per VM — it is a cluster-wide setting in ESA
(Note: In ESA, dedup+compression is either ON or OFF for the entire cluster)
Expected ratios by workload type:
| Workload | Dedup Ratio | Compression Ratio | Combined |
|---|---|---|---|
| General server (Windows/Linux) | 1.2-1.5:1 | 1.5-2.0:1 | 2.0-3.0:1 |
| VDI (linked clones) | 2.0-4.0:1 | 1.5-2.0:1 | 3.0-8.0:1 |
| Database (unique data) | 1.0-1.2:1 | 1.5-2.5:1 | 1.5-3.0:1 |
| Dev/Test (golden images) | 2.0-3.0:1 | 1.5-2.0:1 | 3.0-6.0:1 |
Apply to capacity calculation:
- Conservative estimate: 2.0:1 combined ratio
- 62 TB usable (FTT=1 RAID-5) × 2.0 = 124 TB effective capacity
- This changes the fit analysis: now sufficient for 83 TB requirement with headroom
Design vSAN datastore layout:
VCF architecture: Each workload domain gets its own vSAN datastore automatically
- Management domain: vSAN datastore for management VMs (dedicated)
- Workload Domain 1 (Tier-1): vSAN datastore with Tier-1-Critical policy as default
- Workload Domain 2 (Tier-2/3): vSAN datastore with Tier-2-Business policy as default
Additional datastores:
- vSAN HCI Mesh (if needed): Share unused capacity across workload domains without vMotion
- Use case: Workload Domain 2 has excess capacity; Workload Domain 1 can mount it remotely
- Constraint: Remote datastore performance is lower (network hop); suitable for Tier-3 only
- NFS/VMFS datastores (if external storage exists):
- Use for backup targets, ISO libraries, templates
- Never mix vSAN + external storage for the same workload tier (inconsistent performance)
Document: Default storage policy per datastore, exceptions, and the policy change workflow.
Validation Gate
Check: Storage policies defined for each tier with FTT, encryption, QoS, and dedup settings; datastore layout documented
Expected: 3+ storage policies with distinct settings, IOPS limits for noisy-neighbor prevention, dedup/compression ratio estimates with capacity impact, datastore-per-domain layout
Common Errors
Task 3 vSAN Encryption & Key Management Design
Design end-to-end encryption for data at rest using vSAN native encryption with proper key management architecture.
Design key management hierarchy:
vSAN encryption uses a two-tier key model:
- Key Encryption Key (KEK): Stored in external KMS or vCenter Native Key Provider
- Encrypts/wraps the DEK
- Rotated annually (or per compliance policy)
- If KEK is lost: data is permanently inaccessible
- Data Encryption Key (DEK): Generated per host, stored locally (encrypted by KEK)
- Used for actual data encryption (AES-256-XTS)
- Re-wrapped when KEK rotates (data is NOT re-encrypted)
- Shallow re-key: New KEK wraps existing DEK (fast, no I/O)
- Deep re-key: New DEK generated, data re-encrypted (slow, I/O intensive)
Key Provider options:
A) vCenter Native Key Provider (NKP):
- Built into vCenter 7.0u2+ and 9.0
- No external KMS needed
- Keys backed up via vCenter backup
- Limitation: Single point of failure if vCenter backup is lost
B) External KMS (KMIP 1.1 compliant):
- Examples: HyTrust KeyControl, Thales CipherTrust, Entrust
- HA: Deploy KMS cluster (2+ nodes)
- Stronger compliance posture (separation of key management from infrastructure)
Design decision: External KMS for production (SOC 2 requirement), NKP for dev/test (cost savings).
Configure encryption layers:
Layer 1 — vSAN Data-at-Rest Encryption:
- Enabled at cluster level
- Encrypts all data on vSAN datastore (VM files, swap, snapshots, metadata)
- Performance impact: 5-10% overhead on write IOPS (AES-NI hardware acceleration)
- Verify: Hosts must have AES-NI instruction support (all modern Xeon/EPYC do)
Layer 2 — VM Encryption (vSphere-level, optional):
- Per-VM encryption using storage policy
- Encrypts VMDK at the vSphere layer before reaching vSAN
- Double encryption warning: Enabling both vSAN encryption AND VM encryption creates double encryption overhead with no security benefit
- Recommendation: Use vSAN encryption for all VMs; add VM encryption only for cross-vCenter vMotion scenarios (encrypts data in transit)
Layer 3 — vSAN in-transit encryption:
- Encrypts vSAN data replication traffic between hosts
- Options: Cluster-only (within cluster) or Preferred/Required
- Performance: Adds ~5% latency to vSAN replication
- Recommended: Enable for compliance environments
Design key rotation and lifecycle:
- Key rotation schedule:
- KEK rotation: Annually (shallow re-key, minimal impact)
- DEK rotation: Every 3 years or after suspected compromise (deep re-key, plan maintenance window)
- Post-incident rotation: Immediate shallow re-key after any security event
- Key backup and recovery:
- NKP: Keys included in vCenter file-level backup; backup to separate secure location
- External KMS: KMS cluster handles replication; backup KMS database independently
- Test recovery: Quarterly test of key recovery procedure (restore vCenter from backup, verify data accessibility)
- Failure scenarios:
- KMS unavailable during host boot: Host cannot mount vSAN datastore; VMs do not start
Mitigation: KMS cluster with 2+ nodes across failure domains
- vCenter down: Existing hosts continue running (keys cached in memory); new hosts cannot join
Mitigation: Ensure vCenter HA is configured; documented break-glass procedure for emergency vCenter recovery
- KEK lost permanently: Data is UNRECOVERABLE
Mitigation: Multiple KMS backup copies in geographically separate locations
Create encryption design decision entry:
Decision D-006 — Encryption Architecture:
Decision: vSAN data-at-rest encryption with external KMS for Tier-1/2, NKP for Tier-3
Alternatives:
A) VM-level encryption per policy — Rejected: more complex policy management, double encryption risk, vMotion encryption is separate anyway
B) No encryption, rely on physical security — Rejected: SOC 2 requires encryption at rest, drive disposal risk
C) Full-disk encryption (SED drives) — Rejected: vendor lock-in, limited key management flexibility, no integration with vSphere policy framework
Compliance mapping:
- SOC 2 CC6.1: Encryption of data at rest ✓
- SOC 2 CC6.7: Key management separation ✓ (external KMS)
- NIST 800-171: FIPS 140-2 validated algorithm (AES-256) ✓
Document the key management network design:
- KMS communication: KMIP port 5696 (TCP)
- Network: Dedicated management VLAN, firewalled
- vCenter → KMS: Mutual TLS authentication (client certificate)
Validation Gate
Check: Encryption design covers all three layers with key management hierarchy, rotation schedule, and failure scenarios
Expected: Two-tier key hierarchy documented, KMS HA architecture designed, rotation schedule defined, 3+ failure scenarios analyzed with mitigation, compliance mapping completed
Common Errors
Task 4 vSAN Health Monitoring & Maintenance Design
Design the ongoing operational framework for vSAN health monitoring, maintenance mode behavior, and capacity alerts.
Design vSAN Health monitoring integration:
- vSAN Health Service checks to monitor:
- Cluster health: Overall status, network partition detection
- Disk health: Disk balance, disk latency, congestion
- Data health: Object health, policy compliance, stale components
- Network health: vSAN VMkernel connectivity, multicast/unicast status
- Physical disk health: SMART data, wear indicators for NVMe
- Aria Operations integration:
- Enable vSAN management pack
- Custom dashboards: Capacity trending, IOPS distribution, resync progress
- Alerts: NVMe wear level >80% → Warning; >90% → Critical
- Capacity alert thresholds:
- 70% used → Warning (plan procurement)
- 80% used → Critical (restrict new VM provisioning)
- 85% used → Emergency (begin workload migration or expansion)- Skyline Health (Broadcom):
- Proactive support: Auto-detects known issues against KB articles
- Requires outbound HTTPS — for air-gapped: use Skyline Collector appliance
Design maintenance mode behavior:
Three evacuation modes:
- Ensure data accessibility (fastest):
- Moves data components to maintain accessibility but does not rebuild full protection
- Data is temporarily at reduced FTT (e.g., FTT=2 → FTT=1 during maintenance)
- Use for: Short maintenance (<4 hours), patching, firmware updates
- Full data migration (slowest, safest):
- Rebuilds all components on remaining hosts to maintain full FTT
- Requires sufficient free capacity on other hosts
- Use for: Host decommission, extended maintenance (>24 hours)
- Capacity requirement: Host's capacity × FTT multiplier must fit on remaining hosts
- No data migration (fastest, risk):
- Does not move any data; VMs are migrated but data components remain on offline host
- If another host fails during maintenance, data loss possible
- Use for: Quick reboot only (< 15 minutes)
Design decision: Default to 'Ensure accessibility' for patching; require change approval for 'No data migration'
Plan disk replacement and expansion procedures:
- Disk failure replacement:
- vSAN ESA auto-detects disk failure and begins component rebuild on remaining disks
- Rebuild time depends on data volume and cluster load (~100 GB/hour estimate)
- During rebuild: Performance degraded (~20-30% IOPS reduction cluster-wide)
- After physical replacement: vSAN reclaims new disk automatically (ESA — no disk group concept)
- Capacity expansion options:
a) Add disks to existing hosts:
- Hot-add NVMe supported on most server platforms
- vSAN discovers new disk, adds to storage pool
- Rebalance starts automatically (proactive rebalance in ESA)
b) Add hosts to cluster:
- VCF workflow: SDDC Manager → Add Host to Cluster
- New host contributes storage pool capacity + compute
- Minimum host add: 1 host at a time
- Post-add: Enable vSAN proactive rebalance for even distribution
c) HCI Mesh (share capacity across domains):
- Remote mount of vSAN datastore from another cluster
- Latency: +1-2ms over direct local access
- Use case: Short-term capacity relief while procurement completes
Create operational design documentation:
- Capacity planning cadence:
- Monthly: Review vSAN capacity usage trending in Aria Operations
- Quarterly: Procurement forecast based on growth rate
- Annual: Full capacity audit with lifecycle refresh planning
- Maintenance window design:
- Host maintenance: Serial, one host at a time, minimum 4 hours between hosts
- Reason: vSAN resync must complete before next host enters maintenance
- Monitor: vSAN resync dashboard — proceed only when resync objects = 0
- Sequence: Follow VCF upgrade order (KB 390634) — SDDC Manager first, then management domain, then workload domains
- Design decision entry:
Decision D-007 — Maintenance Mode Default:
Decision: 'Ensure data accessibility' as default evacuation mode
Alternative: 'Full data migration' — Rejected as default: too slow for routine patching (8+ hours per host in large clusters)
Alternative: 'No data migration' — Rejected as default: unacceptable data-loss risk during host reboot failure
Risk: If second host fails during maintenance with 'Ensure accessibility', data at reduced FTT
Mitigation: Health pre-check before maintenance; abort if any components already in degraded state
Validation Gate
Check: Operational design covers health monitoring, maintenance procedures, and capacity management
Expected: Alert thresholds defined, maintenance mode policy documented with alternatives, disk replacement/expansion procedures outlined, capacity planning cadence established
Common Errors
Final Validation
Complete vSAN storage design with capacity math, tiered policies, encryption architecture, and operational procedures
✓ Capacity calculations show raw-to-usable chain → Full math with metadata, slack, overhead, and FTT deductions documented
✓ Storage policies defined per tier → 3+ policies with FTT, QoS, dedup/compression, and encryption settings
✓ Encryption design includes key management → KEK/DEK hierarchy, KMS HA, rotation schedule, and failure recovery documented
✓ Operational procedures documented → Maintenance mode policy, disk replacement, capacity alerts, and monitoring integration
Cleanup / Restore
• Save all design documents and capacity models
• If using Holodeck: revert vSAN configuration to base snapshot
Design Reflection (VCDX)
Storage design separates good architects from great ones. VCDX panelists probe three areas relentlessly: (1) Show the capacity math — raw to usable, including slack space that most candidates forget. (2) Justify your protection scheme — why RAID-5 over mirroring? The answer must include host count, performance impact, and cost trade-off. (3) What happens when things break? Disk failure during maintenance, KMS outage, capacity exhaustion. The operational design is where VCDX candidates differentiate.
Requirements
- 83 TB usable storage across 3 tiers with different protection levels
- Encryption at rest for SOC 2 compliance (AES-256)
- IOPS limits to prevent noisy-neighbor effects between workload tiers
- Capacity for 3-year growth at 20% p.a.
Constraints
- vSAN ESA: minimum 1 NVMe TLC device per storage pool per host
- Deduplication/compression is cluster-wide in ESA (cannot per-VM)
- VCF workload domains have isolated vSAN datastores
- Per-core licensing affects host addition cost decisions
Assumptions
- NVMe devices are 3.84 TB TLC SSDs (HCL-certified)
- Conservative dedup/compression ratio of 2.0:1 for capacity planning
- Mixed workload I/O profile: 70% read, 30% write
- AES-NI hardware acceleration available on all hosts
Risks
- Actual dedup ratio lower than estimated — capacity runs out sooner
- NVMe SSD endurance reached before refresh cycle (high-write workloads)
- KMS failure renders all encrypted data inaccessible until recovery
- Capacity expansion procurement takes 8-12 weeks — insufficient lead time if alerts trigger late
Self-Assessment Discussion Prompts
- When would you choose FTT=2 Mirror over FTT=2 RAID-6, given the 3× vs 1.5× capacity overhead?
- If the customer refuses external KMS due to cost, what is the risk posture of using Native Key Provider for production?
- How would you redesign storage if the customer adds a VDI workload (500 desktops) to the existing cluster?
- What is the impact on resync time if you add a 9th host to the 8-host cluster — does it speed up or slow down existing operations?
Extensions
vSAN Stretched Cluster Storage Considerations
Extend the storage design for a stretched cluster spanning two sites. Address site affinity policies (keep replicas co-located with compute), witness storage requirements, and the capacity impact of cross-site FTT configurations. Calculate the WAN bandwidth needed for synchronous replication of vSAN write traffic.
vSAN HCI Mesh Cross-Cluster Storage Sharing
Design a vSAN HCI Mesh topology where a compute-heavy cluster mounts remote datastores from a storage-rich cluster. Document the performance trade-offs (additional network hop), the fault domain implications, and the use cases where HCI Mesh is preferable to adding local storage.
vSAN Performance Benchmarking with HCIBench
Use HCIBench (VMware's open-source vSAN benchmarking tool) on Holodeck to measure IOPS, latency, and throughput under different FTT and dedup/compression configurations. Compare ESA vs OSA performance profiles and document the results for inclusion in the VCDX design document.
⚠ Known Pitfalls (from Community KB)
References
- vSAN 9.0 Planning and Deployment Guide — ESA Architecture chapter
- vSAN 9.0 Design Guide — Capacity Planning and Sizing
- VMware KB 2150753 — vSAN Capacity Planning considerations
- vSAN Encryption — Native Key Provider vs External KMS comparison
- vSAN Health Service documentation — Alert thresholds and remediation