Academy/VVS/Private AI Ready Infrastructure
This solution targets VCF 5.2

Private AI Ready Infrastructure for VMware Cloud Foundation

VCF 5.2architectvcdxautomationPages 170-236

Design, implementation, and operational guidance for AI GPU-enabled accelerated workload domains running on vSphere with Tanzu. Extends DRI with GPU passthrough/vGPU (NVIDIA A100/L40S/H100), SR-IOV, GPUDirect RDMA, NVLink/NVSwitch support, NVIDIA AI Enterprise (NVAIE) licensing (CLS/DLS), VMware Data Services Manager for database services (PostgreSQL with pgvector, MySQL), and VMware Private AI Foundation with NVIDIA add-on. Supports deep learning VMs, Tanzu Kubernetes Grid clusters running AI containers, and RAG workloads.

Key Components: NSX, SDDC Manager, vCenter, ESXi, Tanzu

External dependencies: VMware Data Services Manager 2.0.2 (part of VMware Private AI Foundation with NVIDIA)

Design Decisions
Implementation
Operations
VCDX Defense
Quiz (15)
Flashcards (15)

7 design decisions

DD-IDDecisionQuality
AIR-TZU-NET-001Set up networking for 100 Gbps or higher if possible.Performance

Decision: Set up networking for 100 Gbps or higher if possible.

Rationale: Enough bandwidth and very low latency for inference and fine-tuning with vSAN ESA.

Implication: Cost increased.

Component: TZU

AIR-TZU-NET-002Add /28 overlay-backed NSX segment for Supervisor control plane nodes.Manageability

Decision: Add /28 overlay-backed NSX segment for Supervisor control plane nodes.

Rationale: Supports Supervisor control plane.

Implication: Must create overlay segment.

Component: TZU

AIR-TZU-NET-003Use /20 subnet for pod networking.Meets 2000 pods design requirement.Private IP behind NAT; reusableRecoverability

Decision: Use /20 subnet for pod networking.Meets 2000 pods design requirement.Private IP behind NAT; reusable across Supervisors.

Rationale: AIR-TZU-NET-004Use /22 subnet for services.Meets 2000 pods design requirement.Private IP behind NAT.

Implication: AIR-TZU-NET-005/24+ subnet for ingress endpoints on corporate network.Sufficient in most cases.Must be routable; evaluate ingress needs.

Component: TZU

AIR-TZU-SEC-003JustificationImplicationSecurity

Decision: JustificationImplication

Rationale: See source document for rationale

Implication: AIR-TZU-SEC-001AD security groups for DevOps admins with Can Edit on namespaces.Auditable RBAC.AD group mgmt.

Component: TZU

AIR-TZU-SEC-004JustificationImplicationSecurity

Decision: JustificationImplication

Rationale: See source document for rationale

Implication: AIR-TZU-SEC-001AD security groups for DevOps admins with Can Edit on namespaces.Auditable RBAC.AD group mgmt.

Component: TZU

AIR-NVD-LIC-002Internet access required for CLS (ports 80, 443 egress).ManageabilitySecurity

Decision: Internet access required for CLS (ports 80, 443 egress).

Rationale: See source document for rationale

Implication: Security risks; enforce FW/IDS/monitoring.

Component: NVD

AIR-NVD-LIC-001Account for extra compute/storage/network for DLS in management domain.Manageability

Decision: Account for extra compute/storage/network for DLS in management domain.

Rationale: See source document for rationale

Implication: Increased mgmt resources; LCM required.

Component: NVD

Prerequisites

  • VCF version in Support Matrix (5.1.0-5.2.1)
  • DRI prerequisites (vSphere with Tanzu activated)
  • SR-IOV enabled globally at BIOS/UEFI on all hosts with GPUs
  • GPU-enabled ESXi hosts (vLCM-managed) with compatible NVIDIA GPU (A100, L40S, H100)
  • Supported GPU sharing mode: Time slicing or MIG
  • NSX Edge cluster large-sized
  • AD/DNS/NTP/CA external services (as per IAM + DRI)
  • Workspace ONE Access optional for Service Broker catalog
  • VMware Data Services Manager license (part of VMware Private AI Foundation with NVIDIA)
  • NVIDIA AI Enterprise license (NVAIE) — per-GPU
  • Tanzu Network account for DSM refresh token
  • Pvf Components
  • NameDesc
  • NVIDIA GPU deviceNVIDIA A100, L40S, H100
  • Supported GPU sharingTime slicing or Multi-Instance GPU (MIG)
  • NVIDIA vGPU driverInstalled on ESXi via VIB
  • NVIDIA AI Enterprise suiteAI software (NGC catalog access)
  • Implementation Steps
  • Configure vSphere for AI GPU-enabled workloads (Global SR-IOV, NVIDIA vGPU VIB)
  • Activate Supervisor for AI-ready workloads (per DRI procedures)

Implementation Procedure

Implementation

Triton Inference ServerCloud + edge inferencing; HTTP/REST and GRPC for remote clients; shared library C API for edge
Generative AI Workflow - RAGReference solution for RAG-augmenting LLM with enterprise knowledge base

Day-2 Operations Tasks

Operations

As needed

Personas

As needed

Persona: Cloud Administrator

As needed

PersonaCloud Administrator

As needed

ResponsibilityFull admin access to vSphere with Tanzu — configure/activate Supervisors and namespaces

As needed

Mapping

As needed

vCenterAdministrator

As needed

SDDC ManagerAdministrator

As needed

Persona: DevOps Engineer

As needed

PersonaDevOps Engineer

As needed

Monitoring Points

  • Configure: PostgreSQL 15.5+vmware.v2.0.0; Replica Mode = Single vSphere Cluster; Topology = 3 (1 Primary, 1 Replica, 1 Monitor)
  • Verify Docker Hub registry access (pull pgvector containers)
  • Verify NVIDIA NGC registry access (pull PyTorch containers)
  • Per VCF standard procedures; additional considerations for DSM (drain DBs gracefully) and deep learning VMs (checkpoint workloads).
  • MonitoringVCF Operations 8.16.1+ monitors Supervisor, TKG clusters, Deep Learning VMs. vSphere Client custom charts show GPU metrics at cluster level.

Likely Panelist Questions

Q: Why did you choose this architecture?

See design decisions for rationale

Failure Scenarios

Not enabling Global SR-IOV at BIOS — vGPU/MIG fail to enumerate
Impact:
Mitigation:
Heterogeneous GPU mix within a cluster — vMotion and placement fail
Impact:
Mitigation:
Under-provisioned NSX Edge (not large) — Supervisor enable fails
Impact:
Mitigation:
Not reserving sufficient IPs in infrastructure policy pools — DSM DB deployments fail (see AIR-DSM-003)
Impact:
Mitigation:
DLS license appliance failure
Impact:
Mitigation:

Trade-off Analysis

Trade-Offs Analysis

Chosen:

Justification:

Running DSM in a workload domain instead of management domain — violates AIR-DSM-001 constraint

Chosen:

Justification:

Quiz — Private AI Ready Infrastructure

0/15
Q1
Which GPU sharing modes are supported in this solution?
  • Time slicing only
  • MIG only
  • Time slicing and Multi-Instance GPU (MIG)
  • Fractional vMotion
Supported GPU sharing modes: Time slicing and MIG (Multi-Instance GPU).
Q2
What must be enabled in BIOS/UEFI on every GPU host before vGPU/MIG works?
  • Hyper-Threading
  • Global SR-IOV
  • TPM 2.0
  • Secure Boot
Global SR-IOV is required; for MIG, a vGPU is associated with a virtual function at boot time.
Q3
What is the minimum subnet size designed for pod networking in PAI?
  • /24
  • /22
  • /20
  • /16
AIR-TZU-NET-003: /20 subnet for pod networking (supports 2000 pods).
Q4
Which NVIDIA license system deployment is hosted on the NVIDIA Licensing Portal?
  • CLS (Cloud License Service)
  • DLS (Delegated License Service)
  • NVL (NVIDIA Virtual License)
  • NGC (NVIDIA GPU Cloud)
CLS is hosted on the NVIDIA Licensing Portal; DLS is on-prem.
Q5
What is the HA topology required for a production PostgreSQL in DSM per this design?
  • Single node
  • 2-node primary/standby
  • 3 or 5 nodes
  • Unlimited
AIR-DSM-002: deploy PostgreSQL in HA mode (3 or 5 nodes) for production.
Q6
How many IP addresses does a 5-node PostgreSQL DSM cluster require?
  • 5
  • 6
  • 7
  • 10
5 node IPs + 1 kube_VIP + 1 DB LB = 7 IPs (per AIR-DSM-003 example).
Q7
Which GPU architecture example is recommended for middle-range (7-13B parameter) LLMs?
  • NVIDIA H100 (Hopper)
  • NVIDIA A100 (Ampere)
  • NVIDIA RTX 4090
  • NVIDIA V100 (Volta)
A100 (40/80 GB HBM2e) is suited for middle-range LLMs (7-13B) and embedding models per design.
Q8
Which DSM decision requires internet access for package downloads?
  • AIR-DSM-001 (deploy in mgmt)
  • AIR-DSM-006 (S3 with TLS)
  • AIR-DSM-007 (Tanzu Network refresh token)
  • AIR-DSM-003 (IP pool)
AIR-DSM-007: Tanzu Network refresh token required in connected env; air-gapped env requires manual upload of repository.
Q9
What technology combines high-speed GPU-to-GPU communication suitable for generative AI scaling?
  • PCIe only
  • GPUDirect RDMA + NVLink + NVSwitch
  • SR-IOV + vMotion
  • VMXNET3
GPUDirect RDMA (direct GPU-to-GPU via network), NVLink (GPU-GPU bridge), NVSwitch (multi-GPU orchestration).
Q10
Which vSphere feature groups complementary hardware devices connected via PCIe switch/interconnect?
  • VM classes
  • Device groups (vSphere 8)
  • Resource pools
  • DRS VM groups
Device groups in vSphere 8 identify hardware-connected devices as unified entity; recognized by DRS and HA; supported with TKG Service.
Q11
What filesystem/protocol supports ReadWriteMany persistent volumes for TKG model repositories?
  • NFS v3
  • vSAN File Service (via CNS file volumes)
  • iSCSI
  • HDFS
vSphere with Tanzu uses CNS file volumes backed by vSAN File Service for RWX persistent volumes (model repos, archives).
Q12
Limit for file shares per vSAN cluster for DSM/RWX volumes?
  • 10
  • 50
  • 100
  • 1000
Maximum of 100 file shares per vSAN cluster; max file share size = max available vSAN capacity.
Q13
Which feature of vSAN ESA reduces network traffic for AI data transfers?
  • Deduplication
  • Advanced compression at top of stack (enabled by default), data transmitted compressed across network
  • Erasure coding
  • Replication
vSAN ESA advanced compression at top of stack = data compressed before network transfer, optimizing bandwidth and storage.
Q14
Which precision is recommended for LLM inference deployment to optimize resources?
  • FP64
  • FP32 only
  • FP16 / BF16 / Int8
  • FP4 only
Convert from FP32 (training) to FP16/BF16/Int8 (inference) for model size/compute reduction without significant accuracy loss.
Q15
Which persona uses VMware Aria Automation Service Broker with Can Edit for AI workload deployment?
  • Cloud Administrator
  • DevOps Engineer
  • Data Scientist
  • Auditor
Data Scientist persona: Can Edit on Service Broker, TKC, vSphere Namespace — deploys VMs via VM Service and Kubernetes API.

Flashcards — Private AI Ready Infrastructure

Card 1 of 15
Supported NVIDIA GPUs (PAI)
NVIDIA A100, L40S, H100 (Ampere, Ada Lovelace, Hopper architectures).

Labs

Lab 1: Enable GPU Passthrough and Deploy a Deep Learning VM

Configure Global SR-IOV on ESXi host, install NVIDIA vGPU VIB, create vGPU-enabled VM class, and deploy a deep learning VM with PyTorch.

Starting State: VCF 5.2 with VI workload domain; host has NVIDIA A100/L40S/H100 installed; NVAIE license available; Tanzu content library subscribed; Supervisor activated.

Lab 2: Deploy PostgreSQL with pgvector via Data Services Manager for RAG

Configure DSM infrastructure policy, deploy 3-node HA PostgreSQL with pgvector extension, verify connectivity.

Starting State: DSM appliance deployed in management domain; Tanzu Network refresh token configured; S3-compatible bucket with TLS (MinIO) provisioned; LDAP integration complete; Storage Policy and VM Classes defined.

Lab 3: Deploy a RAG Workload in Jupyter Notebook Using pgvector

Deploy a Jupyter Notebook that ingests text documents into a pgvector-backed knowledge base (NASA history books) and queries an LLM via RAG.

Starting State: Deep learning VM with PyTorch container from Lab 1; pgvector DB from Lab 2; Docker Hub + NVIDIA NGC registry access.

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.