Private AI Ready Infrastructure for VMware Cloud Foundation
Design, implementation, and operational guidance for AI GPU-enabled accelerated workload domains running on vSphere with Tanzu. Extends DRI with GPU passthrough/vGPU (NVIDIA A100/L40S/H100), SR-IOV, GPUDirect RDMA, NVLink/NVSwitch support, NVIDIA AI Enterprise (NVAIE) licensing (CLS/DLS), VMware Data Services Manager for database services (PostgreSQL with pgvector, MySQL), and VMware Private AI Foundation with NVIDIA add-on. Supports deep learning VMs, Tanzu Kubernetes Grid clusters running AI containers, and RAG workloads.
Key Components: NSX, SDDC Manager, vCenter, ESXi, Tanzu
External dependencies: VMware Data Services Manager 2.0.2 (part of VMware Private AI Foundation with NVIDIA)
7 design decisions
| DD-ID | Decision | Quality |
|---|---|---|
| AIR-TZU-NET-001 | Set up networking for 100 Gbps or higher if possible. | Performance |
Decision: Set up networking for 100 Gbps or higher if possible. Rationale: Enough bandwidth and very low latency for inference and fine-tuning with vSAN ESA. Implication: Cost increased. Component: TZU | ||
| AIR-TZU-NET-002 | Add /28 overlay-backed NSX segment for Supervisor control plane nodes. | Manageability |
Decision: Add /28 overlay-backed NSX segment for Supervisor control plane nodes. Rationale: Supports Supervisor control plane. Implication: Must create overlay segment. Component: TZU | ||
| AIR-TZU-NET-003 | Use /20 subnet for pod networking.Meets 2000 pods design requirement.Private IP behind NAT; reusable | Recoverability |
Decision: Use /20 subnet for pod networking.Meets 2000 pods design requirement.Private IP behind NAT; reusable across Supervisors. Rationale: AIR-TZU-NET-004Use /22 subnet for services.Meets 2000 pods design requirement.Private IP behind NAT. Implication: AIR-TZU-NET-005/24+ subnet for ingress endpoints on corporate network.Sufficient in most cases.Must be routable; evaluate ingress needs. Component: TZU | ||
| AIR-TZU-SEC-003 | JustificationImplication | Security |
Decision: JustificationImplication Rationale: See source document for rationale Implication: AIR-TZU-SEC-001AD security groups for DevOps admins with Can Edit on namespaces.Auditable RBAC.AD group mgmt. Component: TZU | ||
| AIR-TZU-SEC-004 | JustificationImplication | Security |
Decision: JustificationImplication Rationale: See source document for rationale Implication: AIR-TZU-SEC-001AD security groups for DevOps admins with Can Edit on namespaces.Auditable RBAC.AD group mgmt. Component: TZU | ||
| AIR-NVD-LIC-002 | Internet access required for CLS (ports 80, 443 egress). | ManageabilitySecurity |
Decision: Internet access required for CLS (ports 80, 443 egress). Rationale: See source document for rationale Implication: Security risks; enforce FW/IDS/monitoring. Component: NVD | ||
| AIR-NVD-LIC-001 | Account for extra compute/storage/network for DLS in management domain. | Manageability |
Decision: Account for extra compute/storage/network for DLS in management domain. Rationale: See source document for rationale Implication: Increased mgmt resources; LCM required. Component: NVD | ||
Prerequisites
- VCF version in Support Matrix (5.1.0-5.2.1)
- DRI prerequisites (vSphere with Tanzu activated)
- SR-IOV enabled globally at BIOS/UEFI on all hosts with GPUs
- GPU-enabled ESXi hosts (vLCM-managed) with compatible NVIDIA GPU (A100, L40S, H100)
- Supported GPU sharing mode: Time slicing or MIG
- NSX Edge cluster large-sized
- AD/DNS/NTP/CA external services (as per IAM + DRI)
- Workspace ONE Access optional for Service Broker catalog
- VMware Data Services Manager license (part of VMware Private AI Foundation with NVIDIA)
- NVIDIA AI Enterprise license (NVAIE) — per-GPU
- Tanzu Network account for DSM refresh token
- Pvf Components
- NameDesc
- NVIDIA GPU deviceNVIDIA A100, L40S, H100
- Supported GPU sharingTime slicing or Multi-Instance GPU (MIG)
- NVIDIA vGPU driverInstalled on ESXi via VIB
- NVIDIA AI Enterprise suiteAI software (NGC catalog access)
- Implementation Steps
- Configure vSphere for AI GPU-enabled workloads (Global SR-IOV, NVIDIA vGPU VIB)
- Activate Supervisor for AI-ready workloads (per DRI procedures)
Implementation Procedure
Implementation
Triton Inference ServerCloud + edge inferencing; HTTP/REST and GRPC for remote clients; shared library C API for edge
Generative AI Workflow - RAGReference solution for RAG-augmenting LLM with enterprise knowledge base
Day-2 Operations Tasks
Operations
As neededPersonas
As neededPersona: Cloud Administrator
As neededPersonaCloud Administrator
As neededResponsibilityFull admin access to vSphere with Tanzu — configure/activate Supervisors and namespaces
As neededMapping
As neededvCenterAdministrator
As neededSDDC ManagerAdministrator
As neededPersona: DevOps Engineer
As neededPersonaDevOps Engineer
As neededMonitoring Points
- Configure: PostgreSQL 15.5+vmware.v2.0.0; Replica Mode = Single vSphere Cluster; Topology = 3 (1 Primary, 1 Replica, 1 Monitor)
- Verify Docker Hub registry access (pull pgvector containers)
- Verify NVIDIA NGC registry access (pull PyTorch containers)
- Per VCF standard procedures; additional considerations for DSM (drain DBs gracefully) and deep learning VMs (checkpoint workloads).
- MonitoringVCF Operations 8.16.1+ monitors Supervisor, TKG clusters, Deep Learning VMs. vSphere Client custom charts show GPU metrics at cluster level.
Likely Panelist Questions
Q: Why did you choose this architecture?
See design decisions for rationale
Failure Scenarios
Trade-off Analysis
Trade-Offs Analysis
Chosen:
Justification:
Running DSM in a workload domain instead of management domain — violates AIR-DSM-001 constraint
Chosen:
Justification:
Quiz — Private AI Ready Infrastructure
- Time slicing only
- MIG only
- Time slicing and Multi-Instance GPU (MIG)
- Fractional vMotion
- Hyper-Threading
- Global SR-IOV
- TPM 2.0
- Secure Boot
- /24
- /22
- /20
- /16
- CLS (Cloud License Service)
- DLS (Delegated License Service)
- NVL (NVIDIA Virtual License)
- NGC (NVIDIA GPU Cloud)
- Single node
- 2-node primary/standby
- 3 or 5 nodes
- Unlimited
- 5
- 6
- 7
- 10
- NVIDIA H100 (Hopper)
- NVIDIA A100 (Ampere)
- NVIDIA RTX 4090
- NVIDIA V100 (Volta)
- AIR-DSM-001 (deploy in mgmt)
- AIR-DSM-006 (S3 with TLS)
- AIR-DSM-007 (Tanzu Network refresh token)
- AIR-DSM-003 (IP pool)
- PCIe only
- GPUDirect RDMA + NVLink + NVSwitch
- SR-IOV + vMotion
- VMXNET3
- VM classes
- Device groups (vSphere 8)
- Resource pools
- DRS VM groups
- NFS v3
- vSAN File Service (via CNS file volumes)
- iSCSI
- HDFS
- 10
- 50
- 100
- 1000
- Deduplication
- Advanced compression at top of stack (enabled by default), data transmitted compressed across network
- Erasure coding
- Replication
- FP64
- FP32 only
- FP16 / BF16 / Int8
- FP4 only
- Cloud Administrator
- DevOps Engineer
- Data Scientist
- Auditor
Flashcards — Private AI Ready Infrastructure
Labs
Lab 1: Enable GPU Passthrough and Deploy a Deep Learning VM
Configure Global SR-IOV on ESXi host, install NVIDIA vGPU VIB, create vGPU-enabled VM class, and deploy a deep learning VM with PyTorch.
Starting State: VCF 5.2 with VI workload domain; host has NVIDIA A100/L40S/H100 installed; NVAIE license available; Tanzu content library subscribed; Supervisor activated.
Lab 2: Deploy PostgreSQL with pgvector via Data Services Manager for RAG
Configure DSM infrastructure policy, deploy 3-node HA PostgreSQL with pgvector extension, verify connectivity.
Starting State: DSM appliance deployed in management domain; Tanzu Network refresh token configured; S3-compatible bucket with TLS (MinIO) provisioned; LDAP integration complete; Storage Policy and VM Classes defined.
Lab 3: Deploy a RAG Workload in Jupyter Notebook Using pgvector
Deploy a Jupyter Notebook that ingests text documents into a pgvector-backed knowledge base (NASA history books) and queries an LLM via RAG.
Starting State: Deep learning VM with PyTorch container from Lab 1; pgvector DB from Lab 2; Docker Hub + NVIDIA NGC registry access.