Academy/NSX 4.x Network Virtualization Professional (2V0-41.24)/Lab N1: Deploy DPDK-Accelerated NSX Edge Cluster in VCF 9.0
This lab targets VCF 9.0

Lab N1: Deploy DPDK-Accelerated NSX Edge Cluster in VCF 9.0

VCF 9.0Advancedvcp-foundation⏱ 120 min

NSX 9.0.x (feature set inherited from the NSX 4.2 line; NSX 4.2 docs remain a valid technical reference) Edge cluster deployment — Enhanced Data Path (DPDK fastpath), Edge sizing, N-S throughput validation

Objectives

  • Understand NSX Edge node architecture and DPDK Enhanced Data Path
  • Deploy and configure an Edge cluster suitable for high-throughput N-S traffic
  • Configure Edge transport nodes with proper CPU reservation and NUMA alignment
  • Validate Edge cluster health and throughput performance
  • Understand Edge sizing tiers and their impact on supported services

Prerequisites

VCF 9.0 management domain operational, dedicated ESXi hosts for Edge VMs (or shared hosts with sufficient resources), VLAN uplink segment configured for Edge external connectivity, overlay transport zone operational

Prior labs: net-virt-01

Required skills:

  • NSX Edge architecture concepts
  • DPDK/SR-IOV fundamentals
  • vSphere resource management (CPU reservation, NUMA)

Lab Environment

VCF 9.0 with 2+ ESXi hosts designated for Edge workloads. Each host should have sufficient resources for Edge VMs: Large form factor = 8 vCPU, 32 GB RAM per Edge VM. Physical NICs connected to upstream switches for Edge uplink VLANs.

Tasks

Task 1 Deploy and Validate DPDK-Accelerated Edge Cluster

Edge sizing in VCF 9.0 follows predefined form factors: Small (2 vCPU/4GB — lab only), Medium (4 vCPU/8GB — small deployments), Large (8 vCPU/32GB — production standard), Extra Large (16+ vCPU/64GB — high throughput). DPDK fastpath bypasses the kernel network stack, dedicating CPU cores to packet processing via poll-mode drivers. In VCF deployments, Edge nodes are deployed automatically by SDDC Manager/VCF Operations — this lab walks through the manual process for understanding.

Deploy NSX Edge nodes with the Enhanced Data Path (DPDK fastpath) for maximum north-south throughput, configure proper resource reservations, form an Edge cluster, and validate performance under load. Understanding Edge architecture is critical for VCDX — Edge sizing directly impacts north-south bandwidth, stateful service capacity, and failure domain design.

Step 1

Plan Edge host resources. NSX Edge VMs require: (a) CPU reservation — 100% for all vCPUs (DPDK poll-mode drivers consume CPU continuously even at idle); (b) Memory reservation — 100% (no overcommit); (c) NUMA alignment — all vCPUs and memory on a single NUMA node for optimal DPDK performance; (d) Latency sensitivity — set to 'High' in VM Advanced settings. For a 4-Edge cluster on 2 hosts (anti-affinity), each host needs: 2 × 8 vCPU = 16 cores + 2 × 32 GB = 64 GB reserved.

Step 2
Prepare Edge transport node profiles. NSX Manager → System → Fabric → Profiles → Transport Node Profiles. Verify (or create) an Edge transport node profile with: (a) Overlay transport zone membership; (b) VLAN transport zone for uplinks; (c) TEP IP pool assignment; (d) Uplink profile with at least 2 uplinks for redundancy (active-standby or LACP depending on physical topology).
Step 3
Deploy Edge nodes via NSX Manager. System → Fabric → Edge Transport Nodes → Add Edge VM. Configure: Name='edge-node-01', Form Factor=Large, Host=esxi-edge-01 (place on dedicated Edge host), Datastore=select shared or local datastore, Management Network=management port group. Configure uplinks: Edge uplink 1 → VLAN uplink segment (toward physical router), Edge overlay TEP → overlay TEP segment. Repeat for edge-node-02 on esxi-edge-02 (anti-affinity).
Step 4
Verify Edge node deployment. Wait for deployment to complete (5-10 minutes per Edge). In NSX Manager → System → Fabric → Edge Transport Nodes, verify: Configuration State='Success', Node Status='Up', Manager Connectivity='Up', Controller Connectivity='Up'. If any show 'Failed', check: (a) Edge VM console for boot errors, (b) management network connectivity from Edge VM, (c) DNS resolution from Edge (required for Manager/Controller communication).
Step 5
Form Edge cluster. System → Fabric → Edge Clusters → Add Edge Cluster. Name='edge-cluster-prod', Member Type=Edge Node, add edge-node-01 and edge-node-02 (and edge-node-03/04 for production). For DPDK high-availability: set the Edge cluster profile with allocation rules — 'AllocationBasedOnFailureDomain' distributes T0 service instances across failure domains.
Step 6

Verify DPDK fastpath. SSH to edge-node-01: get datapath. Output should show 'enhanced' (DPDK) vs 'standard' (kernel). Run: get datapath cpu-stats — shows per-core packet processing statistics. DPDK cores (typically cores 1-3 for Large form factor) should show high packet counts even at idle (poll mode). Core 0 is reserved for management/control plane.

Step 7

Configure anti-affinity rules. In vSphere Client, create a DRS anti-affinity rule: Name='Edge-Anti-Affinity', Members=edge-node-01, edge-node-02 (and 03/04). This ensures Edge VMs never co-locate on the same host, maintaining N+1 availability. In VCF, SDDC Manager creates these rules automatically during Edge deployment.

Step 8

Performance baseline. Attach iperf3 server on the external network (beyond the physical router). From a VM on the overlay: iperf3 -c <external-iperf-server> -P 8 -t 30 (8 parallel flows, 30 seconds). Record throughput — Large Edge with DPDK should achieve 10+ Gbps for single T0 in Active-Standby, or 20+ Gbps in Active-Active (ECMP across 2 Edge nodes). Monitor Edge CPU during the test: get datapath cpu-stats — DPDK cores should show proportional increase.

Step 9
Edge failure simulation. Gracefully shut down edge-node-01 VM. Monitor: (a) T0 BGP neighbor state — should transition from Established to the backup Edge node within BFD timeout (sub-second with BFD, 12-180s without); (b) iperf3 traffic — brief interruption followed by recovery on edge-node-02; (c) NSX Manager → Networking → Tier-0 → edge-node-01 shows 'Down', edge-node-02 shows 'Active'. Power edge-node-01 back on and verify it rejoins the cluster.
Step 10
Review Edge cluster capacity. NSX Manager → System → Fabric → Edge Clusters → edge-cluster-prod → Overview shows: Member count, Services hosted (T0, T1 with services, load balancers), Deployment status. For capacity planning: each Edge node can host a limited number of T1 service routers — Large form factor supports up to 400 T1 service instances. Monitor: get edge-cluster status from any Edge node CLI.

Validation Gate

Check: Edge cluster operational with DPDK fastpath and validated throughput

Expected: Edge cluster formed with 2+ nodes, DPDK Enhanced Data Path confirmed, anti-affinity rules enforced, baseline throughput measured (10+ Gbps per Edge node), Edge failure/recovery validated with sub-second failover (BFD-enabled)

Common Errors

Edge VM deployment fails with 'Insufficient resources'
Fix: Edge VMs require 100% CPU and memory reservation. Verify the target host has unreserved capacity. For Large form factor: 8 vCPU + 32 GB free on a single NUMA node. Check: vSphere Client → Host → Monitor → Hardware → NUMA Topology.
Edge shows 'Controller Connectivity: Down'
Fix: Edge VM cannot reach NSX Controllers. Check: (1) Edge management IP is correct and on the right VLAN, (2) DNS resolves NSX Manager FQDN from Edge, (3) no firewall blocking port 1235 (Edge-to-Controller). SSH to Edge and run: get managers — should show connected status.
DPDK shows 'standard' instead of 'enhanced' datapath
Fix: The Edge VM may not have been deployed with the DPDK-capable form factor, or the ESXi host NIC doesn't support the required features. Verify: (1) Edge was deployed as Large or XLarge (Small doesn't support DPDK), (2) ESXi host NICs are on the HCL for enhanced datapath.
iperf3 throughput significantly lower than expected
Fix: Check: (1) Edge uplink speed (10GbE vs 25GbE physical NIC), (2) iperf3 using enough parallel flows (-P 8 minimum for ECMP), (3) NUMA misalignment causing cross-NUMA memory access penalties, (4) physical switch port speed/duplex settings.

Final Validation

NSX Edge cluster deployed with DPDK-accelerated data path and validated performance

✓ Edge node health → All Edge nodes: Configuration=Success, Status=Up, Manager+Controller=Connected

✓ DPDK datapath → get datapath returns 'enhanced' on all Edge nodes

✓ Edge cluster formation → Cluster shows correct member count with allocation rules configured

✓ Anti-affinity rules → DRS rule prevents Edge VMs on same host

✓ Throughput baseline → ≥10 Gbps per Edge node with Large form factor

✓ Failover validation → Traffic recovers within BFD timeout after Edge node failure

Cleanup / Restore

• Remove T0/T1 gateways from Edge cluster (if created in this lab)

• Delete Edge cluster (System → Fabric → Edge Clusters → Delete)

• Delete Edge transport nodes (System → Fabric → Edge Transport Nodes → Delete — this powers off and removes Edge VMs)

• Remove anti-affinity DRS rules from vSphere

Design Reflection (VCDX)

Edge cluster design is a high-value VCDX topic. Key decisions: (1) Form factor — Large is the VCF standard; Small/Medium are for lab/dev only. (2) Placement — dedicated Edge hosts vs shared, NUMA alignment, anti-affinity rules. (3) Cluster sizing — N+1 minimum; N+2 for environments requiring zero-downtime maintenance. (4) DPDK vs standard — DPDK is mandatory for any serious throughput requirement. (5) Edge-per-rack vs centralized — affects failure domain and latency. (6) In VCF 9.0, Edge deployment is automated via VCF Operations — understanding the manual process helps troubleshoot automation failures.

Requirements

  • 10+ Gbps north-south throughput per T0 gateway
  • Sub-second failover for stateful services
  • Support for 400+ T1 service instances per Edge cluster

Constraints

  • Large Edge VM requires 8 vCPU + 32 GB with 100% reservation per Edge node
  • DPDK cores are consumed even at idle — cannot overcommit Edge host CPU
  • Maximum 10 Edge nodes per cluster (practical limit based on control plane scalability)

Assumptions

  • Dedicated hosts for Edge VMs (no workload contention)
  • Physical uplinks are 10GbE minimum (25GbE recommended for DPDK throughput)
  • BFD is enabled for all BGP neighbors to achieve sub-second failover

Risks

  • Undersized Edge form factor causes packet drops under load — difficult to diagnose without DPDK CPU stats
  • Single Edge cluster for all T0/T1 creates a blast radius — consider per-domain Edge clusters for isolation
  • NUMA misalignment silently degrades throughput by 30-40% — always verify NUMA placement post-deployment

Self-Assessment Discussion Prompts

  1. How would you design Edge placement for a 3-site VCF deployment with stretched clusters?
  2. What is the impact of running Edge VMs on the same hosts as production workloads?
  3. How do you calculate the required number of Edge nodes for a given north-south bandwidth requirement?
  4. When would you choose Active-Standby over Active-Active for a Tier-0 gateway, and how does that affect Edge sizing?

Extensions

Deploy Extra Large Edge nodes and benchmark throughput improvement vs Large

Configure multi-TEP on Edge nodes and verify ECMP across multiple uplinks

Test Edge maintenance mode — migrate services to partner Edge and validate zero-downtime

Configure Edge node syslog forwarding and SNMP monitoring for production observability

⚠ Known Pitfalls (from Community KB)

Deploying Small form factor Edge for production — Small lacks DPDK support and is limited to ~2 Gbps throughput
Ignoring NUMA alignment — DPDK performance degrades 30-40% with cross-NUMA memory access; verify with esxtop (NUMA home node column)
Placing all Edge nodes on the same host — a single host failure takes down the entire Edge cluster; anti-affinity is mandatory
Not reserving 100% CPU/memory for Edge VMs — contention from other VMs causes packet drops in the DPDK poll-mode driver

References

  • NSX 4.2 Installation Guide — Edge Node Deployment: techdocs.broadcom.com
  • VCF 9.0 Planning and Preparation Guide — Edge Cluster Sizing: techdocs.broadcom.com
  • KB 92345 — Troubleshooting NSX Edge Node Connectivity Issues
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.