Academy/VCAP — VCF VKS (3V0-24.25)/NSX Networking Validation
This lab targets VCF 9.0

NSX Networking Validation

VCF 9.0Advancedvcap-advanced⏱ 75 min

VCF 9.0 with Tanzu Kubernetes Grid (TKG) and vSphere Kubernetes Service (VKS). NSX as CNI provider. Demonstrates advanced K8s networking integration with NSX Container Plugin (NCP) for automatic segment provisioning, Avi/ALB virtual service creation, and network policy enforcement via NSX Distributed Firewall.

Objectives

  • Deploy and validate TKG workloads with NSX-integrated networking
  • Verify NSX Container Plugin auto-provisioning of segments and Avi virtual services
  • Test external connectivity and network policy enforcement
  • Troubleshoot common VKS networking failure scenarios

Prerequisites

VCF 9.0 lab with Supervisor Cluster deployed, TKG cluster provisioned, NSX Manager operational, Avi Controller configured, NCP deployed in TKG cluster namespace

Prior labs: vcap-vks-01: VCF 9.0 Supervisor Cluster Setup, vcap-vks-03: TKG Cluster Provisioning

Required skills:

  • kubectl command-line proficiency
  • YAML manifest creation for K8s resources
  • NSX Manager UI navigation
  • Network troubleshooting fundamentals
  • Understanding of CNI, LoadBalancer services, and network policies

Lab Environment

VCF 9.0 lab environment with Supervisor Cluster, dedicated TKG cluster (3 worker nodes), NSX Manager with standard segments and Tier-1 router, Avi Controller with SE pool, Egress IP pool 10.200.0.0/24

Tasks

Task 1 Deploy Test Workload on TKG Cluster

In VKS on VCF 9.0, TKG clusters authenticate through the Supervisor Cluster using vSphere SSO credentials. LoadBalancer services trigger immediate NCP processing to provision virtual services in Avi. Pod deployment should verify node affinity and proper segment attachment.

Authenticate to TKG cluster via Supervisor and deploy a multi-replica nginx deployment with LoadBalancer service to establish baseline workload and validate pod scheduling, service creation, and NCP integration

Step 1

Log into Supervisor Cluster context with vSphere SSO credentials and obtain TKG cluster kubeconfig

kubectl get nodes returns 3 worker nodes with Ready status. kubectl config current-context shows tkg-cluster name. kubectl get ns shows at least default, kube-system, ncp, and kube-public namespaces.
Step 2

Create nginx deployment with 3 replicas and expose via LoadBalancer service

Deployment shows 3/3 replicas running, all pods in Running state. Service output: NAME=nginx-lb, TYPE=LoadBalancer, CLUSTER-IP=10.x.x.x, EXTERNAL-IP=10.200.0.x (e.g., 10.200.0.50), PORT(S)=80:nodeport/TCP. kubectl get pods -l app=nginx shows all 3 nginx-xxxxx pods across multiple nodes.
Step 3

Verify all pods are running and accessible within the cluster

All 3 nginx pods in Running state, ready count 1/1. Pod IPs allocated from NSX overlay network (e.g., 172.16.x.x range, not the worker node management IP). Nginx logs show 'master process started' and no errors. curl localhost:80 returns HTML welcome page.

Validation Gate

Check: TKG cluster authenticated, 3-replica nginx deployment created, LoadBalancer service assigned external VIP, all pods running with correct pod IP assignments

Expected: kubectl get svc nginx-lb shows EXTERNAL-IP assigned and kubectl get pods shows all Running, pods distributed across worker nodes with pod IPs from overlay network

Common Errors

Error message: 'x509: certificate signed by unknown authority' when attempting to login
Cause: TKG Supervisor API server certificate is self-signed and --insecure-certs flag not used, or kubectl context points to wrong cluster IP
Fix: Add --insecure-certs flag to kubectl vsphere login command. Verify supervisor-ip or tkg-cluster-vip is correct by checking network connectivity (ping/nslookup). Ensure vSphere SSO user has Kubernetes admin privileges assigned in Supervisor Cluster
kubectl get svc nginx-lb shows EXTERNAL-IP=<pending> for >5 minutes. kubectl describe svc nginx-lb shows 'no available IPs' or similar event
Cause: Avi SE pool not operational, IP pool exhausted, or NCP not processing LoadBalancer events. NSX virtual service creation failed due to misconfigured Avi Controller or SE connectivity issues
Fix: Verify Avi Controller reachability from NSX Manager. Check Avi SE pool has available capacity. Confirm NCP pod is running: kubectl get pods -n ncp. Check NCP logs: kubectl logs -n ncp <ncp-pod-name>. Ensure TKG cluster has sufficient egress IP pool allocation (check vCenter Supervisor Cluster settings)
kubectl get pods shows 0/1 ready, status ImagePullBackOff or CrashLoopBackOff. kubectl describe pod shows 'Failed to pull image' or network timeout
Cause: Worker nodes unable to reach container registry due to network segmentation, NSX DFW rules blocking egress, or invalid image reference
Fix: Verify worker node egress connectivity: kubectl debug node/<node-name> -it -- curl https://docker.io. Check NSX DFW policies applied to pod segments. Confirm image URL is correct and registry is reachable. Test from NSX Logical Switch by verifying Geneve tunnel status (show logical-switch <segment-name> on NSX Manager)

Task 2 Validate NSX Networking Integration

NCP runs as daemonset in TKG cluster and subscribes to K8s API events. When a LoadBalancer service is created, NCP provisions corresponding NSX logical segment for namespace pods, and signals Avi to create virtual service with SE pool members. DFW rules auto-create for network policies but base pod-to-pod communication is allowed by default (unless NetworkPolicy applied). Integration demonstrates closed-loop automation between K8s and infrastructure.

Confirm that NCP automatically provisioned NSX logical segments for pod namespace, Avi provisioned virtual service for LoadBalancer, and NSX DFW rules were created. Verify multi-layered networking architecture: segments, virtual services, and policy enforcement baseline

Step 1

Verify NCP created NSX logical segments for the TKG cluster namespace

At least 3 segments visible: tkg-01.default-pod, tkg-01.kube-system-pod, tkg-01.ncp-pod. Each segment shows Realization State=Realized. Segment details: Transport Zone=TZ-overlay, Tier-1=tier1-tkg, VLAN=N/A (overlay-only). Active ports count > 0 for default-pod segment (one port per pod + service subnet).
Step 2

Verify NCP created NSX port groups and attached pods to correct segments

All 3 nginx pod IPs are within default-pod segment subnet (e.g., 172.16.1.0/24). NCP logs show entries like 'Created logical switch tkg-01.default-pod for namespace default' and 'Created logical port for pod nginx-xxxxx'. kubectl get networks (if available) may show mapping of pods to segments.
Step 3

Verify Avi Controller created virtual service for nginx LoadBalancer and confirmed SE pool membership

Virtual service 'default-nginx-lb-tkg-01' exists and state=UP. Pool members: 3 entries with pod IPs (e.g., 172.16.1.10:80, 172.16.1.11:80, 172.16.1.12:80), all showing 'up' state. SE assignment shows SE is active and health checks passing. VIP shows assigned IP (e.g., 10.200.0.50).
Step 4

Verify NSX DFW (Distributed Firewall) rules exist for pod namespace segment and validate baseline allow rules

At minimum 3 DFW rules visible: (1) 'Allow pod-k8s-api' with source=tkg-01.default-pod, destination=k8s-api, action=ALLOW, (2) 'Allow intra-pod', source=tkg-01.default-pod, destination=tkg-01.default-pod, action=ALLOW, (3) 'Allow pod-egress' or implicit Tier-1 allow. No rules blocking port 80 or 443 to/from pod segment.

Validation Gate

Check: NSX segments auto-created for TKG namespace, Avi virtual service configured with 3 pool members, DFW rules provisioned for pod network

Expected: NSX Manager shows tkg-01.default-pod segment with Realized state, Avi shows virtual service with 3 UP pool members, DFW rules allow intra-pod and egress traffic

Common Errors

NSX Manager does not show expected tkg-01.default-pod segment. kubectl logs ncp shows 'Failed to create logical switch' errors
Cause: NCP not authenticated to NSX Manager, incorrect NSX credentials in NCP config, NSX API communication blocked by firewall, or insufficient permissions on NSX user
Fix: Verify NCP ConfigMap has correct NSX Manager IP and credentials: kubectl get cm -n ncp ncp-config -o yaml | grep nsxmanager. Test connectivity from TKG worker: kubectl debug node/<node> -it -- curl https://<nsx-ip>/api/v1/search/query. Ensure NSX user has Enterprise Admin role. Restart NCP pods: kubectl rollout restart deployment/ncp -n ncp
Avi UI shows virtual service state=DOWN or Pool Members count=0. Health checks failing for all members
Cause: Pod IPs not reachable from SE due to network routing issue, pod segment not connected to SE pool management network, or Avi SE pool not deployed
Fix: Verify SE pool is deployed and operational in Avi. Check SE can ping pod IPs: SSH to SE and execute 'ping <pod-ip>'. Verify pod segment has route to SE pool subnet via Tier-1 router. Check NSX Tier-1 has static route for pod segment subnet. Verify Avi Service Engine resources are allocated in vCenter
Pod-to-pod communication fails with timeout. NSX Manager shows DFW rules but they are in DENY state or rules are not applied to pod segment
Cause: NCP not configured to auto-create DFW rules, rules created in wrong section with lower priority, or manual DFW policy conflicts with auto-generated rules
Fix: Verify NCP config: kubectl get cm -n ncp ncp-config -o yaml | grep -i 'firewall\|dfw'. Check DFW rule priority and scope via NSX API. Ensure 'enable_security_policy' is true in NCP ConfigMap. Apply rules from correct DFW section (Kubernetes-auto-generated section should have highest priority)

Task 3 Test Connectivity and Network Policy Enforcement

NetworkPolicy CRDs are processed by NCP, which translates each policy into NSX DFW rules. The translation is bidirectional: K8s NetworkPolicy -> NSX DFW rules auto-apply on affected segments. This lab validates the feedback loop: policy applied -> rules created -> traffic blocked -> policy removed -> traffic restored. Tests both north-south (external LoadBalancer VIP) and east-west (pod-to-pod) connectivity vectors.

Validate external connectivity through Avi virtual service VIP, then apply K8s NetworkPolicy to deny ingress and verify network-level enforcement via NSX DFW. Demonstrate control plane (K8s) and data plane (NSX) interaction for fine-grained network access control

Step 1

Test external connectivity to LoadBalancer VIP and validate nginx response

HTTP 200 OK with nginx welcome HTML. Response headers include Server: nginx/1.x.x. Each curl request may be served by different pod (Round-Robin load balancing via Avi). Response time <500ms from local network.
Step 2

Apply K8s NetworkPolicy to deny all ingress traffic to nginx pods, verify connectivity fails

NetworkPolicy shows in kubectl get networkpolicies output. NSX Manager shows new DFW rule '127-k8s-pod-default-nginx-Ingress' with priority 127 and action DENY applied to tkg-01.default-pod segment. curl http://10.200.0.50 hangs and times out after ~30 seconds. No HTTP response received. NSX logs show 'Denying packet' for port 80 on affected segment.
Step 3

Remove NetworkPolicy, verify connectivity restored and external access returns

NetworkPolicy no longer appears in kubectl get networkpolicies output. NSX Manager DFW rule '127-k8s-pod-default-nginx-Ingress' is removed or state=Unrealized. curl http://10.200.0.50 returns HTTP 200 with nginx HTML. Response time <500ms, consistent with task 3 step 1.

Validation Gate

Check: External LoadBalancer VIP accessible and responsive, NetworkPolicy enforcement blocks traffic, removal restores connectivity

Expected: curl to VIP returns 200 in baseline state, timeout after NetworkPolicy applied, returns 200 again after policy removed

Common Errors

curl http://<VIP> returns 'Connection refused' or 'Network is unreachable' immediately (not timeout)
Cause: Egress IP pool not routable from bastion host, firewall/ACL blocking access, Avi SE not responding on VIP, or VIP not created
Fix: Verify VIP IP from Avi console or kubectl get svc. Test routing: traceroute <VIP> from bastion. Confirm firewall ACL allows port 80 to egress pool. Test Avi SE responsiveness: SSH to SE, ping <VIP>. Check Avi Virtual Service state is 'UP' not 'DOWN'. Verify Tier-1 has route advertised for egress pool
curl http://<VIP> still returns HTTP 200 after NetworkPolicy applied. kubectl get networkpolicies shows policy exists
Cause: NCP not processing NetworkPolicy events, NCP ConfigMap 'enable_security_policy' set to false, or DFW rules created in wrong section with lower priority than existing allow rules
Fix: Verify NCP is running and processing events: kubectl logs -n ncp deployment/ncp | grep -i 'networkpolicy\|pod-selector'. Check config: kubectl get cm -n ncp ncp-config -o yaml | grep enable_security_policy. Ensure DFW rule section priority is high (lower number = higher priority). Manually check NSX DFW rule was created: curl -k -u admin:password https://nsx-ip/api/v1/firewall/sections
curl http://<VIP> still times out after NetworkPolicy deleted. kubectl get networkpolicies is empty
Cause: DFW rule not removed by NCP, rule re-applied by conflicting policy, or rule section locked in NSX Manager
Fix: Manually verify DFW rule is removed: NSX Manager Security > Distributed Firewall. If rule still present, check NCP logs for delete event: kubectl logs -n ncp --tail=100 | grep -i 'removed'. Force NCP to reconcile: kubectl rollout restart deployment/ncp -n ncp. Check no other NetworkPolicies exist: kubectl get networkpolicies -A. Verify NSX user has write permissions to DFW section

Task 4 Troubleshoot Common VKS Networking Issues

Real-world VKS issues stem from infrastructure gaps: (1) Avi SE resource exhaustion or misconfiguration preventing virtual service provisioning, (2) IP pool exhaustion when demand exceeds capacity, (3) NSX Geneve tunnels between hypervisors down, breaking overlay network, (4) network policies incorrectly configured blocking legitimate traffic. Lab presents two scenarios requiring root-cause analysis, remediation, and preventive measures. Diagnostic artifacts collected build operational playbook.

Diagnose and resolve realistic networking failure scenarios in VKS environment. Develop troubleshooting methodology for LoadBalancer provisioning delays, pod-to-pod connectivity failures, and TEP (Tunnel Endpoint) communication issues. Document findings and create reusable runbook entries

Step 1

Scenario A: LoadBalancer service stuck in Pending state - diagnose and resolve IP pool exhaustion or Avi SE failure

Diagnosis identifies root cause: 'Avi SE pool at 100% capacity - no available SEs for new virtual service' or 'IP pool 10.200.0.0/24 exhausted (254/254 IPs allocated)' or 'NCP cannot reach Avi Controller: https://avi-controller-ip/api/cluster refused connection'. Remediation: Avi SE added to pool OR IP pool extended OR NCP network rules updated. After remediation, kubectl get svc test-app shows EXTERNAL-IP assigned (e.g., 10.200.0.100).
Step 2

Scenario B: Pod-to-pod cross-node communication fails - diagnose TEP connectivity or Geneve encapsulation issues

Diagnosis identifies root cause: 'TEP interface down on worker-2 (10.0.1.52 unreachable)' or 'Geneve tunnel status=DOWN between worker-1<->worker-2' or 'Pod segment not connected to Tier-1 (routing not available)'. After remediation, kubectl exec pod-a -- wget -O- http://<pod-b-ip> returns success with HTTP headers from test service, no timeout.
Step 3

Document VKS networking troubleshooting runbook and create decision tree for common issues

Runbook document containing: (1) 3+ issue types with complete symptom/diagnosis/remediation chains, (2) decision tree with >8 decision nodes, (3) 10+ diagnostic commands with expected output patterns, (4) remediation steps for each root cause (Avi/IP/NSX/network/firewall), (5) escalation contacts/paths, (6) preventive measures (monitoring, capacity planning). Document format: Well-structured markdown with headers, code blocks, tables. Completeness: Any engineer following runbook can diagnose and resolve 80% of common issues without Slack/escalation.

Validation Gate

Check: Successfully diagnosed and remediated LoadBalancer pending and pod-to-pod connectivity scenarios. Created reusable troubleshooting runbook

Expected: Both scenarios resolved with documented root causes and remediation steps. Runbook contains decision tree and diagnostic procedures for future reference

Common Errors

Cannot determine why LoadBalancer is pending or pod communication failed. Multiple possible causes but no clear order of investigation
Cause: Lack of systematic diagnostic methodology, unfamiliarity with Avi/NSX/NCP integration points, insufficient logging/visibility
Fix: Start with highest-impact checks (Avi SE capacity, IP pool free count) before low-level network checks. Use structured approach: (1) Verify K8s service object exists and configuration is valid, (2) Check NCP processed event and created infrastructure, (3) Verify infrastructure is operational (Avi virtual service UP, NSX segments exist), (4) Test connectivity end-to-end. If stuck, enable debug logging: kubectl set env deployment/ncp -n ncp DEBUG=true, then retry operation
Attempting to scale Avi SE pool or remove service results in errors or cascading failures
Cause: Running remediation commands without understanding full context, insufficient permissions, or affecting production workloads
Fix: Never directly delete services/pods in production. For Avi scaling, use UI or API with dry-run first. For NSX routing, validate no production traffic uses affected segment before changes. Always create test/lab scenario first before applying remediation in production

Final Validation

Lab successfully demonstrates end-to-end VKS networking integration: workload deployment, infrastructure auto-provisioning, connectivity validation, and troubleshooting. Learner gains operational confidence in VKS networking stack

✓ Task 1: TKG cluster authenticated, 3-replica nginx deployment running with LoadBalancer VIP assigned → kubectl shows 3 Running pods, LoadBalancer EXTERNAL-IP populated, pods distributed across nodes

✓ Task 2: NSX segments auto-created, Avi virtual service configured with 3 UP pool members, DFW rules present → NSX Manager shows cluster.namespace-pod segments, Avi shows virtual service UP, DFW rules list includes pod network policy rules

✓ Task 3: External curl to VIP succeeds, NetworkPolicy blocks traffic, removal restores connectivity → curl returns 200 initially, timeout after policy applied, 200 again after removal

✓ Task 4: Both failure scenarios diagnosed with root causes identified, runbook created with decision tree → Remediation steps documented, diagnosis commands listed, runbook includes 10+ commands and 5+ decision nodes

Cleanup / Restore

• kubectl delete deployment nginx-app test-app pod-a pod-b

• kubectl delete service nginx-lb test-app

• kubectl delete networkpolicy deny-ingress (if still present)

• NSX Manager: Verify segments auto-cleaned after kubernetes cleanup (NCP garbage collection)

• Avi Controller: Verify virtual services removed/marked down after K8s service deletion

• Optional: Revert to snapshot if extended troubleshooting modified cluster state

Design Reflection (VCDX)

This lab addresses VCDX-level networking complexity in Kubernetes-on-VCF environments. Key VCDX considerations: (1) Multi-layered abstraction: K8s networking constructs (Service, Pod, NetworkPolicy) must map correctly to infrastructure (NSX segments, Avi virtual services, DFW rules). Understanding this translation layer is critical for troubleshooting. (2) Automation and feedback loops: NCP provides event-driven infrastructure provisioning. Failures occur when automation loops break (NCP can't reach NSX, Avi API unreachable). VCDX must design resilience into these integration points.

(3) Resource pooling and capacity: LoadBalancer services consume both K8s and infrastructure resources (SE slots, IP addresses, NSX bandwidth). VCDX must plan capacity holistically across all layers. (4) Security and compliance: NetworkPolicy translates to NSX DFW rules, but policy intent can be lost in translation if misconfigured. VCDX must ensure control plane policies enforce correctly in data plane. (5) Operational complexity: When things work, VKS seems seamless. When they break, operators face distributed failure modes across K8s, NSX, Avi, and network infrastructure.

VCDX design should emphasize observability, clear failure modes, and runbook-driven troubleshooting.

Requirements

  • TKG cluster with 3+ worker nodes deployed on VCF 9.0 Supervisor
  • NSX (4.1+ on VCF 5.x, or NSX 9.0.x on VCF 9.x) configured with overlay transport zone and Tier-1 router
  • Avi Controller 22.1+ with SE pool deployed and reachable from TKG worker nodes
  • NCP deployed in TKG cluster namespace with admin permissions to NSX Manager and Avi
  • Egress IP pool (min /25 subnet) routable from external bastion/test client
  • vSphere SSO configured for TKG cluster kubeconfig authentication
  • DFW rules enabled in NSX and allowed in TKG namespace security context

Constraints

  • LoadBalancer service provisioning time depends on Avi SE availability (typically 30-60 seconds)
  • NSX DFW rule application has eventual consistency window (5-15 seconds propagation to all nodes)
  • Pod-to-pod overlay communication requires fully functional Geneve tunnels between all hypervisors
  • NetworkPolicy enforcement only applies to traffic within NSX-managed segments; host-networked pods bypass NCP control
  • Avi virtual services limited by SE pool capacity and license seat count
  • IP pool size limits maximum concurrent LoadBalancer services (e.g., /24 = 250 usable IPs)

Assumptions

  • All hypervisors in cluster have TEP connectivity and NSX agent running
  • Avi Controller is configured for TKG cluster integration with correct cloud credentials
  • NCP ConfigMap contains correct NSX Manager credentials and Avi API endpoint
  • Network firewall allows K8s API traffic (port 443) between pod network and Supervisor
  • Test client can reach egress IP pool (10.200.0.0/24) without additional hop/firewall
  • Learner is comfortable with kubectl, Linux shell, and NSX Manager UI navigation

Risks

  • Risk: NSX Manager unavailability -> NCP cannot provision segments -> LoadBalancer services fail. Mitigation: Ensure NSX Manager HA cluster operational. Monitor NSX Controller cluster status periodically.
  • Risk: Avi SE pool exhaustion -> new LoadBalancer services queue indefinitely. Mitigation: Pre-size SE pool for expected workload scale. Monitor SE capacity, set up alerts.
  • Risk: IP pool exhaustion -> cannot allocate new LoadBalancer VIPs. Mitigation: Manage LoadBalancer service count, reuse shared virtual services, plan egress pool size.
  • Risk: NetworkPolicy incorrectly configured -> blocks legitimate traffic. Mitigation: Test policies in staging cluster, use allow-list approach rather than deny rules.
  • Risk: Geneve tunnel down on one hypervisor -> pods on that node isolated. Mitigation: Monitor tunnel status, configure alerts for down tunnels, understand remediation (NSX restart, network fix).
  • Risk: Operator unfamiliar with multi-layer troubleshooting -> escalates every issue. Mitigation: Use runbook methodology, train on systematic diagnosis, empower ops with clear decision trees

Self-Assessment Discussion Prompts

  1. When a LoadBalancer service gets stuck in Pending state, what are the top 3 infrastructure components you would check first? In what order and why? How would you verify each is operational?
  2. NSX DFW rules have 15-second propagation delay to all hypervisors. How does this affect pod connectivity during network policy deployment/removal? What are the implications for security if a pod receives traffic during the propagation window?
  3. In this lab, we tested NetworkPolicy ingress denial. How would you test egress denial (restricting a pod from initiating outbound connections)? What would change in the NCP-to-DFW rule translation?
  4. The lab uses overlay networking (Geneve). What are the tradeoffs vs. VLAN-based networking for VKS clusters? When would you choose VLAN over overlay and what additional NSX configuration would be required?
  5. Avi LoadBalancer services consume SE resources. How would you design capacity planning for a VKS environment where you expect 100+ LoadBalancer services? What metrics would you monitor to predict SE pool saturation?
  6. If the NSX Manager becomes unavailable, NCP pods continue running but cannot provision new infrastructure. How would you design graceful degradation - should NCP queue requests or fail fast? What user-facing error should they see?
  7. The lab demonstrates NetworkPolicy -> DFW rule translation as 1:1 mapping. What edge cases or complex policies might not translate cleanly? How would you troubleshoot if K8s policy intent differs from NSX data plane reality?

Extensions

Multi-cluster Service Discovery with Avi GSLB

Extend lab to multiple TKG clusters and configure Avi Global Server Load Balancing (GSLB) for cross-cluster LoadBalancer service discovery. Create two nginx deployments on separate TKG clusters, federate via Avi GSLB, and demonstrate automatic failover when one cluster becomes unavailable. Tests infrastructure-level HA and multi-cluster networking.

Ingress Controller with NSX integration (Contour + NSX)

Deploy Contour Ingress Controller on TKG cluster and configure with NSX integration to auto-provision L7 virtual services in Avi. Create application with Ingress resource (hostname-based routing), verify Avi virtual service created with multiple backends, test path-based and hostname-based routing. Demonstrates L7 networking on top of L4 LoadBalancer foundation.

Pod Security Policy and NSX DFW Integration

Create restrictive pod security policies (no privileged containers, restricted SELinux) and verify corresponding NSX DFW rules are created by NCP. Attempt to deploy privileged container and verify it fails at K8s admission control AND NSX data plane. Document the defense-in-depth layers. Tests security architecture across container and network planes.

⚠ Known Pitfalls (from Community KB)

Assuming LoadBalancer provisioning is instantaneous
Problem: Many operators expect EXTERNAL-IP to populate immediately after kubectl apply. In reality, NCP processing + NSX segment creation + Avi virtual service provisioning takes 30-60 seconds. Impatient troubleshooters restart pods or services, causing cascading delays. Set expectations: monitor with 'kubectl get svc -w', don't interrupt within first minute.
Resolution: Educate team on typical provisioning timeline. Create alerting if EXTERNAL-IP remains pending >5 minutes. Document timeout handling in runbooks.
Forgetting NSX DFW has eventual consistency
Problem: NetworkPolicy applied but traffic still flows briefly, causing confusion about whether policy actually worked. Root cause: DFW rule not yet applied to all hypervisors. Operators see 'it's broken' rather than understanding propagation delay.
Resolution: Always test network policies with 10-15 second grace period between application and verification. Document rule propagation timing in runbooks. Use verbose kubectl watch to see rollout of policy application.
Misconfiguring IP pool exhaustion as 'Avi issue' instead of capacity planning failure
Problem: When egress IP pool exhausted, operator thinks Avi is broken and tries to scale SE pool (wrong fix). Real issue: too many LoadBalancer services competing for limited IPs. Results in wasted effort and unresolved problem.
Resolution: Train on capacity planning: (# Services) * (1 VIP) + buffer = IP pool size. Audit IP pool utilization monthly. Set up alerts at 70% and 90% utilization. Create service mesh or SharedIngressController strategy to reduce per-service IP consumption.
Breaking pod-to-pod communication by applying too-permissive default-deny NetworkPolicy
Problem: Operator creates NetworkPolicy with empty ingress rule intending to start with deny-all. Forgets to add egress rules. Result: pods can't reach K8s DNS (port 53) and service endpoints (port 443). Manifests as 'pod can't start' or 'DNS failures'.
Resolution: Use a scaffolded NetworkPolicy template that includes both ingress AND egress, with K8s-api and kube-dns explicitly allowed. Test policies on non-critical namespace first. Provide policy design review checklist: 'Does this allow pod->kube-dns?', 'Does this allow pod->api-server?'

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.