NSX Connectivity Troubleshooting
Objectives
- Map complete NSX data plane path (VM, segment, T1 router, T0 router, external gateway)
- Understand NSX control plane dependencies (Manager cluster, CCP, LCP on transport nodes)
- Use NSX Traceflow tool to trace packet path and identify drop reasons
- Interpret nsxcli commands for router, firewall, and logical switch state validation
- Diagnose transport endpoint (TEP) connectivity failures and BFD session health
- Correlate control plane (Manager, CCP/LCP) issues with data plane connectivity loss
- Document RCA (root cause analysis) and create troubleshooting runbooks for common NSX issues
Prerequisites
VCF lab with NSX 4.x deployed and operational (at minimum, 2 T1 routers, 1 T0 router, 2+ transport nodes)
Prior labs: NSX Fundamentals (basic segment creation, static routing)
Required skills:
- NSX Manager UI navigation
- SSH/CLI access to transport nodes (ESXi hosts or Edge nodes)
- Basic network troubleshooting (IP routing, packet flow concepts)
- Firewall rule logic and stateful vs. stateless packet handling
Lab Environment
Multi-tier NSX topology: VM-A on Tier-1-app, VM-B on Tier-1-db, both connected to Tier-0 router. Transport nodes: 2+ ESXi hosts or 2 Edge VMs. NSX Manager cluster (at least 1 node, 3+ recommended for HA).
Tasks
Task 1 Map the NSX Network Path and Identify Control Plane Dependencies
Develop a complete understanding of the data plane (packet forwarding) and control plane (configuration sync, state replication) paths. Demonstrate the ability to trace a packet from source VM through all routing, switching, and firewall layers to destination VM.
Draw a detailed network topology diagram showing: VM-A on logical segment app-segment (connected to T1-app), VM-B on logical segment db-segment (connected to T1-db), both T1 routers connected to T0 router. Include physical transport nodes (ESXi hosts or Edge VMs) where each logical component resides.
Identify all components in the data plane path: (1) Distributed Firewall (DFW) at source and destination segments, (2) Logical switches (app-segment, db-segment), (3) T1 Service Router (T1-SR) and Distributed Router (T1-DR), (4) T0 Service Router (T0-SR) and Distributed Router (T0-DR), (5) Tunnel between transport nodes (if applicable).
Identify control plane components: (1) NSX Manager Cluster (controls policy distribution), (2) Central Control Plane (CCP, typically co-located with NSX Manager nodes), (3) Local Control Plane (LCP, runs on each transport node), (4) Edge Control Plane (if using Edge VMs). Document the communication path: Manager → CCP → LCP on each transport node.
For each layer in the path (DFW, logical switch, T1 DR, T1 SR, T0 DR, T0 SR), document: (a) Where it runs (which transport node), (b) What it does (function), (c) What can fail there (common failure modes).
In NSX Manager UI, navigate to Networking > Connectivity > Segments. Verify that both app-segment and db-segment exist and are connected to their respective T1 routers. Verify T1 and T0 routers are connected. Document the state (READY, DOWN, DEGRADED).
From NSX Manager, navigate to System > Fabric > Transport Nodes. Verify both ESXi hosts (or Edge VMs) show 'READY' state. Click each node and check: (a) CCP connectivity (should be READY), (b) VNI pool status (should have available VNIs), (c) TEP (tunnel endpoint) IP address and reachability.
Document your complete topology map and control plane architecture in a written summary. Include: (1) ASCII diagram or reference to your saved diagram, (2) Data plane component list and function, (3) Control plane component list and communication path, (4) Baseline state of all components (READY), (5) Potential failure points and impact (if DFW is DOWN, which VMs are affected? If T0 SR fails, which traffic is rerouted?).
Validation Gate
Check: Complete topology map created with identified data plane and control plane components
Expected: Lab notes contain detailed topology diagram, component-by-component analysis, documented baseline state, and written summary of NSX architecture
Common Errors
Task 2 Troubleshoot with NSX Traceflow
Master the Traceflow tool to diagnose connectivity issues and identify the exact point where packets are dropped (DFW, T1 router, T0 router, TEP failure, etc.). Learn to interpret Traceflow output to correlate drops with configuration errors.
Open NSX Manager and navigate to Troubleshooting > Traceflow. Click 'New Trace' to create a new trace from VM-A to VM-B.
Configure the trace: Source = VM-A (on app-segment), Destination = VM-B (on db-segment). Protocol = TCP. Source Port = any, Destination Port = 443 (or any port where VM-B is listening). Action = 'Run Trace'.
If trace is DELIVERED, examine the detailed hop-by-hop results: (1) Packet enters DFW at source (rule applied, allowed). (2) Packet switches through app-segment (MAC learned, forwarded to T1 DR port). (3) T1-app DR routes to T0 (route lookup succeeded, next-hop is T0 DR). (4) T0 DR routes to T1-db (return route lookup succeeded). (5) T1-db DR routes to db-segment (MAC lookup, forwarded to VM-B port). (6) DFW at destination checks rule (allowed). (7) Packet delivered to VM-B.
Simulate a failure scenario: In NSX Manager, navigate to Networking > Security > Distributed Firewall. Find the DFW rule allowing app-to-db traffic. Create a NEW rule ABOVE it that explicitly DENIES the same traffic (to simulate misconfiguration). Save the rule.
Re-run Traceflow (VM-A to VM-B, TCP port 443). Trace will now show BLOCKED at the DFW ingress layer. Examine the detailed output: 'DFW@app: BLOCK (rule <ID>, policy <POLICY-NAME>)'.
Fix the DFW rule by deleting the deny rule or moving the allow rule above it. Re-run Traceflow to confirm the trace now shows DELIVERED.
Simulate a routing failure: In NSX Manager, navigate to Networking > Connectivity > Tier-0 > Routing > Static Routes. Find the route to db-segment (if using static routing) or use dynamic routing to examine the OSPF/BGP table. Delete the route temporarily to simulate routing failure.
Re-run Traceflow. Observe the result: [T1-app DR: DROP (no route to <db-segment-subnet>)] → [DROPPED]. This identifies the routing issue. Restore the route, re-run Traceflow to confirm recovery.
Document your Traceflow methodology: (1) Always run Traceflow bidirectionally (A→B and B→A) if traffic is bi-directional. (2) Test common protocols (TCP/UDP, common ports 80, 443, 3306). (3) If Traceflow succeeds but real traffic fails, suspect MTU issues, TCP option stripping, or asymmetric routing. (4) Traceflow does NOT account for packet size; large packets may fail in real world due to MTU but succeed in Traceflow.
Validation Gate
Check: Traceflow used to diagnose multiple connectivity scenarios (baseline, DFW block, routing failure)
Expected: Lab notes document Traceflow output for each scenario, interpreted drop reasons, and remediation steps
Common Errors
Task 3 CLI-Based Troubleshooting and Transport Node Diagnostics
Master nsxcli commands and lower-layer diagnostics to validate control plane sync, routing tables, firewall state, and TEP connectivity when UI-based tools are unavailable or slow.
SSH to one of your transport nodes (ESXi host or Edge VM). Log in as root or a user with admin privileges.
Run 'get managers' to verify the host can reach NSX Manager cluster. Output should show Manager node(s) with status CONNECTED.
Run 'get logical-switches' to list logical switches on this host. Output should show app-segment and db-segment (or whichever segments are connected to this host's VMs).
Run 'get logical-routers' to verify T1 and T0 routers are visible to this host. Output should list all T1 routers and the T0 router.
Run 'get logical-router-port T1-app' (or similar) to examine routing table entries for the T1-app router. Look for routes to db-segment subnet and upstream routes to T0.
Run 'get firewall-rules' to list DFW rules installed on this transport node. Output should show the security policy (e.g., policy-app-db) and its rules.
Run 'get tunnel-interfaces' or 'get vtep-interfaces' to list tunnel endpoints (TEP) on this host and their operational status. Output should show TEP IP address and adjacency count (how many other hosts this host is tunneled to).
Run 'get bfd-sessions' to check Bidirectional Forwarding Detection (BFD) sessions to other transport nodes. BFD is NSX's keep-alive mechanism for tunnel health. Output should show all BFD sessions as UP.
Run 'vmkping <other-host-TEP-IP>' to verify TEP network connectivity. Example: 'vmkping 192.168.100.11'. Output should show successful pings (0% loss).
For Edge VMs (if used for T0 or T1 services), run 'get controllers' to verify the Edge VM's connection to NSX Manager's Control Plane. Output should show CONNECTED.
Document your nsxcli diagnostics workflow: (1) Check Manager connectivity (get managers). (2) List logical switches/routers (get logical-switches, get logical-routers). (3) Inspect routing tables (get logical-router-port). (4) Verify firewall rules (get firewall-rules, check install status). (5) Validate TEP connectivity (get tunnel-interfaces, vmkping). (6) Confirm BFD health (get bfd-sessions).
Validation Gate
Check: CLI commands executed on transport node(s) with documented output and interpretation
Expected: Lab notes contain nsxcli command output for logical switches, routers, firewall rules, TEP interfaces, BFD sessions, and documented health assessment
Common Errors
Task 4 Remediation, Validation, and RCA Documentation
Apply fixes to identified issues, validate remediation with Traceflow, and document root cause analysis in a structured format suitable for runbooks and incident post-mortems.
Based on your diagnostics from Tasks 2-3, identify the root cause of the connectivity issue (e.g., DFW rule misconfiguration, routing table incomplete, TEP network unreachable). Document the exact issue in a brief RCA statement.
Apply the fix. Example: If DFW rule is misconfigured, correct the rule in NSX Manager > Networking > Security > Distributed Firewall. Edit rule, fix CIDR block, save and publish.
Monitor the rule push process. Run 'get firewall-rules' on transport nodes to confirm rule status changes from PENDING → INSTALLED. Allow 30-60 seconds for sync.
Re-run Traceflow from VM-A to VM-B with the same parameters as before (TCP, port 443). Trace should now show DELIVERED (all hops pass, no blocks).
Validate the fix at the application layer. SSH to VM-B and verify the service is listening on port 443. From VM-A, attempt to connect to VM-B: 'telnet <VM-B-IP> 443' or 'curl https://<VM-B-IP>' (if web service). Connection should succeed.
Document the RCA in a structured format: (1) Incident Date/Time. (2) Symptom (user report or automated alert). (3) Timeline (when incident was detected, when troubleshooting started, when fix was applied). (4) Root Cause (technical explanation). (5) Impact (how many VMs/users affected, downtime duration). (6) Fix Applied (specific configuration change). (7) Validation (Traceflow + application test). (8) Prevention (how to avoid this in the future).
Create a troubleshooting runbook entry based on this incident. Runbook should be a step-by-step guide that a junior engineer or NOC (Network Operations Center) tech can follow to diagnose the same issue in the future.
If the issue was a design flaw (not a one-off misconfiguration), recommend architectural changes. Example: If T1 routers lack redundancy, design and propose a HA T1 setup. If TEP network is single-home, design a dual-homed TEP network.
Validation Gate
Check: Issue fixed, remediation validated with Traceflow, RCA documented, runbook created
Expected: Lab notes contain complete RCA document (symptom, root cause, fix, validation, prevention) and troubleshooting runbook entry suitable for team reference
Common Errors
Final Validation
Lab completed when all 4 tasks are executed and fully documented with topology maps, Traceflow tests, CLI diagnostics, RCA, and runbooks.
✓ Complete NSX topology mapped with data plane and control plane identified → Task 1 complete: diagram shows all components (VM, DFW, logical switches, T1/T0 routers, transport nodes, Manager, CCP, LCP)
✓ Traceflow used to diagnose baseline, failure, and remediation scenarios → Task 2 complete: Traceflow traces documented showing DELIVERED (healthy), BLOCKED (DFW misconfiguration), DROPPED (routing failure), and remediated DELIVERED
✓ nsxcli commands executed to validate control and data plane state → Task 3 complete: Output from get managers, get logical-switches, get logical-routers, get firewall-rules, get tunnel-interfaces, get bfd-sessions, vmkping documented and interpreted
✓ Issue remediated and validated with Traceflow + application testing → Task 4 complete: RCA document (symptom, root cause, fix, impact, prevention), troubleshooting runbook entry, and architectural recommendations for future prevention
Cleanup / Restore
• Remove any temporary DFW rules created during testing (deny rules added for failure simulation)
• Restore any deleted logical switches, routers, or routes (use NSX Manager snapshots if available)
• Verify VM-A and VM-B are back to baseline state (connectivity DELIVERED, no policy blocks)
• Document any permanent architectural changes recommended (for infrastructure team to implement)
Design Reflection (VCDX)
This lab emphasizes NSX troubleshooting methodology and operational excellence. VCDX panelists will probe: (1) How do you debug control plane issues when UI is slow or unavailable? (2) What is your mental model of NSX data flow and where failures occur? (3) Can you prioritize between control plane (Manager, CCP/LCP) failures and data plane (DFW, routing, TEP) failures? (4) How do you design NSX monitoring to detect these issues proactively? The RCAR framework emphasizes root cause analysis and architectural thinking, not just tactical fixes.
Requirements
- NSX 4.x Manager cluster (minimum 1 node, 3+ recommended for HA testing)
- 2+ transport nodes (ESXi hosts or Edge VMs) for tunnel/TEP testing
- 2 logical segments (app-segment, db-segment)
- 2 tier-1 routers (T1-app, T1-db) and 1 tier-0 router (T0)
- Test VMs: VM-A on app-segment, VM-B on db-segment
- Firewall rules defining app-to-db policy
- SSH/CLI access to transport nodes for nsxcli commands
Constraints
- NSX Traceflow only works for VMs registered with NSX; bare-metal or external devices require packet capture (tcpdump) instead
- DFW rule changes have 5-15 second propagation latency from Manager to transport nodes; tests must account for this sync time
- TEP network must be routable; if TEP network is not L2-adjacent, tunneling requires L3 routing setup (may not be available in lab)
- BFD/tunnel health is binary (UP/DOWN); subtle issues like high latency or packet loss may not show as DOWN immediately
- nsxcli commands vary by host type (ESXi vs. Edge VM); command syntax differs between NSX versions
Assumptions
- NSX Manager and transport nodes have network connectivity (at baseline)
- Lab operator has admin-level access to NSX Manager and transport node CLIs
- VMs can reach each other if no security policies interfere (i.e., baseline connectivity works before introducing deliberate failures)
- Firewall policies are deployed and synchronized before the lab starts (Manager → LCP sync is operational)
- RCA and runbook documentation are stored in a centralized repository (Confluence, GitHub, etc.) for team access
Risks
- Risk: Breaking actual connectivity during failure simulation. If you delete a real route or firewall rule, production traffic could fail. — Mitigation: Use a dedicated lab environment, not shared infrastructure. Snapshot the NSX Manager config before destructive tests. Have a revert plan (snapshot or config backup).
- Risk: Misinterpreting Traceflow results. A DELIVERED trace does not guarantee real traffic will succeed (MTU, TCP options, asymmetric paths). — Mitigation: Always follow Traceflow with application-level testing (telnet, curl, packet capture). Document limitations of Traceflow in your runbook.
- Risk: Control plane issues (Manager down, LCP disconnected) cascading to data plane. If Manager is unavailable, LCP goes read-only; new policies cannot be deployed. — Mitigation: Test with Manager deliberately unavailable (simulate network partition). Understand graceful degradation: data plane uses cached rules, services degrade but do not fail completely.
- Risk: TEP network partition causing split-brain. Host A can reach Host B, but Host B cannot reach Host A due to asymmetric routing. Traceflow from A→B succeeds, but real traffic fails both ways. — Mitigation: Test bidirectional paths. Use vmkping and BFD status to confirm tunnel health, not just one-way connectivity.
Self-Assessment Discussion Prompts
- In the DFW misconfiguration scenario, you identified the issue via Traceflow. But in a real incident, what would prevent you from using Traceflow (e.g., VMs not NSX-registered, NSX Manager down)? What is your backup diagnostic method?
- You validated the fix with Traceflow DELIVERED + application test. But a week later, user reports intermittent failures (traffic works sometimes, fails sometimes). Traceflow always shows DELIVERED. What could cause transient failures that Traceflow doesn't catch?
- In your RCA, you recommended adding a second T1 SR Edge VM for HA. What is the cost/benefit? What are the design trade-offs? How would you validate that HA failover actually works under load?
- Your lab uses static routing (T1 → T0). In production, would you use dynamic routing (OSPF, BGP)? What are the pros and cons? How does dynamic routing change your troubleshooting approach?
- You ran Traceflow and got DELIVERED, but nsxcli get firewall-rules showed the rule status as PENDING (not yet installed). Can Traceflow deliver packets if firewall rules are not yet installed? What is the behavior?
- If NSX Manager becomes unavailable (network partition), you said data plane goes read-only. What happens to new traffic if a new VM is spun up? Is that VM reachable without Manager? When does Manager become critical again?
- You used vmkping to validate TEP network. But vmkping is ICMP; NSX uses VXLAN (UDP 4789). Can network policies block VXLAN while allowing ICMP? How would you detect that scenario?