Academy/VCF 9.0 Support (2V0-15.25)/NSX Connectivity Troubleshooting
This lab targets VCF 9.0

NSX Connectivity Troubleshooting

VCF 9.0Advancedvcp-foundationvcp-supportnsx-advanced⏱ 90 min

Advanced NSX data plane and control plane diagnostics.

Objectives

  • Map complete NSX data plane path (VM, segment, T1 router, T0 router, external gateway)
  • Understand NSX control plane dependencies (Manager cluster, CCP, LCP on transport nodes)
  • Use NSX Traceflow tool to trace packet path and identify drop reasons
  • Interpret nsxcli commands for router, firewall, and logical switch state validation
  • Diagnose transport endpoint (TEP) connectivity failures and BFD session health
  • Correlate control plane (Manager, CCP/LCP) issues with data plane connectivity loss
  • Document RCA (root cause analysis) and create troubleshooting runbooks for common NSX issues

Prerequisites

VCF lab with NSX 4.x deployed and operational (at minimum, 2 T1 routers, 1 T0 router, 2+ transport nodes)

Prior labs: NSX Fundamentals (basic segment creation, static routing)

Required skills:

  • NSX Manager UI navigation
  • SSH/CLI access to transport nodes (ESXi hosts or Edge nodes)
  • Basic network troubleshooting (IP routing, packet flow concepts)
  • Firewall rule logic and stateful vs. stateless packet handling

Lab Environment

Multi-tier NSX topology: VM-A on Tier-1-app, VM-B on Tier-1-db, both connected to Tier-0 router. Transport nodes: 2+ ESXi hosts or 2 Edge VMs. NSX Manager cluster (at least 1 node, 3+ recommended for HA).

Tasks

Task 1 Map the NSX Network Path and Identify Control Plane Dependencies

Task requires multi-layer system thinking. Panelists will ask: (1) What happens to control plane if NSX Manager becomes unavailable? (2) How does CCP (Central Control Plane) differ from LCP (Local Control Plane)? (3) What is the role of the East-West Gateway (EWG) in T1-to-T1 traffic? (4) Can you articulate the difference between a T0 data plane issue vs. a control plane issue?

Develop a complete understanding of the data plane (packet forwarding) and control plane (configuration sync, state replication) paths. Demonstrate the ability to trace a packet from source VM through all routing, switching, and firewall layers to destination VM.

Step 1

Draw a detailed network topology diagram showing: VM-A on logical segment app-segment (connected to T1-app), VM-B on logical segment db-segment (connected to T1-db), both T1 routers connected to T0 router. Include physical transport nodes (ESXi hosts or Edge VMs) where each logical component resides.

Diagram shows: [VM-A on ESXi Host 1] → [app-segment (DFW attached)] → [T1-app SR (Service Router)] → [T1-app DR (Distributed Router)] → [T0 SR] → [T0 DR] → [T1-db DR] → [T1-db SR] → [db-segment (DFW attached)] → [VM-B on ESXi Host 2]. Label each arrow with component function (switching, routing, firewall).
Use a tool like draw.io or Visio to create a professional diagram. Include physical host locations to show which components run on which transport node.
Step 2

Identify all components in the data plane path: (1) Distributed Firewall (DFW) at source and destination segments, (2) Logical switches (app-segment, db-segment), (3) T1 Service Router (T1-SR) and Distributed Router (T1-DR), (4) T0 Service Router (T0-SR) and Distributed Router (T0-DR), (5) Tunnel between transport nodes (if applicable).

Path documented: Packet from VM-A ingresses app-segment → hits DFW at ingress (security policy applied) → logical switch forwards to T1-app DR → T1-app DR routes to T0 DR (local path if on same host, or tunneled path if cross-host) → T0 DR routes to T1-db DR → T1-db DR routes to logical switch db-segment → DFW at egress (security policy applied) → delivers to VM-B. Total components in path: 7 (2 DFWs, 2 logical switches, 2 T1 routers, 1 T0 router).
Not all components are guaranteed to be in the path. If VM-A and VM-B are on the same T1 (not common in multi-tier), traffic may bypass T0 entirely.
Step 3
Identify control plane components: (1) NSX Manager Cluster (controls policy distribution), (2) Central Control Plane (CCP, typically co-located with NSX Manager nodes), (3) Local Control Plane (LCP, runs on each transport node), (4) Edge Control Plane (if using Edge VMs). Document the communication path: Manager → CCP → LCP on each transport node.
Control plane architecture documented: NSX Manager cluster (3 nodes) publishes DFW rules and routing tables to CCP → CCP replicates to LCP on ESXi-1 and ESXi-2 → Each LCP caches rules locally and pushes to kernel driver (nsx-mpa for DFW, nsx-ra for routing). If control plane link is broken, data plane uses cached rules but cannot accept new policies.
Control plane latency is typically 100-500ms. If a new DFW rule is added, there may be a short window where some packets use old rule. Document this for panel discussion.
Step 4

For each layer in the path (DFW, logical switch, T1 DR, T1 SR, T0 DR, T0 SR), document: (a) Where it runs (which transport node), (b) What it does (function), (c) What can fail there (common failure modes).

Detailed layer analysis: (1) DFW at source: runs on ESXi-1, filters egress traffic, can drop packets if rule denies. (2) app-segment: logical switch on ESXi-1, learns MAC addresses, can lose MAC entries if control plane disconnects. (3) T1-app DR: distributed on ESXi-1, handles local routing, can have stale routing table if CCP link is down. (4) T1-app SR: centralized on Edge VM, used for NAT/services, can be unavailable if Edge VM crashes. (5) T0 components: similar failure modes at cluster interconnect layer.
Most NSX failures occur at control plane boundaries: rule caching inconsistency, routing table divergence, tunnel down.
Step 5

In NSX Manager UI, navigate to Networking > Connectivity > Segments. Verify that both app-segment and db-segment exist and are connected to their respective T1 routers. Verify T1 and T0 routers are connected. Document the state (READY, DOWN, DEGRADED).

UI shows: app-segment (READY, connected to T1-app, DFW attached), db-segment (READY, connected to T1-db, DFW attached), T1-app (READY, connected to T0), T1-db (READY, connected to T0), T0 (READY, uplink connected). All components green = healthy baseline.
If any component shows DEGRADED or DOWN, do not proceed with subsequent tasks. Resolve first.
Step 6

From NSX Manager, navigate to System > Fabric > Transport Nodes. Verify both ESXi hosts (or Edge VMs) show 'READY' state. Click each node and check: (a) CCP connectivity (should be READY), (b) VNI pool status (should have available VNIs), (c) TEP (tunnel endpoint) IP address and reachability.

Transport nodes show: ESXi-1 (READY, CCP: READY, TEP: 192.168.100.10/24), ESXi-2 (READY, CCP: READY, TEP: 192.168.100.11/24). Both TEPs should be on the same subnet and reachable via vmkping.
TEP reachability is prerequisite for NSX tunneling. If TEP is unreachable, all NSX traffic fails. Note TEP IP and VLAN for later CLI diagnostics.
Step 7

Document your complete topology map and control plane architecture in a written summary. Include: (1) ASCII diagram or reference to your saved diagram, (2) Data plane component list and function, (3) Control plane component list and communication path, (4) Baseline state of all components (READY), (5) Potential failure points and impact (if DFW is DOWN, which VMs are affected? If T0 SR fails, which traffic is rerouted?).

Written summary (500+ words) describing the complete NSX topology. Example: 'If NSX Manager CCP link is lost, the LCP on each transport node continues operating with cached state. New DFW rules cannot be pushed, but existing rules remain active. This is a data plane resilience feature: control plane is HA, but data plane degrades gracefully if control plane is unavailable (read-only mode).'
This summary demonstrates to VCDX panelists that you understand NSX as a system, not just a checklist of components.

Validation Gate

Check: Complete topology map created with identified data plane and control plane components

Expected: Lab notes contain detailed topology diagram, component-by-component analysis, documented baseline state, and written summary of NSX architecture

Common Errors

Treating all NSX routers as equivalent. T1 routers are per-tenant or per-app-tier; T0 routers are per-cluster backbone. They have different roles and failure implications.
Fix: Remember: T1 is local (VPC-like), T0 is global (core router-like). A failed T1 affects one app; a failed T0 affects all traffic.
Ignoring the control plane entirely. Focus solely on data plane 'packet flow' without understanding how rules are distributed and cached.
Fix: Control plane issues often manifest as data plane symptoms. Before troubleshooting packet drops, always confirm control plane is healthy (Manager UP, CCP READY, LCP synchronized).
Assuming all traffic traverses T0. In NSX, T1-to-T1 traffic on the same T0 may use an East-West Gateway (EWG) and never touch the T0 data plane. This is an optimization but can be a surprise.
Fix: Verify T1 router configuration: Does it have a T0 uplink? Is EWG enabled? Document the actual path, not the assumed path.

Task 2 Troubleshoot with NSX Traceflow

Traceflow is the most powerful NSX troubleshooting tool for data plane issues. Panelists will ask: (1) How does Traceflow simulate a real packet? (2) What is the difference between 'DELIVERED' and 'BLOCKED'? (3) If Traceflow succeeds but real traffic fails, what could be the issue? (4) Can Traceflow detect transient connectivity issues or only sustained failures?

Master the Traceflow tool to diagnose connectivity issues and identify the exact point where packets are dropped (DFW, T1 router, T0 router, TEP failure, etc.). Learn to interpret Traceflow output to correlate drops with configuration errors.

Step 1

Open NSX Manager and navigate to Troubleshooting > Traceflow. Click 'New Trace' to create a new trace from VM-A to VM-B.

Traceflow UI opens with source/destination VM selector. Both your test VMs (VM-A, VM-B) should be available in the dropdown.
Traceflow requires VMs to be registered with NSX. If VMs don't appear, they may not be on NSX-managed segments. Verify VM placement first.
Step 2

Configure the trace: Source = VM-A (on app-segment), Destination = VM-B (on db-segment). Protocol = TCP. Source Port = any, Destination Port = 443 (or any port where VM-B is listening). Action = 'Run Trace'.

Traceflow executes and returns one of: (1) DELIVERED - packet traversed the full path successfully, (2) BLOCKED - packet blocked at a layer (DFW, routing, etc.), (3) DROPPED - packet dropped at a component.
DELIVERED means all layers passed the packet. BLOCKED means a security policy explicitly denied it. DROPPED is less common but indicates unexpected behavior (e.g., routing failure).
Step 3

If trace is DELIVERED, examine the detailed hop-by-hop results: (1) Packet enters DFW at source (rule applied, allowed). (2) Packet switches through app-segment (MAC learned, forwarded to T1 DR port). (3) T1-app DR routes to T0 (route lookup succeeded, next-hop is T0 DR). (4) T0 DR routes to T1-db (return route lookup succeeded). (5) T1-db DR routes to db-segment (MAC lookup, forwarded to VM-B port). (6) DFW at destination checks rule (allowed). (7) Packet delivered to VM-B.

Traceflow hops document: [DFW@app: ALLOW] → [Switch@app: FORWARD] → [T1-app DR: ROUTE to T0] → [T0 DR: ROUTE to T1-db] → [T1-db DR: ROUTE to db-segment] → [Switch@db: FORWARD] → [DFW@db: ALLOW] → [DELIVERED]. Latency estimate shown (typical: 2-5ms for NSX logic).
Each hop is logged with rule ID, route entry, MAC address. If a hop fails, note the exact layer (e.g., 'DFW rejected at ingress due to rule <security-policy-ID>').
Step 4

Simulate a failure scenario: In NSX Manager, navigate to Networking > Security > Distributed Firewall. Find the DFW rule allowing app-to-db traffic. Create a NEW rule ABOVE it that explicitly DENIES the same traffic (to simulate misconfiguration). Save the rule.

DFW rule added in priority order: Rule-1 (DENY app-to-db), Rule-2 (ALLOW app-to-db). Rules are evaluated top-to-bottom; Rule-1 matches first and drops traffic.
DFW rule order matters. This is a common misconfiguration: admin adds a blanket DENY rule and forgets to adjust priorities.
Step 5

Re-run Traceflow (VM-A to VM-B, TCP port 443). Trace will now show BLOCKED at the DFW ingress layer. Examine the detailed output: 'DFW@app: BLOCK (rule <ID>, policy <POLICY-NAME>)'.

Traceflow output: [DFW@app: BLOCK (rule vcp-support-07.deny-app-db, policy app-deny-policy)] → [BLOCKED]. The trace identifies the exact rule blocking traffic. Rule ID can be used to locate the misconfigured rule in NSX Manager.
Real troubleshooting: When a user reports 'can't reach VM-B from VM-A,' run Traceflow. It immediately identifies the blocking layer and rule. Fix the rule (delete or reprioritize) and re-run Traceflow to confirm.
Step 6

Fix the DFW rule by deleting the deny rule or moving the allow rule above it. Re-run Traceflow to confirm the trace now shows DELIVERED.

Trace reverts to DELIVERED. This demonstrates the fix and validates the remediation.
In a real incident, you'd communicate with the app team: 'The DFW rule policy-XYZ was blocking traffic. I've corrected the rule priority. Traffic should now flow. Please test from your application and confirm.'
Step 7

Simulate a routing failure: In NSX Manager, navigate to Networking > Connectivity > Tier-0 > Routing > Static Routes. Find the route to db-segment (if using static routing) or use dynamic routing to examine the OSPF/BGP table. Delete the route temporarily to simulate routing failure.

Route removed. Next Traceflow attempt from VM-A to VM-B will show DROPPED or BLOCKED at the T1-app DR routing layer.
Routing failures are less common in NSX (OSPF/BGP is dynamic) but can occur if routes are manually deleted or exported incorrectly.
Step 8
Re-run Traceflow. Observe the result: [T1-app DR: DROP (no route to <db-segment-subnet>)] → [DROPPED]. This identifies the routing issue. Restore the route, re-run Traceflow to confirm recovery.
Traceflow shows DROP at routing layer. Pinpoints exact layer (T1 vs. T0 router). After restoring route, trace shows DELIVERED.
Traceflow is deterministic: it simulates an exact packet and shows exactly where it would drop. Use this to build confidence in your configuration changes before deploying to production.
Step 9
Document your Traceflow methodology: (1) Always run Traceflow bidirectionally (A→B and B→A) if traffic is bi-directional. (2) Test common protocols (TCP/UDP, common ports 80, 443, 3306). (3) If Traceflow succeeds but real traffic fails, suspect MTU issues, TCP option stripping, or asymmetric routing. (4) Traceflow does NOT account for packet size; large packets may fail in real world due to MTU but succeed in Traceflow.
Documented methodology: 'Traceflow Test Procedure: (1) Bidirectional traces (A→B, B→A). (2) Common protocols: TCP 80, 443; UDP 53, 123. (3) Success criteria: all traces DELIVERED and rules are expected. (4) If real traffic fails despite DELIVERED trace, run: vmkping (TEP connectivity), tcpdump on VTEP (packet capture), guest OS routing table (ensure guest has return route).'
This methodology is your Traceflow troubleshooting SOP (Standard Operating Procedure). Document it for your lab reference and panel discussion.

Validation Gate

Check: Traceflow used to diagnose multiple connectivity scenarios (baseline, DFW block, routing failure)

Expected: Lab notes document Traceflow output for each scenario, interpreted drop reasons, and remediation steps

Common Errors

Relying solely on Traceflow for diagnosis. Traceflow is a simulator; it does NOT account for MTU, packet fragmentation, TCP options, or asymmetric return paths.
Fix: If Traceflow succeeds but real traffic fails, use tcpdump/Wireshark to capture actual packets. Compare packet size, TCP flags, and return path in real network.
Assuming DELIVERED means the application is working. Traceflow validates L3-L7 policy and routing, but doesn't test application-level protocols (is port 443 actually listening? Is TLS handshake working?).
Fix: DELIVERED trace is necessary but not sufficient. Follow up with: SSH to VM-B, verify listening ports (netstat -an), test connectivity (telnet <VM-B-IP> 443, curl, nc).
Not using Traceflow bidirectionally. Traffic A→B succeeds but B→A fails due to asymmetric routing or reverse DFW rule misconfiguration.
Fix: Always test both directions. Document which direction fails. Asymmetric failures often indicate return path issues or missing stateful DFW rules.

Task 3 CLI-Based Troubleshooting and Transport Node Diagnostics

CLI troubleshooting is the gold standard for deep diagnostics. Panelists will ask: (1) How do you validate that LCP has synchronized the DFW rules from Manager? (2) What does a healthy BFD session look like? (3) How would you diagnose a TEP network partition? (4) Can you correlate an nsxcli 'get logical-switch' output with a Traceflow drop?

Master nsxcli commands and lower-layer diagnostics to validate control plane sync, routing tables, firewall state, and TEP connectivity when UI-based tools are unavailable or slow.

Step 1

SSH to one of your transport nodes (ESXi host or Edge VM). Log in as root or a user with admin privileges.

SSH session established. Prompt shows: '[root@<host>]# ' or similar, ready for commands.
Ensure you have nsxcli available. On ESXi, nsxcli is in the default PATH. On Edge VMs, it may be in /opt/vmware/nsx/bin/ or bundled with the system.
Step 2

Run 'get managers' to verify the host can reach NSX Manager cluster. Output should show Manager node(s) with status CONNECTED.

Output: 'Manager IP: 192.168.100.200, Status: CONNECTED, Last Update: 30s ago'. If status is DISCONNECTED or UNKNOWN, the LCP lost connection to Manager. This is a critical control plane issue.
If Manager is DISCONNECTED, check network connectivity: vmkping <Manager-IP>. If vmkping fails, network path to Manager is down. If vmkping succeeds, suspect Manager process issue (NSX Manager app may be down or overloaded).
Step 3

Run 'get logical-switches' to list logical switches on this host. Output should show app-segment and db-segment (or whichever segments are connected to this host's VMs).

Output: '1. app-segment (UUID: <UUID>, VLAN: 0, VNI: 5000, State: UP, MAC Entries: 3). 2. db-segment (UUID: <UUID>, VLAN: 0, VNI: 5001, State: UP, MAC Entries: 2)'. Each segment shows VNI (virtual network ID assigned by NSX), state (UP=healthy, DOWN=control plane issue), and MAC table size.
MAC Entries = number of MAC addresses learned on this segment. If you have 3 VMs on the segment, expect ~3 MAC entries. If 0, segment is not forwarding traffic.
Step 4

Run 'get logical-routers' to verify T1 and T0 routers are visible to this host. Output should list all T1 routers and the T0 router.

Output: '1. T1-app (UUID: <UUID>, Type: TIER1, Status: UP). 2. T1-db (UUID: <UUID>, Type: TIER1, Status: UP). 3. T0 (UUID: <UUID>, Type: TIER0, Status: UP)'. Each router shows its type and operational status.
If a T1 router shows status DOWN, its LCP is not operational on this host. This would indicate control plane sync failure or a LCP process crash.
Step 5

Run 'get logical-router-port T1-app' (or similar) to examine routing table entries for the T1-app router. Look for routes to db-segment subnet and upstream routes to T0.

Output: 'T1-app routing table: 1. Network: 10.2.0.0/24 (db-segment), Next Hop: T0 (via <T0-DR-IP>), Status: UP. 2. Network: 192.168.1.0/24 (app-segment), Next Hop: Direct (local), Status: UP'. Each route shows destination subnet and next-hop, validating routing configuration.
If a route shows status DOWN or missing (e.g., route to db-segment not listed), it indicates a control plane sync failure. The LCP cache is stale or incomplete.
Step 6

Run 'get firewall-rules' to list DFW rules installed on this transport node. Output should show the security policy (e.g., policy-app-db) and its rules.

Output: 'Firewall Policy: app-db-policy (UUID: <UUID>). Rule 1: ALLOW TCP/443 from app-segment to db-segment (Status: INSTALLED). Rule 2: DENY all (Status: INSTALLED)'. Each rule shows its match criteria and install status (INSTALLED = active and cached locally, PENDING = waiting for Manager push, ERROR = sync failure).
If rules show PENDING, the host is waiting for Manager to push the rule. If PENDING persists for > 5 minutes, suspect Manager-to-LCP communication issue. If ERROR, the rule is malformed or incompatible with this host version.
Step 7

Run 'get tunnel-interfaces' or 'get vtep-interfaces' to list tunnel endpoints (TEP) on this host and their operational status. Output should show TEP IP address and adjacency count (how many other hosts this host is tunneled to).

Output: 'VTEP Interface: vmk10 (IP: 192.168.100.10, Netmask: 255.255.255.0, MTU: 1500, Active Tunnels: 1 (to 192.168.100.11))'. Active Tunnels > 0 indicates connectivity to other hosts. MTU 1500 is standard (ESXi may use 1600 with jumbo frames).
If Active Tunnels = 0, this host is isolated from other hosts in the NSX cluster. Data plane traffic between VMs on different hosts will fail. Check TEP network connectivity: vmkping 192.168.100.11 (the other host's TEP).
Step 8

Run 'get bfd-sessions' to check Bidirectional Forwarding Detection (BFD) sessions to other transport nodes. BFD is NSX's keep-alive mechanism for tunnel health. Output should show all BFD sessions as UP.

Output: 'BFD Session 1: Peer 192.168.100.11 (T0 tunnel), Status: UP, Tx Interval: 1000ms, Rx Interval: 1000ms, Multiplier: 3. BFD Session 2: Peer 192.168.100.12 (T1 tunnel), Status: UP'. Status UP = tunnel is healthy and keep-alives are exchanged every 1000ms.
If BFD shows DOWN, the tunnel to that peer is broken. This is a critical data plane failure. Correlate with vmkping to TEP and network diagnostics (network partition, firewall blocking UDP 3784).
Step 9

Run 'vmkping <other-host-TEP-IP>' to verify TEP network connectivity. Example: 'vmkping 192.168.100.11'. Output should show successful pings (0% loss).

Output: 'PING 192.168.100.11 (192.168.100.11) 56(84) bytes of data. 64 bytes from 192.168.100.11: icmp_seq=1 ttl=64 time=2.34ms. 0% loss, avg latency 2.3ms'. Successful vmkping confirms network path is operational.
If vmkping fails (100% loss), the TEP network is partitioned. Check ESXi network stack: esxcli network nic list (verify vmk10 is UP), esxcli network route ipv4 list (verify route to TEP subnet). If network is healthy but vmkping fails, suspect firewall policy on network switch blocking traffic.
Step 10

For Edge VMs (if used for T0 or T1 services), run 'get controllers' to verify the Edge VM's connection to NSX Manager's Control Plane. Output should show CONNECTED.

Output: 'Controller: 192.168.100.200 (nsxmanager-node-1), Status: CONNECTED, Role: central-control-plane'. If DISCONNECTED, the Edge VM cannot receive configuration updates from Manager.
Edge VMs are mission-critical. If Edge-to-Manager link fails, all T0/T1 services on that Edge VM stop functioning. Always keep at least 2 Edge VMs for HA.
Step 11

Document your nsxcli diagnostics workflow: (1) Check Manager connectivity (get managers). (2) List logical switches/routers (get logical-switches, get logical-routers). (3) Inspect routing tables (get logical-router-port). (4) Verify firewall rules (get firewall-rules, check install status). (5) Validate TEP connectivity (get tunnel-interfaces, vmkping). (6) Confirm BFD health (get bfd-sessions).

Documented workflow in lab notes. Example: 'NSX CLI Troubleshooting SOP: (1) Manager UP? (2) Switches/Routers UP? (3) Routes valid? (4) Firewall rules INSTALLED? (5) TEP reachable (vmkping)? (6) BFD sessions UP? If any step fails, escalate to that layer's issue.'
This workflow is your CLI troubleshooting checklist. Use it during incidents to systematically validate control and data plane health.

Validation Gate

Check: CLI commands executed on transport node(s) with documented output and interpretation

Expected: Lab notes contain nsxcli command output for logical switches, routers, firewall rules, TEP interfaces, BFD sessions, and documented health assessment

Common Errors

Misinterpreting 'get firewall-rules' status. INSTALLED means the rule is active on this host, but doesn't mean it was pushed from Manager recently. A cached rule from days ago still shows INSTALLED until the policy is updated.
Fix: Check Manager logs to confirm the rule version matches. Use 'get firewall-rules verbose' (if available) to see rule push timestamp. If rule timestamp is old, suspect Manager is not pushing updates.
Assuming TEP network is OK if vmkping to one peer succeeds. TEP might have asymmetric paths (A→B works, B→A fails).
Fix: vmkping is ICMP-based. NSX uses VXLAN (UDP 4789) for tunnels and UDP 3784 for BFD. Firewall may block BFD but allow ICMP. Use 'get bfd-sessions' to verify tunnel health is actually UP, not just assumed.
Not distinguishing between 'Manager disconnected' (control plane failure) and 'Rule sync failure' (data plane stale but operational). A disconnected Manager is critical; stale rules are recoverable.
Fix: If Manager is disconnected, data plane uses cached rules (read-only mode). If Manager is connected but 'get firewall-rules' shows PENDING for hours, suspect Manager → LCP link issue or Manager is overloaded.

Task 4 Remediation, Validation, and RCA Documentation

This task emphasizes the full troubleshooting cycle: diagnose → fix → validate → document. Panelists will ask: (1) How do you ensure a fix doesn't break something else? (2) How long do you test the fix before declaring the incident closed? (3) What is the difference between a temporary workaround and a permanent fix? (4) How would you automate this troubleshooting in a CI/CD pipeline?

Apply fixes to identified issues, validate remediation with Traceflow, and document root cause analysis in a structured format suitable for runbooks and incident post-mortems.

Step 1

Based on your diagnostics from Tasks 2-3, identify the root cause of the connectivity issue (e.g., DFW rule misconfiguration, routing table incomplete, TEP network unreachable). Document the exact issue in a brief RCA statement.

RCA statement: 'Traffic from VM-A to VM-B fails due to missing DFW allow rule in policy app-to-db. Traceflow showed BLOCKED at DFW ingress, rule ID <XXXX>. nsxcli get firewall-rules shows rule status ERROR (malformed CIDR block). Root cause: DFW rule syntax error during initial deployment.'
Distinguish root cause from symptom. Symptom: 'User says app is slow.' Root cause: 'DFW rule drop rate is 50% due to misconfigured range.'
Step 2

Apply the fix. Example: If DFW rule is misconfigured, correct the rule in NSX Manager > Networking > Security > Distributed Firewall. Edit rule, fix CIDR block, save and publish.

DFW rule updated and published. NSX Manager confirms rule is ACTIVE (published to Manager cluster). LCP on transport nodes will receive updated rule within seconds.
Always test rule changes in a staging environment first (if available). In a VCDX interview, emphasize change control: 'I would create a change request, have peer review the rule, test in staging, then deploy to production during maintenance window.'
Step 3
Monitor the rule push process. Run 'get firewall-rules' on transport nodes to confirm rule status changes from PENDING → INSTALLED. Allow 30-60 seconds for sync.
nsxcli output shows rule now INSTALLED on transport nodes. Latency between fix publication and INSTALLED status is typically 5-15 seconds.
If INSTALLED status doesn't appear within 60 seconds, suspect Manager-to-LCP link issue. Check 'get managers' on transport node; if status is DISCONNECTED, diagnose Manager connectivity first.
Step 4

Re-run Traceflow from VM-A to VM-B with the same parameters as before (TCP, port 443). Trace should now show DELIVERED (all hops pass, no blocks).

Traceflow output: [DFW@app: ALLOW] → [... (hops)] → [DELIVERED]. Compare this output to the pre-fix trace to confirm the fix resolved the specific drop point.
If trace still shows BLOCKED at a different layer (e.g., T1 DR routing), there are multiple issues. Continue troubleshooting to resolve all identified problems before considering the incident closed.
Step 5

Validate the fix at the application layer. SSH to VM-B and verify the service is listening on port 443. From VM-A, attempt to connect to VM-B: 'telnet <VM-B-IP> 443' or 'curl https://<VM-B-IP>' (if web service). Connection should succeed.

Connection succeeds. Example: 'Connected to 192.168.2.10:443. HTTP 200 OK' or similar application-level success. This confirms not just network connectivity (Traceflow), but end-to-end functionality.
Some teams rely solely on Traceflow validation. Always do application-level testing to confirm the user-reported issue is actually resolved.
Step 6

Document the RCA in a structured format: (1) Incident Date/Time. (2) Symptom (user report or automated alert). (3) Timeline (when incident was detected, when troubleshooting started, when fix was applied). (4) Root Cause (technical explanation). (5) Impact (how many VMs/users affected, downtime duration). (6) Fix Applied (specific configuration change). (7) Validation (Traceflow + application test). (8) Prevention (how to avoid this in the future).

RCA document (500+ words): 'Incident: VM-A → VM-B connectivity loss on 2024-12-15 10:30 UTC. User Report: 'Application server cannot reach database.' Symptom: Traceflow VM-A → VM-B showed BLOCKED at DFW (rule XYZ). Root Cause: DFW rule deployed with typo in destination CIDR (192.168.2.0/24 instead of 192.168.20.0/24). Impact: 1 application server, 5000 users affected, ~45-minute downtime. Fix: Corrected CIDR in DFW rule, re-published. Validation: Traceflow DELIVERED, application connectivity restored. Prevention: Implement DFW rule deployment pipeline with syntax validation and staging environment testing.'
This RCA format is industry-standard (used by Google SRE, AWS, etc.). Panelists will recognize and appreciate the structure.
Step 7

Create a troubleshooting runbook entry based on this incident. Runbook should be a step-by-step guide that a junior engineer or NOC (Network Operations Center) tech can follow to diagnose the same issue in the future.

Runbook entry: 'NSX Connectivity Troubleshooting Runbook > DFW Rule Misconfiguration. Symptom: Ping fails between VMs on different T1s. Step 1: Run Traceflow. If BLOCKED at DFW ingress, proceed to Step 2. Step 2: Get rule ID from Traceflow. Step 3: SSH to transport node, run `get firewall-rules | grep <rule-ID>`. Check rule status and CIDR. Step 4: If CIDR is wrong, escalate to Network Security team to correct. Step 5: After fix, re-run Traceflow to confirm DELIVERED. Step 6: Test application connectivity (ping, telnet, curl). Success = incident closed.'
Runbooks are VCDX-gold. Panelists want to see evidence that you're building organizational knowledge from incidents, not just fixing today's problem.
Step 8

If the issue was a design flaw (not a one-off misconfiguration), recommend architectural changes. Example: If T1 routers lack redundancy, design and propose a HA T1 setup. If TEP network is single-home, design a dual-homed TEP network.

Architectural recommendation: 'To prevent future T1 router failures, deploy 2 T1 routers in HA pair with stateful failover. Add a second T1 SR Edge VM. Implement BFD health checks to trigger automatic failover within 3 seconds. Estimated cost: $10K hardware, 2 days implementation. Benefit: 99.9% uptime improvement for tier-1 critical apps.'
This is where you distinguish yourself as a senior engineer. VCDX panelists want to see you think beyond immediate fixes to systemic improvements.

Validation Gate

Check: Issue fixed, remediation validated with Traceflow, RCA documented, runbook created

Expected: Lab notes contain complete RCA document (symptom, root cause, fix, validation, prevention) and troubleshooting runbook entry suitable for team reference

Common Errors

Fixing the symptom without addressing the root cause. Example: Adding a static ARP entry to work around a network discovery issue, instead of fixing the underlying DHCP or VLAN misconfiguration.
Fix: Always ask 'Why?' three times before applying a fix. Why is the DFW rule misconfigured? (Typo in rule creation.) Why did the typo make it through QA? (No rule syntax validation in deployment pipeline.) True fix: Implement validation automation.
Not testing the fix thoroughly. A DFW rule fix may unblock the intended traffic but inadvertently open a security hole. Always test both positive (intended traffic flows) and negative (unintended traffic is still blocked) cases.
Fix: For DFW rules, test: (1) Correct traffic flows (bidirectional). (2) Incorrect traffic is blocked (e.g., wrong IP range still blocked). (3) No unintended side effects (other rules still working).
Declaring the incident closed too early. Fix is applied, but validation is incomplete (only Traceflow tested, not actual user traffic). Weeks later, user reports issue persists.
Fix: Validation checklist: (1) Traceflow DELIVERED. (2) Application-level test successful (telnet, curl, etc.). (3) Full workload test (run user workflow end-to-end). (4) Monitor for 24 hours (verify fix is stable, no anomalies). Only then close incident.

Final Validation

Lab completed when all 4 tasks are executed and fully documented with topology maps, Traceflow tests, CLI diagnostics, RCA, and runbooks.

✓ Complete NSX topology mapped with data plane and control plane identified → Task 1 complete: diagram shows all components (VM, DFW, logical switches, T1/T0 routers, transport nodes, Manager, CCP, LCP)

✓ Traceflow used to diagnose baseline, failure, and remediation scenarios → Task 2 complete: Traceflow traces documented showing DELIVERED (healthy), BLOCKED (DFW misconfiguration), DROPPED (routing failure), and remediated DELIVERED

✓ nsxcli commands executed to validate control and data plane state → Task 3 complete: Output from get managers, get logical-switches, get logical-routers, get firewall-rules, get tunnel-interfaces, get bfd-sessions, vmkping documented and interpreted

✓ Issue remediated and validated with Traceflow + application testing → Task 4 complete: RCA document (symptom, root cause, fix, impact, prevention), troubleshooting runbook entry, and architectural recommendations for future prevention

Cleanup / Restore

• Remove any temporary DFW rules created during testing (deny rules added for failure simulation)

• Restore any deleted logical switches, routers, or routes (use NSX Manager snapshots if available)

• Verify VM-A and VM-B are back to baseline state (connectivity DELIVERED, no policy blocks)

• Document any permanent architectural changes recommended (for infrastructure team to implement)

Design Reflection (VCDX)

This lab emphasizes NSX troubleshooting methodology and operational excellence. VCDX panelists will probe: (1) How do you debug control plane issues when UI is slow or unavailable? (2) What is your mental model of NSX data flow and where failures occur? (3) Can you prioritize between control plane (Manager, CCP/LCP) failures and data plane (DFW, routing, TEP) failures? (4) How do you design NSX monitoring to detect these issues proactively? The RCAR framework emphasizes root cause analysis and architectural thinking, not just tactical fixes.

Requirements

  • NSX 4.x Manager cluster (minimum 1 node, 3+ recommended for HA testing)
  • 2+ transport nodes (ESXi hosts or Edge VMs) for tunnel/TEP testing
  • 2 logical segments (app-segment, db-segment)
  • 2 tier-1 routers (T1-app, T1-db) and 1 tier-0 router (T0)
  • Test VMs: VM-A on app-segment, VM-B on db-segment
  • Firewall rules defining app-to-db policy
  • SSH/CLI access to transport nodes for nsxcli commands

Constraints

  • NSX Traceflow only works for VMs registered with NSX; bare-metal or external devices require packet capture (tcpdump) instead
  • DFW rule changes have 5-15 second propagation latency from Manager to transport nodes; tests must account for this sync time
  • TEP network must be routable; if TEP network is not L2-adjacent, tunneling requires L3 routing setup (may not be available in lab)
  • BFD/tunnel health is binary (UP/DOWN); subtle issues like high latency or packet loss may not show as DOWN immediately
  • nsxcli commands vary by host type (ESXi vs. Edge VM); command syntax differs between NSX versions

Assumptions

  • NSX Manager and transport nodes have network connectivity (at baseline)
  • Lab operator has admin-level access to NSX Manager and transport node CLIs
  • VMs can reach each other if no security policies interfere (i.e., baseline connectivity works before introducing deliberate failures)
  • Firewall policies are deployed and synchronized before the lab starts (Manager → LCP sync is operational)
  • RCA and runbook documentation are stored in a centralized repository (Confluence, GitHub, etc.) for team access

Risks

  • Risk: Breaking actual connectivity during failure simulation. If you delete a real route or firewall rule, production traffic could fail. — Mitigation: Use a dedicated lab environment, not shared infrastructure. Snapshot the NSX Manager config before destructive tests. Have a revert plan (snapshot or config backup).
  • Risk: Misinterpreting Traceflow results. A DELIVERED trace does not guarantee real traffic will succeed (MTU, TCP options, asymmetric paths). — Mitigation: Always follow Traceflow with application-level testing (telnet, curl, packet capture). Document limitations of Traceflow in your runbook.
  • Risk: Control plane issues (Manager down, LCP disconnected) cascading to data plane. If Manager is unavailable, LCP goes read-only; new policies cannot be deployed. — Mitigation: Test with Manager deliberately unavailable (simulate network partition). Understand graceful degradation: data plane uses cached rules, services degrade but do not fail completely.
  • Risk: TEP network partition causing split-brain. Host A can reach Host B, but Host B cannot reach Host A due to asymmetric routing. Traceflow from A→B succeeds, but real traffic fails both ways. — Mitigation: Test bidirectional paths. Use vmkping and BFD status to confirm tunnel health, not just one-way connectivity.

Self-Assessment Discussion Prompts

  1. In the DFW misconfiguration scenario, you identified the issue via Traceflow. But in a real incident, what would prevent you from using Traceflow (e.g., VMs not NSX-registered, NSX Manager down)? What is your backup diagnostic method?
  2. You validated the fix with Traceflow DELIVERED + application test. But a week later, user reports intermittent failures (traffic works sometimes, fails sometimes). Traceflow always shows DELIVERED. What could cause transient failures that Traceflow doesn't catch?
  3. In your RCA, you recommended adding a second T1 SR Edge VM for HA. What is the cost/benefit? What are the design trade-offs? How would you validate that HA failover actually works under load?
  4. Your lab uses static routing (T1 → T0). In production, would you use dynamic routing (OSPF, BGP)? What are the pros and cons? How does dynamic routing change your troubleshooting approach?
  5. You ran Traceflow and got DELIVERED, but nsxcli get firewall-rules showed the rule status as PENDING (not yet installed). Can Traceflow deliver packets if firewall rules are not yet installed? What is the behavior?
  6. If NSX Manager becomes unavailable (network partition), you said data plane goes read-only. What happens to new traffic if a new VM is spun up? Is that VM reachable without Manager? When does Manager become critical again?
  7. You used vmkping to validate TEP network. But vmkping is ICMP; NSX uses VXLAN (UDP 4789). Can network policies block VXLAN while allowing ICMP? How would you detect that scenario?
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.