NSX Micro-Segmentation Design
Objectives
- Design NSX network segments with proper CIDR allocation and Tier-0/Tier-1 routing
- Create a Distributed Firewall rule set using a zero-trust, allow-list approach
- Design BGP peering between NSX Edge and physical network
- Plan east-west and north-south traffic flow with security enforcement points
- Design logging, audit, and compliance verification using Traceflow and flow monitoring
Prerequisites
Access to NSX 9.0 and VCF networking documentation
Prior labs: vcp-architect-01, vcp-architect-02
Required skills:
- TCP/IP networking and subnetting
- Firewall rule concepts (stateful inspection, rule ordering)
- BGP fundamentals
- NSX overlay networking basics (Geneve)
Lab Environment
Design exercise — validate DFW rules on Holodeck NSX environment if available
Tasks
Task 1 Network Segment Design & IP Address Planning
Design the overlay network topology with proper segment isolation, CIDR allocation, and Tier-0/Tier-1 gateway architecture.
Design Tier-0 and Tier-1 gateway topology:
NSX gateway hierarchy in VCF 9.0:
- Tier-0 Gateway (Provider Router): Connects NSX overlay to physical network
- 1 per workload domain (or shared across domains via VRF)
- Provides north-south routing, NAT, VPN, BGP peering
- HA mode: Active-Standby (stateful services) or Active-Active (ECMP, no stateful NAT)
- Tier-1 Gateway (Tenant Router): Connects segments to Tier-0
- 1+ per application tier or tenant
- Provides segment interconnection, distributed routing, DFW attachment point
- Can run as Distributed Router (DR) for east-west only, or Service Router (SR) for NAT/LB
Design for 600-VM environment:
- Tier-0 (Management Domain): 1 Active-Standby (provides stateful NAT for management VMs)
- Tier-0 (Workload Domain): 1 Active-Active ECMP (high throughput, no NAT needed)
- Tier-1 (Web): Connected to Workload Tier-0, segments for web servers
- Tier-1 (App): Connected to Workload Tier-0, segments for application servers
- Tier-1 (DB): Connected to Workload Tier-0, segments for database servers
- Tier-1 (Dev): Connected to Workload Tier-0, segments for dev/test
Design CIDR addressing plan:
Overlay network supernet: 172.16.0.0/12 (avoid RFC 1918 conflicts with existing networks)
| Segment | CIDR | Gateway | Tier-1 | VMs | Purpose |
|---|---|---|---|---|---|
| Web-Prod-A | 172.16.1.0/24 | .1 | T1-Web | 30 | Production web servers |
| Web-Prod-B | 172.16.2.0/24 | .1 | T1-Web | 20 | Production web servers (overflow) |
| App-Prod | 172.16.10.0/24 | .1 | T1-App | 80 | Application servers |
| App-API | 172.16.11.0/24 | .1 | T1-App | 40 | API gateway servers |
| DB-Prod | 172.16.20.0/24 | .1 | T1-DB | 30 | Database primary + replica |
| DB-Archive | 172.16.21.0/24 | .1 | T1-DB | 10 | Archive/backup databases |
| Dev-General | 172.16.100.0/23 | .1 | T1-Dev | 150 | Development environment |
| Test-QA | 172.16.102.0/24 | .1 | T1-Dev | 70 | QA/staging |
| Mgmt-Infra | 172.16.200.0/24 | .1 | T1-Mgmt | 20 | Infrastructure management VMs |
Design principles:
- /24 per segment unless >254 VMs needed (then /23)
- Leave gaps between segment ranges for future expansion
- Aggregate at Tier-1 level for BGP advertisement efficiency
- Example: T1-Web advertises 172.16.0.0/20 (covers .1.0-.15.0 range)
Design transport zone and VTEP configuration:
Transport zones (VCF 9.0):
- Overlay TZ (default): All host TEPs participate; carries workload segments
- VLAN TZ: For edge uplinks connecting to physical network
VTEP (Virtual Tunnel Endpoint) configuration:
- VTEP subnet: 192.168.250.0/24 (dedicated, not routable externally)
- VTEP IP pool: 192.168.250.10 - .200 (supports up to 191 host TEPs)
- Teaming policy: Load Balance Source (for multi-pNIC hosts)
Geneve encapsulation:
- Overhead: ~54 bytes (Outer Ethernet 14 + Outer IP 20 + Outer UDP 8 + Geneve header 8 + options ~4)
- Required MTU: Inner frame 1500 + 54 = 1554 minimum; set 1700 for headroom
- Physical switch MTU: 9000 (jumbo frames) recommended for all transport network paths
Verification:
- Ping test between VTEPs: vmkping ++netstack=vxlan -d -s 1572 <remote_VTEP_IP>
- Expected: No fragmentation at MTU 1700
Design BGP peering between NSX Edge and physical network:
- Edge node deployment:
- 2 Edge VMs per Tier-0 gateway (HA pair)
- Edge VM sizing: Large (8 vCPU, 32 GB RAM) for production routing
- Placement: Anti-affinity rule — Edge VMs on separate hosts
- BGP configuration:
- NSX ASN: 65001 (private ASN range)
- Physical router ASN: 65000
- Peering: Each Edge node peers with 2 physical ToR switches (4 peering sessions total)
- BFD: Enabled by default in NSX; timer 500ms × 3 = 1.5s failover detection
- Route advertisement:
- Tier-0 advertises: Connected segments, Tier-1 connected routes, NAT IPs, LB VIPs
- Route aggregation: Advertise summary 172.16.0.0/12 (not individual /24s)
- Physical network: Accepts NSX routes, provides default route to Tier-0
- ECMP design (Active-Active Tier-0):
- Both Edge nodes advertise same routes with equal metrics
- Physical routers distribute traffic across both edges
- Throughput: 2× single edge bandwidth
- Limitation: Stateful services (NAT, LB, VPN) require Active-Standby, not ECMP
| Edge | Physical Router | NSX ASN | Phys ASN | BFD | Status |
|---|---|---|---|---|---|
| Edge-1 | ToR-A1 | 65001 | 65000 | Yes | Active |
| Edge-1 | ToR-A2 | 65001 | 65000 | Yes | Active |
| Edge-2 | ToR-B1 | 65001 | 65000 | Yes | Active |
| Edge-2 | ToR-B2 | 65001 | 65000 | Yes | Active |
Validation Gate
Check: Network design covers segment topology, CIDR plan, transport zone config, and BGP peering
Expected: Tier-0/Tier-1 hierarchy defined, 8+ segments with CIDR addressing, VTEP configuration with MTU validation, BGP peering with BFD and route advertisement plan
Common Errors
Task 2 Distributed Firewall Rule Set Design (Zero-Trust)
Design a comprehensive DFW rule set using zero-trust principles with explicit allow rules and default-deny baseline.
Design security group taxonomy:
NSX Groups (tag-based, not IP-based):
- Group: Web-Servers — Criteria: Tag = 'tier:web' AND Tag = 'env:prod'
- Group: App-Servers — Criteria: Tag = 'tier:app' AND Tag = 'env:prod'
- Group: DB-Servers — Criteria: Tag = 'tier:db' AND Tag = 'env:prod'
- Group: Dev-All — Criteria: Tag = 'env:dev'
- Group: Mgmt-Infra — Criteria: Tag = 'tier:mgmt'
- Group: PCI-Scope — Criteria: Tag = 'compliance:pci' (VMs handling payment data)
- Group: Internet-Facing — Criteria: Tag = 'exposure:internet'
Tag assignment strategy:
- Automated: vSphere tags applied during VM provisioning (Aria Automation blueprint or SDDC Manager)
- Policy: Every VM must have at minimum: tier, env, and compliance tags
- No tag = no access (default-deny catches untagged VMs)
Design DFW rule categories and policies:
NSX DFW evaluates rules top-to-bottom across categories:
1. Ethernet (L2) → 2. Emergency → 3. Infrastructure → 4. Environment → 5. Application → Default
Category 1 — Emergency (highest priority):
| # | Name | Source | Dest | Service | Action | Log |
|---|---|---|---|---|---|---|
| E-001 | Block-Quarantined | Group:Quarantined | Any | Any | Drop | Yes |
| E-002 | Allow-Emergency-Admin | Group:Emergency-Admin-IPs | Any | SSH/RDP | Allow | Yes |
Category 2 — Infrastructure:
| # | Name | Source | Dest | Service | Action | Log |
|---|---|---|---|---|---|---|
| I-001 | Allow-DNS | Any | Group:DNS-Servers | UDP/53, TCP/53 | Allow | No |
| I-002 | Allow-NTP | Any | Group:NTP-Servers | UDP/123 | Allow | No |
| I-003 | Allow-AD-Auth | Any | Group:AD-Controllers | TCP/389,636,88,445 | Allow | No |
| I-004 | Allow-vCenter-Mgmt | Group:Mgmt-Infra | Group:vCenter | TCP/443 | Allow | Yes |
| I-005 | Allow-Monitoring | Any | Group:Monitoring | TCP/9090,9100,514 | Allow | No |
Category 3 — Environment:
| # | Name | Source | Dest | Service | Action | Log |
|---|---|---|---|---|---|---|
| V-001 | Block-Dev-to-Prod | Group:Dev-All | Group:*-Prod | Any | Drop | Yes |
| V-002 | Block-Prod-to-Dev | Any (env:prod) | Group:Dev-All | Any | Drop | Yes |
Design application-tier DFW rules:
Category 4 — Application (bulk of rules):
3-Tier Web Application:
| # | Name | Source | Dest | Service | Action | Log | Notes |
|---|---|---|---|---|---|---|---|
| A-001 | External-to-Web | Group:Load-Balancer-VIP | Group:Web-Servers | TCP/443 | Allow | Yes | HTTPS only |
| A-002 | Web-to-App | Group:Web-Servers | Group:App-Servers | TCP/8080,8443 | Allow | Yes | App APIs |
| A-003 | App-to-DB | Group:App-Servers | Group:DB-Servers | TCP/3306,5432 | Allow | Yes | MySQL/Postgres |
| A-004 | DB-Replication | Group:DB-Servers | Group:DB-Servers | TCP/3306 | Allow | No | DB cluster |
| A-005 | App-to-Cache | Group:App-Servers | Group:Cache-Servers | TCP/6379 | Allow | No | Redis |
| A-006 | Web-to-CDN | Group:Web-Servers | Group:CDN-Egress | TCP/443 | Allow | No | Static assets |
PCI-scoped rules (stricter logging):
| # | Name | Source | Dest | Service | Action | Log | Notes |
|---|---|---|---|---|---|---|---|
| P-001 | PCI-Web-to-Payment | Group:PCI-Scope AND Web | Group:PCI-Scope AND App | TCP/8443 | Allow | Yes | PCI traffic |
| P-002 | PCI-App-to-TokenDB | Group:PCI-Scope AND App | Group:PCI-Scope AND DB | TCP/5432 | Allow | Yes | Tokenized data |
| P-003 | Block-PCI-Egress | Group:PCI-Scope | NOT Group:PCI-Scope | Any | Drop | Yes | No PCI data leakage |
Default rule (last in each category):
| # | Name | Source | Dest | Service | Action | Log |
|---|---|---|---|---|---|---|
| D-001 | Default-Deny-All | Any | Any | Any | Drop | Yes |
Document rule management best practices:
- Rule count management:
- Target: <200 rules total across all categories
- Use groups extensively (tag-based) rather than individual VM/IP rules
- Each new application: 3-5 rules (inbound, inter-tier, outbound, logging)
- Rule lifecycle:
- New rule request: Change management approval required
- Rule review: Quarterly audit of all rules for relevance
- Unused rule detection: NSX Flow Monitoring → rules with 0 hits in 90 days → decommission candidate
- Rule versioning: Export DFW config before every change (NSX API backup)
- Applied-to scope:
- CRITICAL: Always set 'Applied To' on each rule or section
- Without 'Applied To': Rule is applied to EVERY vNIC in the environment (performance impact)
- Example: Web-to-App rule 'Applied To' = Group:Web-Servers + Group:App-Servers
- This pushes the rule only to hosts running those VMs, not all hosts
- Exclusion list:
- VMs that should bypass DFW: NSX Manager, vCenter, SDDC Manager
- Reason: DFW interruption could lock out management access
- Configure: NSX Manager UI → Security → Distributed Firewall → Exclusion List
Validation Gate
Check: DFW rule set designed with zero-trust approach across all categories
Expected: 15+ rules organized by NSX categories (Emergency/Infrastructure/Environment/Application), tag-based groups, default-deny, PCI isolation rules, <200 total rule count target
Common Errors
Task 3 Traffic Flow Analysis & Security Verification
Map all traffic flows through the network and verify DFW enforcement using Traceflow and flow monitoring.
Map east-west traffic flows:
Traffic flow diagram for 3-tier web application:
1. Client → Load Balancer (north-south, enters via Tier-0 Edge): - Ingress: Physical network → BGP → Tier-0 → LB VIP
- DFW enforcement: None (traffic not yet on overlay)
2. Load Balancer → Web Server (east-west within segment):
- LB distributes to web-prod-a segment VMs
- DFW: Rule A-001 allows TCP/443 from LB VIP group to Web-Servers group
- Enforcement point: DFW on destination host's dvPort (kernel module)
3. Web → App (east-west across segments): - Traffic: 172.16.1.x → 172.16.10.x (different segments, same Tier-1)
- Routing: Distributed Router (DR) on source host routes to App segment
- DFW: Rule A-002 allows TCP/8080 from Web-Servers to App-Servers
- Note: DR routing is in-kernel — traffic does NOT hairpin through Edge
4. App → DB (east-west across Tier-1 boundaries): - Traffic: 172.16.10.x → 172.16.20.x (different Tier-1 gateways) - Routing: Source DR → Tier-0 → Destination DR (inter-Tier-1 routing)
- DFW: Rule A-003 allows TCP/3306 from App-Servers to DB-Servers
5. DB → DB Replication (east-west within segment): - Traffic: 172.16.20.x → 172.16.20.y (same segment)
- No routing needed (L2 within segment)
- DFW: Rule A-004 allows TCP/3306 within DB-Servers group
Document each flow with: source, destination, port, routing path, DFW rule hit, and enforcement host.
Design Traceflow verification plan:
Traceflow: NSX built-in diagnostic tool that injects a tagged packet and traces its path through the overlay network.
Test plan:
| # | Source VM | Dest VM | Port | Expected Result | DFW Rule | Purpose |
|---|---|---|---|---|---|---|
| T-001 | Web-01 | App-01 | TCP/8080 | Delivered | A-002 | Verify web→app allowed |
| T-002 | Web-01 | DB-01 | TCP/3306 | Dropped (DFW) | D-001 | Verify web cannot reach DB directly |
| T-003 | App-01 | DB-01 | TCP/3306 | Delivered | A-003 | Verify app→DB allowed |
| T-004 | Dev-01 | App-01 | TCP/8080 | Dropped (DFW) | V-001 | Verify dev→prod blocked |
| T-005 | PCI-Web-01 | Non-PCI-App | TCP/8443 | Dropped (DFW) | P-003 | Verify PCI isolation |
| T-006 | Web-01 | DNS-01 | UDP/53 | Delivered | I-001 | Verify infrastructure access |
| T-007 | Quarantined-VM | Any | Any | Dropped (DFW) | E-001 | Verify quarantine effective |
For each test:
1. Run Traceflow in NSX Manager (Security → Traceflow)
- Verify packet reaches intended hop (delivered or dropped at correct enforcement point)
- Confirm DFW rule hit matches expected rule ID
- Document any unexpected results for remediation
Design flow monitoring and logging:
- NSX Distributed Firewall logging:
- Log level: Per-rule (not global — global logging creates excessive log volume)
- Log rules: All Emergency, Environment blocks, Application allows/blocks for PCI scope
- Do NOT log: High-frequency infrastructure rules (DNS, NTP) — creates log flood
- Log format: Syslog to centralized SIEM (Splunk, Aria Operations for Logs)
- Syslog destination: 172.16.200.10:514 (Aria Operations for Logs appliance)
- Flow monitoring:
- IPFIX export: Enable on selected segments (not all — performance impact)
- Collector: Aria Operations for Networks (vRNI) or third-party IPFIX collector
- Use case: Discover undocumented traffic flows before tightening DFW rules
- Recommendation: Enable IPFIX for 2-4 weeks in monitoring mode before enforcing default-deny
- NSX Intelligence (if licensed):
- Automated micro-segmentation recommendations based on observed traffic
- Visualizes all east-west flows between VMs
- Identifies VMs communicating outside their expected groups
- Use during initial deployment to refine DFW rules before enforcement
Design compliance verification:
- SOC 2 requirements mapping:
- CC6.1 (Logical access): DFW default-deny + allow-list per application ✓
- CC6.6 (Network security): Micro-segmentation isolates PCI from non-PCI ✓
- CC7.1 (Detection): DFW logging + SIEM integration ✓
- CC7.2 (Monitoring): Flow monitoring + anomaly detection ✓
- PCI-DSS compliance:
- Requirement 1: Firewall between PCI and non-PCI (DFW rule P-003) ✓
- Requirement 2: No default credentials (NSX admin password changed) ✓
- Requirement 10: Logging all access to PCI data (rules P-001, P-002 logged) ✓
- Audit-ready documentation:
- Export: NSX DFW rule set as JSON/CSV (API: GET /policy/api/v1/infra/domains/default/security-policies)
- Mapping: Cross-reference each DFW rule to compliance requirement
- Review cadence: Quarterly rule review with security team
- Evidence: Traceflow results + flow monitoring exports for audit package
Validation Gate
Check: Traffic flows mapped, Traceflow tests planned, logging configured, compliance verified
Expected: 5+ east-west flows documented with DFW rule hit, 7+ Traceflow tests planned, logging strategy defined (per-rule, not global), SOC 2 and PCI compliance requirements mapped
Common Errors
Task 4 NSX Security Design Decision & VCDX Defense
Consolidate the NSX security design into defensible design decisions and prepare for VCDX panel challenges on network security.
Create design decision D-009 — Network Security Model:
Decision: NSX Distributed Firewall with zero-trust, tag-based micro-segmentation
Alternatives:
A) Perimeter firewall only (Palo Alto / Fortinet at edge):
- Blocks north-south threats
- Cannot inspect east-west traffic between VMs on same host
- Rejected: 80% of data center traffic is east-west; perimeter-only is insufficient for SOC 2
B) Agent-based host firewall (iptables / Windows Firewall):
- Per-VM enforcement
- Difficult to manage at scale (600 VMs × individual rule sets)
- No centralized visibility or logging
- Rejected: Operational burden, no group-based policy, limited audit capability
C) NSX DFW + physical perimeter firewall (defense-in-depth):
- NSX DFW for east-west micro-segmentation
- Physical firewall for north-south inspection (IDS/IPS, SSL inspection)
- SELECTED: Best of both — DFW for internal isolation, physical for perimeter defense
Justification: NSX DFW provides kernel-level enforcement on every host without network hairpinning. Combined with physical perimeter firewall, this creates defense-in-depth aligned with zero-trust principles.
Prepare VCDX defense responses:
Challenge 1: 'How does DFW handle a rule with 10,000 VMs in the group?'
Response: DFW compiles rules into optimized filter tables per host. A group with 10,000 members is resolved once and pushed to all relevant hosts. Performance impact is minimal because the filter is evaluated per-packet at kernel level. The bottleneck would be group membership updates, not rule evaluation — membership changes trigger incremental updates, not full recompilation.
Challenge 2: 'What happens if NSX Manager goes down — do DFW rules stop working?'
Response: DFW rules are pushed to and cached on each host's kernel module. If NSX Manager becomes unavailable, existing rules continue to enforce. New rules cannot be pushed, and group membership changes are queued. VMs migrated via vMotion carry their DFW rules with them. This is a critical availability design — DFW operates independently of the management plane.
Challenge 3: 'Why tag-based groups instead of IP-based rules?'
Response: IP-based rules break when VMs get new IPs (DHCP renewal, migration, reprovisioning). Tag-based groups are dynamic — when a VM is tagged 'tier:web', it automatically joins the Web-Servers group and all applicable DFW rules apply immediately. This is essential for VCF automation where VMs are provisioned and decommissioned frequently.
Challenge 4: 'How would you implement a quarantine for a compromised VM?'
Response: Tag the VM with 'security:quarantine'. The Quarantined group in Emergency category (rule E-001) immediately drops all traffic. Because Emergency category rules evaluate first, they override all other allow rules. The VM is isolated within seconds without network changes. Forensics team can then add a targeted allow rule for their investigation tools only.
Document DFW operational procedures:
- Day-0: Initial deployment:
- Deploy in monitor-only mode (rules set to Allow with logging)
- Enable IPFIX flow monitoring for 2-4 weeks
- Use NSX Intelligence to discover actual traffic patterns
- Build groups and rules based on observed traffic
- Day-1: Enforcement:
- Switch default rule from Allow → Drop (with logging)
- Validate: Traceflow for every documented flow
- Rollback plan: Revert default to Allow if critical traffic blocked
- Cutover window: Low-traffic period with full team on standby
- Day-2: Ongoing operations:
- New application onboarding: Provide DFW rule template (3-5 rules per app)
- Rule change: Change management process → test in dev → promote to prod
- Incident response: Quarantine procedure documented (tag + Emergency rule)
- Audit: Quarterly rule review, unused rule cleanup, compliance evidence export
- Monitoring:
- Aria Operations for Logs: DFW log aggregation and correlation
- Alert: >100 DFW drops/minute from production VMs → potential misconfiguration
- Dashboard: Top blocked flows, top talkers, rule hit counts
Create NSX security architecture summary diagram (text-based):
[Physical Network]
|
| BGP + BFD (ASN 65000 ↔ 65001)
|
[Tier-0 Gateway] (Active-Active ECMP on 2 Edge VMs)
|
+--[Tier-1: Web]--[Seg: Web-Prod-A]--[DFW: A-001, A-002]
| [Seg: Web-Prod-B]--[DFW: A-001, A-002]
|
+--[Tier-1: App]--[Seg: App-Prod]----[DFW: A-002, A-003, A-005]
| [Seg: App-API]-----[DFW: A-002, A-003]
|
+--[Tier-1: DB]---[Seg: DB-Prod]-----[DFW: A-003, A-004]
| [Seg: DB-Archive]--[DFW: A-003]
|
+--[Tier-1: Dev]--[Seg: Dev-General]--[DFW: V-001 (blocked from prod)]
[Seg: Test-QA]-----[DFW: V-001]
[DFW Category Order]: Emergency → Infrastructure → Environment → Application → Default-Deny
[Groups]: Tag-based (tier, env, compliance) — dynamic membership
[Logging]: Per-rule to Aria Operations for Logs via syslog
[Verification]: Traceflow + IPFIX flow monitoringThis diagram should be in the VCDX design document as the logical network security architecture.
Validation Gate
Check: NSX security design consolidated with design decision, defense responses, operational procedures, and architecture diagram
Expected: Design decision with alternatives, 4+ defense responses, day-0/1/2 operational plan, and architectural diagram
Common Errors
Final Validation
Complete NSX micro-segmentation design with segment topology, DFW rules, traffic flow analysis, and defensible security architecture
✓ Network segment design with CIDR plan → 8+ segments with addressing, Tier-0/Tier-1 hierarchy, BGP peering design
✓ DFW rules follow zero-trust approach → 15+ rules across categories, tag-based groups, default-deny, PCI isolation
✓ Traffic flows mapped and verified → 5+ flows documented, 7+ Traceflow tests, logging strategy, compliance mapping
✓ Design defensible in VCDX context → Design decision with alternatives, 4+ defense responses, operational procedures
Cleanup / Restore
• Save all DFW rule documentation and traffic flow diagrams
• If using Holodeck: Remove test DFW rules and revert to base NSX configuration
Design Reflection (VCDX)
NSX micro-segmentation is the centerpiece of modern VCF security architecture. VCDX panelists will test: (1) Can you explain DFW rule evaluation order? Category → Section → Rule, top-to-bottom, first match wins. (2) Why groups over IPs? Dynamic membership, automation-friendly, no stale rules. (3) What is the blast radius of a misconfigured rule? With Applied-To scope set correctly — limited to affected VMs. Without it — every VM in the environment. (4) Draw me the traffic flow. Be able to trace any packet from source to destination through every enforcement point.
Requirements
- Zero-trust network architecture — no implicit trust between any VMs
- PCI-DSS scope isolation — PCI VMs cannot communicate with non-PCI VMs
- SOC 2 CC6.1 and CC6.6 — logical access controls and network security
- All DFW events logged to centralized SIEM for 90-day retention
Constraints
- DFW rules cannot exceed ~10,000 per host (performance limit)
- Applied-To scope must be set per rule or section — omission impacts all hosts
- NSX Manager unavailability prevents new rule pushes (existing rules persist)
- Geneve overlay adds ~54 bytes overhead — MTU must be 1700+
Assumptions
- All VMs will be tagged during provisioning (automation enforced)
- Physical perimeter firewall handles north-south IDS/IPS inspection
- Aria Operations for Logs deployed for DFW log aggregation
- Security team performs quarterly DFW rule review
Risks
- Misconfigured DFW rule blocks critical application traffic — mitigate with monitoring phase before enforcement
- Tag assignment automation failure leaves VMs untagged — default-deny blocks all traffic, mitigate with provisioning validation
- DFW rule sprawl over time degrades manageability — mitigate with quarterly cleanup and rule count monitoring
- NSX Manager compromise exposes DFW configuration — mitigate with RBAC, MFA, and API access controls
Self-Assessment Discussion Prompts
- If you had to implement micro-segmentation without NSX (e.g., customer uses a different overlay), what alternatives exist?
- How would you handle a DFW rule that needs to allow traffic from an external IP range that changes frequently?
- What is the performance impact of enabling DFW on a host with 50 VMs, each with 5 vNICs?
- If a developer accidentally deploys a VM without tags, what happens and how do you detect it?
Extensions
NSX Gateway Firewall with IDS/IPS
Extend the security design by enabling NSX Gateway Firewall IDS/IPS on the Tier-0 and Tier-1 gateways. Configure signature profiles for common CVEs, set up alert integration with SIEM, and document the performance impact on Edge VM throughput. Compare NSX IDS/IPS with physical firewall IDS/IPS for defense-in-depth.
NSX Distributed IDS/IPS
Enable NSX Distributed IDS/IPS on the DFW for east-west traffic inspection. Design signature profiles for the 3-tier application, configure signature exclusions for false positives, and evaluate the CPU overhead on ESXi hosts. Document the difference between DFW packet filtering and IDS/IPS deep packet inspection.
Automated Micro-Segmentation with Aria Automation
Integrate DFW rule creation with Aria Automation blueprints. Design a workflow where deploying a new application tier automatically creates NSX groups, DFW rules, and Traceflow verification tests. Document the tagging convention, rule template, and approval workflow.
⚠ Known Pitfalls (from Community KB)
References
- NSX 9.0 Administration Guide — Distributed Firewall chapter
- NSX 9.0 Design Guide — Micro-Segmentation best practices
- VMware KB 212073 — NSX DFW Rule Processing Order
- NSX Traceflow Troubleshooting Guide
- NIST SP 800-207 — Zero Trust Architecture (ZTA) reference framework