Academy/VCP-VCF 9.0 Architect (2V0-13.25)/NSX Micro-Segmentation Design
This lab targets VCF 9.0

NSX Micro-Segmentation Design

VCF 9.0Advancedvcp-foundation⏱ 105 min

VCF 9.0 NSX security architecture — DFW, segments, Tier-0/Tier-1 routing, and zero-trust design

Objectives

  • Design NSX network segments with proper CIDR allocation and Tier-0/Tier-1 routing
  • Create a Distributed Firewall rule set using a zero-trust, allow-list approach
  • Design BGP peering between NSX Edge and physical network
  • Plan east-west and north-south traffic flow with security enforcement points
  • Design logging, audit, and compliance verification using Traceflow and flow monitoring

Prerequisites

Access to NSX 9.0 and VCF networking documentation

Prior labs: vcp-architect-01, vcp-architect-02

Required skills:

  • TCP/IP networking and subnetting
  • Firewall rule concepts (stateful inspection, rule ordering)
  • BGP fundamentals
  • NSX overlay networking basics (Geneve)

Lab Environment

Design exercise — validate DFW rules on Holodeck NSX environment if available

Tasks

Task 1 Network Segment Design & IP Address Planning

VCDX panelists evaluate network design for proper isolation, scalable addressing, and clear separation between management, workload, and edge networks. Sloppy CIDR planning reveals weak network fundamentals.

Design the overlay network topology with proper segment isolation, CIDR allocation, and Tier-0/Tier-1 gateway architecture.

Step 1

Design Tier-0 and Tier-1 gateway topology:

NSX gateway hierarchy in VCF 9.0:

  • Tier-0 Gateway (Provider Router): Connects NSX overlay to physical network
    • 1 per workload domain (or shared across domains via VRF)
    • Provides north-south routing, NAT, VPN, BGP peering
    • HA mode: Active-Standby (stateful services) or Active-Active (ECMP, no stateful NAT)
  • Tier-1 Gateway (Tenant Router): Connects segments to Tier-0
    • 1+ per application tier or tenant
    • Provides segment interconnection, distributed routing, DFW attachment point
    • Can run as Distributed Router (DR) for east-west only, or Service Router (SR) for NAT/LB

Design for 600-VM environment:

  • Tier-0 (Management Domain): 1 Active-Standby (provides stateful NAT for management VMs)
  • Tier-0 (Workload Domain): 1 Active-Active ECMP (high throughput, no NAT needed)
  • Tier-1 (Web): Connected to Workload Tier-0, segments for web servers
  • Tier-1 (App): Connected to Workload Tier-0, segments for application servers
  • Tier-1 (DB): Connected to Workload Tier-0, segments for database servers
  • Tier-1 (Dev): Connected to Workload Tier-0, segments for dev/test
Step 2

Design CIDR addressing plan:

Overlay network supernet: 172.16.0.0/12 (avoid RFC 1918 conflicts with existing networks)

SegmentCIDRGatewayTier-1VMsPurpose
Web-Prod-A172.16.1.0/24.1T1-Web30Production web servers
Web-Prod-B172.16.2.0/24.1T1-Web20Production web servers (overflow)
App-Prod172.16.10.0/24.1T1-App80Application servers
App-API172.16.11.0/24.1T1-App40API gateway servers
DB-Prod172.16.20.0/24.1T1-DB30Database primary + replica
DB-Archive172.16.21.0/24.1T1-DB10Archive/backup databases
Dev-General172.16.100.0/23.1T1-Dev150Development environment
Test-QA172.16.102.0/24.1T1-Dev70QA/staging
Mgmt-Infra172.16.200.0/24.1T1-Mgmt20Infrastructure management VMs

Design principles:

  • /24 per segment unless >254 VMs needed (then /23)
  • Leave gaps between segment ranges for future expansion
  • Aggregate at Tier-1 level for BGP advertisement efficiency
  • Example: T1-Web advertises 172.16.0.0/20 (covers .1.0-.15.0 range)
Step 3

Design transport zone and VTEP configuration:

Transport zones (VCF 9.0):

  • Overlay TZ (default): All host TEPs participate; carries workload segments
  • VLAN TZ: For edge uplinks connecting to physical network

VTEP (Virtual Tunnel Endpoint) configuration:

  • VTEP subnet: 192.168.250.0/24 (dedicated, not routable externally)
  • VTEP IP pool: 192.168.250.10 - .200 (supports up to 191 host TEPs)
  • Teaming policy: Load Balance Source (for multi-pNIC hosts)

Geneve encapsulation:

  • Overhead: ~54 bytes (Outer Ethernet 14 + Outer IP 20 + Outer UDP 8 + Geneve header 8 + options ~4)
  • Required MTU: Inner frame 1500 + 54 = 1554 minimum; set 1700 for headroom
  • Physical switch MTU: 9000 (jumbo frames) recommended for all transport network paths

Verification:

  • Ping test between VTEPs: vmkping ++netstack=vxlan -d -s 1572 <remote_VTEP_IP>
  • Expected: No fragmentation at MTU 1700
Step 4

Design BGP peering between NSX Edge and physical network:

  1. Edge node deployment:
  • 2 Edge VMs per Tier-0 gateway (HA pair)
  • Edge VM sizing: Large (8 vCPU, 32 GB RAM) for production routing
  • Placement: Anti-affinity rule — Edge VMs on separate hosts
  1. BGP configuration:
  • NSX ASN: 65001 (private ASN range)
  • Physical router ASN: 65000
  • Peering: Each Edge node peers with 2 physical ToR switches (4 peering sessions total)
  • BFD: Enabled by default in NSX; timer 500ms × 3 = 1.5s failover detection
  1. Route advertisement:
  • Tier-0 advertises: Connected segments, Tier-1 connected routes, NAT IPs, LB VIPs
  • Route aggregation: Advertise summary 172.16.0.0/12 (not individual /24s)
  • Physical network: Accepts NSX routes, provides default route to Tier-0
  1. ECMP design (Active-Active Tier-0):
  • Both Edge nodes advertise same routes with equal metrics
  • Physical routers distribute traffic across both edges
  • Throughput: 2× single edge bandwidth
  • Limitation: Stateful services (NAT, LB, VPN) require Active-Standby, not ECMP
EdgePhysical RouterNSX ASNPhys ASNBFDStatus
Edge-1ToR-A16500165000YesActive
Edge-1ToR-A26500165000YesActive
Edge-2ToR-B16500165000YesActive
Edge-2ToR-B26500165000YesActive

Validation Gate

Check: Network design covers segment topology, CIDR plan, transport zone config, and BGP peering

Expected: Tier-0/Tier-1 hierarchy defined, 8+ segments with CIDR addressing, VTEP configuration with MTU validation, BGP peering with BFD and route advertisement plan

Common Errors

Using overlapping CIDR ranges between overlay and underlay networks — causes routing black holes
Forgetting Geneve overhead (~54 bytes) when setting MTU — results in fragmentation and vSAN/overlay performance degradation
Not enabling BFD on BGP peers — without BFD, BGP failover takes 180 seconds (3× keepalive timer) instead of 1.5 seconds
Deploying both Edge VMs on the same host — defeats HA purpose; use anti-affinity DRS rule

Task 2 Distributed Firewall Rule Set Design (Zero-Trust)

The DFW design is where security architects differentiate. VCDX panelists look for: rule ordering logic, group-based policies (not IP-based), and evidence that the candidate understands the enforcement pipeline.

Design a comprehensive DFW rule set using zero-trust principles with explicit allow rules and default-deny baseline.

Step 1

Design security group taxonomy:

NSX Groups (tag-based, not IP-based):

  • Group: Web-Servers — Criteria: Tag = 'tier:web' AND Tag = 'env:prod'
  • Group: App-Servers — Criteria: Tag = 'tier:app' AND Tag = 'env:prod'
  • Group: DB-Servers — Criteria: Tag = 'tier:db' AND Tag = 'env:prod'
  • Group: Dev-All — Criteria: Tag = 'env:dev'
  • Group: Mgmt-Infra — Criteria: Tag = 'tier:mgmt'
  • Group: PCI-Scope — Criteria: Tag = 'compliance:pci' (VMs handling payment data)
  • Group: Internet-Facing — Criteria: Tag = 'exposure:internet'

Tag assignment strategy:

  • Automated: vSphere tags applied during VM provisioning (Aria Automation blueprint or SDDC Manager)
  • Policy: Every VM must have at minimum: tier, env, and compliance tags
  • No tag = no access (default-deny catches untagged VMs)
Step 2

Design DFW rule categories and policies:

NSX DFW evaluates rules top-to-bottom across categories:

1. Ethernet (L2) → 2. Emergency → 3. Infrastructure → 4. Environment → 5. Application → Default

Category 1 — Emergency (highest priority):

#NameSourceDestServiceActionLog
E-001Block-QuarantinedGroup:QuarantinedAnyAnyDropYes
E-002Allow-Emergency-AdminGroup:Emergency-Admin-IPsAnySSH/RDPAllowYes

Category 2 — Infrastructure:

#NameSourceDestServiceActionLog
I-001Allow-DNSAnyGroup:DNS-ServersUDP/53, TCP/53AllowNo
I-002Allow-NTPAnyGroup:NTP-ServersUDP/123AllowNo
I-003Allow-AD-AuthAnyGroup:AD-ControllersTCP/389,636,88,445AllowNo
I-004Allow-vCenter-MgmtGroup:Mgmt-InfraGroup:vCenterTCP/443AllowYes
I-005Allow-MonitoringAnyGroup:MonitoringTCP/9090,9100,514AllowNo

Category 3 — Environment:

#NameSourceDestServiceActionLog
V-001Block-Dev-to-ProdGroup:Dev-AllGroup:*-ProdAnyDropYes
V-002Block-Prod-to-DevAny (env:prod)Group:Dev-AllAnyDropYes
Step 3

Design application-tier DFW rules:

Category 4 — Application (bulk of rules):

3-Tier Web Application:

#NameSourceDestServiceActionLogNotes
A-001External-to-WebGroup:Load-Balancer-VIPGroup:Web-ServersTCP/443AllowYesHTTPS only
A-002Web-to-AppGroup:Web-ServersGroup:App-ServersTCP/8080,8443AllowYesApp APIs
A-003App-to-DBGroup:App-ServersGroup:DB-ServersTCP/3306,5432AllowYesMySQL/Postgres
A-004DB-ReplicationGroup:DB-ServersGroup:DB-ServersTCP/3306AllowNoDB cluster
A-005App-to-CacheGroup:App-ServersGroup:Cache-ServersTCP/6379AllowNoRedis
A-006Web-to-CDNGroup:Web-ServersGroup:CDN-EgressTCP/443AllowNoStatic assets

PCI-scoped rules (stricter logging):

#NameSourceDestServiceActionLogNotes
P-001PCI-Web-to-PaymentGroup:PCI-Scope AND WebGroup:PCI-Scope AND AppTCP/8443AllowYesPCI traffic
P-002PCI-App-to-TokenDBGroup:PCI-Scope AND AppGroup:PCI-Scope AND DBTCP/5432AllowYesTokenized data
P-003Block-PCI-EgressGroup:PCI-ScopeNOT Group:PCI-ScopeAnyDropYesNo PCI data leakage

Default rule (last in each category):

#NameSourceDestServiceActionLog
D-001Default-Deny-AllAnyAnyAnyDropYes
Step 4

Document rule management best practices:

  1. Rule count management:
  • Target: <200 rules total across all categories
  • Use groups extensively (tag-based) rather than individual VM/IP rules
  • Each new application: 3-5 rules (inbound, inter-tier, outbound, logging)
  1. Rule lifecycle:
  • New rule request: Change management approval required
  • Rule review: Quarterly audit of all rules for relevance
   - Unused rule detection: NSX Flow Monitoring → rules with 0 hits in 90 days → decommission candidate
  • Rule versioning: Export DFW config before every change (NSX API backup)
  1. Applied-to scope:
  • CRITICAL: Always set 'Applied To' on each rule or section
  • Without 'Applied To': Rule is applied to EVERY vNIC in the environment (performance impact)
  • Example: Web-to-App rule 'Applied To' = Group:Web-Servers + Group:App-Servers
  • This pushes the rule only to hosts running those VMs, not all hosts
  1. Exclusion list:
  • VMs that should bypass DFW: NSX Manager, vCenter, SDDC Manager
  • Reason: DFW interruption could lock out management access
   - Configure: NSX Manager UI → Security → Distributed Firewall → Exclusion List

Validation Gate

Check: DFW rule set designed with zero-trust approach across all categories

Expected: 15+ rules organized by NSX categories (Emergency/Infrastructure/Environment/Application), tag-based groups, default-deny, PCI isolation rules, <200 total rule count target

Common Errors

Using IP addresses instead of groups — IPs change, groups are dynamic and self-maintaining
Not setting 'Applied To' scope — without it, every rule is evaluated on every host for every vNIC, causing CPU overhead
Putting application rules in the Infrastructure category — rule ordering matters; application rules in higher categories bypass environmental isolation
Forgetting to exclude management VMs from DFW — a misconfigured rule could lock out vCenter/NSX Manager access

Task 3 Traffic Flow Analysis & Security Verification

A design without verification is just a theory. VCDX candidates must show they can validate their security design, not just define it.

Map all traffic flows through the network and verify DFW enforcement using Traceflow and flow monitoring.

Step 1

Map east-west traffic flows:

Traffic flow diagram for 3-tier web application:

1. Client → Load Balancer (north-south, enters via Tier-0 Edge):
   - Ingress: Physical network → BGP → Tier-0 → LB VIP
  • DFW enforcement: None (traffic not yet on overlay)
2. Load Balancer → Web Server (east-west within segment):
  • LB distributes to web-prod-a segment VMs
  • DFW: Rule A-001 allows TCP/443 from LB VIP group to Web-Servers group
  • Enforcement point: DFW on destination host's dvPort (kernel module)
3. Web → App (east-west across segments):
   - Traffic: 172.16.1.x → 172.16.10.x (different segments, same Tier-1)
  • Routing: Distributed Router (DR) on source host routes to App segment
  • DFW: Rule A-002 allows TCP/8080 from Web-Servers to App-Servers
  • Note: DR routing is in-kernel — traffic does NOT hairpin through Edge
4. App → DB (east-west across Tier-1 boundaries):
   - Traffic: 172.16.10.x → 172.16.20.x (different Tier-1 gateways)
   - Routing: Source DR → Tier-0 → Destination DR (inter-Tier-1 routing)
  • DFW: Rule A-003 allows TCP/3306 from App-Servers to DB-Servers
5. DB → DB Replication (east-west within segment):
   - Traffic: 172.16.20.x → 172.16.20.y (same segment)
  • No routing needed (L2 within segment)
  • DFW: Rule A-004 allows TCP/3306 within DB-Servers group

Document each flow with: source, destination, port, routing path, DFW rule hit, and enforcement host.

Step 2

Design Traceflow verification plan:

Traceflow: NSX built-in diagnostic tool that injects a tagged packet and traces its path through the overlay network.

Test plan:

#Source VMDest VMPortExpected ResultDFW RulePurpose
T-001Web-01App-01TCP/8080DeliveredA-002Verify web→app allowed
T-002Web-01DB-01TCP/3306Dropped (DFW)D-001Verify web cannot reach DB directly
T-003App-01DB-01TCP/3306DeliveredA-003Verify app→DB allowed
T-004Dev-01App-01TCP/8080Dropped (DFW)V-001Verify dev→prod blocked
T-005PCI-Web-01Non-PCI-AppTCP/8443Dropped (DFW)P-003Verify PCI isolation
T-006Web-01DNS-01UDP/53DeliveredI-001Verify infrastructure access
T-007Quarantined-VMAnyAnyDropped (DFW)E-001Verify quarantine effective

For each test:

1. Run Traceflow in NSX Manager (Security → Traceflow)
  1. Verify packet reaches intended hop (delivered or dropped at correct enforcement point)
  2. Confirm DFW rule hit matches expected rule ID
  3. Document any unexpected results for remediation
Step 3

Design flow monitoring and logging:

  1. NSX Distributed Firewall logging:
  • Log level: Per-rule (not global — global logging creates excessive log volume)
  • Log rules: All Emergency, Environment blocks, Application allows/blocks for PCI scope
  • Do NOT log: High-frequency infrastructure rules (DNS, NTP) — creates log flood
  • Log format: Syslog to centralized SIEM (Splunk, Aria Operations for Logs)
  • Syslog destination: 172.16.200.10:514 (Aria Operations for Logs appliance)
  1. Flow monitoring:
  • IPFIX export: Enable on selected segments (not all — performance impact)
  • Collector: Aria Operations for Networks (vRNI) or third-party IPFIX collector
  • Use case: Discover undocumented traffic flows before tightening DFW rules
  • Recommendation: Enable IPFIX for 2-4 weeks in monitoring mode before enforcing default-deny
  1. NSX Intelligence (if licensed):
  • Automated micro-segmentation recommendations based on observed traffic
  • Visualizes all east-west flows between VMs
  • Identifies VMs communicating outside their expected groups
  • Use during initial deployment to refine DFW rules before enforcement
Step 4

Design compliance verification:

  1. SOC 2 requirements mapping:
  • CC6.1 (Logical access): DFW default-deny + allow-list per application ✓
  • CC6.6 (Network security): Micro-segmentation isolates PCI from non-PCI ✓
  • CC7.1 (Detection): DFW logging + SIEM integration ✓
  • CC7.2 (Monitoring): Flow monitoring + anomaly detection ✓
  1. PCI-DSS compliance:
  • Requirement 1: Firewall between PCI and non-PCI (DFW rule P-003) ✓
  • Requirement 2: No default credentials (NSX admin password changed) ✓
  • Requirement 10: Logging all access to PCI data (rules P-001, P-002 logged) ✓
  1. Audit-ready documentation:
  • Export: NSX DFW rule set as JSON/CSV (API: GET /policy/api/v1/infra/domains/default/security-policies)
  • Mapping: Cross-reference each DFW rule to compliance requirement
  • Review cadence: Quarterly rule review with security team
  • Evidence: Traceflow results + flow monitoring exports for audit package

Validation Gate

Check: Traffic flows mapped, Traceflow tests planned, logging configured, compliance verified

Expected: 5+ east-west flows documented with DFW rule hit, 7+ Traceflow tests planned, logging strategy defined (per-rule, not global), SOC 2 and PCI compliance requirements mapped

Common Errors

Enabling global DFW logging — generates massive log volume and impacts DFW performance; log selectively per rule
Not running Traceflow for blocked paths — verifying only 'allow' rules misses misconfigured 'deny' rules
Forgetting that inter-Tier-1 traffic routes through Tier-0 — DFW rules must account for this routing path
Assuming micro-segmentation alone satisfies PCI — DFW is one control; logging, monitoring, and access management are also required

Task 4 NSX Security Design Decision & VCDX Defense

Network security is a VCDX panel favorite. Be ready to draw traffic flows on demand, explain why group-based policy beats IP-based, and articulate the DFW rule evaluation pipeline.

Consolidate the NSX security design into defensible design decisions and prepare for VCDX panel challenges on network security.

Step 1

Create design decision D-009 — Network Security Model:

Decision: NSX Distributed Firewall with zero-trust, tag-based micro-segmentation
Alternatives:
A) Perimeter firewall only (Palo Alto / Fortinet at edge):

  • Blocks north-south threats
  • Cannot inspect east-west traffic between VMs on same host
  • Rejected: 80% of data center traffic is east-west; perimeter-only is insufficient for SOC 2

B) Agent-based host firewall (iptables / Windows Firewall):

  • Per-VM enforcement
  • Difficult to manage at scale (600 VMs × individual rule sets)
  • No centralized visibility or logging
  • Rejected: Operational burden, no group-based policy, limited audit capability

C) NSX DFW + physical perimeter firewall (defense-in-depth):

  • NSX DFW for east-west micro-segmentation
  • Physical firewall for north-south inspection (IDS/IPS, SSL inspection)
  • SELECTED: Best of both — DFW for internal isolation, physical for perimeter defense

Justification: NSX DFW provides kernel-level enforcement on every host without network hairpinning. Combined with physical perimeter firewall, this creates defense-in-depth aligned with zero-trust principles.

Step 2

Prepare VCDX defense responses:

Challenge 1: 'How does DFW handle a rule with 10,000 VMs in the group?'
Response: DFW compiles rules into optimized filter tables per host. A group with 10,000 members is resolved once and pushed to all relevant hosts. Performance impact is minimal because the filter is evaluated per-packet at kernel level. The bottleneck would be group membership updates, not rule evaluation — membership changes trigger incremental updates, not full recompilation.

Challenge 2: 'What happens if NSX Manager goes down — do DFW rules stop working?'
Response: DFW rules are pushed to and cached on each host's kernel module. If NSX Manager becomes unavailable, existing rules continue to enforce. New rules cannot be pushed, and group membership changes are queued. VMs migrated via vMotion carry their DFW rules with them. This is a critical availability design — DFW operates independently of the management plane.

Challenge 3: 'Why tag-based groups instead of IP-based rules?'
Response: IP-based rules break when VMs get new IPs (DHCP renewal, migration, reprovisioning). Tag-based groups are dynamic — when a VM is tagged 'tier:web', it automatically joins the Web-Servers group and all applicable DFW rules apply immediately. This is essential for VCF automation where VMs are provisioned and decommissioned frequently.

Challenge 4: 'How would you implement a quarantine for a compromised VM?'
Response: Tag the VM with 'security:quarantine'. The Quarantined group in Emergency category (rule E-001) immediately drops all traffic. Because Emergency category rules evaluate first, they override all other allow rules. The VM is isolated within seconds without network changes. Forensics team can then add a targeted allow rule for their investigation tools only.

Step 3

Document DFW operational procedures:

  1. Day-0: Initial deployment:
  • Deploy in monitor-only mode (rules set to Allow with logging)
  • Enable IPFIX flow monitoring for 2-4 weeks
  • Use NSX Intelligence to discover actual traffic patterns
  • Build groups and rules based on observed traffic
  1. Day-1: Enforcement:
   - Switch default rule from Allow → Drop (with logging)
  • Validate: Traceflow for every documented flow
  • Rollback plan: Revert default to Allow if critical traffic blocked
  • Cutover window: Low-traffic period with full team on standby
  1. Day-2: Ongoing operations:
  • New application onboarding: Provide DFW rule template (3-5 rules per app)
   - Rule change: Change management process → test in dev → promote to prod
  • Incident response: Quarantine procedure documented (tag + Emergency rule)
  • Audit: Quarterly rule review, unused rule cleanup, compliance evidence export
  1. Monitoring:
  • Aria Operations for Logs: DFW log aggregation and correlation
   - Alert: >100 DFW drops/minute from production VMs → potential misconfiguration
  • Dashboard: Top blocked flows, top talkers, rule hit counts
Step 4

Create NSX security architecture summary diagram (text-based):

[Physical Network]
     |
     | BGP + BFD (ASN 65000 ↔ 65001)
     |
[Tier-0 Gateway] (Active-Active ECMP on 2 Edge VMs)
     |
     +--[Tier-1: Web]--[Seg: Web-Prod-A]--[DFW: A-001, A-002]
     |                  [Seg: Web-Prod-B]--[DFW: A-001, A-002]
     |
     +--[Tier-1: App]--[Seg: App-Prod]----[DFW: A-002, A-003, A-005]
     |                  [Seg: App-API]-----[DFW: A-002, A-003]
     |
     +--[Tier-1: DB]---[Seg: DB-Prod]-----[DFW: A-003, A-004]
     |                  [Seg: DB-Archive]--[DFW: A-003]
     |
     +--[Tier-1: Dev]--[Seg: Dev-General]--[DFW: V-001 (blocked from prod)]
                        [Seg: Test-QA]-----[DFW: V-001]

[DFW Category Order]: Emergency → Infrastructure → Environment → Application → Default-Deny
[Groups]: Tag-based (tier, env, compliance) — dynamic membership
[Logging]: Per-rule to Aria Operations for Logs via syslog
[Verification]: Traceflow + IPFIX flow monitoring

This diagram should be in the VCDX design document as the logical network security architecture.

Validation Gate

Check: NSX security design consolidated with design decision, defense responses, operational procedures, and architecture diagram

Expected: Design decision with alternatives, 4+ defense responses, day-0/1/2 operational plan, and architectural diagram

Common Errors

Deploying DFW in enforcement mode without monitoring phase — will block unknown traffic flows and cause outages
Not having a quarantine procedure — security incidents require rapid VM isolation capability
Forgetting that DFW rules survive NSX Manager failure — this is a common panel question and a design strength to highlight
Not including the DFW rule evaluation order in the design document — Emergency → Infrastructure → Environment → Application → Default is critical for rule design

Final Validation

Complete NSX micro-segmentation design with segment topology, DFW rules, traffic flow analysis, and defensible security architecture

✓ Network segment design with CIDR plan → 8+ segments with addressing, Tier-0/Tier-1 hierarchy, BGP peering design

✓ DFW rules follow zero-trust approach → 15+ rules across categories, tag-based groups, default-deny, PCI isolation

✓ Traffic flows mapped and verified → 5+ flows documented, 7+ Traceflow tests, logging strategy, compliance mapping

✓ Design defensible in VCDX context → Design decision with alternatives, 4+ defense responses, operational procedures

Cleanup / Restore

• Save all DFW rule documentation and traffic flow diagrams

• If using Holodeck: Remove test DFW rules and revert to base NSX configuration

Design Reflection (VCDX)

NSX micro-segmentation is the centerpiece of modern VCF security architecture. VCDX panelists will test: (1) Can you explain DFW rule evaluation order? Category → Section → Rule, top-to-bottom, first match wins. (2) Why groups over IPs? Dynamic membership, automation-friendly, no stale rules. (3) What is the blast radius of a misconfigured rule? With Applied-To scope set correctly — limited to affected VMs. Without it — every VM in the environment. (4) Draw me the traffic flow. Be able to trace any packet from source to destination through every enforcement point.

Requirements

  • Zero-trust network architecture — no implicit trust between any VMs
  • PCI-DSS scope isolation — PCI VMs cannot communicate with non-PCI VMs
  • SOC 2 CC6.1 and CC6.6 — logical access controls and network security
  • All DFW events logged to centralized SIEM for 90-day retention

Constraints

  • DFW rules cannot exceed ~10,000 per host (performance limit)
  • Applied-To scope must be set per rule or section — omission impacts all hosts
  • NSX Manager unavailability prevents new rule pushes (existing rules persist)
  • Geneve overlay adds ~54 bytes overhead — MTU must be 1700+

Assumptions

  • All VMs will be tagged during provisioning (automation enforced)
  • Physical perimeter firewall handles north-south IDS/IPS inspection
  • Aria Operations for Logs deployed for DFW log aggregation
  • Security team performs quarterly DFW rule review

Risks

  • Misconfigured DFW rule blocks critical application traffic — mitigate with monitoring phase before enforcement
  • Tag assignment automation failure leaves VMs untagged — default-deny blocks all traffic, mitigate with provisioning validation
  • DFW rule sprawl over time degrades manageability — mitigate with quarterly cleanup and rule count monitoring
  • NSX Manager compromise exposes DFW configuration — mitigate with RBAC, MFA, and API access controls

Self-Assessment Discussion Prompts

  1. If you had to implement micro-segmentation without NSX (e.g., customer uses a different overlay), what alternatives exist?
  2. How would you handle a DFW rule that needs to allow traffic from an external IP range that changes frequently?
  3. What is the performance impact of enabling DFW on a host with 50 VMs, each with 5 vNICs?
  4. If a developer accidentally deploys a VM without tags, what happens and how do you detect it?

Extensions

NSX Gateway Firewall with IDS/IPS

Extend the security design by enabling NSX Gateway Firewall IDS/IPS on the Tier-0 and Tier-1 gateways. Configure signature profiles for common CVEs, set up alert integration with SIEM, and document the performance impact on Edge VM throughput. Compare NSX IDS/IPS with physical firewall IDS/IPS for defense-in-depth.

NSX Distributed IDS/IPS

Enable NSX Distributed IDS/IPS on the DFW for east-west traffic inspection. Design signature profiles for the 3-tier application, configure signature exclusions for false positives, and evaluate the CPU overhead on ESXi hosts. Document the difference between DFW packet filtering and IDS/IPS deep packet inspection.

Automated Micro-Segmentation with Aria Automation

Integrate DFW rule creation with Aria Automation blueprints. Design a workflow where deploying a new application tier automatically creates NSX groups, DFW rules, and Traceflow verification tests. Document the tagging convention, rule template, and approval workflow.

⚠ Known Pitfalls (from Community KB)

Not setting 'Applied To' scope on DFW rules — without it, every rule is pushed to every host and evaluated for every vNIC, causing unnecessary CPU overhead and potential performance degradation on large clusters
Using IP-based rules instead of tag-based groups — IP addresses change with DHCP, migration, and reprovisioning, leading to stale rules that either block legitimate traffic or allow unauthorized access
Enabling DFW enforcement without a monitoring phase — unknown traffic flows will be blocked by default-deny, causing application outages that could have been prevented with 2-4 weeks of flow discovery
Forgetting to exclude NSX Manager and vCenter from DFW — a misconfigured rule that blocks management traffic can lock out all administrative access to the entire environment

References

  • NSX 9.0 Administration Guide — Distributed Firewall chapter
  • NSX 9.0 Design Guide — Micro-Segmentation best practices
  • VMware KB 212073 — NSX DFW Rule Processing Order
  • NSX Traceflow Troubleshooting Guide
  • NIST SP 800-207 — Zero Trust Architecture (ZTA) reference framework
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.