Academy/NSX 4.x Network Virtualization Professional (2V0-41.24)/Lab N2: NSX Federation Across Two VCF 9.0 Instances
This lab targets VCF 9.0

Lab N2: NSX Federation Across Two VCF 9.0 Instances

VCF 9.0Advancedvcp-foundation⏱ 150 min

NSX 9.0.x (feature set inherited from the NSX 4.2 line; NSX 4.2 docs remain a valid technical reference) Federation — Global Manager, cross-site policy, stretched segments, location-aware networking

Objectives

  • Understand NSX Federation architecture — Global Manager vs Local Manager
  • Configure federation peering between two VCF instances
  • Create stretched segments and gateways spanning both sites
  • Implement location-aware DFW policies for cross-site security
  • Test site failure scenarios and understand split-brain handling
  • Compare federated vs non-federated multi-site designs

Prerequisites

Two VCF 9.0 instances (or simulated with two NSX Manager clusters in the same lab), each with Edge clusters and overlay transport zones operational. Inter-site L3 connectivity established between the two environments.

Prior labs: net-virt-01, net-virt-02, net-virt-05

Required skills:

  • NSX Manager administration
  • Multi-site networking concepts
  • BGP/routing between sites

Lab Environment

Two VCF sites: Site-A (NSX Manager cluster-A, Edge cluster-A) and Site-B (NSX Manager cluster-B, Edge cluster-B). Both sites connected via L3 routed link (WAN or simulated inter-site VLAN). Global Manager appliance deployed on management network accessible from both sites.

Tasks

Task 1 Configure NSX Federation with Stretched Segments and Cross-Site Policy

NSX Federation uses a hierarchical model: Global Manager (GM) defines global objects (stretched T0/T1, segments, DFW policies) that are synchronized to Local Managers (LM) at each site. Each LM maintains an independent control plane — if inter-site connectivity is lost, each site continues to operate locally. Federation is not a replacement for disaster recovery — it enables policy consistency and workload mobility across sites. In VCF 9.0, Federation is supported across VCF instances in the same or different Private Clouds.

Build a federated NSX topology spanning two VCF instances, demonstrating how NSX Federation enables consistent policy across sites while maintaining local control plane independence. This is critical for disaster recovery, multi-site workload mobility, and consistent security posture across geographically distributed infrastructure.

Step 1

Deploy Global Manager. Deploy the NSX Global Manager appliance (OVA) on the management network. The GM is a standalone appliance (not part of either site's NSX Manager cluster). Size: Medium for lab (4 vCPU, 16 GB). Configure: FQDN='gm.lab.local', VIP for cluster access. The GM must be reachable from both Site-A and Site-B NSX Manager clusters.

Step 2
Register Local Managers with Global Manager. Log into the Global Manager UI → System → Location Manager → Add Location. For Site-A: Name='Site-A', NSX Manager FQDN/IP=<site-a-nsx-mgr>, Credentials=admin. Repeat for Site-B. After registration, the GM discovers each site's fabric: Edge clusters, transport zones, segments. Verify: Location Manager shows both sites with status 'Active'.
Step 3
Create stretched Tier-0 gateway. On the Global Manager → Networking → Tier-0 Gateways → Add. Name='GM-T0-Stretched', HA Mode=Active-Standby (recommended for stretched — stateful services needed for cross-site NAT). Assign Edge clusters: Site-A Edge cluster as Primary, Site-B Edge cluster as Secondary. Configure uplink interfaces at each site to peer with local physical routers. The GM creates the T0 on both LMs simultaneously.
Step 4
Create stretched Tier-1 and overlay segment. On the GM → Networking → Tier-1 Gateways → Add. Name='GM-T1-App', Linked T0='GM-T0-Stretched'. Create a stretched segment: Name='seg-app-stretched', Connected Gateway='GM-T1-App', Subnet=10.20.1.1/24, Span=Both sites. This segment is realized on transport nodes at BOTH sites — VMs at Site-A and Site-B can communicate on the same L2 domain.
Step 5
Verify stretched segment realization. On each site's Local Manager (NOT the GM), navigate to Networking → Segments — the stretched segment should appear as a 'Read-Only' object (managed by GM, cannot be edited locally). Deploy a VM at each site on the stretched segment: vm-site-a (10.20.1.11) at Site-A, vm-site-b (10.20.1.12) at Site-B. Ping between them — traffic traverses the inter-site GENEVE tunnel (TEP-to-TEP across sites).
Step 6
Configure location-aware DFW policies. On the GM → Security → Distributed Firewall → Add Policy. Name='Cross-Site-Policy'. Create rules with location awareness: (a) Rule 1: Source=SG-Web (Site-A only), Dest=SG-App (any site), Service=HTTPS, Action=Allow, Applied-To=Site-A only — this rule is pushed ONLY to Site-A hosts; (b) Rule 2: Source=SG-DB (Site-B only), Dest=<backup-server>, Service=NFS, Action=Allow, Applied-To=Site-B only. Location-scoped rules enable site-specific policies within a global policy framework.
Step 7
Verify policy distribution. On Site-A's Local Manager → Security → DFW — the global policies appear as read-only entries above local policies. Local admins can still create site-local DFW policies below the global ones. On Site-B, verify the same global policy appears but with the Site-B-specific rules. This demonstrates the hierarchy: Global policies (from GM) → Local policies (from LM).
Step 8

Test ingress optimization (north-south path). With a stretched segment, a VM at Site-A could receive traffic via Site-B's T0 (suboptimal). Configure route advertisement with locality: on GM-T0-Stretched, enable 'Ingress Routing Optimization' — this ensures each site advertises routes for VMs currently at that site only, preventing traffic tromboning through the remote site's Edge. Verify with BGP table inspection at each site's physical router.

Step 9

Simulate site failure. Disconnect inter-site connectivity (or shut down Site-B's NSX Manager). Observe: (a) GM shows Site-B as 'Disconnected'; (b) Site-A continues operating normally — all local VMs maintain connectivity and DFW enforcement; (c) Site-B also continues independently — its LM takes autonomous control of local fabric. This is the key Federation design principle: independent failure domains. Cross-site stretched segment traffic is disrupted, but local operations are unaffected.

Step 10

Recovery and reconciliation. Restore inter-site connectivity. Observe: (a) GM reconnects to Site-B and shows 'Active' status; (b) Any local changes made during the partition are flagged — GM reconciliation process merges non-conflicting changes and flags conflicts for admin review; (c) Stretched segment connectivity resumes — VMs at both sites can communicate again. Monitor the GM Events log for reconciliation details.

Validation Gate

Check: NSX Federation operational with stretched resources and cross-site policy

Expected: Global Manager managing both sites, stretched T0/T1/segments operational with cross-site VM connectivity, location-aware DFW policies distributed correctly, site failure handled gracefully with independent local operations, recovery reconciliation successful

Common Errors

Local Manager registration fails with 'Certificate error'
Fix: GM and LM must trust each other's certificates. If using self-signed certs (lab), import the LM cert into GM trust store and vice versa. In production, use CA-signed certificates for all NSX Manager appliances.
Stretched segment shows 'Partial Success' — one site not realized
Fix: The overlay transport zone at the failing site may not have the same name or configuration. Ensure both sites have compatible overlay TZs. Also verify: inter-site TEP connectivity — TEPs at Site-A must be able to reach TEPs at Site-B via the WAN/inter-site link.
Cross-site VM traffic fails despite segment being stretched
Fix: Inter-site GENEVE tunnel requires: (a) L3 reachability between TEP IPs at each site, (b) MTU ≥ 9000 on the inter-site path (or GENEVE MTU adjusted to match WAN MTU), (c) UDP 6081 not blocked by WAN firewalls.
Global policy not appearing on Local Manager
Fix: Check GM-to-LM synchronization status: GM → System → Location Manager → site status. If 'Sync Pending', the LM may be under heavy load or network latency between GM and LM is high. Allow time for sync (up to 5 minutes for large policy sets).

Final Validation

NSX Federation operational with cross-site networking and security policy

✓ GM location status → Both sites show 'Active' in Location Manager

✓ Stretched segment connectivity → VMs at Site-A and Site-B communicate on the stretched segment

✓ Location-aware DFW → Site-specific rules only appear on hosts at the target site

✓ Ingress optimization → Each site advertises routes only for local VMs

✓ Site failure independence → Surviving site operates normally during partition

✓ Recovery reconciliation → GM reconnects and reconciles after partition healing

Cleanup / Restore

• Delete stretched segments and VMs from GM

• Remove stretched T1 and T0 gateways from GM

• Deregister Local Managers from Global Manager (Location Manager → Remove)

• Power off Global Manager appliance

Design Reflection (VCDX)

Federation is a premium VCDX topic — it addresses multi-site design, disaster recovery, and operational complexity. Key design decisions: (1) GM placement — should be on a management segment accessible from all sites, ideally in a third 'witness' site for partition tolerance; (2) Stretched vs local objects — only stretch what needs cross-site L2 mobility; keep everything else local to minimize blast radius; (3) Location-aware DFW — enables global security baseline with site-specific exceptions; (4) Inter-site MTU — GENEVE over WAN requires careful MTU planning or IP fragmentation tolerance.

Requirements

  • Consistent security policy across 2+ VCF sites
  • Workload mobility between sites without IP change
  • Independent site operations during WAN failure

Constraints

  • Global Manager is a single point of policy management — its failure prevents global changes (but doesn't affect local operations)
  • Stretched segments require inter-site TEP connectivity with ≥ 9000 MTU or adjusted GENEVE MTU
  • Maximum 4 locations per federation (NSX 4.2 limit)

Assumptions

  • Inter-site WAN provides reliable L3 connectivity with < 150ms RTT
  • Both sites have compatible NSX versions (federation requires same major version)
  • GM is deployed in a highly available configuration (3-node cluster for production)

Risks

  • GM failure during a partition leaves both sites operating independently — any changes diverge and require manual reconciliation
  • Stretched segment traffic trombone through remote site Edge if ingress optimization is not configured
  • Over-stretching resources (every segment, every policy) increases inter-site dependency and blast radius

Self-Assessment Discussion Prompts

  1. When would you choose NSX Federation over independent NSX domains with API-driven policy sync?
  2. How do you handle MTU constraints on a 1500-byte WAN for GENEVE stretched segments?
  3. What is the GM recovery procedure if the Global Manager cluster is permanently lost?
  4. How does Federation interact with VCF's Private Cloud → Fleet → Instance hierarchy?

Extensions

Deploy a 3-node GM cluster (Active-Standby-Witness) and test GM failover

Configure Avi (NSX ALB) GSLB to steer traffic to the closest site for a stretched application

Implement a DR runbook: planned failover of all workloads from Site-A to Site-B using stretched segments

Test the maximum number of stretched segments and DFW rules to understand federation scale limits

⚠ Known Pitfalls (from Community KB)

Placing the Global Manager at one site only — if that site is lost, global changes are impossible until GM recovery; use a witness site or 3-node cluster
Stretching all segments by default — creates unnecessary inter-site dependency; only stretch segments that truly need cross-site L2 mobility
Ignoring WAN MTU for GENEVE overhead — stretched segment traffic uses GENEVE encapsulation across the WAN; if WAN MTU is 1500, either enable fragmentation or reduce inner MTU to ~1400
Not testing site partition behavior before production — split-brain handling is complex; validate reconciliation procedures in the lab first

References

  • NSX 4.2 Administration Guide — NSX Federation: techdocs.broadcom.com
  • VCF 9.0 Multi-Site Design Guide: techdocs.broadcom.com
  • KB 93267 — NSX Federation Troubleshooting and Reconciliation
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.