Avi Load Balancer Troubleshooting
Objectives
- Describe the Avi Load Balancer architecture within VCF 9.0 including Controller cluster, Service Engine placement, and NSX integration
- Configure and validate health monitors (TCP, HTTP, HTTPS, custom) for backend pool members and interpret health check results
- Diagnose cascading pool member failures by correlating SE resource utilization, backend application health, and health monitor configuration
- Perform root cause analysis on Virtual Service failures using Avi analytics, application logs, and end-to-end timing data
- Apply immediate remediation actions including SE scaling, pool member tuning, and health check threshold adjustment
- Design proactive monitoring dashboards and alert policies to detect load balancer degradation before SLA breach
- Estimate SLA recovery timelines and articulate load balancer design tradeoffs in a VCDX defense context
Prerequisites
VCF 9.0 workload domain operational with Avi Controller cluster (3-node) deployed and integrated with NSX. At least one Virtual Service configured with a pool of 3+ backend web servers. NSX overlay segments available for SE placement. Avi Controller UI accessible via HTTPS from management workstation.
Prior labs: vcp-support-07, avi-alb-01
Required skills:
- Basic understanding of L4/L7 load balancing concepts (VIP, pool, health monitor)
- Familiarity with NSX overlay networking and segment configuration
- vSphere resource monitoring (CPU, memory, network) for virtual appliances
- HTTP/HTTPS fundamentals including status codes, SSL/TLS handshake, and connection lifecycle
- CLI and REST API interaction for troubleshooting (curl, Avi CLI shell)
Lab Environment
Three-node Avi Controller cluster deployed on the management domain. Two Service Engines (SE-1, SE-2) deployed in elastic HA (N+M) mode on the workload domain, connected to NSX overlay segments. One L7 Virtual Service (vs-web-prod) distributing HTTPS traffic across a pool of 4 backend web servers (web-01 through web-04) on an NSX segment. A second L4 Virtual Service (vs-db-prod) proxying TCP/3306 to 2 database servers (db-01, db-02). Health monitors configured: HTTP GET /health for web pool, TCP connect for DB pool. Simulated failure conditions pre-staged: web-03 returning HTTP 503, web-04 with expired SSL certificate, SE-1 approaching 85% CPU utilization.
graph TB CLIENT[Client Traffic] -->|HTTPS 443| VIP_WEB[VIP: vs-web-prod 10.50.1.100] CLIENT -->|TCP 3306| VIP_DB[VIP: vs-db-prod 10.50.1.101] VIP_WEB --> SE1[SE-1 10.50.2.10] VIP_WEB --> SE2[SE-2 10.50.2.11] VIP_DB --> SE1 SE1 --> WEB01[web-01 10.50.3.11] SE1 --> WEB02[web-02 10.50.3.12] SE2 --> WEB03[web-03 10.50.3.13 - HTTP 503] SE2 --> WEB04[web-04 10.50.3.14 - SSL Expired] SE1 --> DB01[db-01 10.50.4.11] SE1 --> DB02[db-02 10.50.4.12] CTRL1[Avi Controller-1 10.50.0.21] --- CTRL2[Avi Controller-2 10.50.0.22] CTRL2 --- CTRL3[Avi Controller-3 10.50.0.23] CTRL1 -.->|Manages| SE1 CTRL1 -.->|Manages| SE2
IP Addressing
| Network | Purpose | VLAN |
|---|---|---|
10.50.0.0/24 | Avi Controller management network | VLAN 1644 |
10.50.1.0/24 | Virtual Service VIP network (client-facing) | VLAN 1650 |
10.50.2.0/24 | Service Engine data network | VLAN 1651 |
10.50.3.0/24 | Web server backend pool network (NSX overlay) | NSX Segment web-seg |
10.50.4.0/24 | Database server backend pool network (NSX overlay) | NSX Segment db-seg |
Credentials
| System | Username | Password |
|---|---|---|
| Avi Controller UI/API | admin | Set during Avi Controller initial setup. Minimum 8 characters, mixed case, special characters required. |
| Avi CLI Shell | admin | Same as Avi Controller admin password. Access via SSH to Controller node. |
| Backend Web Servers | root | Standard lab credentials for web-01 through web-04. |
| NSX Manager | admin | NSX Manager admin credentials for verifying SE segment placement. |
Tasks
Task 1 Avi Architecture & SE Deployment Model in VCF
availabilityBefore troubleshooting failures, you must understand the Avi distributed architecture in VCF 9.0. The Controller cluster is the control plane; Service Engines are the data plane. Misunderstanding SE placement, HA modes, or Controller-SE communication causes misdiagnosis. VCDX panelists expect candidates to articulate why the architecture is split-brain resilient and how SE placement on NSX overlay segments affects failure domains.
Log in to the Avi Controller UI (https://10.50.0.21). Navigate to Infrastructure > Service Engine Group. Document the SE Group configuration: HA mode (elastic HA N+M, dedicated, or legacy HA), minimum and maximum number of SEs, memory and vCPU allocation per SE, and the vCenter/NSX cloud connector binding.
Navigate to Infrastructure > Service Engines. For each SE (SE-1, SE-2), record: power state, vCPU count, memory allocation, connected networks (management, VIP, backend segments), and the host on which the SE VM is deployed. Verify that SEs are placed on different ESXi hosts for anti-affinity.
Verify Controller cluster health. Navigate to Administration > Controller > Nodes. Confirm all 3 Controller nodes show 'CLUSTER_UP' state. Check the cluster VIP is reachable: ping 10.50.0.20 (cluster VIP). Also verify the leader election status by checking which node is the current leader.
Examine the NSX Cloud connector configuration. Navigate to Infrastructure > Clouds. Select the NSX cloud and verify: NSX Manager IP, transport zone mapping, SE management network, and the content library used for SE image deployment. Document any warnings or errors on the cloud status.
From the Avi CLI shell (SSH to Controller), run: show serviceengine detail | grep -A 5 'se_name\|oper_status\|se_group_ref\|host_ref'. Cross-reference the CLI output with the UI findings to confirm consistency. Also run: show cloud status to verify cloud connector health from the CLI perspective.
Validation Gate
Check: Confirm: (1) SE Group HA mode is Elastic HA N+M, (2) Both SEs are on different ESXi hosts, (3) All 3 Controller nodes are CLUSTER_UP, (4) NSX Cloud connector status is GREEN, (5) CLI and UI data are consistent.
Expected: All 5 architecture validation points confirmed. You have a complete picture of the Avi control plane and data plane topology before proceeding to troubleshooting.
Common Errors
Task 2 Health Monitor & Pool Configuration Troubleshooting
recoverabilityHealth monitors are the primary mechanism for detecting backend failures. Misconfigured health checks cause false positives (marking healthy servers as down) or false negatives (continuing to send traffic to failed servers). This task builds the skill of diagnosing health monitor behavior, which is the most common source of Avi pool member issues in VCF deployments and a frequent VCDX troubleshooting scenario.
Navigate to Applications > Virtual Services > vs-web-prod > Pool (web-pool). Examine each pool member's health status. Identify which members are UP, which are DOWN, and which are in ERROR state. For each DOWN/ERROR member, click the health monitor icon to view the last health check result and failure reason.
Examine the health monitor configuration. Navigate to Templates > Profiles > Health Monitors. Open the HTTP health monitor (http-health-check) used by web-pool. Document: type, send interval, receive timeout, successful checks, failed checks, HTTP request method, expected response code, and send/receive strings.
Diagnose web-03 (HTTP 503). SSH to web-03 (10.50.3.13) and check the application status: systemctl status nginx (or the appropriate web service). Check the application log for errors. Test the health endpoint locally: curl -v http://localhost/health. Determine why the backend is returning 503.
Diagnose web-04 (SSL handshake error). From the Avi Controller CLI, run: debug virtualservice vs-web-prod level debug. Then check the health monitor logs: show pool web-pool detail | grep -A 10 'web-04'. Also SSH to web-04 and check the SSL certificate: openssl s_client -connect localhost:443 2>&1 | grep -i 'verify\|expire\|not after'. Identify the certificate issue.
Remediate both failures. For web-03: work with the application team to resolve the upstream issue, or temporarily disable the member with 'graceful disable' (which drains existing connections). For web-04: renew the SSL certificate (or for lab purposes, disable SSL verification on the health monitor temporarily). After remediation, monitor the health check results: the member should transition from DOWN to UP after 3 consecutive successful checks (30 seconds with default timing). Navigate back to the pool view and verify all members are UP.
Validation Gate
Check: Confirm: (1) Root cause identified for both web-03 (HTTP 503 from upstream issue) and web-04 (expired SSL certificate), (2) Health monitor configuration documented with timing parameters, (3) At least one member remediated and transitioned to UP state, (4) Pool capacity percentage calculated correctly.
Expected: Both root causes identified and documented. At least web-04 restored to UP after certificate remediation. Health monitor behavior (timing, thresholds, SSL verification) fully understood.
Common Errors
Task 3 Virtual Service Failure Diagnosis & Traffic Analysis
performanceVirtual Service failures impact end users directly. This task teaches systematic diagnosis using Avi's built-in analytics engine, which provides real-time metrics, application logs, and client insight data that most traditional load balancers lack. The ability to correlate client-side errors with server-side metrics and SE resource utilization is a differentiating skill for VCDX candidates defending load balancer design decisions.
Navigate to Applications > Virtual Services > vs-web-prod. Open the Analytics tab. Set the time range to the last 1 hour. Review the key metrics: End-to-End Timing (client RTT, server RTT, application response time, data transfer time), Throughput (requests/sec, bandwidth), and Error Rate (percentage of 4xx and 5xx responses). Identify any anomalies or spikes.
Open the Logs tab for vs-web-prod. Filter by: HTTP Status >= 500. Examine 5-10 error entries. For each, note: timestamp, client IP, server IP (which pool member handled the request), response code, response time, and the full request URI. Look for patterns: are errors concentrated on specific pool members, specific URIs, or specific time windows?
Check SE resource utilization. Navigate to Infrastructure > Service Engines > SE-1. Review CPU, memory, network throughput, and connection count metrics for the last hour. Note that SE-1 was pre-staged at 85% CPU utilization. Determine whether SE resource exhaustion is contributing to Virtual Service degradation. Also check: SE-1 > Connected Virtual Services to see how many VS are sharing this SE.
Simulate a client request and trace it end-to-end. From a test client (or the Avi Controller CLI), run: curl -v -H 'Host: app.lab.local' https://10.50.1.100/api/data --resolve app.lab.local:443:10.50.1.100. Observe the response. Then in the Avi UI, find this specific request in the Logs tab (filter by timestamp or URI /api/data). Click the log entry to see the full request trace: which SE handled it, which pool member was selected, load balancing algorithm used, and timing breakdown.
Trigger and verify SE scale-out. If SE-1 CPU remains above 80%, the Controller should initiate an automatic scale-out. Navigate to Administration > Events. Filter for 'scale' events. If no scale-out has occurred, manually trigger it: navigate to Infrastructure > SE Group > Default-SEG > Edit, and temporarily lower the max_vs_per_se or adjust the buffer_se count to force a new SE deployment. Monitor the SE deployment in vCenter (new VM being created from content library). Once SE-3 is deployed, verify that Virtual Services are redistributed.
Validation Gate
Check: Confirm: (1) End-to-End Timing breakdown analyzed and server-side latency identified, (2) Application logs correlated error entries with specific pool members, (3) SE resource exhaustion identified as contributing factor, (4) Client request traced end-to-end through the Avi data path, (5) SE scale-out mechanism understood and verified.
Expected: Complete picture of Virtual Service failure: root causes include pool member failures (web-03 app issue, web-04 cert issue) compounded by SE-1 CPU exhaustion. Remediation includes pool member fixes and SE scale-out.
Common Errors
Task 4 Avi Operational Monitoring & VCDX Defense
manageabilityTroubleshooting is reactive; monitoring is proactive. This task transitions from break-fix to designing ongoing operational monitoring that prevents recurrence. VCDX candidates must demonstrate that their load balancer design includes observability, alerting, and SLA compliance tracking. This task also prepares you to defend Avi design decisions in a VCDX panel scenario.
Design an alert policy for pool member health degradation. Navigate to Operations > Alerts > Alert Actions. Create a new alert action that sends notifications when pool capacity drops below 75% (i.e., more than 1 of 4 members is DOWN). Configure: Alert Rule (pool_member_down count >= 2), Alert Action (email/syslog/SNMP trap), and Throttle (no more than 1 alert per 5 minutes to avoid alert storms). Document your alert policy design.
Configure SE resource monitoring thresholds. Navigate to Infrastructure > SE Group > Default-SEG > Edit > Advanced. Set the following thresholds: SE CPU high-water mark: 80%, SE memory high-water mark: 90%, connection high-water mark: 80% of max connections. Verify that auto-scale triggers are configured to respond when these thresholds are breached. Document the auto-scale policy: what triggers scale-out, what triggers scale-in, and what are the cooldown periods.
Build an SLA recovery estimate. Based on the scenario in this lab (50% pool capacity loss, elevated SE CPU), calculate: (1) Mean Time to Detect (MTTD): how long from failure onset to alert trigger, (2) Mean Time to Respond (MTTR): estimated time for an engineer to diagnose and begin remediation, (3) Mean Time to Recover (MTTR-full): time from remediation start to full service restoration, (4) Total SLA impact: estimated error rate during the incident window and time to return to baseline (<0.1% error rate). Document your calculations.
Prepare a VCDX defense summary for your Avi Load Balancer design. Document the following design decisions and their justifications: (1) Why Elastic HA N+M over Dedicated mode, (2) Why 3-node Controller cluster instead of single Controller, (3) Health monitor timing choices (interval, timeout, threshold), (4) SE sizing and auto-scale policy rationale, (5) Monitoring and alerting strategy. For each decision, state the requirement it satisfies, the constraint it operates within, and the risk if the decision were different.
Integrate Avi monitoring with the VCF observability stack. Navigate to Administration > Settings > Analytics. Verify that metrics export is configured to send data to Aria Operations (formerly vRealize Operations). Check the integration endpoint, export interval, and which metrics are being exported. Also verify syslog export under Administration > Settings > Syslog for sending application logs to Aria Operations for Logs. Document the end-to-end observability pipeline: Avi metrics/logs -> Aria Operations/Logs -> Dashboards -> Alerts.
Validation Gate
Check: Confirm: (1) Alert policy created for pool capacity degradation with escalation tiers, (2) SE auto-scale thresholds and cooldown periods configured, (3) SLA recovery timeline calculated with MTTD/MTTR, (4) VCDX defense document prepared with 5 design decisions and RCAR justification, (5) Avi monitoring integrated with VCF observability stack (Aria Operations).
Expected: Complete operational monitoring framework established. VCDX defense preparation document created demonstrating tradeoff awareness and architectural justification for all Avi design decisions.
Common Errors
Final Validation
This lab validated your ability to understand Avi Load Balancer architecture in VCF 9.0, diagnose health monitor and pool member failures, analyze Virtual Service performance using Avi analytics, and design proactive monitoring to prevent SLA breaches. You also prepared VCDX defense material articulating load balancer design tradeoffs with RCAR justification.
✓ Avi architecture documented: Controller cluster (3-node), SE deployment mode (Elastic HA N+M), NSX cloud connector status, SE anti-affinity verified → All architecture components validated and consistent between UI and CLI
✓ Pool member failures diagnosed: web-03 root cause (upstream application returning HTTP 503), web-04 root cause (expired SSL certificate causing health check failure) → Both root causes identified with evidence from health monitor logs and backend investigation
✓ Virtual Service analytics analyzed: End-to-End Timing breakdown, application log correlation, SE resource utilization impact on latency → Server-side latency spike correlated with pool capacity loss and SE CPU exhaustion
✓ SE scale-out mechanism verified: auto-scale trigger, SE deployment from content library, Virtual Service redistribution → Scale-out completed successfully, SE-1 CPU reduced below high-water mark
✓ Operational monitoring designed: alert policy with escalation tiers, SLA recovery timeline (MTTD + MTTR), VCDX defense document with 5 design decisions → Complete monitoring framework and defense preparation document with RCAR justification for each design decision
Cleanup / Restore
• Remove debug logging on vs-web-prod: navigate to the Virtual Service and disable debug-level analytics to reduce Controller storage consumption
• Revert SE Group max_scaleout to original value if it was modified during Task 3 to force scale-out
• Restore web-04 SSL certificate to a valid certificate (if a temporary health monitor workaround was applied, re-enable SSL verification)
• Revert to the pre-avi-troubleshooting snapshot if a clean state is required for subsequent labs
Design Reflection (VCDX)
Avi Load Balancer in VCF 9.0 represents a significant architectural shift from the NSX-T built-in load balancer. VCDX candidates must defend why the distributed control-plane/data-plane architecture (Controller cluster + Service Engines) is superior for enterprise workloads: independent scaling of management and data path, headless SE operation during Controller outages, and deep application-layer analytics. The key design tension is between resource efficiency (Elastic HA sharing SEs across Virtual Services) and isolation (Dedicated mode guaranteeing SE resources per VS). Your defense should articulate how SE sizing, auto-scale policy, health monitor timing, and monitoring integration form a cohesive availability strategy that meets specific RPO/RTO requirements.
Requirements
- L7 load balancing with SSL offload for production web applications on VCF 9.0 workload domains
- 99.95% availability SLA for load-balanced services (maximum 22 minutes downtime per month)
- Sub-second failover detection for backend pool member failures
- Centralized monitoring and alerting integrated with VCF observability stack (Aria Operations)
Constraints
- Maximum 4 Service Engines per SE Group due to ESXi host resource limits in the workload domain
- Avi Controller cluster must share the management domain with SDDC Manager, vCenter, and NSX Manager
- Health monitor traffic must traverse NSX overlay segments, subject to NSX transport node performance
- Certificate lifecycle managed by enterprise PKI with 90-day rotation policy
Assumptions
- Backend application teams maintain the /health endpoint and return accurate HTTP status codes
- NSX overlay network provides consistent sub-millisecond latency between SE and backend servers
- SE OVA images in the vSphere content library are updated during VCF lifecycle manager patches
- Aria Operations management pack for Avi is installed and maintained by the monitoring team
Risks
- SE scale-out blocked during peak demand if ESXi resources exhausted or content library unavailable — mitigate with buffer_se warm standby and resource reservation
- Health monitor false positives during NSX control plane maintenance causing unnecessary pool member flapping — mitigate with increased failed_checks threshold during maintenance windows
- Controller cluster quorum loss (2 of 3 nodes down) leaves SEs in headless mode with stale configuration — mitigate with Controller node placement across fault domains and automated health monitoring
- Certificate expiration cascading to pool member failures if PKI renewal automation fails — mitigate with certificate expiry alerting at 30, 14, and 7 days before expiration
Self-Assessment Discussion Prompts
- A VCDX panelist asks: 'Why did you choose Elastic HA over Dedicated mode for your SE Group? What would happen to your SLA if two Virtual Services on the same SE both experience traffic spikes simultaneously?' Defend your decision with resource utilization data and auto-scale policy.
- The panel challenges: 'Your health monitor detection window is 30 seconds. A competing design uses a 3-second detection window. Why is yours better or worse?' Discuss the tradeoff between detection speed and false positive risk, and how your SLA calculation accounts for the detection window.
- A panelist asks: 'What happens to your load-balanced applications if all three Avi Controller nodes fail simultaneously?' Explain headless SE mode, its limitations (no new configuration, no scale-out, no analytics), and your mitigation strategy for Controller availability.
- The panel asks: 'How do you handle certificate rotation across 50 pool members without causing a rolling outage?' Discuss health monitor SSL verification settings, certificate staging strategies, and how Avi's connection draining interacts with certificate updates.
- A panelist probes: 'Your SE scale-out takes 2-5 minutes. What happens to user experience during that window?' Discuss the impact of SE resource exhaustion on latency, how existing SEs absorb traffic during scale-out, and whether your buffer_se configuration adequately mitigates the gap.
Extensions
Avi WAF Policy Troubleshooting
Extend the L7 Virtual Service with a Web Application Firewall (WAF) policy. Configure WAF in detection mode, generate test traffic with common attack patterns (SQL injection, XSS), analyze WAF logs for false positives, and tune the WAF policy to reduce false positive rate while maintaining protection. This extension adds security observability to the load balancer design.
Multi-Site GSLB Failover with Avi DNS
Configure Global Server Load Balancing (GSLB) across two simulated sites using Avi DNS. Create a GSLB service that routes traffic based on health and proximity. Simulate a site failure and verify automatic DNS failover to the secondary site. Measure DNS TTL impact on failover time. This extension tests disaster recovery design for load-balanced services.
Avi REST API Automation for Incident Response
Use the Avi REST API to automate common troubleshooting tasks: query pool member health status, trigger SE scale-out, export application logs for a specific time window, and create automated health reports. Build a Python script that could be integrated into a runbook for on-call engineers. This extension demonstrates operational automation maturity for VCDX defense.
⚠ Known Pitfalls (from Community KB)
References
- Avi Load Balancer Documentation - VCF 9.0 Integration GuideTier 1 — Official
- Avi Networks Knowledge Base - Health Monitor ConfigurationTier 1 — Official
- VMware Cloud Foundation 9.0 Architecture Guide - Networking and Load BalancingTier 1 — Official
- Avi Service Engine Sizing and Scaling Best PracticesTier 1 — Official
- VCDX Application Design - Load Balancer Section GuidelinesTier 1 — Official