Academy/VCF 9.0 Support (2V0-15.25)/Avi Load Balancer Troubleshooting
This lab targets VCF 9.0

Avi Load Balancer Troubleshooting

VCF 9.0Intermediatevcp-foundation⏱ 90 min

VCF 9.0 integrates Avi Load Balancer (formerly NSX Advanced Load Balancer) as the standard load balancing solution, replacing the deprecated NSX-T built-in load balancer. Avi Controller 30.2.x is the aligned version for VCF 9.0 deployments.

Objectives

  • Describe the Avi Load Balancer architecture within VCF 9.0 including Controller cluster, Service Engine placement, and NSX integration
  • Configure and validate health monitors (TCP, HTTP, HTTPS, custom) for backend pool members and interpret health check results
  • Diagnose cascading pool member failures by correlating SE resource utilization, backend application health, and health monitor configuration
  • Perform root cause analysis on Virtual Service failures using Avi analytics, application logs, and end-to-end timing data
  • Apply immediate remediation actions including SE scaling, pool member tuning, and health check threshold adjustment
  • Design proactive monitoring dashboards and alert policies to detect load balancer degradation before SLA breach
  • Estimate SLA recovery timelines and articulate load balancer design tradeoffs in a VCDX defense context

Prerequisites

VCF 9.0 workload domain operational with Avi Controller cluster (3-node) deployed and integrated with NSX. At least one Virtual Service configured with a pool of 3+ backend web servers. NSX overlay segments available for SE placement. Avi Controller UI accessible via HTTPS from management workstation.

Prior labs: vcp-support-07, avi-alb-01

Required skills:

  • Basic understanding of L4/L7 load balancing concepts (VIP, pool, health monitor)
  • Familiarity with NSX overlay networking and segment configuration
  • vSphere resource monitoring (CPU, memory, network) for virtual appliances
  • HTTP/HTTPS fundamentals including status codes, SSL/TLS handshake, and connection lifecycle
  • CLI and REST API interaction for troubleshooting (curl, Avi CLI shell)

Lab Environment

Three-node Avi Controller cluster deployed on the management domain. Two Service Engines (SE-1, SE-2) deployed in elastic HA (N+M) mode on the workload domain, connected to NSX overlay segments. One L7 Virtual Service (vs-web-prod) distributing HTTPS traffic across a pool of 4 backend web servers (web-01 through web-04) on an NSX segment. A second L4 Virtual Service (vs-db-prod) proxying TCP/3306 to 2 database servers (db-01, db-02). Health monitors configured: HTTP GET /health for web pool, TCP connect for DB pool. Simulated failure conditions pre-staged: web-03 returning HTTP 503, web-04 with expired SSL certificate, SE-1 approaching 85% CPU utilization.

graph TB
  CLIENT[Client Traffic] -->|HTTPS 443| VIP_WEB[VIP: vs-web-prod 10.50.1.100]
  CLIENT -->|TCP 3306| VIP_DB[VIP: vs-db-prod 10.50.1.101]
  VIP_WEB --> SE1[SE-1 10.50.2.10]
  VIP_WEB --> SE2[SE-2 10.50.2.11]
  VIP_DB --> SE1
  SE1 --> WEB01[web-01 10.50.3.11]
  SE1 --> WEB02[web-02 10.50.3.12]
  SE2 --> WEB03[web-03 10.50.3.13 - HTTP 503]
  SE2 --> WEB04[web-04 10.50.3.14 - SSL Expired]
  SE1 --> DB01[db-01 10.50.4.11]
  SE1 --> DB02[db-02 10.50.4.12]
  CTRL1[Avi Controller-1 10.50.0.21] --- CTRL2[Avi Controller-2 10.50.0.22]
  CTRL2 --- CTRL3[Avi Controller-3 10.50.0.23]
  CTRL1 -.->|Manages| SE1
  CTRL1 -.->|Manages| SE2

IP Addressing

NetworkPurposeVLAN
10.50.0.0/24Avi Controller management networkVLAN 1644
10.50.1.0/24Virtual Service VIP network (client-facing)VLAN 1650
10.50.2.0/24Service Engine data networkVLAN 1651
10.50.3.0/24Web server backend pool network (NSX overlay)NSX Segment web-seg
10.50.4.0/24Database server backend pool network (NSX overlay)NSX Segment db-seg

Credentials

SystemUsernamePassword
Avi Controller UI/APIadminSet during Avi Controller initial setup. Minimum 8 characters, mixed case, special characters required.
Avi CLI ShelladminSame as Avi Controller admin password. Access via SSH to Controller node.
Backend Web ServersrootStandard lab credentials for web-01 through web-04.
NSX ManageradminNSX Manager admin credentials for verifying SE segment placement.

Tasks

Task 1 Avi Architecture & SE Deployment Model in VCF

availability

Before troubleshooting failures, you must understand the Avi distributed architecture in VCF 9.0. The Controller cluster is the control plane; Service Engines are the data plane. Misunderstanding SE placement, HA modes, or Controller-SE communication causes misdiagnosis. VCDX panelists expect candidates to articulate why the architecture is split-brain resilient and how SE placement on NSX overlay segments affects failure domains.

Step 1

Log in to the Avi Controller UI (https://10.50.0.21). Navigate to Infrastructure > Service Engine Group. Document the SE Group configuration: HA mode (elastic HA N+M, dedicated, or legacy HA), minimum and maximum number of SEs, memory and vCPU allocation per SE, and the vCenter/NSX cloud connector binding.

SE Group 'Default-SEG' configured with Elastic HA (N+M) mode, min_scaleout=2, max_scaleout=4, memory_per_se=4 GB, vcpus_per_se=2. Cloud connector bound to the VCF workload domain vCenter and NSX Manager.
Elastic HA N+M is the recommended mode for VCF 9.0 because it allows SEs to scale out automatically when traffic increases. Dedicated mode pins one SE per Virtual Service which wastes resources in multi-tenant environments.
Step 2

Navigate to Infrastructure > Service Engines. For each SE (SE-1, SE-2), record: power state, vCPU count, memory allocation, connected networks (management, VIP, backend segments), and the host on which the SE VM is deployed. Verify that SEs are placed on different ESXi hosts for anti-affinity.

SE-1 on esxi-host-01, SE-2 on esxi-host-02, both powered on with 2 vCPU, 4 GB RAM, connected to management, VIP, and backend NSX segments. Anti-affinity confirmed (different hosts).
SE anti-affinity is critical for availability. If both SEs land on the same host, a single host failure takes down all Virtual Services. Avi enforces anti-affinity via vSphere DRS rules when integrated with vCenter.
Step 3

Verify Controller cluster health. Navigate to Administration > Controller > Nodes. Confirm all 3 Controller nodes show 'CLUSTER_UP' state. Check the cluster VIP is reachable: ping 10.50.0.20 (cluster VIP). Also verify the leader election status by checking which node is the current leader.

All 3 nodes in CLUSTER_UP state. Cluster VIP 10.50.0.20 responding. One node identified as leader, two as followers. Leader election timestamp visible.
The Avi Controller cluster uses a Raft-based consensus protocol. If one Controller goes down, the remaining two maintain quorum. If two go down, the cluster becomes read-only and SEs continue forwarding traffic with their last-known configuration (headless mode).
Step 4

Examine the NSX Cloud connector configuration. Navigate to Infrastructure > Clouds. Select the NSX cloud and verify: NSX Manager IP, transport zone mapping, SE management network, and the content library used for SE image deployment. Document any warnings or errors on the cloud status.

NSX Cloud connector status: GREEN. NSX Manager 10.50.0.10 connected. Transport zone mapped to overlay-tz. SE management network: avi-mgmt-seg. Content library: Avi-SE-Images. No warnings or errors.
When Avi deploys a new SE, it pulls the SE OVA from the vSphere content library and places it on an NSX segment. If the content library is missing or the SE image version mismatches the Controller version, SE deployment fails silently. Always verify the content library contains the correct SE image version matching your Controller.
Step 5

From the Avi CLI shell (SSH to Controller), run: show serviceengine detail | grep -A 5 'se_name\|oper_status\|se_group_ref\|host_ref'. Cross-reference the CLI output with the UI findings to confirm consistency. Also run: show cloud status to verify cloud connector health from the CLI perspective.

CLI output matches UI: 2 SEs operational, both in Default-SEG, deployed on separate hosts. Cloud status shows connected to NSX Manager and vCenter with no errors.
Always cross-reference UI and CLI data during troubleshooting. Discrepancies between UI and CLI often indicate a stale UI cache or Controller-SE communication issues. The CLI pulls real-time data from the Controller database.

Validation Gate

Check: Confirm: (1) SE Group HA mode is Elastic HA N+M, (2) Both SEs are on different ESXi hosts, (3) All 3 Controller nodes are CLUSTER_UP, (4) NSX Cloud connector status is GREEN, (5) CLI and UI data are consistent.

Expected: All 5 architecture validation points confirmed. You have a complete picture of the Avi control plane and data plane topology before proceeding to troubleshooting.

Common Errors

SE shows 'OPER_DOWN' in UI but VM is powered on in vCenter
Cause: SE has lost communication with the Controller cluster. Common causes: management network connectivity loss, Controller cluster VIP unreachable, or SE process crashed.
Fix: Check SE management network connectivity (ping from SE to Controller VIP). Verify NSX segment connectivity for SE management interface. If SE process crashed, restart the SE VM from vCenter. Check Controller logs: /var/log/se_agent.log on the Controller.
NSX Cloud connector shows 'DISCONNECTED' or 'ERROR'
Cause: NSX Manager credentials expired, NSX Manager unreachable, or certificate trust issue between Avi Controller and NSX Manager.
Fix: Re-enter NSX Manager credentials in Infrastructure > Clouds > NSX Cloud > Edit. Verify network connectivity between Controller and NSX Manager. Check certificate validity: the Avi Controller must trust the NSX Manager certificate.
Only 1 SE deployed despite min_scaleout=2 in SE Group
Cause: Insufficient resources on ESXi hosts, DRS anti-affinity rule preventing placement, or content library SE image missing.
Fix: Check vCenter Events for SE VM deployment failures. Verify ESXi host resources (CPU, memory, datastore). Confirm content library has the correct SE OVA image. Check DRS rules for conflicts.

Task 2 Health Monitor & Pool Configuration Troubleshooting

recoverability

Health monitors are the primary mechanism for detecting backend failures. Misconfigured health checks cause false positives (marking healthy servers as down) or false negatives (continuing to send traffic to failed servers). This task builds the skill of diagnosing health monitor behavior, which is the most common source of Avi pool member issues in VCF deployments and a frequent VCDX troubleshooting scenario.

Step 1

Navigate to Applications > Virtual Services > vs-web-prod > Pool (web-pool). Examine each pool member's health status. Identify which members are UP, which are DOWN, and which are in ERROR state. For each DOWN/ERROR member, click the health monitor icon to view the last health check result and failure reason.

web-01: UP, web-02: UP, web-03: DOWN (HTTP health check returned 503 instead of expected 200), web-04: DOWN (health check connection failed - SSL handshake error). Two of four pool members are unhealthy, reducing pool capacity by 50%.
The pool member detail view shows the exact health check response. For HTTP monitors, you can see the full response code and body. For TCP monitors, you see connection state. This is your first diagnostic data point.
Step 2

Examine the health monitor configuration. Navigate to Templates > Profiles > Health Monitors. Open the HTTP health monitor (http-health-check) used by web-pool. Document: type, send interval, receive timeout, successful checks, failed checks, HTTP request method, expected response code, and send/receive strings.

Health monitor: http-health-check, Type: HTTP, Send Interval: 10s, Receive Timeout: 4s, Successful Checks: 3, Failed Checks: 3, Method: GET, Path: /health, Expected Response: 200, No send/receive string matching configured.
The ratio of send_interval to receive_timeout determines how aggressively the monitor detects failures. With 10s interval and 4s timeout, a server must fail 3 consecutive checks (30 seconds) before being marked DOWN. For production SLA requirements, consider reducing failed_checks to 2 for faster detection.
Step 3

Diagnose web-03 (HTTP 503). SSH to web-03 (10.50.3.13) and check the application status: systemctl status nginx (or the appropriate web service). Check the application log for errors. Test the health endpoint locally: curl -v http://localhost/health. Determine why the backend is returning 503.

Nginx is running but the upstream application is overloaded or misconfigured. curl localhost/health returns HTTP 503 with body indicating upstream connection refused or resource exhaustion. Application log shows connection pool exhaustion to the database tier or memory pressure.
A 503 from the backend means the web server itself is healthy but it cannot fulfill requests. This is different from a connection refused (server process down) or timeout (network issue). Understanding the distinction is critical for choosing the right remediation: restart the app vs. scale the database vs. fix the network.
Step 4

Diagnose web-04 (SSL handshake error). From the Avi Controller CLI, run: debug virtualservice vs-web-prod level debug. Then check the health monitor logs: show pool web-pool detail | grep -A 10 'web-04'. Also SSH to web-04 and check the SSL certificate: openssl s_client -connect localhost:443 2>&1 | grep -i 'verify\|expire\|not after'. Identify the certificate issue.

Certificate on web-04 has expired (Not After date is in the past). The Avi health monitor performs SSL verification by default and rejects expired certificates. Health check log shows: SSL_ERROR_EXPIRED_CERT or similar handshake failure.
If the health monitor has SSL verification enabled (verify_server_cert=true), expired or self-signed certificates cause health check failures even if the application is serving traffic correctly. You can temporarily disable verification on the health monitor, but the correct fix is to renew the certificate. In VCDX design, always document your certificate lifecycle management strategy.
Step 5

Remediate both failures. For web-03: work with the application team to resolve the upstream issue, or temporarily disable the member with 'graceful disable' (which drains existing connections). For web-04: renew the SSL certificate (or for lab purposes, disable SSL verification on the health monitor temporarily). After remediation, monitor the health check results: the member should transition from DOWN to UP after 3 consecutive successful checks (30 seconds with default timing). Navigate back to the pool view and verify all members are UP.

After remediation: web-03 gracefully disabled (draining) or restored to UP after upstream fix. web-04 restored to UP after certificate renewal or health monitor SSL verification adjustment. Pool capacity restored to 75% or 100% depending on web-03 resolution approach.
Graceful disable vs. force disable: graceful disable stops new connections but allows existing ones to complete (connection draining). Force disable immediately drops all connections. Always use graceful disable in production to avoid client-visible errors during maintenance windows.

Validation Gate

Check: Confirm: (1) Root cause identified for both web-03 (HTTP 503 from upstream issue) and web-04 (expired SSL certificate), (2) Health monitor configuration documented with timing parameters, (3) At least one member remediated and transitioned to UP state, (4) Pool capacity percentage calculated correctly.

Expected: Both root causes identified and documented. At least web-04 restored to UP after certificate remediation. Health monitor behavior (timing, thresholds, SSL verification) fully understood.

Common Errors

Health monitor shows member as DOWN but curl from SE to backend succeeds
Cause: Health monitor is checking a different port, path, or protocol than the manual curl test. Or the health monitor expects a specific response string that the backend does not return.
Fix: Compare the health monitor configuration (port, path, expected response) with the actual backend endpoint. Use the Avi CLI to see the exact health check request/response: debug healthmonitor web-pool.
Pool member flaps between UP and DOWN repeatedly
Cause: Backend server is borderline unhealthy (intermittent 503s or slow responses near the timeout threshold). The receive_timeout is too close to the server's actual response time.
Fix: Increase receive_timeout to accommodate normal response time variance. Alternatively, increase successful_checks to require more consecutive passes before marking UP, which dampens flapping.
All pool members marked DOWN simultaneously after health monitor change
Cause: New health monitor configuration is incompatible with backend (wrong path, wrong port, wrong expected response code).
Fix: Revert the health monitor change immediately. Test the new health monitor configuration against one member first using the Avi CLI debug before applying to the pool.

Task 3 Virtual Service Failure Diagnosis & Traffic Analysis

performance

Virtual Service failures impact end users directly. This task teaches systematic diagnosis using Avi's built-in analytics engine, which provides real-time metrics, application logs, and client insight data that most traditional load balancers lack. The ability to correlate client-side errors with server-side metrics and SE resource utilization is a differentiating skill for VCDX candidates defending load balancer design decisions.

Step 1

Navigate to Applications > Virtual Services > vs-web-prod. Open the Analytics tab. Set the time range to the last 1 hour. Review the key metrics: End-to-End Timing (client RTT, server RTT, application response time, data transfer time), Throughput (requests/sec, bandwidth), and Error Rate (percentage of 4xx and 5xx responses). Identify any anomalies or spikes.

End-to-End Timing shows elevated application response time (>500ms, normally <100ms). Error rate shows a spike correlating with web-03 and web-04 going DOWN. Throughput may show a dip as pool capacity dropped by 50%. Client RTT is normal, indicating the issue is server-side, not network.
Avi's End-to-End Timing is one of its most powerful diagnostic tools. It breaks down latency into 5 components: Client RTT, Server RTT, App Response, Data Transfer, and Avi processing. This immediately tells you whether the problem is client-side, network, application, or load balancer. Traditional load balancers only show total latency.
Step 2

Open the Logs tab for vs-web-prod. Filter by: HTTP Status >= 500. Examine 5-10 error entries. For each, note: timestamp, client IP, server IP (which pool member handled the request), response code, response time, and the full request URI. Look for patterns: are errors concentrated on specific pool members, specific URIs, or specific time windows?

Error entries show HTTP 503 responses, all originating from web-03 before it was marked DOWN. After web-03 was marked DOWN by the health monitor, 503 errors stop because traffic is redistributed to healthy members. Some entries may show connection reset errors for web-04 due to SSL failures.
The gap between when a server starts failing and when the health monitor marks it DOWN is the 'detection window.' During this window, real client requests hit the failing server. With the default health monitor (10s interval, 3 failed checks), this window is up to 30 seconds. Reducing this window is a key design tradeoff: faster detection vs. more health check traffic and higher false positive risk.
Step 3

Check SE resource utilization. Navigate to Infrastructure > Service Engines > SE-1. Review CPU, memory, network throughput, and connection count metrics for the last hour. Note that SE-1 was pre-staged at 85% CPU utilization. Determine whether SE resource exhaustion is contributing to Virtual Service degradation. Also check: SE-1 > Connected Virtual Services to see how many VS are sharing this SE.

SE-1 CPU at 85% utilization, approaching the default high-water mark (80%). Memory utilization at 60%. SE-1 is handling both vs-web-prod and vs-db-prod. High CPU on SE-1 may cause increased latency for all Virtual Services on this SE. SE-2 shows normal utilization (40% CPU).
When SE CPU exceeds the high-water mark, the Avi Controller should trigger a scale-out event in Elastic HA mode, deploying a new SE (SE-3) and redistributing Virtual Services. If scale-out does not occur, check: (1) max_scaleout limit in SE Group, (2) available ESXi host resources, (3) SE image in content library. The scale-out decision is logged in Controller events.
Step 4

Simulate a client request and trace it end-to-end. From a test client (or the Avi Controller CLI), run: curl -v -H 'Host: app.lab.local' https://10.50.1.100/api/data --resolve app.lab.local:443:10.50.1.100. Observe the response. Then in the Avi UI, find this specific request in the Logs tab (filter by timestamp or URI /api/data). Click the log entry to see the full request trace: which SE handled it, which pool member was selected, load balancing algorithm used, and timing breakdown.

Request successfully served by one of the healthy pool members (web-01 or web-02). Log entry shows: SE that handled the request, pool member selected via Round Robin (or configured algorithm), total response time, and timing breakdown. If SE-1 handled it, response time may be elevated due to SE CPU pressure.
The Avi application log for each request includes a unique request_id that can be correlated with backend application logs. This end-to-end tracing capability is invaluable for troubleshooting intermittent issues. In VCDX defense, articulating this observability advantage over traditional load balancers demonstrates depth of understanding.
Step 5

Trigger and verify SE scale-out. If SE-1 CPU remains above 80%, the Controller should initiate an automatic scale-out. Navigate to Administration > Events. Filter for 'scale' events. If no scale-out has occurred, manually trigger it: navigate to Infrastructure > SE Group > Default-SEG > Edit, and temporarily lower the max_vs_per_se or adjust the buffer_se count to force a new SE deployment. Monitor the SE deployment in vCenter (new VM being created from content library). Once SE-3 is deployed, verify that Virtual Services are redistributed.

Scale-out event logged. New SE-3 deploying from content library (takes 2-5 minutes). After SE-3 is operational, vs-web-prod or vs-db-prod migrates to SE-3, reducing load on SE-1. SE-1 CPU drops below 80%.
SE scale-out in VCF involves multiple integrations: Avi Controller requests a new VM from vCenter, vCenter deploys the OVA from the content library, NSX attaches the SE to the correct overlay segments, and the Controller programs the SE with Virtual Service configuration. Any failure in this chain (content library, NSX segment, ESXi resources) blocks scale-out silently. Always check vCenter Events and Controller Events together.

Validation Gate

Check: Confirm: (1) End-to-End Timing breakdown analyzed and server-side latency identified, (2) Application logs correlated error entries with specific pool members, (3) SE resource exhaustion identified as contributing factor, (4) Client request traced end-to-end through the Avi data path, (5) SE scale-out mechanism understood and verified.

Expected: Complete picture of Virtual Service failure: root causes include pool member failures (web-03 app issue, web-04 cert issue) compounded by SE-1 CPU exhaustion. Remediation includes pool member fixes and SE scale-out.

Common Errors

Avi Logs tab shows 'Log collection disabled' or no entries
Cause: Application logging is not enabled on the Virtual Service or the log aggregation policy is set to only collect significant logs.
Fix: Navigate to vs-web-prod > Analytics Policy. Enable 'Full Client Logs' and set the log duration. Note: full logging increases Controller storage usage. For production, use 'Significant Logs' (errors and slow responses only).
SE scale-out event logged but new SE never appears in vCenter
Cause: Content library missing the SE OVA, insufficient ESXi resources, or vCenter permissions issue for the Avi service account.
Fix: Check vCenter > Content Libraries for the SE image. Verify ESXi host resources. Check the Avi Controller service account permissions in vCenter (requires VM creation, network assignment, datastore access).
Virtual Service shows 'OPER_RESOURCES' warning after scale-out
Cause: Virtual Service requires more SE capacity than available. Common when max_scaleout limit is reached or all SEs are at capacity.
Fix: Increase max_scaleout in SE Group configuration, or increase per-SE resources (vCPU, memory). Alternatively, move less critical Virtual Services to a separate SE Group to free capacity.

Task 4 Avi Operational Monitoring & VCDX Defense

manageability

Troubleshooting is reactive; monitoring is proactive. This task transitions from break-fix to designing ongoing operational monitoring that prevents recurrence. VCDX candidates must demonstrate that their load balancer design includes observability, alerting, and SLA compliance tracking. This task also prepares you to defend Avi design decisions in a VCDX panel scenario.

Step 1

Design an alert policy for pool member health degradation. Navigate to Operations > Alerts > Alert Actions. Create a new alert action that sends notifications when pool capacity drops below 75% (i.e., more than 1 of 4 members is DOWN). Configure: Alert Rule (pool_member_down count >= 2), Alert Action (email/syslog/SNMP trap), and Throttle (no more than 1 alert per 5 minutes to avoid alert storms). Document your alert policy design.

Alert policy created: Rule triggers when 2+ pool members are DOWN. Action sends syslog to central logging server (e.g., vRealize Log Insight / Aria Operations for Logs). Throttle set to 5 minutes. Alert priority: HIGH.
Alert fatigue is a real operational risk. Design alerts with escalation tiers: WARNING at 25% capacity loss (1 of 4 members down), CRITICAL at 50% capacity loss (2 of 4 members down), EMERGENCY at 75% capacity loss (3 of 4 members down). Each tier should have different notification channels (dashboard vs. email vs. PagerDuty).
Step 2

Configure SE resource monitoring thresholds. Navigate to Infrastructure > SE Group > Default-SEG > Edit > Advanced. Set the following thresholds: SE CPU high-water mark: 80%, SE memory high-water mark: 90%, connection high-water mark: 80% of max connections. Verify that auto-scale triggers are configured to respond when these thresholds are breached. Document the auto-scale policy: what triggers scale-out, what triggers scale-in, and what are the cooldown periods.

SE thresholds configured: CPU 80%, Memory 90%, Connections 80%. Auto-scale policy: scale-out triggers at high-water mark, scale-in triggers when utilization drops below 30% for 10+ minutes (cooldown). Minimum SEs: 2, Maximum SEs: 4. Scale-out cooldown: 5 minutes. Scale-in cooldown: 30 minutes.
Asymmetric cooldown periods (short scale-out, long scale-in) are a best practice. You want to add capacity quickly during spikes but remove capacity slowly to avoid oscillation. A 30-minute scale-in cooldown prevents the situation where an SE is removed and traffic immediately spikes again, requiring another scale-out.
Step 3

Build an SLA recovery estimate. Based on the scenario in this lab (50% pool capacity loss, elevated SE CPU), calculate: (1) Mean Time to Detect (MTTD): how long from failure onset to alert trigger, (2) Mean Time to Respond (MTTR): estimated time for an engineer to diagnose and begin remediation, (3) Mean Time to Recover (MTTR-full): time from remediation start to full service restoration, (4) Total SLA impact: estimated error rate during the incident window and time to return to baseline (<0.1% error rate). Document your calculations.

MTTD: ~30 seconds (health monitor detection window: 10s interval x 3 failed checks). MTTR-respond: 5-15 minutes (alert to engineer engagement). MTTR-full: 15-30 minutes (certificate renewal + application fix + health check convergence). Total incident window: ~45-75 minutes. During the incident, error rate was approximately 50% (2 of 4 members down, round-robin distribution). Time to <0.1% error rate: within 2 minutes of pool member restoration (3 successful health checks).
In a VCDX defense, panelists often ask: 'What is your Recovery Time Objective for this component?' Your answer should include MTTD + MTTR + validation time. The health monitor timing directly affects MTTD, which is a design decision. Faster detection (lower interval, fewer failed checks) trades off against false positive risk.
Step 4

Prepare a VCDX defense summary for your Avi Load Balancer design. Document the following design decisions and their justifications: (1) Why Elastic HA N+M over Dedicated mode, (2) Why 3-node Controller cluster instead of single Controller, (3) Health monitor timing choices (interval, timeout, threshold), (4) SE sizing and auto-scale policy rationale, (5) Monitoring and alerting strategy. For each decision, state the requirement it satisfies, the constraint it operates within, and the risk if the decision were different.

Design defense document covering 5 decisions with RCAR justification. Example: Elastic HA N+M chosen because: Requirement = multi-tenant environment with variable traffic, Constraint = limited ESXi host resources (cannot dedicate SEs per VS), Assumption = traffic patterns are bursty not constant, Risk = if max_scaleout is too low, scale-out may be blocked during peak demand.
VCDX panelists look for tradeoff awareness, not perfect answers. For example: 'I chose Elastic HA because it optimizes resource utilization, but I accept the tradeoff that SE scale-out takes 2-5 minutes during which performance may degrade. To mitigate this, I set a buffer_se of 1 to keep a warm standby SE ready.' This demonstrates architectural maturity.
Step 5

Integrate Avi monitoring with the VCF observability stack. Navigate to Administration > Settings > Analytics. Verify that metrics export is configured to send data to Aria Operations (formerly vRealize Operations). Check the integration endpoint, export interval, and which metrics are being exported. Also verify syslog export under Administration > Settings > Syslog for sending application logs to Aria Operations for Logs. Document the end-to-end observability pipeline: Avi metrics/logs -> Aria Operations/Logs -> Dashboards -> Alerts.

Metrics export configured to Aria Operations endpoint (https://aria-ops.lab.local/api). Export interval: 60 seconds. Metrics exported: VS performance, SE health, pool member status, error rates. Syslog export configured to Aria Operations for Logs (10.50.0.30:514, UDP). Full observability pipeline documented.
In VCF 9.0, the recommended observability stack is Aria Operations + Aria Operations for Logs. Avi's native analytics are excellent for deep-dive troubleshooting, but Aria provides the single-pane-of-glass view across the entire VCF stack (compute, storage, network, load balancing). VCDX candidates should demonstrate how Avi monitoring integrates into the broader VCF operations model.

Validation Gate

Check: Confirm: (1) Alert policy created for pool capacity degradation with escalation tiers, (2) SE auto-scale thresholds and cooldown periods configured, (3) SLA recovery timeline calculated with MTTD/MTTR, (4) VCDX defense document prepared with 5 design decisions and RCAR justification, (5) Avi monitoring integrated with VCF observability stack (Aria Operations).

Expected: Complete operational monitoring framework established. VCDX defense preparation document created demonstrating tradeoff awareness and architectural justification for all Avi design decisions.

Common Errors

Alerts not triggering despite pool members being DOWN
Cause: Alert rule is configured on the wrong object (Virtual Service vs. Pool) or the threshold condition does not match the actual metric name in Avi.
Fix: Verify the alert rule references the correct metric: 'pool.num_servers_down' not 'virtualservice.num_se_down'. Test the alert by temporarily lowering the threshold to trigger on the current state.
Aria Operations shows no Avi data despite export being configured
Cause: Firewall blocking the metrics export port, Aria Operations management pack for Avi not installed, or authentication credentials for the export endpoint are incorrect.
Fix: Verify network connectivity from Avi Controller to Aria Operations endpoint. Install the Avi management pack in Aria Operations. Re-enter export credentials and test the connection from the Avi UI.
SE scale-in removes an SE during business hours causing brief traffic redistribution
Cause: Scale-in cooldown period too short or scale-in not restricted to maintenance windows.
Fix: Increase scale-in cooldown to 60+ minutes. Optionally, configure scale-in to only occur during defined maintenance windows using the Avi scheduler. In production, many teams disable automatic scale-in entirely and handle it manually.

Final Validation

This lab validated your ability to understand Avi Load Balancer architecture in VCF 9.0, diagnose health monitor and pool member failures, analyze Virtual Service performance using Avi analytics, and design proactive monitoring to prevent SLA breaches. You also prepared VCDX defense material articulating load balancer design tradeoffs with RCAR justification.

✓ Avi architecture documented: Controller cluster (3-node), SE deployment mode (Elastic HA N+M), NSX cloud connector status, SE anti-affinity verified → All architecture components validated and consistent between UI and CLI

✓ Pool member failures diagnosed: web-03 root cause (upstream application returning HTTP 503), web-04 root cause (expired SSL certificate causing health check failure) → Both root causes identified with evidence from health monitor logs and backend investigation

✓ Virtual Service analytics analyzed: End-to-End Timing breakdown, application log correlation, SE resource utilization impact on latency → Server-side latency spike correlated with pool capacity loss and SE CPU exhaustion

✓ SE scale-out mechanism verified: auto-scale trigger, SE deployment from content library, Virtual Service redistribution → Scale-out completed successfully, SE-1 CPU reduced below high-water mark

✓ Operational monitoring designed: alert policy with escalation tiers, SLA recovery timeline (MTTD + MTTR), VCDX defense document with 5 design decisions → Complete monitoring framework and defense preparation document with RCAR justification for each design decision

Cleanup / Restore

• Remove debug logging on vs-web-prod: navigate to the Virtual Service and disable debug-level analytics to reduce Controller storage consumption

• Revert SE Group max_scaleout to original value if it was modified during Task 3 to force scale-out

• Restore web-04 SSL certificate to a valid certificate (if a temporary health monitor workaround was applied, re-enable SSL verification)

• Revert to the pre-avi-troubleshooting snapshot if a clean state is required for subsequent labs

Design Reflection (VCDX)

Avi Load Balancer in VCF 9.0 represents a significant architectural shift from the NSX-T built-in load balancer. VCDX candidates must defend why the distributed control-plane/data-plane architecture (Controller cluster + Service Engines) is superior for enterprise workloads: independent scaling of management and data path, headless SE operation during Controller outages, and deep application-layer analytics. The key design tension is between resource efficiency (Elastic HA sharing SEs across Virtual Services) and isolation (Dedicated mode guaranteeing SE resources per VS). Your defense should articulate how SE sizing, auto-scale policy, health monitor timing, and monitoring integration form a cohesive availability strategy that meets specific RPO/RTO requirements.

Requirements

  • L7 load balancing with SSL offload for production web applications on VCF 9.0 workload domains
  • 99.95% availability SLA for load-balanced services (maximum 22 minutes downtime per month)
  • Sub-second failover detection for backend pool member failures
  • Centralized monitoring and alerting integrated with VCF observability stack (Aria Operations)

Constraints

  • Maximum 4 Service Engines per SE Group due to ESXi host resource limits in the workload domain
  • Avi Controller cluster must share the management domain with SDDC Manager, vCenter, and NSX Manager
  • Health monitor traffic must traverse NSX overlay segments, subject to NSX transport node performance
  • Certificate lifecycle managed by enterprise PKI with 90-day rotation policy

Assumptions

  • Backend application teams maintain the /health endpoint and return accurate HTTP status codes
  • NSX overlay network provides consistent sub-millisecond latency between SE and backend servers
  • SE OVA images in the vSphere content library are updated during VCF lifecycle manager patches
  • Aria Operations management pack for Avi is installed and maintained by the monitoring team

Risks

  • SE scale-out blocked during peak demand if ESXi resources exhausted or content library unavailable — mitigate with buffer_se warm standby and resource reservation
  • Health monitor false positives during NSX control plane maintenance causing unnecessary pool member flapping — mitigate with increased failed_checks threshold during maintenance windows
  • Controller cluster quorum loss (2 of 3 nodes down) leaves SEs in headless mode with stale configuration — mitigate with Controller node placement across fault domains and automated health monitoring
  • Certificate expiration cascading to pool member failures if PKI renewal automation fails — mitigate with certificate expiry alerting at 30, 14, and 7 days before expiration

Self-Assessment Discussion Prompts

  1. A VCDX panelist asks: 'Why did you choose Elastic HA over Dedicated mode for your SE Group? What would happen to your SLA if two Virtual Services on the same SE both experience traffic spikes simultaneously?' Defend your decision with resource utilization data and auto-scale policy.
  2. The panel challenges: 'Your health monitor detection window is 30 seconds. A competing design uses a 3-second detection window. Why is yours better or worse?' Discuss the tradeoff between detection speed and false positive risk, and how your SLA calculation accounts for the detection window.
  3. A panelist asks: 'What happens to your load-balanced applications if all three Avi Controller nodes fail simultaneously?' Explain headless SE mode, its limitations (no new configuration, no scale-out, no analytics), and your mitigation strategy for Controller availability.
  4. The panel asks: 'How do you handle certificate rotation across 50 pool members without causing a rolling outage?' Discuss health monitor SSL verification settings, certificate staging strategies, and how Avi's connection draining interacts with certificate updates.
  5. A panelist probes: 'Your SE scale-out takes 2-5 minutes. What happens to user experience during that window?' Discuss the impact of SE resource exhaustion on latency, how existing SEs absorb traffic during scale-out, and whether your buffer_se configuration adequately mitigates the gap.

Extensions

Avi WAF Policy Troubleshooting

Extend the L7 Virtual Service with a Web Application Firewall (WAF) policy. Configure WAF in detection mode, generate test traffic with common attack patterns (SQL injection, XSS), analyze WAF logs for false positives, and tune the WAF policy to reduce false positive rate while maintaining protection. This extension adds security observability to the load balancer design.

Multi-Site GSLB Failover with Avi DNS

Configure Global Server Load Balancing (GSLB) across two simulated sites using Avi DNS. Create a GSLB service that routes traffic based on health and proximity. Simulate a site failure and verify automatic DNS failover to the secondary site. Measure DNS TTL impact on failover time. This extension tests disaster recovery design for load-balanced services.

Avi REST API Automation for Incident Response

Use the Avi REST API to automate common troubleshooting tasks: query pool member health status, trigger SE scale-out, export application logs for a specific time window, and create automated health reports. Build a Python script that could be integrated into a runbook for on-call engineers. This extension demonstrates operational automation maturity for VCDX defense.

⚠ Known Pitfalls (from Community KB)

Assuming NSX-T LB Configuration Applies to Avi
Problem: VCF 9.0 replaced the NSX-T built-in load balancer with Avi. Configuration concepts (pools, monitors, VIPs) are similar but the implementation, CLI, API, and operational model are completely different. Do not reference NSX-T LB documentation or commands when working with Avi.
Ignoring SE-to-Controller Version Alignment
Problem: Service Engine image version must match the Controller version exactly. After a Controller upgrade, existing SEs must be upgraded (rolling or disruptive). Version mismatch causes SE communication failures and potential data plane outages. Always verify SE version alignment after any Controller maintenance.
Overlooking Health Monitor Impact on Backend Servers
Problem: Aggressive health monitor settings (1-second interval across 50 pool members) generate significant synthetic traffic to backend servers. In resource-constrained environments, health check traffic itself can contribute to backend overload. Calculate total health check requests per second across all monitors and pools before finalizing timing.
Confusing SE Data Plane Availability with Controller Availability
Problem: If all three Controller nodes fail, Service Engines continue forwarding traffic using their last-known configuration (headless mode). This is a strength of the architecture but has limitations: no new VS configuration, no auto-scale, no analytics collection. Design your Controller availability strategy separately from your data plane availability strategy.

References

Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.