Skip to content
Guide

Reliability Friday 09: How Load Balancer Retries Multiply Backend Traffic

A user sends one request; with two retries at the client and two at the load balancer, the backend sees nine attempts. Count the layers that retry this week.

Oğuzhan Gerçek··3 min read
Reliability Friday 09: How Load Balancer Retries Multiply Backend Traffic — 1/8Reliability Friday 09: How Load Balancer Retries Multiply Backend Traffic — 2/8Reliability Friday 09: How Load Balancer Retries Multiply Backend Traffic — 3/8Reliability Friday 09: How Load Balancer Retries Multiply Backend Traffic — 4/8Reliability Friday 09: How Load Balancer Retries Multiply Backend Traffic — 5/8Reliability Friday 09: How Load Balancer Retries Multiply Backend Traffic — 6/8Reliability Friday 09: How Load Balancer Retries Multiply Backend Traffic — 7/8Reliability Friday 09: How Load Balancer Retries Multiply Backend Traffic — 8/8

Short answer: Your users and your traffic did not grow, yet the number of backend requests can multiply by nine. The user thinks they sent a single request; that request passes through the client, the CDN / WAF, the load balancer, the ingress and the service, and every one of those layers can carry a retry policy of its own. Client: 1 initial attempt + 2 retries = 3 attempts. Load balancer, for each client attempt: 1 initial attempt + 2 retries = 3 attempts. Result: 3 × 3 = 9. One request does not necessarily stay one request at the backend.

Don't Deploy on Friday. Verify on Friday.

This week's check

Find the layers in your system that retry: client / SDK, CDN / WAF, load balancer, ingress / service mesh, application.

# NGINX
nginx -T 2>&1 | grep -E 'proxy_next_upstream|proxy_next_upstream_tries|proxy_next_upstream_timeout'

# HAProxy
grep -E 'retries|retry-on|redispatch' /etc/haproxy/haproxy.cfg

Do you know how many layers retry?

Expected output

In NGINX, a proxy_next_upstream directive means the load balancer retries; proxy_next_upstream_tries and proxy_next_upstream_timeout say where that retry stops. If no limit is written, NGINX tries every server in the upstream group in turn. In HAProxy, retries sets the number of attempts, retry-on the failures that trigger them, and redispatch whether an attempt may be sent to a different server.

The output itself is not the problem. The problem is that nobody has counted how many layers these lines appear in at once. Run the same query against the settings of the client SDK, the CDN / WAF and the ingress / service mesh.

Measure the hidden traffic. The ratio to watch:

Upstream Attempts / Incoming Requests   target ≈ 1

If this ratio drifts away from 1, the invisible traffic in your system is growing. Watch it together with: retry rate and retry reason, exhausted retries, backend saturation, timeouts and p99 latency, duplicate transactions, and the success rate after a retry.

When to act

The acceptable value of the ratio depends on your own targets; but if it drifts away from 1 while backend saturation and p99 latency rise together, the retries are no longer rescuing requests, they are feeding the outage. The same holds when the success rate after a retry falls: a retry that fails only adds new traffic to the incident.

Risk

The failure loop runs like this: the backend slows down, timeouts begin, the client and the load balancer retry, the backend receives more requests, latency and timeouts climb further, and the loop starts again. A retry can turn a partial failure into a full outage.

The second risk is not about traffic. If the retried operation is POST /orders, POST /payments or PATCH /inventory, the second attempt may not merely create traffic; it may perform the same operation again: a duplicate order, a double payment, wrong inventory, a repeated notification. On non-idempotent requests a retry is not an availability setting, it is a data consistency risk.

Automate it

A weekly manual check is a starting point, not the destination. Make retry policies visible in code and configuration:

    Limit them to transient errors only.Set a total retry budget.Use exponential backoff with jitter.Coordinate them across layers.Protect non-idempotent operations.

Then put Upstream Attempts / Incoming Requests on a dashboard and alert when it drifts away from 1; that is how invisible traffic becomes visible. The layer inventory and that dashboard are part of our observability and APM and network engineering services.

Eclit note

Retry is not free capacity. A retry can rescue a request; an uncontrolled retry can grow an outage.

In the previous episode on this site we filled in the incident escalation matrix; that check was about who holds the decision. This week we are chasing traffic that nobody decided to send. The preventive maintenance side belongs to our system reliability services.

Every Friday. One Production Check.