Retry storms: why retries take your service down
Short answer: a retry is one more request, and retries happen exactly when the dependency is already failing. With 3 retries, a dependency that rejects half of its calls receives 1.9× the traffic; one that rejects all of them receives 4×. Stack two layers of retries and the multiplier is 16×.
Retries help against rare, random failures. Against overload or a hard limit they add load, not successes. Cap them with a retry budget, spread them with backoff and jitter, and retry at one layer only.
The arithmetic
If each attempt fails with probability p and a client retries up to r times, one call
turns into 1 + p + p² + … + pʳ attempts on average. With the common default of 3 retries:
| Share of attempts that fail | Attempts per call | Load on the dependency |
|---|---|---|
| 1 % (a healthy service) | 1.01 | +1 % |
| 10 % | 1.11 | +11 % |
| 50 % | 1.88 | +88 % |
| 90 % | 3.44 | +244 % |
| 100 % (down or fully overloaded) | 4.00 | +300 % |
Layers multiply. If a gateway retries 3 times, the service behind it retries 3 times and the database client retries 3 times, one user request can become 4 × 4 × 4 = 64 attempts at the bottom when the database is down.
Why the storm outlives its trigger
The dangerous part is the feedback loop. A short trigger (a deploy, a slow disk, a cache flush, a spike) pushes a service past its capacity. Timeouts and errors trigger retries, the retries add load, the added load causes more timeouts. When the trigger goes away, the retry traffic alone can keep the service above capacity. Researchers call this a metastable failure: the system stays broken in a stable way until someone sheds load (Bronson et al., HotOS 2021).
Timeouts shorter than the dependency's actual latency make it worse: the caller gives up, the dependency still does the work, and the retry asks for the same work again.
A simulated example: retries against a rate limit
Stackrig's API with rate limit template is a gateway that wraps a paid geocoding API with a hard limit of 40 calls/s. 30 requests/s arrive, a response cache answers half of them, so 15 calls/s reach the partner. Every call retries up to 3 times with backoff. These are Stackrig model results for that design, not measurements:
| Scenario | Calls/s at the partner | p99 | Error rate |
|---|---|---|---|
| 1× traffic, 3 retries | 15 | 0.32 s | 0 % |
| 3× traffic, 3 retries | 95 (58 % rejected) | ~1.0 s | 5.6 % |
| 3× traffic, no retries | 45 | 0.25 s | 5.6 % |
| 3× traffic, cache hit ratio 0.5 → 0.8 | 18 | 0.29 s | 0 % |
At 3× traffic, 45 calls/s want through a 40/s limit. Retries more than double the calls the partner receives, make p99 about four times worse, and do not save a single request: the limit lets 40 calls/s succeed either way, so the error rate is 5.6 % with or without them. The fix that works reduces the demand (a better cache hit ratio), not the reaction to it. It is also the cost lever: the partner bills $0.50 per 1,000 accepted calls, so every cache hit saves money, while the rejected retries only cost time.
Fixes that work
- A retry budget. Allow retries only while they are a small share of the traffic, for example 10 % of requests per dependency over the last minute. In a healthy system that never triggers; in an outage it caps the multiplier at 1.1× instead of 4×.
- Exponential backoff with jitter. Backoff alone makes every client retry on the same schedule, in waves. A random delay spreads the retries out (AWS Architecture Blog).
- Retry at one layer. Usually at the edge closest to the user, or at the one layer that knows whether the operation is idempotent. Everything below fails fast.
- Circuit breakers. After a burst of failures, stop calling the dependency for a while and fail fast, so it can recover without your traffic on top.
- Don't retry what can't succeed. A 429 with
Retry-After, a 400, or a dependency that is known to be down: retrying just adds load. - Timeouts from the dependency's real latency. Set them above its p99 under normal load, and give the whole request a deadline so inner retries stop when the caller has given up.
The Google SRE book's chapter on cascading failures covers the same ground from production experience.
See it before it happens
In Stackrig every connection has its own timeout, retries with backoff, retry budget and circuit breaker. Turn the traffic up or slow a dependency down and the retry storm shows up in the numbers: the calls per second on the connection, p99 and the errors. Stackrig publishes where its model is exact and where it is not; the exact tipping point of a retry storm is one of the listed limits on How accurate is Stackrig?
The live demo on the start page shows a smaller cousin of the storm: at 3× traffic the API's threads fill up waiting on an overloaded database.
Open “API with rate limit” in the playground Watch the live demo Get early access