Metastable failure
A server at 80% utilisation, briefly disturbed and then fully restored, can serve nobody at all for hours - because the retries meant to hide the disturbance are now the only thing holding it down.
Most outages end when their cause ends. You restart the process, the dependency comes back, the traffic spike passes, and service returns. That is the model of failure almost everyone carries around, and it is the reason this class of outage takes teams by surprise for hours: the cause was fixed a long time ago and nothing recovered.
Let it settle. It serves 8 requests per tick against a capacity of 10, comfortably. Now stall the server for a moment, then press Put everything back - full capacity, normal traffic, exactly the configuration that was working a minute ago.
Goodput stays at zero.
The three states
The framing comes from Bronson and colleagues, who named the pattern after reviewing a decade of outages at hyperscale.[1]paperMetastable Failures in Distributed Systems A system moves through three states, and the middle one is where almost everything in production actually lives:
Stable. Load is low enough that no fault can push the system somewhere it cannot climb back from.
Vulnerable. The system is healthy. It is serving every request, it is not overloaded, and its dashboards are green. What it has lost is the guarantee that a transient fault stays transient.
Metastable. A trigger has pushed it over, and a sustaining effect now keeps it there independently of whatever caused it.
The vulnerable state is not an overloaded state; a system can run for months or years in the vulnerable state and then get stuck in a metastable state without any increase in load.
Unverified quote - not yet checked word-for-word against the source.
The paper is blunt about why this is so common: production systems choose to run vulnerable, because the alternative is buying capacity you almost never use. That is a defensible trade. It is only a bad one when nobody has worked out where the boundary is.
Why removing the trigger does not help
The mechanism in the widget is the one the paper calls out as the most common sustaining effect: request retries.[1]paperMetastable Failures in Distributed Systems
Follow one request. It is queued. The client waits for its timeout and then gives up - but the server has no way to know that, so the request stays in the queue and is eventually served at full cost, producing nothing. Meanwhile the client has sent a retry, which joins the same queue behind the same backlog and meets the same fate.
That is the whole loop, and it has two properties that make it vicious:
- The server looks healthy. It is at 100% utilisation, completing requests at exactly its rated capacity. Every server-side metric except goodput looks like a busy afternoon.
- The work is self-generating. Each timeout produces a retry, and each retry produces another timeout. The original trigger is not needed for any of it.
The number nobody computes
Here is the uncomfortable arithmetic. You size capacity against offered load, so you think in terms of utilisation: 8 against 10 is 80%, which feels fine.
But once everything is timing out, every request is retried up to maxRetries times, so
the load actually arriving is up to:
offered load x (retries + 1)
Which means the load you can absorb a trigger at is not your capacity. It is:
safe load = capacity / (retries + 1)
With a capacity of 10 and 3 retries, that is 2.5 requests per tick. A system at 25% utilisation is already vulnerable. Set the widget’s load to 3 - 30% utilisation, 70% headroom - hit it with the same brief stall, and it never comes back.
The empirical follow-up work found the same shapes across a range of real systems and workloads, which is some comfort: the pattern generalises, so the reasoning transfers even when your sustaining effect is not retries.[2]paperMetastable Failures in the Wild
Two ways out, and one that does not work
Try the verbs in this order.
Add capacity. There is no verb for it because it is not a fix - go and raise the capacity slider instead, while the system is stuck. Unless you raise it past the retry load rather than the offered load, nothing changes. This is the mitigation teams reach for first, because it is the one that works for ordinary overload.
Shed the queue. Press Shed the queue and goodput returns immediately. Note what this actually did: it threw away work that had already been paid for, and told a large number of clients to go away. Recovery in a metastable failure is not free and it is not graceful, which is exactly why it rarely happens by accident.
Remove the second equilibrium. The better answer is to make the failed state impossible to sustain in the first place. Turn on Retry budget - the client-side throttle from the SRE book, capping retries at 10% of traffic[3]bookSite Reliability Engineering: Handling Overload - and run the same trigger.
Unbounded retries
Every timeout becomes up to three more requests, which is enough to keep the server saturated on its own.
Retry budget
Retries capped at 10% of fresh traffic. The failed state has nothing left to feed on.
| Measure | Unbounded retries | Retry budget | Gap |
|---|---|---|---|
| Goodput | 8.00 | 8.00 | same |
| Wasted share of served work | 0.00 | 0.00 | same |
| Queue depth | 0.00 | 0.00 | same |
Fresh requests per tick, before any retries.
How long a client waits before giving up. The server carries on regardless.
A queue short enough to drain inside the timeout cannot abandon anything.
Both arms take the same hit. Only one of them is still down afterwards, and the difference is not capacity, speed, or how the outage began. With retries capped at 10%, the worst load they can generate is 1.1x offered - 8.8 against a capacity of 10 - which is not enough to keep the server saturated. The failed state has nothing to feed on, so it cannot persist.
What to carry away
Goodput, not throughput. A metastable system is fully busy and completely useless. Any alerting built on utilisation, queue depth or request rate will show you a system under healthy load. Measure work that somebody is still waiting for.
The trigger is not the incident. By the time you are paged, the thing that started it is very likely over. Hunting for it is the natural instinct and it wastes the hour that matters.
Recovery means shedding, not scaling. Adding capacity to a system whose load is self-generated raises the load too. The corrective push has to be a big one - drop the queue, cut the traffic hard, or turn off the retries.[1]paperMetastable Failures in Distributed Systems
The dial
The line - when someone asks
A metastable failure is one that keeps itself alive: a trigger pushes the system into a state where the work it does stops producing anything useful, and something about that state - usually retries - generates enough extra load to keep it there. The tell is that removing the trigger does not help. Capacity is normal, traffic is normal, and goodput is zero, because the server is busy completing requests whose clients gave up long ago. Getting out requires dropping work, not adding capacity.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]
Secondary
- [3]