Backpressure and bufferbloat
A queue is not a solution to being overloaded; it is a place to store the evidence that you are, and a bigger one only means you find out later and hurt more.
The instinct when a system falls behind is to add a buffer. It feels like giving the system room to breathe. It is usually the opposite: you have converted a visible, immediate failure into an invisible, delayed one, and made every request slower on the way.
Utilisation is the whole story
Let λ be the arrival rate and μ the service rate. Their ratio ρ = λ/μ is utilisation,
and it decides everything.
Below 1, the queue is stable: it wanders up and down but keeps returning to empty. At or above 1, it grows without bound - not slowly, not eventually, but by construction, because work is arriving faster than it can possibly leave.
The part that surprises people is how badly things degrade before reaching 1. Waiting
time climbs roughly as 1/(1-ρ), so going from 50% to 85% utilisation does not make things
70% worse - it makes them several times worse. Drag arrivals up slowly in the widget and
watch the mean wait; the curve is not a straight line.
Little’s Law
L = λW: the average number of items in a system equals the arrival rate times the average
time each spends there.[3]paperA Proof for the Queuing Formula L = λW
It is almost embarrassingly simple and it holds for essentially any stable system,
regardless of the arrival distribution or service discipline. Rearranged as W = L/λ, it
gives you the practical form: if you know your queue depth and your arrival rate, you
know your latency. A queue holding 500 items at 100 requests per second is adding five
seconds, and no amount of optimising the workers will change that arithmetic.
Why a bigger buffer makes it worse
This is the bufferbloat result.[1]paperBufferbloat: Dark Buffers in the Internet
Throughput is capped by the service rate. A buffer cannot raise it. So once you are saturated, every extra slot of buffer contributes exactly zero throughput and exactly one slot’s worth of extra waiting for everything behind it.
Set arrivals above service capacity in the widget and compare a queue capacity of 20 with 400. The completed-work counter barely moves. The mean wait explodes.
The fix that came out of the bufferbloat work was CoDel, which does not measure queue length at all - it measures how long items have been waiting, and starts dropping when that persists above a target.[2]paperControlling Queue Delay Length is a proxy that varies with rate; delay is the thing you actually care about.
Backpressure versus load shedding
There are only three honest responses to more work than you can do.
- Backpressure - refuse to accept more, and let the pressure propagate to whoever is producing. Bounded queues, blocking writes, TCP’s receive window, gRPC flow control.
- Load shedding - accept the connection and reject the work immediately, cheaply, with a clear error.[4]bookSite Reliability Engineering - Handling Overload
- Buffering - pretend you can do it, and fail later, slower, and less legibly.
Toggle Shed load when full in the widget. The same amount of work is lost either way. The difference is that shedding tells the client now, while it can still retry elsewhere or degrade gracefully, instead of after a timeout on a request that was doomed the moment it was queued.
Where you have already met this
TCP’s receive window is backpressure. So is gRPC and HTTP/2 flow control, a bounded
ArrayBlockingQueue, a thread pool with a bounded work queue, Kafka consumer lag as a
signal to slow producers, and every reactive-streams request(n).
The systems that get this wrong are usually the ones with an unbounded queue somewhere in the middle - an unbounded channel, an executor with an unbounded backlog, a message consumer with no prefetch limit. It works perfectly until the day it does not, and then it fails as an out-of-memory kill rather than as a slowdown you could have seen coming.
The dial
The line - when someone asks
Backpressure is letting a slow consumer tell a fast producer to slow down, instead of buffering the difference. Without it, a system whose arrival rate exceeds its service rate grows an unbounded queue, and because throughput is capped by the service rate, all that queue adds is delay. The counter-intuitive part is that a larger buffer makes things worse, not better - it defers the failure and inflates the latency of everything that survives.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]
- [3]
Secondary
- [4]