yield point
Performance & Systemsintermediateupdated 2026-08-22

Backpressure and bufferbloat

A queue is not a solution to being overloaded; it is a place to store the evidence that you are, and a bigger one only means you find out later and hurt more.

The instinct when a system falls behind is to add a buffer. It feels like giving the system room to breathe. It is usually the opposite: you have converted a visible, immediate failure into an invisible, delayed one, and made every request slower on the way.

Push arrivals just past service capacity and watch the queue go unbounded. Then make the buffer bigger and notice that nothing gets fixed - the wait just gets worse.

Queue0/120
in 0out 0
Utilisation
0.78
Mean wait
0.0 ticks
Completed
0
tick 0 / 2000
Break it
Do

The bigger this is, the longer the queue can hide the problem before anything is rejected.

On: reject immediately when the queue is full. Off: let it grow to capacity and buffer.

Measured vs theory
MetricSimulatedFormulaError
utilisation0.7780.7780.0%

Utilisation is the whole story

Let λ be the arrival rate and μ the service rate. Their ratio ρ = λ/μ is utilisation, and it decides everything.

Below 1, the queue is stable: it wanders up and down but keeps returning to empty. At or above 1, it grows without bound - not slowly, not eventually, but by construction, because work is arriving faster than it can possibly leave.

The part that surprises people is how badly things degrade before reaching 1. Waiting time climbs roughly as 1/(1-ρ), so going from 50% to 85% utilisation does not make things 70% worse - it makes them several times worse. Drag arrivals up slowly in the widget and watch the mean wait; the curve is not a straight line.

Little’s Law

L = λW: the average number of items in a system equals the arrival rate times the average time each spends there.[3]paperA Proof for the Queuing Formula L = λWLittle, J. D. C., Operations Research 9(3), 1961

It is almost embarrassingly simple and it holds for essentially any stable system, regardless of the arrival distribution or service discipline. Rearranged as W = L/λ, it gives you the practical form: if you know your queue depth and your arrival rate, you know your latency. A queue holding 500 items at 100 requests per second is adding five seconds, and no amount of optimising the workers will change that arithmetic.

Why a bigger buffer makes it worse

This is the bufferbloat result.[1]paperBufferbloat: Dark Buffers in the InternetGettys, J. & Nichols, K., ACM Queue 9(11), 2011

Throughput is capped by the service rate. A buffer cannot raise it. So once you are saturated, every extra slot of buffer contributes exactly zero throughput and exactly one slot’s worth of extra waiting for everything behind it.

Set arrivals above service capacity in the widget and compare a queue capacity of 20 with 400. The completed-work counter barely moves. The mean wait explodes.

The fix that came out of the bufferbloat work was CoDel, which does not measure queue length at all - it measures how long items have been waiting, and starts dropping when that persists above a target.[2]paperControlling Queue DelayNichols, K. & Jacobson, V., ACM Queue 10(5), 2012 Length is a proxy that varies with rate; delay is the thing you actually care about.

Backpressure versus load shedding

There are only three honest responses to more work than you can do.

  • Backpressure - refuse to accept more, and let the pressure propagate to whoever is producing. Bounded queues, blocking writes, TCP’s receive window, gRPC flow control.
  • Load shedding - accept the connection and reject the work immediately, cheaply, with a clear error.[4]bookSite Reliability Engineering - Handling OverloadGoogle
  • Buffering - pretend you can do it, and fail later, slower, and less legibly.

Toggle Shed load when full in the widget. The same amount of work is lost either way. The difference is that shedding tells the client now, while it can still retry elsewhere or degrade gracefully, instead of after a timeout on a request that was doomed the moment it was queued.

Where you have already met this

TCP’s receive window is backpressure. So is gRPC and HTTP/2 flow control, a bounded ArrayBlockingQueue, a thread pool with a bounded work queue, Kafka consumer lag as a signal to slow producers, and every reactive-streams request(n).

The systems that get this wrong are usually the ones with an unbounded queue somewhere in the middle - an unbounded channel, an executor with an unbounded backlog, a message consumer with no prefetch limit. It works perfectly until the day it does not, and then it fails as an out-of-memory kill rather than as a slowdown you could have seen coming.

The dial

You gainBuffering absorbs bursts, so a brief spike does not become dropped work
You payEvery item of buffer is latency, and past the point of saturation it is pure latency with no throughput gained

The line - when someone asks

Backpressure is letting a slow consumer tell a fast producer to slow down, instead of buffering the difference. Without it, a system whose arrival rate exceeds its service rate grows an unbounded queue, and because throughput is capped by the service rate, all that queue adds is delay. The counter-intuitive part is that a larger buffer makes things worse, not better - it defers the failure and inflates the latency of everything that survives.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary

Secondary