yield point
Performance & Systemsintermediateupdated 2026-08-22

Tail latency amplification

A service where 99% of calls are fast becomes a service where most calls are slow, the moment one user request has to touch a hundred machines.

Most latency intuition is built on a single server answering a single request. Fan-out breaks that intuition completely, and it breaks it in a direction nobody expects: the system gets slower for the typical user precisely because you made each server’s job smaller.

Drag fan-out up and watch the user distribution (green) walk away from the single-call distribution (grey). At N=100 the p99 of one call becomes the typical user's experience.

110latency (ms)
Fan-out
100
One call p50 / p99
NaN / NaN ms
User p50 / p99
NaN / NaN ms
User p50 ÷ call p50
-
tick 0 / 3000
Break it

How many servers one user request must wait on.

Send a second copy when the first is late. A call is only slow if both are.

Why the tail eats the median

A user request that fans out to N servers in parallel finishes only when the last response arrives. If any one server independently exceeds some latency threshold with probability p, then the user request exceeds it with probability 1 − (1 − p)^N.

That expression is unremarkable until you put real numbers in it. At p = 0.01 - a server that is slow for just one call in a hundred - a fan-out of N = 100 gives 1 − 0.99¹⁰⁰ ≈ 0.634. A “1% slow” backend produces a service that is slow for roughly two out of every three users.[1]paperThe Tail at ScaleDean, J. & Barroso, L. A., Communications of the ACM 56(2), 2013

If a user request must collect responses from 100 such servers in parallel, then 63% of user requests will take more than one second

Dean, J. & Barroso, L. A., The Tail at Scale, Communications of the ACM 56(2), 2013p. 76[1]

The consequence is that p99 of a component is not a rare event at the system level. Once fan-out is wide, the component’s tail is the system’s common case, and optimising the component’s median buys you almost nothing.

Why your measurements hide it

Load generators typically issue a request, wait for the response, then issue the next one. When the system stalls, the generator stalls with it - so the requests that would have landed during the stall are never sent, and never recorded. The result is a histogram that omits precisely the samples that mattered.[2]talkHow NOT to Measure LatencyTene, G., 2015

Gil Tene named this coordinated omission, and it is why so many p99 numbers in production dashboards are quietly fictional. Recording latency against intended send time rather than actual send time is the correction; HdrHistogram implements it directly.[3]source codeHdrHistogram - a High Dynamic Range histogramTene, G.

What actually works

The paper’s framing is that tail tolerance should be a property of the system rather than an absence of slow components - you will never eliminate the tail, so route around it.[1]paperThe Tail at ScaleDean, J. & Barroso, L. A., Communications of the ACM 56(2), 2013

  • Hedged requests - send to one replica, and if it has not answered by the 95th percentile, send a second. Small extra load, large tail reduction.
  • Tied requests - send to two replicas immediately, each told about the other, so whichever dequeues first cancels its twin.
  • Good-enough responses - return with 95 of 100 shards answered rather than waiting on the straggler, when the workload tolerates it.
  • Micro-partitioning and probation - many small partitions per machine so a slow machine can be shed quickly.

All four are the same move: stop making the user wait on the maximum.

The dial

You gainFanning out cuts the work any single server does, so median latency falls
You payThe user waits on the slowest of N responses, so your p99 becomes their typical

The line - when someone asks

Tail latency amplification is what happens when one user request fans out to many servers: the user waits for the slowest response, so a rare slow call stops being rare. With 100 parallel calls that each have a 1% chance of exceeding a second, about 63% of user requests exceed a second. You fix it by not waiting on everything - hedged and tied requests, and returning good-enough partial results.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary

Secondary