yield point
Performance & Systemsintermediateupdated 2026-08-31

Hedged requests

Sending 5% of your requests twice cuts p99 by a quarter. The same setting, on a busy day, can make it forty times worse.

A request that has already taken longer than 95% of requests take is, on the evidence, having a bad time. It landed on a machine that is garbage collecting, or behind a slow neighbour, or on a disk that is busy.

So send another one.[1]paperThe Tail at ScaleDean, J. & Barroso, L. A., Communications of the ACM 56(2), 2013

Hedging at the 95th percentile costs 5% more traffic and cuts p99 by about a quarter. Now turn off 'Cancel the loser' and drag utilisation past 0.6, and watch the same setting make p99 forty times worse.

110second copy sentp50p99latency (ticks)
p99
0 ticks
One attempt, no queue
129 ticks
Theory says
98 ticks
p50
0 ticks
Extra attempts
0%
Second copies sent
after 75 ticks
Loser cancelled
yes
Waiting for a worker
0
tick 0 / 20000
Break it
Do

Share of the workers the ordinary traffic already commits. Hedges are extra.

Log-normal sigma. Zero would mean every request takes exactly the same time, which is the one case hedging cannot help.

100 disables hedging entirely. Lower means more copies, sent sooner.

Stop the other copy as soon as one answers, instead of letting it run to completion.

Safety properties
  • ✓
    No second copy is sent before its first has outlived the percentile

    held on every tick so far

  • ✓
    Every attempt is running, queued, finished, or cancelled

    held on every tick so far

The cost is exactly the tail

The second copy goes out only when the first has already outlived the threshold, so the extra traffic is the tail beyond it. Hedge at the 95th percentile and you send 5% more requests. At the 99th, 1%.

The benefit is a product, not a sum

The client waits only if both copies are slow, and the copies are independent, so the two tail probabilities multiply:

P(still waiting at t) = P(L > t) x P(L > t - d)

Squaring a small number is what does the work. At the default settings that takes p99 from about 129 ticks to about 98.

Set Tail heaviness near zero and the benefit vanishes. That is worth understanding rather than memorising: hedging pays when being slow is bad luck, and buys nothing when it is the normal case. If your tail comes from one consistently overloaded shard rather than from variance, the second copy lands on the same shard and you have paid 5% for nothing.

The threshold is calibrated on a system that is fine

Identical arrivals, workers and service times. The right arm sends a second copy of any request that outlives the 95th percentile, and takes whichever answers first. Then drag utilisation up, and turn off 'Cancel the loser' to see the version of this that takes a service down.

No hedging

One attempt per request. The tail is whatever the service-time distribution says it is.

10100p50p99latency (ticks)
p99
123 ticks
One attempt, no queue
129 ticks
Theory says
129 ticks
p50
19 ticks
Extra attempts
0%
Second copies sent
never
Loser cancelled
yes
Waiting for a worker
0

Hedge at p95

A second copy after the 95th percentile, so both copies must be slow for the client to wait.

10100second copy sentp50p99latency (ticks)
p99
96 ticks
One attempt, no queue
129 ticks
Theory says
98 ticks
p50
19 ticks
Extra attempts
5%
Second copies sent
after 75 ticks
Loser cancelled
yes
Waiting for a worker
0
Measured at tick 6000 - both sides, same seed, same inputs
MeasureNo hedgingHedge at p95Gap
p99 latency123.096.01.3×
Extra attempts sent0.000.05∞
tick 6000
Break it
Do

Share of the workers the ordinary traffic already commits. Hedges are extra.

Log-normal sigma. Zero would mean every request takes exactly the same time, which is the one case hedging cannot help.

Stop the other copy as soon as one answers, instead of letting it run to completion.

Now turn off Cancel the loser and drag Utilisation before hedging past 0.6.

p99 does not degrade. It explodes, to tens of times worse than not hedging at all.

The mechanism is worth stating slowly, because the number that made hedging affordable is the one that fails. The threshold is a fixed percentile of the healthy latency distribution. Once queueing pushes ordinary requests past it, every request qualifies for a hedge. The extra load goes from 5% to nearly 100%. That extra load is what caused the queueing. The system now has two equilibria and has been shoved into the bad one, which is exactly the shape described on the metastable failure page.

What makes it safe

Cancel the loser. This is not an optimisation, it is the thing that bounds the feedback loop. Without it a hedge costs a whole extra unit of work; with it the wasted work is capped at the delay. Turn cancellation on in the widget and the same 85% utilisation that produced the explosion is fine.[3]docsgRPC request hedging configuration

Cap the hedge rate. Budget hedges as a share of traffic - say 5% - and stop issuing them when the budget is spent. This converts the runaway into a graceful loss of benefit, and it is the single change that turns the technique from conditionally dangerous into safe.

Never hedge below the percentile you measured. A hedge at p50 is not a hedge, it is sending everything twice. A retry loop with no delay is a hedge with the threshold at zero, which is the same bug arriving through a different door.

Do not hedge a write, or anything else that is not idempotent. Two copies means the work may happen twice, and the redundancy literature is about reads for this reason.[2]paperLow Latency via RedundancyVulimiri, A. et al., CoNEXT 2013, 2013

What to do with this

Hedge reads, above p95, with cancellation and a rate cap, on services whose tail comes from variance rather than from a hot shard. That combination is one of the best latency improvements available for the effort.

Any of those four missing turns it into a way to convert a slow afternoon into an outage.

The dial

You gainA tail that requires two independent pieces of bad luck instead of one, for a few percent more traffic
You payA threshold calibrated on a healthy system, which stops being a few percent exactly when you need it most

The line - when someone asks

A request that has already outlived the 95th percentile is having a bad time, so send a second copy and take whichever answers first. The extra traffic is exactly the tail beyond the threshold, so 5% by construction, and the client now waits only if both copies are slow - two independent tails multiplying, which is what collapses the p99. The trap is that the threshold is a percentile of the healthy distribution: once queueing pushes ordinary requests past it, every request earns a hedge, the cost goes to 100%, and that is the load which caused the queueing. Cancelling the loser is what stops the loop.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary

Secondary