Rate limiting and the boundary burst
A fixed-window limiter set to 60 per minute will let 120 through in an instant, and the admitted-percentage graph you would use to check it reads exactly the same as one that works.
Rate limiting is one of those things everyone implements and almost nobody tests against an adversary. The default implementation - a counter and a timestamp per client, reset every minute - is wrong in a specific, exploitable way, and the metric you would naturally use to check it cannot see the problem.
Under steady traffic it looks perfect. Now press Time the burst to the boundary.
Twice the limit, by construction
A fixed window keeps a count and resets it when the clock rolls over. It has no memory
across its own reset. So a client that sends its full allowance in the last moment of one
window and its full allowance in the first moment of the next puts 2 x limit through an
interval that can be arbitrarily short.
window N | window N+1
............#####|#####............
60 60
\_______/
120 in an instant
This is not an implementation bug you can patch. It is what a fixed window is. And the boundary is trivial to find: reset at the top of the minute and any client with a clock is aimed at it, whether they meant to be or not - a fleet that retries on a schedule produces the same pattern by accident.
Why it survives in production
Fixed window
Resets on a wall-clock boundary with no memory of what came before it.
Sliding window
Counts what is genuinely inside the trailing window, so there is no boundary to aim at.
| Measure | Fixed window | Sliding window | Gap |
|---|---|---|---|
| Most admitted in one window | 120.0 | 60.0 | 2.0× |
| Requests admitted | 0.50 | 0.50 | same |
| Rejected | 60.0 | 120.0 | 2.0× |
Both arms have the same limit and face the same client. Look at the admitted row.
It is the same on both. Around half, because both limiters admit their configured allowance per window against a client offering twice it.
The sliding window fixes it by having no boundary to aim at: it counts what is genuinely inside the trailing window. That costs you a timestamp per request per client instead of a single integer, which is the entire reason people reach for the fixed window in the first place.
The token bucket is better, and not exempt
A token bucket refills continuously, so there is no reset - which is a real improvement, and it is the formalism the traffic-shaping RFCs standardised, with a committed rate and a committed burst size as two separate parameters.[1]RFCRFC 2697: A Single Rate Three Color Marker[2]RFCRFC 2698: A Two Rate Three Color Marker
But its ceiling over any window is:
burst + limit
Set Algorithm to token bucket and leave the capacity at 60 - one window’s worth, which is the natural default. The invariant still goes red, and the worst case is still 120. A bucket sized at one window’s worth is exactly as permissive as the fixed window it replaced.
Now drop the capacity to 10 and run the same attack. Clean.
Worth keeping the two numbers apart: burst + limit is a bound over any arrival pattern,
while the two-tick boundary attack only extracts burst plus a trickle of refill. The bound
is what you size against; the attack is what you observe.
What to carry away
Measure the guarantee, not the rate. Keep the timestamps of admitted requests and report the largest count in any trailing window. If you only graph the percentage rejected, you cannot see this class of bug at all.
A fixed window’s real limit is twice its stated one. If that is acceptable, configure it at half and say so. If it is not, you need a sliding window or a small bucket.
Your burst is part of your limit. burst + limit is the number to put in the
documentation, because it is the number a client can actually reach.
Boundaries get aimed at by accident. You do not need an attacker. A fleet of clients retrying on the minute is the same traffic shape.
The dial
The line - when someone asks
Every rate limiter promises the same thing - no more than the limit in any window - and a fixed window cannot keep it, because it resets on a boundary with no memory of what came before. A client that spends its allowance either side of that reset gets twice the limit in an arbitrarily short interval. A token bucket is better but not exempt: its real ceiling is burst plus limit, so sizing the bucket at one window's worth doubles your limit in exactly the same way.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]