yield point
AI / LLM Internalsintermediateupdated 2026-08-28

Continuous batching

A static batch of 16 leaves about a third of the GPU generating padding, and doubling the batch to fix it makes that share worse.

Generating a token requires reading every weight in the model out of memory and then doing a trivial amount of arithmetic with it. So a forward pass over a batch of sixteen sequences costs about what a forward pass over one costs: the weights get read either way.

That single fact makes batching the biggest lever there is in inference throughput, and it is why the obvious implementation is so tempting. Collect requests, run them together, return them together.

A static batch of 16 leaves about a third of the GPU generating padding, and the slots are full the whole time. Tick 'Continuous batching' and watch occupancy, throughput and latency all improve at once.

Waiting for a slot0/400
0.16/tick
Slots generating tokens0/16
0.00/tick
in 0out 0
Slot-time producing tokens
0%
Theory says
65%
Slots held by finished work
0
Tokens per tick
0.0
Mean latency
0 ticks
Lengths
48 to 192 tokens
Longest of a batch
184
Scheduler
static
Wasted by design
35%
tick 0 / 4000
Break it
Do

Half-width of the length window, as a share of the mean. Zero means every request generates exactly the same number of tokens.

Retire each sequence the tick it finishes and admit a waiting request into the slot.

Safety properties
  • ✓
    Every slot-tick is either producing a token or admitted to be waste

    held on every tick so far

  • ✓
    No more sequences in flight than there are slots

    held on every tick so far

The slots are full of nothing

Watch the Slots held by finished work line. Requests do not all generate the same number of tokens: one answers in a sentence, another writes an essay. A static batch runs until its longest member is done, so the short ones sit in their slots producing padding while they wait for the long one.

The useful share is the mean length divided by the expected longest in the batch:

occupancy = E[L] / E[max of B draws]

The number worth sitting with is that the slots are full the entire time. Every dashboard that reports GPU utilisation or batch occupancy shows a healthy green line. They are full of sequences that finished.

Making the batch bigger makes it worse

Identical arrivals, identical batch size, identical output lengths. The only difference is whether a slot is freed when its sequence finishes or when the whole batch does. Drag the batch size from 4 up to 32: the static arm loses about 13 points of occupancy on the way, so the gap widens as you pull the throughput lever harder.

Static batching

The batch runs until its longest member is done, and finished sequences hold their slots until then.

Waiting for a slot287/400
0.34/tick
Slots generating tokens1/4
1.00/tick
peak 1
in 23out 2571
Slot-time producing tokens
71%
Theory says
74%
Slots held by finished work
3
Tokens per tick
2.9
Mean latency
425 ticks
Lengths
48 to 192 tokens
Longest of a batch
163
Scheduler
static
Wasted by design
26%

Continuous batching

A slot is freed the tick its sequence ends and a waiting request moves in.

Waiting for a slot276/400
0.34/tick
Slots generating tokens4/4
4.00/tick
peak 1
in 31out 3572
Slot-time producing tokens
99%
Theory says
100%
Slots held by finished work
0
Tokens per tick
4.0
Mean latency
410 ticks
Lengths
48 to 192 tokens
Longest of a batch
163
Scheduler
continuous
Wasted by design
0%
Measured at tick 900 - both sides, same seed, same inputs
MeasureStatic batchingContinuous batchingGap
Slot-time producing tokens0.710.991.4×
Tokens per tick2.863.971.4×
tick 900
Break it
Do

Half-width of the length window, as a share of the mean. Zero means every request generates exactly the same number of tokens.

Drag Batch size from 4 up to 32 on both arms.

Total throughput rises, which is the thing everyone measures and the reason this survives. But E[max of B] climbs with B while E[L] does not move at all, so the share of each slot doing useful work falls: about 83% at a batch of 2, 74% at 4, 65% at 16, 63% at 64.

The curve is steep and then flat. It is heading for E[L] / hi - the mean over the longest length that exists at all - which for this spread is 62.5%, and it never comes back up. So the honest summary is not that huge batches are uniquely bad. It is that the waste arrives almost immediately, is a third of the machine by the time the batch is any useful size, and no batch size you can choose recovers it.

Set Spread of output lengths to zero and the waste vanishes completely. That is worth knowing because it is the shape of every benchmark that makes static batching look fine: fixed-length generation is the one workload where a static batch is optimal, and it is not a workload anybody actually serves.

The fix that is not a trade

Continuous batching retires a sequence the tick it finishes and admits a waiting request into the free slot.[1]paperOrca: A Distributed Serving System for Transformer-Based Generative ModelsYu, G. et al., OSDI 2022, 2022 That is the entire idea. Tick the toggle and occupancy goes to essentially one at any batch size.

It is rare for something to improve throughput and latency together, and worth being precise about why this does. The slots were being held by sequences that had already finished, so freeing them takes nothing from anyone: the long request still gets its forward pass every tick, and the short request behind it starts immediately instead of waiting for a batch boundary. Nobody is being starved to pay for it.

What it costs instead

The cost is not in the trade, it is in the scheduler. A static batch checks once whether B sequences fit in memory. A continuous one has to decide on every tick whether the next waiting request fits alongside whatever is currently running, and the answer changes constantly because the KV cache of every live sequence is still growing.

That admission decision is the hard part of a modern inference server, and it is why PagedAttention matters: paging the KV cache instead of reserving one contiguous block per sequence is what makes the memory budget answerable at all.[2]paperEfficient Memory Management for Large Language Model Serving with PagedAttentionKwon, W. et al., SOSP 2023, 2023

What to do with this

Do not read batch occupancy as utilisation. Full slots and useful slots are different numbers, and only one of them is on your dashboard. Instrument tokens generated against slot-ticks paid for.

Suspect any throughput benchmark run at a fixed output length. It is measuring the one distribution on which the naive scheduler is optimal.

If you are choosing a serving stack, this is the feature to check for. It is worth more than most model-level optimisations, it applies to every workload with variable output lengths, and it is not something you can bolt on afterwards.

If you are building one, the scheduler is the product. The batching idea is a paragraph. Deciding what fits in memory, on every tick, without ever being wrong, is the rest of it.

The dial

You gainEssentially all of the hardware doing work somebody asked for, and lower latency at the same time
You payA scheduler that admits against a memory budget on every tick instead of once per batch

The line - when someone asks

Decoding is memory-bound, so a forward pass over a batch costs about what a pass over one costs, which makes batching the largest lever in inference throughput. A static batch runs until its longest member finishes, so every sequence that ends early holds its slot producing nothing: the useful share is the mean length over the expected longest, and since the expected longest grows with the batch while the mean does not, bigger batches waste a larger fraction. Continuous batching frees each slot the tick its sequence ends, which raises throughput and lowers latency together.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary