Continuous batching
A static batch of 16 leaves about a third of the GPU generating padding, and doubling the batch to fix it makes that share worse.
Generating a token requires reading every weight in the model out of memory and then doing a trivial amount of arithmetic with it. So a forward pass over a batch of sixteen sequences costs about what a forward pass over one costs: the weights get read either way.
That single fact makes batching the biggest lever there is in inference throughput, and it is why the obvious implementation is so tempting. Collect requests, run them together, return them together.
The slots are full of nothing
Watch the Slots held by finished work line. Requests do not all generate the same number of tokens: one answers in a sentence, another writes an essay. A static batch runs until its longest member is done, so the short ones sit in their slots producing padding while they wait for the long one.
The useful share is the mean length divided by the expected longest in the batch:
occupancy = E[L] / E[max of B draws]
The number worth sitting with is that the slots are full the entire time. Every dashboard that reports GPU utilisation or batch occupancy shows a healthy green line. They are full of sequences that finished.
Making the batch bigger makes it worse
Static batching
The batch runs until its longest member is done, and finished sequences hold their slots until then.
Continuous batching
A slot is freed the tick its sequence ends and a waiting request moves in.
| Measure | Static batching | Continuous batching | Gap |
|---|---|---|---|
| Slot-time producing tokens | 0.71 | 0.99 | 1.4× |
| Tokens per tick | 2.86 | 3.97 | 1.4× |
Half-width of the length window, as a share of the mean. Zero means every request generates exactly the same number of tokens.
Drag Batch size from 4 up to 32 on both arms.
Total throughput rises, which is the thing everyone measures and the reason this survives. But
E[max of B] climbs with B while E[L] does not move at all, so the share of each slot doing
useful work falls: about 83% at a batch of 2, 74% at 4, 65% at 16, 63% at 64.
The curve is steep and then flat. It is heading for E[L] / hi - the mean over the longest
length that exists at all - which for this spread is 62.5%, and it never comes back up. So the
honest summary is not that huge batches are uniquely bad. It is that the waste arrives almost
immediately, is a third of the machine by the time the batch is any useful size, and no batch
size you can choose recovers it.
Set Spread of output lengths to zero and the waste vanishes completely. That is worth knowing because it is the shape of every benchmark that makes static batching look fine: fixed-length generation is the one workload where a static batch is optimal, and it is not a workload anybody actually serves.
The fix that is not a trade
Continuous batching retires a sequence the tick it finishes and admits a waiting request into the free slot.[1]paperOrca: A Distributed Serving System for Transformer-Based Generative Models That is the entire idea. Tick the toggle and occupancy goes to essentially one at any batch size.
It is rare for something to improve throughput and latency together, and worth being precise about why this does. The slots were being held by sequences that had already finished, so freeing them takes nothing from anyone: the long request still gets its forward pass every tick, and the short request behind it starts immediately instead of waiting for a batch boundary. Nobody is being starved to pay for it.
What it costs instead
The cost is not in the trade, it is in the scheduler. A static batch checks once whether B sequences fit in memory. A continuous one has to decide on every tick whether the next waiting request fits alongside whatever is currently running, and the answer changes constantly because the KV cache of every live sequence is still growing.
That admission decision is the hard part of a modern inference server, and it is why PagedAttention matters: paging the KV cache instead of reserving one contiguous block per sequence is what makes the memory budget answerable at all.[2]paperEfficient Memory Management for Large Language Model Serving with PagedAttention
What to do with this
Do not read batch occupancy as utilisation. Full slots and useful slots are different numbers, and only one of them is on your dashboard. Instrument tokens generated against slot-ticks paid for.
Suspect any throughput benchmark run at a fixed output length. It is measuring the one distribution on which the naive scheduler is optimal.
If you are choosing a serving stack, this is the feature to check for. It is worth more than most model-level optimisations, it applies to every workload with variable output lengths, and it is not something you can bolt on afterwards.
If you are building one, the scheduler is the product. The batching idea is a paragraph. Deciding what fits in memory, on every tick, without ever being wrong, is the rest of it.
The dial
The line - when someone asks
Decoding is memory-bound, so a forward pass over a batch costs about what a pass over one costs, which makes batching the largest lever in inference throughput. A static batch runs until its longest member finishes, so every sequence that ends early holds its slot producing nothing: the useful share is the mean length over the expected longest, and since the expected longest grows with the batch while the mean does not, bigger batches waste a larger fraction. Continuous batching frees each slot the tick its sequence ends, which raises throughput and lowers latency together.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]
- [3]