Prompt caching economics
A cached prefix costs a tenth of an uncached one to read and a quarter more to write - so one badly placed token turns an 88% saving into paying 25% over the list price.
Almost every agent sends the same few thousand tokens on every single request: the tool definitions, the system prompt, the instructions nobody has touched in months. Caching that prefix is the largest cost lever available and it usually takes one line to enable.
It is also the only optimisation we know of that can make your bill go up, and the thing that flips it costs one token.
Three rates, one sum
There is no model here. A request is a sum of three token counts against three published rates:[3]docsPrompt caching
| What it is | Billed at |
|---|---|
| Read back from cache | 0.1x the base input price |
| Written to cache, 5-minute TTL | 1.25x |
| Written to cache, 1-hour TTL | 2x |
| Plain uncached input | 1x |
With a 20,000 token prefix and a 400 token turn, that is 2,400 tokens billed against the 20,400 you would pay uncached. An 88% saving, and it needs no cleverness at all - just a prefix that does not change.
The word that ruins it is “exact”
A cache matches on an exact prefix. Not a similar one, not a mostly-identical one. So the question is never “how much of my prompt changed”, it is “where is the first token that changed” - because everything after it is a miss.
Press Put a clock in the prompt at position 0 and watch the bill.
uncached: 20,000 x 1.00 + 400 = 20,400
cached, stable: 20,000 x 0.10 + 400 = 2,400
cached, clock on top: 20,000 x 1.25 + 400 = 25,400
The order the provider concatenates in is tools, then system, then messages, so the things most likely to be regenerated are also the things sitting furthest forward.[3]docsPrompt caching A tool list built by iterating a hash map, a system prompt with today’s date, a session id in the header - each of these is a one-line change that silently reprices everything behind it.
Clock on the first line
Nothing after it matches, so the whole prefix is rewritten every request at 1.25x.
Clock on the last line
Everything cacheable comes first, so the whole prefix is read back at 0.1x.
| Measure | Clock on the first line | Clock on the last line | Gap |
|---|---|---|---|
| Cost per request | 25400 | 4700 | 5.4× |
| Saved against no caching | -0.25 | 0.77 | same |
| Prefix served from cache | 0.00 | 1.00 | ∞ |
Tools, then system prompt, then history - in that order, and a change anywhere invalidates everything after it.
The break-even nobody computes
Partial stability has a threshold, and it is higher than people guess. Caching pays only
when what you save on the reused part covers what you lose on the rewritten part. Setting
those equal, with W the write multiplier and R the read multiplier:
break-even stable share = (W - 1) / (W - R)
At the five-minute rate that is 0.25 / 1.15, about 22%. At the one-hour rate it is
1.00 / 1.90, about 53%.
The other way to lose it: nobody is talking
A cache entry has a TTL, and the default is five minutes. Press Traffic goes quiet on a perfectly stable prompt.
Every request becomes a cold write at 1.25x. The prompt never changed; the rate did. This is why low-traffic paths - a nightly job, a webhook, an internal tool three people use - are so often the ones quietly paying a premium. Nobody looks, because the absolute numbers are small.
The one-hour TTL fixes it, and it is not free: writes cost 2x rather than 1.25x. It wins only across gaps that would have killed the five-minute entry. If it also misses, you are 100% over list rather than 25% over.
Where the mechanism comes from
None of this is a billing invention. Reusing attention state across prompts that share a prefix is a real systems technique with real literature behind it: precomputing the attention states of frequently occurring segments and reusing them across requests,[1]paperPrompt Cache: Modular Attention Reuse for Low-Latency Inference made practical at serving scale by memory management that lets separate requests share the same physical KV blocks.[2]paperEfficient Memory Management for Large Language Model Serving with PagedAttention The pricing is a fairly direct reflection of that: a read is cheap because the work was already done, and a write costs a premium because somebody has to hold the memory.
Knowing that is what makes the rules predictable rather than arbitrary. The cache is a prefix-keyed lookup of computed state, so anything that changes the prefix changes the key, and there is no partial credit for a prompt that is nearly the same.
What to carry away
Find the first changing token. Not the percentage that changed - the position. That single number decides your bill.
Everything volatile goes last. Timestamps, session ids, retrieved documents, anything regenerated per request. This is usually a ten-minute change and it is the highest-value ten minutes in agent engineering.
Below 22% stable, do not cache. Compute it before you enable it, and compute it again after somebody adds a field to the system prompt.
Check your quiet paths. A stable prompt on a slow schedule is paying 25% over list and nothing in your dashboards will say so.
The dial
The line - when someone asks
Prompt caching bills a cache read at 0.1x the base input rate and a five-minute cache write at 1.25x, matching on an exact prefix. That last word is the whole thing: one token that changes near the top invalidates everything after it, and because a miss still writes, you end up paying 25% more than you would with no cache at all. The number to know is the break-even, (W - 1) / (W - R), which is about 22% of your prefix staying byte-identical at the five-minute rate and about 53% at the one-hour rate.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]
Secondary
- [3]