yield point
Agentsintermediateupdated 2026-08-24

Prompt caching economics

A cached prefix costs a tenth of an uncached one to read and a quarter more to write - so one badly placed token turns an 88% saving into paying 25% over the list price.

Almost every agent sends the same few thousand tokens on every single request: the tool definitions, the system prompt, the instructions nobody has touched in months. Caching that prefix is the largest cost lever available and it usually takes one line to enable.

It is also the only optimisation we know of that can make your bill go up, and the thing that flips it costs one token.

Let it settle at 88% saved, then put a clock in the prompt at position 0. The bill goes 25% above what you would have paid with no cache at all. Now find the position where it starts paying for itself again.

This request, billed against the same request uncached0 / 20,400
read from cache (0.1x)written to cache (1.25x)plain input (1x)
This request
0 tok
Settles at
88% saved
Stable share
100%
Break-even
22%
Mean so far
0 tok (0%)
Requests
0
tick 0 / 400
Break it
Do

Tools, then system prompt, then history - in that order, and a change anywhere invalidates everything after it.

How much of the prefix is byte-identical to last time. One changing token invalidates everything after it, so this is a position, not a percentage of edits.

Safety properties
  • ✓
    Every token in the request is billed exactly once

    held on every tick so far

  • ✓
    No request costs less than a full cache hit

    held on every tick so far

Three rates, one sum

There is no model here. A request is a sum of three token counts against three published rates:[3]docsPrompt cachingAnthropic, 2026

What it is Billed at
Read back from cache 0.1x the base input price
Written to cache, 5-minute TTL 1.25x
Written to cache, 1-hour TTL 2x
Plain uncached input 1x

With a 20,000 token prefix and a 400 token turn, that is 2,400 tokens billed against the 20,400 you would pay uncached. An 88% saving, and it needs no cleverness at all - just a prefix that does not change.

The word that ruins it is “exact”

A cache matches on an exact prefix. Not a similar one, not a mostly-identical one. So the question is never “how much of my prompt changed”, it is “where is the first token that changed” - because everything after it is a miss.

Press Put a clock in the prompt at position 0 and watch the bill.

uncached:              20,000 x 1.00  + 400  =  20,400
cached, stable:        20,000 x 0.10  + 400  =   2,400
cached, clock on top:  20,000 x 1.25  + 400  =  25,400

The order the provider concatenates in is tools, then system, then messages, so the things most likely to be regenerated are also the things sitting furthest forward.[3]docsPrompt cachingAnthropic, 2026 A tool list built by iterating a hash map, a system prompt with today’s date, a session id in the header - each of these is a one-line change that silently reprices everything behind it.

Identical prompts with identical content at an identical request rate. The only difference is where the timestamp sits. Raise the prefix size and watch the gap widen in absolute tokens while the percentages stay exactly where they are.

Clock on the first line

Nothing after it matches, so the whole prefix is rewritten every request at 1.25x.

This request, billed against the same request uncached25,400 / 20,400
read from cache (0.1x)written to cache (1.25x)plain input (1x)
This request
25,400 tok
Settles at
-25% saved
Stable share
0%
Break-even
22%
Mean so far
25,400 tok (-25%)
Requests
10

Clock on the last line

Everything cacheable comes first, so the whole prefix is read back at 0.1x.

This request, billed against the same request uncached2,400 / 20,400
read from cache (0.1x)written to cache (1.25x)plain input (1x)
This request
2,400 tok
Settles at
88% saved
Stable share
100%
Break-even
22%
Mean so far
4,700 tok (77%)
Requests
10
Measured at tick 10 - both sides, same seed, same inputs
MeasureClock on the first lineClock on the last lineGap
Cost per request2540047005.4×
Saved against no caching-0.250.77same
Prefix served from cache0.001.00∞
tick 10
Break it
Do

Tools, then system prompt, then history - in that order, and a change anywhere invalidates everything after it.

The break-even nobody computes

Partial stability has a threshold, and it is higher than people guess. Caching pays only when what you save on the reused part covers what you lose on the rewritten part. Setting those equal, with W the write multiplier and R the read multiplier:

break-even stable share = (W - 1) / (W - R)

At the five-minute rate that is 0.25 / 1.15, about 22%. At the one-hour rate it is 1.00 / 1.90, about 53%.

The other way to lose it: nobody is talking

A cache entry has a TTL, and the default is five minutes. Press Traffic goes quiet on a perfectly stable prompt.

Every request becomes a cold write at 1.25x. The prompt never changed; the rate did. This is why low-traffic paths - a nightly job, a webhook, an internal tool three people use - are so often the ones quietly paying a premium. Nobody looks, because the absolute numbers are small.

The one-hour TTL fixes it, and it is not free: writes cost 2x rather than 1.25x. It wins only across gaps that would have killed the five-minute entry. If it also misses, you are 100% over list rather than 25% over.

Where the mechanism comes from

None of this is a billing invention. Reusing attention state across prompts that share a prefix is a real systems technique with real literature behind it: precomputing the attention states of frequently occurring segments and reusing them across requests,[1]paperPrompt Cache: Modular Attention Reuse for Low-Latency InferenceGim, I. et al., MLSys 2024, 2024 made practical at serving scale by memory management that lets separate requests share the same physical KV blocks.[2]paperEfficient Memory Management for Large Language Model Serving with PagedAttentionKwon, W. et al., SOSP '23, 2023 The pricing is a fairly direct reflection of that: a read is cheap because the work was already done, and a write costs a premium because somebody has to hold the memory.

Knowing that is what makes the rules predictable rather than arbitrary. The cache is a prefix-keyed lookup of computed state, so anything that changes the prefix changes the key, and there is no partial credit for a prompt that is nearly the same.

What to carry away

Find the first changing token. Not the percentage that changed - the position. That single number decides your bill.

Everything volatile goes last. Timestamps, session ids, retrieved documents, anything regenerated per request. This is usually a ten-minute change and it is the highest-value ten minutes in agent engineering.

Below 22% stable, do not cache. Compute it before you enable it, and compute it again after somebody adds a field to the system prompt.

Check your quiet paths. A stable prompt on a slow schedule is paying 25% over list and nothing in your dashboards will say so.

The dial

You gainA stable prefix is read back at a tenth of the base input price, which is most of the bill on any agent with a long system prompt
You payA write costs more than a plain input token, so below a break-even share of stability the cache is a surcharge rather than a discount

The line - when someone asks

Prompt caching bills a cache read at 0.1x the base input rate and a five-minute cache write at 1.25x, matching on an exact prefix. That last word is the whole thing: one token that changes near the top invalidates everything after it, and because a miss still writes, you end up paying 25% more than you would with no cache at all. The number to know is the break-even, (W - 1) / (W - R), which is about 22% of your prefix staying byte-identical at the five-minute rate and about 53% at the one-hour rate.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary

Secondary