yield point
Agentsintermediateupdated 2026-08-24

Lost in the middle

Retrieval worked, the answer is sitting in the context, and the model still gets it wrong - because a fact in the middle of a long context is read far less reliably than the same fact at either end.

Your retrieval is working. You can prove it: the answer chunk is in the context, every time, and your recall@k dashboard is green.

The model still gets a third of the questions wrong, and it will keep getting them wrong until you move where the answer sits.

The answer is always in the context - never anywhere else. Drag it from chunk 1 to chunk 10 and watch a third of your answers disappear without a single retrieval failure.

Answers found
-
This position
48%
Reliably read
4 of 20
Edge over middle
40 pts
Context
5,000 tok
Trials
0 / 0 (0%)
tick 0 / 3000
Break it
Do

How many chunks the retriever puts in the context. More is better recall and a worse position.

Used when the ordering is pinned. The answer is always genuinely in the context.

How far primacy and recency extend, as a share of the context.

Length hurts on its own, independently of position.

Safety properties
  • ✓
    The answer is always genuinely in the context

    held on every tick so far

  • ✓
    Hits never exceed trials

    held on every tick so far

Each cell is one retrieved chunk, shaded by how likely the model is to use what is in it. The highlighted cell is the one holding the answer. Drag Answer is chunk number from 1 to 10 and watch the hit rate collapse - with retrieval succeeding on every single trial.

The curve is a U

Liu et al. measured this across multi-document question answering and key-value retrieval: performance is highest when the relevant information is at the beginning or the end of the context, and degrades significantly when the model has to reach into the middle - including in models explicitly built and sold for long contexts.[1]paperLost in the Middle: How Language Models Use Long ContextsLiu, N. F. et al., TACL, 2024

Two separate effects stack, and it is worth keeping them apart:

Position. Primacy and recency both help, and they help from opposite ends, so the penalty is worst dead centre. At the default shape that is about 40 percentage points between the edge and the middle.

Length. Even holding position constant, a longer context reads worse. Set Answer is chunk number to 1 - the strongest position there is - and raise the chunk count. Accuracy still falls.

Only a fifth of your context is reliably read

Look at the reliably read figure: at the default shape, 4 of 20 positions clear the bar. Now raise the chunk count to 50 and look again. It becomes 10 of 50.

The share does not improve. A bigger context does not buy you more reliable positions, it buys you more unreliable ones - and you pay for every one of them in tokens.

This is also the honest reading of the long-context numbers on model cards. The RULER benchmark exists more or less for this reason: the context length a model claims and the length over which it actually performs are different numbers, and the gap widens as the task gets harder than pure retrieval.[2]paperRULER: What's the Real Context Size of Your Long-Context Language Models?Hsieh, C.-P. et al., 2024

The largest free lever is ordering

Identical retrieval, identical context size, and the answer genuinely present in both. The only difference is whether the pipeline puts its best chunk first. Now raise the chunk count on both and watch which one degrades.

Arbitrary order

Chunks in whatever order the store returned them. The answer lands anywhere, so you get the average of the curve.

Answers found
73%
This position
49%
Reliably read
4 of 20
Edge over middle
40 pts
Context
5,000 tok
Trials
141 / 200 (71%)

Best chunk first

Sorted by relevance, which puts the best chunk at position one - the strongest position in the context.

Answers found
96%
This position
88%
Reliably read
4 of 20
Edge over middle
40 pts
Context
5,000 tok
Trials
185 / 200 (93%)
Measured at tick 200 - both sides, same seed, same inputs
MeasureArbitrary orderBest chunk firstGap
Answers found0.700.931.3×
Answer position0.420.00∞
Context size50005000same
tick 200
Break it
Do

How many chunks the retriever puts in the context. More is better recall and a worse position.

How far primacy and recency extend, as a share of the context.

Length hurts on its own, independently of position.

Both arms retrieve the same chunks, into the same context size, with the answer genuinely present. The only difference is whether the pipeline sorts by relevance.

Sorting best-first is worth roughly 40 points here, and almost nobody counts it as a design decision - it is just what the vector store happened to return. Which means the reverse is a live risk: someone changes the ordering to document order, or by date, or to group by source, and every retrieval metric stays exactly where it was while answer quality falls off a cliff.

Two traps this creates

Raising k looks free and is not. Retrieve more because you are worried about recall, and every retrieval metric improves - they measure whether the answer made it into the set. The context got longer and the answer moved inward, which no retrieval metric measures at all. You will read the resulting failures as a model problem. The honest version of this trade needs both curves plotted together, and most teams only have the first.

A split answer multiplies the penalty. Press Split the answer in two - now the model needs two chunks, and both have to be in a position it reads. The probabilities multiply, so two chunks at 60% each is 36%, not 60%. Chunking strategy is therefore a context-position decision and not only a retrieval one: a chunk size that splits the fact you need in half has doubled your exposure to this curve.

What to carry away

“Retrieved” and “used” are different. Recall@k cannot see this failure. If your evaluation stops at whether the answer was in the context, it is blind to the thing that is actually costing you.

Order by relevance, deliberately. Best chunk first. Write it down as a decision so nobody undoes it while tidying up.

k has a cost you are not measuring. More chunks is more recall, more tokens, a longer context and a worse position. Only the first of those is on your dashboard.

Keep facts whole. A fact split across two chunks needs two lucky positions instead of one.

The dial

You gainMore retrieved chunks means a better chance the answer is in the context at all
You payA longer context reads worse everywhere, and pushes whatever you retrieved further from the two positions the model reliably uses

The line - when someone asks

Language models do not read a context uniformly. Accuracy is highest when the relevant information sits at the very beginning or the very end, degrades substantially in the middle, and degrades again simply for the context being long. That makes "did we retrieve it" and "will the model use it" two different questions, and almost every retrieval evaluation only measures the first. The cheapest fix is ordering: sorting chunks by relevance puts your best one in the strongest position in the context, which is a large accuracy effect that nobody ever wrote down as a design decision.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary