Lost in the middle
Retrieval worked, the answer is sitting in the context, and the model still gets it wrong - because a fact in the middle of a long context is read far less reliably than the same fact at either end.
Your retrieval is working. You can prove it: the answer chunk is in the context, every time, and your recall@k dashboard is green.
The model still gets a third of the questions wrong, and it will keep getting them wrong until you move where the answer sits.
Each cell is one retrieved chunk, shaded by how likely the model is to use what is in it. The highlighted cell is the one holding the answer. Drag Answer is chunk number from 1 to 10 and watch the hit rate collapse - with retrieval succeeding on every single trial.
The curve is a U
Liu et al. measured this across multi-document question answering and key-value retrieval: performance is highest when the relevant information is at the beginning or the end of the context, and degrades significantly when the model has to reach into the middle - including in models explicitly built and sold for long contexts.[1]paperLost in the Middle: How Language Models Use Long Contexts
Two separate effects stack, and it is worth keeping them apart:
Position. Primacy and recency both help, and they help from opposite ends, so the penalty is worst dead centre. At the default shape that is about 40 percentage points between the edge and the middle.
Length. Even holding position constant, a longer context reads worse. Set Answer is chunk number to 1 - the strongest position there is - and raise the chunk count. Accuracy still falls.
Only a fifth of your context is reliably read
Look at the reliably read figure: at the default shape, 4 of 20 positions clear the bar. Now raise the chunk count to 50 and look again. It becomes 10 of 50.
The share does not improve. A bigger context does not buy you more reliable positions, it buys you more unreliable ones - and you pay for every one of them in tokens.
This is also the honest reading of the long-context numbers on model cards. The RULER benchmark exists more or less for this reason: the context length a model claims and the length over which it actually performs are different numbers, and the gap widens as the task gets harder than pure retrieval.[2]paperRULER: What's the Real Context Size of Your Long-Context Language Models?
The largest free lever is ordering
Arbitrary order
Chunks in whatever order the store returned them. The answer lands anywhere, so you get the average of the curve.
Best chunk first
Sorted by relevance, which puts the best chunk at position one - the strongest position in the context.
| Measure | Arbitrary order | Best chunk first | Gap |
|---|---|---|---|
| Answers found | 0.70 | 0.93 | 1.3× |
| Answer position | 0.42 | 0.00 | ∞ |
| Context size | 5000 | 5000 | same |
How many chunks the retriever puts in the context. More is better recall and a worse position.
How far primacy and recency extend, as a share of the context.
Length hurts on its own, independently of position.
Both arms retrieve the same chunks, into the same context size, with the answer genuinely present. The only difference is whether the pipeline sorts by relevance.
Sorting best-first is worth roughly 40 points here, and almost nobody counts it as a design decision - it is just what the vector store happened to return. Which means the reverse is a live risk: someone changes the ordering to document order, or by date, or to group by source, and every retrieval metric stays exactly where it was while answer quality falls off a cliff.
Two traps this creates
Raising k looks free and is not. Retrieve more because you are worried about recall, and every retrieval metric improves - they measure whether the answer made it into the set. The context got longer and the answer moved inward, which no retrieval metric measures at all. You will read the resulting failures as a model problem. The honest version of this trade needs both curves plotted together, and most teams only have the first.
A split answer multiplies the penalty. Press Split the answer in two - now the model needs two chunks, and both have to be in a position it reads. The probabilities multiply, so two chunks at 60% each is 36%, not 60%. Chunking strategy is therefore a context-position decision and not only a retrieval one: a chunk size that splits the fact you need in half has doubled your exposure to this curve.
What to carry away
“Retrieved” and “used” are different. Recall@k cannot see this failure. If your evaluation stops at whether the answer was in the context, it is blind to the thing that is actually costing you.
Order by relevance, deliberately. Best chunk first. Write it down as a decision so nobody undoes it while tidying up.
k has a cost you are not measuring. More chunks is more recall, more tokens, a longer context and a worse position. Only the first of those is on your dashboard.
Keep facts whole. A fact split across two chunks needs two lucky positions instead of one.
The dial
The line - when someone asks
Language models do not read a context uniformly. Accuracy is highest when the relevant information sits at the very beginning or the very end, degrades substantially in the middle, and degrades again simply for the context being long. That makes "did we retrieve it" and "will the model use it" two different questions, and almost every retrieval evaluation only measures the first. The cheapest fix is ordering: sorting chunks by relevance puts your best one in the strongest position in the context, which is a large accuracy effect that nobody ever wrote down as a design decision.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]