Context compaction in agent loops
An agent that runs out of context does not crash - it summarises, forgets the one detail the task depended on, and continues with complete confidence.
An agent loop is a simple thing: call the model, run whatever tool it asks for, append the result, repeat.[3]paperReAct: Synergizing Reasoning and Acting in Language Models The complication is that every iteration makes the context longer, and the context is finite.
What happens when it fills is the single biggest determinant of whether a long-running agent works, and it is almost always an afterthought.
Compaction is lossy, and quiet about it
When context approaches the window, something has to go. The usual approach keeps the most recent turns verbatim and replaces older ones with a summary.
That is a compression step with an uncontrolled loss ratio. A summary of forty tool calls into two paragraphs necessarily discards nearly everything - and what it discards is chosen by a model that does not know which detail the task will need in twenty turns’ time.
The failure mode is not an exception. It is an agent that established a fact on turn 8, compacted at turn 40, and on turn 60 confidently proceeds as though that fact never existed.
Position matters too
Even facts that survive compaction are not equally usable. Models attend unevenly across long contexts, retrieving information at the beginning and end far more reliably than material buried in the middle.[1]paperLost in the Middle: How Language Models Use Long Contexts
So a summary that faithfully preserves a fact, and places it in the middle of a long context, has preserved it in a form the model may not actually use. Retention as measured by “is the text still there” overstates retention as measured by “does the model act on it”.
The tool-output flood
The most common way this goes wrong in practice is not gradual. A tool returns something enormous - a large file, a verbose log, an unfiltered API response - and a single turn consumes a large fraction of the window.
Press Flood with output in the widget. Compaction fires far more often, facts are evicted much faster, and recall drops sharply. Nothing reports a problem.
The mitigation is unglamorous and works: truncate tool output at the tool boundary, before it ever reaches the context. A tool that can return a hundred thousand tokens should return a summary and a handle to fetch more.
Compact at a boundary, not mid-task
Compaction triggered purely by a token threshold will eventually fire in the middle of a multi-step operation, discarding the intermediate state that operation depends on.
Press Compact now during a run to see the shape of it. The fix is to make compaction aware of the loop rather than only of the token count: prefer a boundary between tasks, and explicitly carry forward the current goal and the facts established for it.
The stronger version of this idea is to stop treating the window as the memory at all - MemGPT’s framing is an explicit tiered memory with the model deciding what to page in and out, rather than a summariser making that decision implicitly and irreversibly.[2]paperMemGPT: Towards LLMs as Operating Systems
The dial
The line - when someone asks
Agent loops accumulate context - tool results, reasoning, messages - until they approach the model's window, at which point older turns must be summarised or dropped. Compaction is lossy by definition, and the loss is silent: nothing errors, the agent simply proceeds without a fact it established earlier. The engineering problem is deciding what is load bearing before you compress it, and doing it at a boundary rather than mid-task.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]
- [3]