yield point
Agentsintermediateupdated 2026-08-22

Context compaction in agent loops

An agent that runs out of context does not crash - it summarises, forgets the one detail the task depended on, and continues with complete confidence.

An agent loop is a simple thing: call the model, run whatever tool it asks for, append the result, repeat.[3]paperReAct: Synergizing Reasoning and Acting in Language ModelsYao, S. et al., ICLR 2023, 2023 The complication is that every iteration makes the context longer, and the context is finite.

What happens when it fills is the single biggest determinant of whether a long-running agent works, and it is almost always an afterthought.

Drop 'facts kept per compaction' to 0.2 and watch the recall rate fall. Nothing errors - the agent simply forgets, and carries on confidently.

Context window0 / 2,000
fact, verbatimfact, compressed by a summaryordinary turn
Compactions
0
Facts lost
0 of 0
Recall rate
-
tick 0 / 1500
Break it

Fraction of the window that triggers compaction.

How much a summary actually preserves. This is the number nobody measures.

Safety properties
  • ✓
    Context never exceeds the window

    held on every tick so far

Compaction is lossy, and quiet about it

When context approaches the window, something has to go. The usual approach keeps the most recent turns verbatim and replaces older ones with a summary.

That is a compression step with an uncontrolled loss ratio. A summary of forty tool calls into two paragraphs necessarily discards nearly everything - and what it discards is chosen by a model that does not know which detail the task will need in twenty turns’ time.

The failure mode is not an exception. It is an agent that established a fact on turn 8, compacted at turn 40, and on turn 60 confidently proceeds as though that fact never existed.

Position matters too

Even facts that survive compaction are not equally usable. Models attend unevenly across long contexts, retrieving information at the beginning and end far more reliably than material buried in the middle.[1]paperLost in the Middle: How Language Models Use Long ContextsLiu, N. F. et al., TACL, 2023

So a summary that faithfully preserves a fact, and places it in the middle of a long context, has preserved it in a form the model may not actually use. Retention as measured by “is the text still there” overstates retention as measured by “does the model act on it”.

The tool-output flood

The most common way this goes wrong in practice is not gradual. A tool returns something enormous - a large file, a verbose log, an unfiltered API response - and a single turn consumes a large fraction of the window.

Press Flood with output in the widget. Compaction fires far more often, facts are evicted much faster, and recall drops sharply. Nothing reports a problem.

The mitigation is unglamorous and works: truncate tool output at the tool boundary, before it ever reaches the context. A tool that can return a hundred thousand tokens should return a summary and a handle to fetch more.

Compact at a boundary, not mid-task

Compaction triggered purely by a token threshold will eventually fire in the middle of a multi-step operation, discarding the intermediate state that operation depends on.

Press Compact now during a run to see the shape of it. The fix is to make compaction aware of the loop rather than only of the token count: prefer a boundary between tasks, and explicitly carry forward the current goal and the facts established for it.

The stronger version of this idea is to stop treating the window as the memory at all - MemGPT’s framing is an explicit tiered memory with the model deciding what to page in and out, rather than a summariser making that decision implicitly and irreversibly.[2]paperMemGPT: Towards LLMs as Operating SystemsPacker, C. et al., 2023

The dial

You gainLong-running tasks that would otherwise hit the window and fail outright
You paySilent, unmeasured information loss at every compaction, biased toward the oldest facts

The line - when someone asks

Agent loops accumulate context - tool results, reasoning, messages - until they approach the model's window, at which point older turns must be summarised or dropped. Compaction is lossy by definition, and the loss is silent: nothing errors, the agent simply proceeds without a fact it established earlier. The engineering problem is deciding what is load bearing before you compress it, and doing it at a boundary rather than mid-task.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary