yield point
Agentsintermediateupdated 2026-08-23

The agent loop and where it burns your budget

An agent that is stuck looks exactly like an agent that is working. It keeps calling tools, keeps producing turns, and spends the entire context window finding out it was wrong on the first one.

Strip away the framing and an agent is a while-loop. The model proposes a tool call, the harness runs it, the result is appended to the conversation, and round it goes until the task is finished or you run out of room.[1]paperReAct: Synergizing Reasoning and Acting in Language ModelsYao, S. et al., ICLR 2023, 2023

Everything that makes agents unreliable in production follows from two facts about that loop, and neither is about the model being clever enough.

Every iteration costs context whether or not it accomplished anything. A failed tool call costs roughly the same tokens as a successful one - the reasoning, the call, and the error message all land in the transcript.

The loop cannot tell that it is stuck. From the inside, retrying a broken tool looks identical to making progress: there is a call, there is a result, the turn completes.

Run it once and watch the task finish. Now wedge a tool and run it again: the loop looks exactly as busy as before, completes nothing, and spends the entire window finding out. Then drop the retry cap to 2 and see what changes.

Context0 / 32,000
Made progressTool errorRepeated the same failure
Steps done
0 / 30
Iterations
0
Context per step
-
Wasted retries
0
Status
running
tick 0 / 400
Break it

What the loop may spend before the window is gone.

The single biggest lever, and the one nobody measures.

Set this high and a wedged tool will consume the entire budget.

Safety properties
  • ✓
    Context never exceeds the budget

    held on every tick so far

The number that actually governs it

People tune the wrong dial. The question is not “how big is my context window” and not “how reliable are my tools”, it is how many tokens you spend per unit of real progress.

Each attempt costs the same whether it succeeds or fails. With failure probability p, the expected number of attempts per success is 1/(1-p), so:

context per completed step ≈ (reasoning + result) / (1 - p)

The consequence is worse than linear, and that surprises people. A 20% failure rate does not cost 20% more context, it costs 25% more. At 50% it costs double.

Result size is the lever nobody measures

Look again at the numerator. Reasoning per turn is a few hundred tokens and you do not control it. The tool result is however many tokens your tool decided to return, and you control it completely.

A tool that returns 6000 tokens where 200 would do has not made your agent slightly less efficient. It has cut the number of steps your agent can take before the window closes by roughly thirty times.

The failure that eats everything

Now press Wedge a tool in the widget.

A wedged tool returns the same error every single time. The model sees an error, reasons about it, tries again - possibly with a small variation that changes nothing - and gets the identical error back. There is no signal anywhere in the loop that says this class of attempt will never work.

With a generous retry budget, the widget spends 98% of its context window on repeated failures and completes zero further steps. Thirty-three consecutive wasted iterations, every one of which looked from the inside like the agent doing its job.

This is why every serious harness ships an iteration cap rather than trusting the loop to notice. LangChain’s AgentExecutor defaults max_iterations to 15 and checks it before every step.[2]source codeLangChain AgentExecutor: max_iterations and the _should_continue check The cap is not a performance tuning knob, it is the only thing standing between a wedged dependency and a fully consumed budget.[4]articleBuilding Effective AgentsAnthropic, 2024

Why the budget is checked before the append

One detail in the widget is easy to miss and is a real bug in real harnesses. The loop checks whether a result will fit before appending it, and stops if it will not.

Appending first and noticing afterwards produces a conversation longer than the model’s window. The next call fails at the provider, and what you get back is not a partial result but an error about token limits, with the transcript already unusable. The widget asserts this as an invariant on every tick, so ballooning the tool results mid-run stops the loop cleanly rather than corrupting it.

What to take away

Agents fail on arithmetic far more often than on reasoning. Three numbers are worth carrying into any design review.

Context per completed step, which is (reasoning + result) / (1 - p) and which is the budget that actually matters. Tool result size, which you control and almost certainly have not measured. And the iteration cap, which is the difference between a failure you can still debug and one that consumed everything on the way down.[3]paperAgentBench: Evaluating LLMs as AgentsLiu, X. et al., 2023

The dial

You gainA model that can act on the world instead of only describing it, and recover from steps it could not have planned
You payEvery iteration costs context whether or not it accomplished anything, and the loop cannot tell the difference

The line - when someone asks

An agent is a while-loop: the model proposes a tool call, the harness runs it, the result is appended, and round it goes until the task is done or the context window is gone. The number that decides whether a long task finishes is not the failure rate and not the window size, it is the ratio between them - tokens burned per unit of real progress, which is roughly the cost of one attempt divided by one minus the failure rate. The failure mode nobody plans for is a tool that fails the same way every time, because the model cannot tell that retrying is pointless and will happily spend the whole budget proving it.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary

Secondary