The agent loop and where it burns your budget
An agent that is stuck looks exactly like an agent that is working. It keeps calling tools, keeps producing turns, and spends the entire context window finding out it was wrong on the first one.
Strip away the framing and an agent is a while-loop. The model proposes a tool call, the harness runs it, the result is appended to the conversation, and round it goes until the task is finished or you run out of room.[1]paperReAct: Synergizing Reasoning and Acting in Language Models
Everything that makes agents unreliable in production follows from two facts about that loop, and neither is about the model being clever enough.
Every iteration costs context whether or not it accomplished anything. A failed tool call costs roughly the same tokens as a successful one - the reasoning, the call, and the error message all land in the transcript.
The loop cannot tell that it is stuck. From the inside, retrying a broken tool looks identical to making progress: there is a call, there is a result, the turn completes.
The number that actually governs it
People tune the wrong dial. The question is not “how big is my context window” and not “how reliable are my tools”, it is how many tokens you spend per unit of real progress.
Each attempt costs the same whether it succeeds or fails. With failure probability p,
the expected number of attempts per success is 1/(1-p), so:
context per completed step ≈ (reasoning + result) / (1 - p)
The consequence is worse than linear, and that surprises people. A 20% failure rate does not cost 20% more context, it costs 25% more. At 50% it costs double.
Result size is the lever nobody measures
Look again at the numerator. Reasoning per turn is a few hundred tokens and you do not control it. The tool result is however many tokens your tool decided to return, and you control it completely.
A tool that returns 6000 tokens where 200 would do has not made your agent slightly less efficient. It has cut the number of steps your agent can take before the window closes by roughly thirty times.
The failure that eats everything
Now press Wedge a tool in the widget.
A wedged tool returns the same error every single time. The model sees an error, reasons about it, tries again - possibly with a small variation that changes nothing - and gets the identical error back. There is no signal anywhere in the loop that says this class of attempt will never work.
With a generous retry budget, the widget spends 98% of its context window on repeated failures and completes zero further steps. Thirty-three consecutive wasted iterations, every one of which looked from the inside like the agent doing its job.
This is why every serious harness ships an iteration cap rather than trusting the loop to
notice. LangChain’s AgentExecutor defaults max_iterations to 15 and checks it before
every step.[2]source codeLangChain AgentExecutor: max_iterations and the _should_continue check The cap is not a performance tuning knob, it is
the only thing standing between a wedged dependency and a fully consumed
budget.[4]articleBuilding Effective Agents
Why the budget is checked before the append
One detail in the widget is easy to miss and is a real bug in real harnesses. The loop checks whether a result will fit before appending it, and stops if it will not.
Appending first and noticing afterwards produces a conversation longer than the model’s window. The next call fails at the provider, and what you get back is not a partial result but an error about token limits, with the transcript already unusable. The widget asserts this as an invariant on every tick, so ballooning the tool results mid-run stops the loop cleanly rather than corrupting it.
What to take away
Agents fail on arithmetic far more often than on reasoning. Three numbers are worth carrying into any design review.
Context per completed step, which is (reasoning + result) / (1 - p) and which is the
budget that actually matters. Tool result size, which you control and almost certainly have
not measured. And the iteration cap, which is the difference between a failure you can
still debug and one that consumed everything on the way down.[3]paperAgentBench: Evaluating LLMs as Agents
The dial
The line - when someone asks
An agent is a while-loop: the model proposes a tool call, the harness runs it, the result is appended, and round it goes until the task is done or the context window is gone. The number that decides whether a long task finishes is not the failure rate and not the window size, it is the ratio between them - tokens burned per unit of real progress, which is roughly the cost of one attempt divided by one minus the failure rate. The failure mode nobody plans for is a tool that fails the same way every time, because the model cannot tell that retrying is pointless and will happily spend the whole budget proving it.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]
- [3]
Secondary
- [4]