Error compounding
An agent that gets 95% of its steps right finishes a twenty-step task about a third of the time, and the number that fixes it is not the one everybody measures.
Ninety-five percent sounds like a good agent. It is the number that gets reported, and it is not the number that decides whether the work gets done.
Twenty steps at 95% each is 0.95^20, which is about 36%. Two runs in three fail, and
every individual step was fine 95% of the time, exactly as advertised.
The horizon is decided by the last nine
Set Errors that announce themselves to zero so nothing is recovered, and read off the 90% horizon as you move per-step accuracy:
| Per-step accuracy | Longest task that finishes 90% of the time |
|---|---|
| 90% | 1 step |
| 95% | 2 steps |
| 99% | 10 steps |
| 99.9% | 105 steps |
The relationship is ln(target) / ln(accuracy), so accuracy enters through a logarithm and
every additional nine multiplies your reachable task length by about ten.
This explains two things at once. It explains why a model that feels only slightly better can be transformatively better inside an agent loop - and why a small regression in step reliability shortens the horizon far more than it looks like it should. The published work on task horizons measures the same relationship from the outside: rather than reporting accuracy, it reports the length of task a model completes at a given success rate, which is the practical form of this arithmetic.[1]paperMeasuring AI Ability to Complete Long Software Tasks
The lever you actually have
You cannot usually move per-step accuracy. You can almost always move detection.
A silent error can never be recovered, because nothing knows it happened. That is why detection multiplies into the recovery term rather than sitting beside it:
effective per-step = accuracy + (1 - accuracy) x loud x recovery
Errors return plausible values
Nothing throws, nothing validates. An error cannot be recovered because nothing knows it happened.
Errors throw
Schemas validated, tools raise. Almost every error gets a chance at repair.
| Measure | Errors return plausible values | Errors throw | Gap |
|---|---|---|---|
| Runs completed | 0.19 | 0.58 | 3.1× |
| Steps before failure | 11.7 | 15.3 | 1.3× |
| Effective per-step | 0.95 | 0.99 | same |
Tool calls, edits, decisions - anything that has to be right for the run to be right.
Both arms use the same model at the same per-step accuracy. One validates its tool results so failures throw; the other lets them return something plausible.
In practice this is unglamorous:
- Tools raise instead of returning an error string. An error string is a plausible value; the model will summarise it and carry on.
- Validate every structured result against its schema, at the boundary, before the model sees it.
- Make “I could not do this” a first-class tool result rather than something the model has to infer from an empty list.
The empirical work on why these systems fail lands in the same place: the largest failure categories are not the model being wrong about something hard, they are specification and inter-step misalignment - agents proceeding confidently on an understanding that stopped being correct several steps ago.[2]paperWhy Do Multi-Agent LLM Systems Fail?
Checkpoints are not free
Press Add checkpoints. Recovery improves, and the run gets longer - a verification step is still a step, and it can still go wrong.
That is the real shape of the trade, and it is why “just add more verification” stops paying
at some point: you are buying a better q with more n, and both are in the same
exponent. The widget charges you the steps so you can find where that turns over for your
own numbers.
What to carry away
Per-step accuracy is not run accuracy. It is run accuracy to the power of one over the step count, and nobody’s dashboard does that arithmetic for them.
Compute your horizon. ln(0.9) / ln(q). It takes ten seconds and it usually explains
the last month of confusing failures.
Detection beats accuracy. You can rarely move the model. You can nearly always make failures loud, and it is worth more.
Fit the task to the horizon. Segment long work rather than betting a whole run on a tail you have already measured.
The dial
The line - when someone asks
Independent steps multiply, so a run's success rate is per-step accuracy raised to the number of steps: 95% over twenty steps is about 36%. The horizon that follows is ln(target) / ln(accuracy), which means every extra nine of per-step reliability multiplies the length of task you can finish by roughly ten. The lever most teams have available is not accuracy at all but detection, because a silent error can never be recovered - making failures loud is usually worth more than a whole point of raw per-step accuracy.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]