yield point
Agentsintermediateupdated 2026-08-24

Error compounding

An agent that gets 95% of its steps right finishes a twenty-step task about a third of the time, and the number that fixes it is not the one everybody measures.

Ninety-five percent sounds like a good agent. It is the number that gets reported, and it is not the number that decides whether the work gets done.

Ninety-five percent per step sounds like a good agent. Over twenty steps it finishes a third of the time. Now find the change that doubles completions without touching per-step accuracy at all.

Steps taken this run0 / 20
step correcterror caught and fixeduncorrected - run over
Runs completed
-
Predicted
59%
Effective per-step
97.4%
90% horizon
4.0 steps
Task length
20 steps
Errors that are loud
60%
tick 0 / 3000
Break it
Do

Tool calls, edits, decisions - anything that has to be right for the run to be right.

A tool that throws, a schema that fails validation. The rest return something plausible and wrong.

Safety properties
  • ✓
    Every step lands in exactly one outcome

    held on every tick so far

  • ✓
    A run stops at its first uncorrected error

    held on every tick so far

Twenty steps at 95% each is 0.95^20, which is about 36%. Two runs in three fail, and every individual step was fine 95% of the time, exactly as advertised.

The horizon is decided by the last nine

Set Errors that announce themselves to zero so nothing is recovered, and read off the 90% horizon as you move per-step accuracy:

Per-step accuracy Longest task that finishes 90% of the time
90% 1 step
95% 2 steps
99% 10 steps
99.9% 105 steps

The relationship is ln(target) / ln(accuracy), so accuracy enters through a logarithm and every additional nine multiplies your reachable task length by about ten.

This explains two things at once. It explains why a model that feels only slightly better can be transformatively better inside an agent loop - and why a small regression in step reliability shortens the horizon far more than it looks like it should. The published work on task horizons measures the same relationship from the outside: rather than reporting accuracy, it reports the length of task a model completes at a given success rate, which is the practical form of this arithmetic.[1]paperMeasuring AI Ability to Complete Long Software TasksKwa, T. et al., 2025

The lever you actually have

You cannot usually move per-step accuracy. You can almost always move detection.

A silent error can never be recovered, because nothing knows it happened. That is why detection multiplies into the recovery term rather than sitting beside it:

effective per-step = accuracy + (1 - accuracy) x loud x recovery

Identical per-step accuracy on both sides - the model is exactly as good. The only difference is whether a failing step announces itself or returns something that looks fine. Now stretch the task and watch which one still finishes.

Errors return plausible values

Nothing throws, nothing validates. An error cannot be recovered because nothing knows it happened.

Steps taken this run2 / 20
step correcterror caught and fixeduncorrected - run over
Runs completed
19%
Predicted
37%
Effective per-step
95.2%
90% horizon
2.1 steps
Task length
20 steps
Errors that are loud
5%

Errors throw

Schemas validated, tools raise. Almost every error gets a chance at repair.

Steps taken this run11 / 20
step correcterror caught and fixeduncorrected - run over
Runs completed
58%
Predicted
79%
Effective per-step
98.8%
90% horizon
8.7 steps
Task length
20 steps
Errors that are loud
95%
Measured at tick 200 - both sides, same seed, same inputs
MeasureErrors return plausible valuesErrors throwGap
Runs completed0.190.583.1×
Steps before failure11.715.31.3×
Effective per-step0.950.99same
tick 200
Break it
Do

Tool calls, edits, decisions - anything that has to be right for the run to be right.

Both arms use the same model at the same per-step accuracy. One validates its tool results so failures throw; the other lets them return something plausible.

In practice this is unglamorous:

  • Tools raise instead of returning an error string. An error string is a plausible value; the model will summarise it and carry on.
  • Validate every structured result against its schema, at the boundary, before the model sees it.
  • Make “I could not do this” a first-class tool result rather than something the model has to infer from an empty list.

The empirical work on why these systems fail lands in the same place: the largest failure categories are not the model being wrong about something hard, they are specification and inter-step misalignment - agents proceeding confidently on an understanding that stopped being correct several steps ago.[2]paperWhy Do Multi-Agent LLM Systems Fail?Cemri, M. et al., 2025

Checkpoints are not free

Press Add checkpoints. Recovery improves, and the run gets longer - a verification step is still a step, and it can still go wrong.

That is the real shape of the trade, and it is why “just add more verification” stops paying at some point: you are buying a better q with more n, and both are in the same exponent. The widget charges you the steps so you can find where that turns over for your own numbers.

What to carry away

Per-step accuracy is not run accuracy. It is run accuracy to the power of one over the step count, and nobody’s dashboard does that arithmetic for them.

Compute your horizon. ln(0.9) / ln(q). It takes ten seconds and it usually explains the last month of confusing failures.

Detection beats accuracy. You can rarely move the model. You can nearly always make failures loud, and it is worth more.

Fit the task to the horizon. Segment long work rather than betting a whole run on a tail you have already measured.

The dial

You gainLonger agent runs do more work per invocation and need less orchestration from you
You payEvery added step multiplies into the success rate, so reach is bounded by a per-step number that improves painfully slowly

The line - when someone asks

Independent steps multiply, so a run's success rate is per-step accuracy raised to the number of steps: 95% over twenty steps is about 36%. The horizon that follows is ln(target) / ln(accuracy), which means every extra nine of per-step reliability multiplies the length of task you can finish by roughly ten. The lever most teams have available is not accuracy at all but detection, because a silent error can never be recovered - making failures loud is usually worth more than a whole point of raw per-step accuracy.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary