yield point
AI / LLM Internalsintermediateupdated 2026-08-25

Reading an eval result

A 3-point win on a 200-item eval ships the worse system about one time in five, and the report will not say so.

An eval reports a number. The number has three or four digits after the decimal point, it appears in a table next to last week’s, and nothing anywhere on the page says how much of it is real.

Each bar is one complete eval run of a candidate that is genuinely 3 points better. Watch how much of the distribution sits left of zero, where the eval reports that your improvement made things worse. Then tick 'Score both on the same items'.

0.0035810ships the wrong onetrue improvementmeasured improvement (percentage points)
Worse system won
no runs yet
Predicted for this run
20%
95% interval
+/- 7.6 pts
Looking for
3.0 pts
Items to see it
2,626
Eval set
200 items
Design
separate runs
Runs
0
tick 0 / 4000
Break it
Do

How much better the candidate genuinely is. The simulation knows; the eval does not.

Two systems that differ by 3 points must disagree on at least 3% of items.

Off means two independent samples, which is what separate eval runs give you.

Safety properties
  • ✓
    The improvement is small enough to fit in the disagreements

    held on every tick so far

  • ✓
    No count exceeds the eval set, and no item is discordant twice

    held on every tick so far

Every bar in that chart is one complete evaluation of a candidate that is genuinely three points better. The rule at zero is the decision you are about to make. Everything left of it is a run where the improvement measured negative and the better system lost, and each one of those looked, from the inside, exactly like the run you are about to do.

What a score is

An accuracy on n items is a proportion, and its standard error is sqrt(p(1-p)/n). At 80% on 200 items that is 2.8 points, so the 95% interval on a single score is roughly plus or minus 5.5 points. Comparing two scores is worse than that, because both are noisy.

Drag Eval set size down to 50 and watch the distribution spread until most of it is on the wrong side of zero. Drag it up to 3000 and watch it pull clear. The item count is not an implementation detail of the eval; it is the resolution of the instrument.

Score both on the same items

Identical settings: the candidate really is better by the amount on the slider, and both sides get the same eval budget. The only difference is whether the two systems were scored on the same items or on separate samples.

Separate runs

Two independent samples, which is what you get from two eval jobs run a week apart.

-8.50-1.7551219ships the wrong onetrue improvementmeasured improvement (percentage points)
Worse system won
16% of runs
Predicted for this run
20%
95% interval
+/- 7.6 pts
Looking for
3.0 pts
Items to see it
2,626
Eval set
200 items
Design
separate runs
Runs
600

Same items

One sample, both systems, difference taken per item. Only the disagreements count.

-4.00-0.633610ships the wrong onetrue improvementmeasured improvement (percentage points)
Worse system won
6% of runs
Predicted for this run
7%
95% interval
+/- 4.4 pts
Looking for
3.0 pts
Items to see it
865
Eval set
200 items
Design
same items
Runs
600
Measured at tick 600 - both sides, same seed, same inputs
MeasureSeparate runsSame itemsGap
Runs where the worse system won0.160.062.6×
Half-width of the 95% interval0.070.041.7×
tick 600
Break it
Do

How much better the candidate genuinely is. The simulation knows; the eval does not.

Two systems that differ by 3 points must disagree on at least 3% of items.

Here is the change that costs nothing. Both arms have the same true improvement and the same number of items. The left one scores the two systems on separate samples, which is what you get from two eval jobs run a week apart. The right one scores both on the same items and takes the difference per item.

The error rate falls by roughly a factor of three.

The reason is that most items in a real eval are answered the same way by both systems, and those items carry no information about which is better. Scoring separately leaves every item’s own difficulty sitting in the noise; scoring on the same items cancels it exactly, and only the disagreements count.[1]paperNote on the sampling error of the difference between correlated proportionsMcNemar, Q., Psychometrika 12(2), 1947 That is McNemar’s test, and it is the right default for comparing two systems on the same benchmark.[2]paperApproximate Statistical Tests for Comparing Supervised Classification Learning AlgorithmsDietterich, T. G., Neural Computation 10(7), 1998

There is an arithmetic floor underneath this, and the widget enforces it. Two systems that differ by three points must disagree on at least three percent of items, because there is nowhere else for the difference to live. Push True improvement above Share of items they answer differently and the invariant goes red: you have asked for a comparison that cannot exist.

Do not report the number you chose on

Press Keep the best of twenty variants. Every one of those twenty is exactly as good as the baseline. The reported winner is about five points ahead.

Taking a maximum over noisy measurements does not find the best variant, it finds the luckiest one, and the bias grows with how many things you tried. The score is real in the sense that it was measured. It is simply not a property of the prompt, and it will not survive the next eval set.

What to do with this

Put an interval on every eval number you publish. It is one line of arithmetic and it changes the conversation from “83.1 versus 80.4” to “somewhere between 5 points worse and 11 points better”, which is the honest version.[4]articleAdding Error Bars to Evals - a statistical approach to language model evaluationsMiller, E., 2024

Pair by default. Score both systems on the same items and difference per item. It costs nothing, it is worth roughly a tripled eval set, and the only reason it is not universal is that separate eval jobs are the more convenient thing to build.

Size the eval to the effect you care about, before you run it. If the smallest improvement worth shipping is two points, work out how many items that needs and get them. An eval that cannot resolve the decision it is being used for will still return a number, on time, every time.

Report variance from more than the sample. Seeds, item order, and judge temperature all move the score, and an interval computed from sampling error alone is an underestimate.[3]paperAccounting for Variance in Machine Learning BenchmarksBouthillier, X. et al., MLSys 2021, 2021

The dial

You gainA number for every change, immediately, from an eval you already have
You payMost of those numbers cannot resolve the effect you are asking about, and they never say which

The line - when someone asks

An eval score is an estimate with a standard error, and on 200 items that error is about 2.8 points, so a 3-point improvement sits well inside the noise and the worse system posts the higher number roughly a fifth of the time. Two cheap changes fix most of it: score both systems on the same items, so item difficulty cancels and only the disagreements count, which is worth about a tripled eval set; and never report the number you selected on, because the best of twenty variants is flattered by around 5 points even when all twenty are identical.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary

Secondary