Reading an eval result
A 3-point win on a 200-item eval ships the worse system about one time in five, and the report will not say so.
An eval reports a number. The number has three or four digits after the decimal point, it appears in a table next to last week’s, and nothing anywhere on the page says how much of it is real.
Every bar in that chart is one complete evaluation of a candidate that is genuinely three points better. The rule at zero is the decision you are about to make. Everything left of it is a run where the improvement measured negative and the better system lost, and each one of those looked, from the inside, exactly like the run you are about to do.
What a score is
An accuracy on n items is a proportion, and its standard error is sqrt(p(1-p)/n). At 80%
on 200 items that is 2.8 points, so the 95% interval on a single score is roughly plus or
minus 5.5 points. Comparing two scores is worse than that, because both are noisy.
Drag Eval set size down to 50 and watch the distribution spread until most of it is on the wrong side of zero. Drag it up to 3000 and watch it pull clear. The item count is not an implementation detail of the eval; it is the resolution of the instrument.
Score both on the same items
Separate runs
Two independent samples, which is what you get from two eval jobs run a week apart.
Same items
One sample, both systems, difference taken per item. Only the disagreements count.
| Measure | Separate runs | Same items | Gap |
|---|---|---|---|
| Runs where the worse system won | 0.16 | 0.06 | 2.6× |
| Half-width of the 95% interval | 0.07 | 0.04 | 1.7× |
How much better the candidate genuinely is. The simulation knows; the eval does not.
Two systems that differ by 3 points must disagree on at least 3% of items.
Here is the change that costs nothing. Both arms have the same true improvement and the same number of items. The left one scores the two systems on separate samples, which is what you get from two eval jobs run a week apart. The right one scores both on the same items and takes the difference per item.
The error rate falls by roughly a factor of three.
The reason is that most items in a real eval are answered the same way by both systems, and those items carry no information about which is better. Scoring separately leaves every item’s own difficulty sitting in the noise; scoring on the same items cancels it exactly, and only the disagreements count.[1]paperNote on the sampling error of the difference between correlated proportions That is McNemar’s test, and it is the right default for comparing two systems on the same benchmark.[2]paperApproximate Statistical Tests for Comparing Supervised Classification Learning Algorithms
There is an arithmetic floor underneath this, and the widget enforces it. Two systems that differ by three points must disagree on at least three percent of items, because there is nowhere else for the difference to live. Push True improvement above Share of items they answer differently and the invariant goes red: you have asked for a comparison that cannot exist.
Do not report the number you chose on
Press Keep the best of twenty variants. Every one of those twenty is exactly as good as the baseline. The reported winner is about five points ahead.
Taking a maximum over noisy measurements does not find the best variant, it finds the luckiest one, and the bias grows with how many things you tried. The score is real in the sense that it was measured. It is simply not a property of the prompt, and it will not survive the next eval set.
What to do with this
Put an interval on every eval number you publish. It is one line of arithmetic and it changes the conversation from “83.1 versus 80.4” to “somewhere between 5 points worse and 11 points better”, which is the honest version.[4]articleAdding Error Bars to Evals - a statistical approach to language model evaluations
Pair by default. Score both systems on the same items and difference per item. It costs nothing, it is worth roughly a tripled eval set, and the only reason it is not universal is that separate eval jobs are the more convenient thing to build.
Size the eval to the effect you care about, before you run it. If the smallest improvement worth shipping is two points, work out how many items that needs and get them. An eval that cannot resolve the decision it is being used for will still return a number, on time, every time.
Report variance from more than the sample. Seeds, item order, and judge temperature all move the score, and an interval computed from sampling error alone is an underestimate.[3]paperAccounting for Variance in Machine Learning Benchmarks
The dial
The line - when someone asks
An eval score is an estimate with a standard error, and on 200 items that error is about 2.8 points, so a 3-point improvement sits well inside the noise and the worse system posts the higher number roughly a fifth of the time. Two cheap changes fix most of it: score both systems on the same items, so item difficulty cancels and only the disagreements count, which is worth about a tripled eval set; and never report the number you selected on, because the best of twenty variants is flattered by around 5 points even when all twenty are identical.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]
- [3]
Secondary
- [4]