Essay · Post-Training Evaluation · Interactive

You Didn’t Pick the Best Checkpoint. You Picked the Luckiest One.

Post-training evaluation is usually treated as measurement: run the eval, read the number, keep the winner. It is actually a selection problem, and selection has statistics of its own — ones that quietly guarantee the number you report is too high.

Series: Knowing if they’re any good (part 6 of 6)
Also in this series: Long-Horizon Eval · The Wrong Test · Goodhart in the Act · The Verification Gap · The Autonomy Dial
The short version

Every post-training run ends with the same ritual: a table of checkpoints, a column of eval scores, and a decision to ship the top row. The ritual has a statistical structure almost nobody accounts for, and it runs in exactly one direction — against you.

Here is a number worth sitting with before anything else. Take a reasonable eval — 200 items, binary pass/fail, and a model that gets about 70% of them right. The standard error on that score is ±3.2 percentage points. Which means a checkpoint that scores 71.5 and a checkpoint that scores 68.0 are, as far as that eval can tell, the same checkpoint. Most post-training decisions I have watched get made — including a good number of my own — were made on differences inside that band.

This is not an argument that evals are useless. It is an argument that post-training evaluation is a different statistical activity than it looks like, and that the standard workflow quietly violates the assumptions of the thing it is pretending to be. Once you see it as a selection problem the fixes are cheap, mostly free, and a couple of them are just changes to how you already spend compute.

01Your ruler is wider than the thing you are measuring

The standard error on a pass-rate eval is √(p(1−p)/N). It is worth having the actual numbers in your head, because the intuition most people carry is badly off:

Eval items1 standard errorTwo checkpoints are distinguishable only if they differ by about
100±4.6 pts18.1 points
200±3.2 pts12.8 points
500±2.1 pts8.1 points
1,000±1.5 pts5.7 points
2,000±1.0 pts4.1 points
8,000±0.5 pts2.0 points

At p = 0.70. The third column is the difference you could detect with 80% power at the usual threshold, comparing two checkpoints on independent eval runs.

Read the right-hand column against what post-training actually produces. A good DPO run might move your target metric three to six points. The checkpoints within that run differ by one or two. To separate a two-point difference on independent evals you need roughly eight thousand items, and almost nobody has eight thousand items of the thing they actually care about — because the expensive, carefully-graded evals are precisely the small ones.

Why this is worse than it looks for agentic evals

Everything above assumes a cheap binary outcome per item. Agentic evals — multi-step tasks with a real environment — are expensive enough that 200 items is often the ceiling, not the floor. They also have higher per-item variance, because a single task can fail for reasons that have nothing to do with the checkpoint: a flaky tool, a timeout, a sampling accident.

So the evals that best reflect what you ship are the ones with the least statistical power, and the evals with real power are the ones furthest from the product. That tension does not resolve. It just has to be managed honestly.

02The winner’s curse

Now the part that turns a measurement problem into a selection problem. You do not evaluate one checkpoint. You evaluate ten or twenty — epochs, learning rates, data mixes, β values — and you keep the best. That verb is the whole issue.

If every measurement is the truth plus noise, then the maximum of a set of measurements is biased upward, because the way to land on top is to be good or to be lucky, and with enough candidates luck wins. The checkpoint you ship is disproportionately one whose noise happened to be positive, and the score you report includes that luck, which does not come to production with it.

Below: every dot is a checkpoint. Horizontal position is what it would truly deliver; vertical is what your eval said. The dashed line is perfect measurement. Green is genuinely the best; red is the one you would ship.

Press draw another run a dozen times at the default settings and watch how rarely green and red are the same dot. The summary statistics run four thousand repeats, so they are stable: at a 200-item eval over twelve checkpoints, you pick the genuinely best one 23% of the time. Random choice would give you 8%. Your eval is doing real work — it is roughly three times better than guessing — and it is still wrong three times in four.

Two more things in that chart deserve attention. The red dot is almost always above the dashed line, which is the bias made visible: you ship checkpoints that got lucky. And the gap between the green dot and the red dot is real quality you paid for in compute and then threw away at the last step, for free, by misreading your own instrument.

The number you put in the launch doc

At the default settings the shipped checkpoint’s reported score runs about 4.8 points above what it will actually deliver. Not because anyone fabricated anything — the eval was run correctly and the number was copied accurately. The inflation is a property of having chosen the maximum, and it is baked in before anyone opens a spreadsheet.

This is why post-training progress so often fails to show up downstream. The gains were partly real and partly selection, the real part survives contact with production and the selected part does not, and the difference shows up as the familiar complaint that the eval moved but nothing feels better.

Drag the eval-size slider up and watch all four numbers improve together. That is the honest relationship: statistical power is the only thing that fixes this, and everything else in this post is a way of buying power without buying items.

03The fix that costs nothing: pair everything

Here is the single highest-leverage change, and it requires no additional compute. Evaluate every checkpoint on exactly the same items, and compare differences rather than levels.

The reason it works is that most of the variance in an eval score is not about the checkpoint at all — it is about which items you happened to sample. Some items are hard, some are easy, and that draw is shared by every checkpoint you score on it. When you compare two checkpoints on the same items, that shared difficulty cancels, and all that remains is the handful of items where the two models actually disagree.

How often the two checkpoints agreeError on the differenceSmallest detectable gap
Independent eval runs (no pairing)±4.6 pts12.8 pts
Paired, 50% correlated±3.2 pts9.1 pts
Paired, 80% correlated±2.1 pts5.7 pts
Paired, 90% correlated±1.5 pts4.1 pts
Paired, 95% correlated±1.0 pts2.9 pts

Same 200-item eval throughout. Only the comparison design changes.

Two checkpoints from the same training run agree on the overwhelming majority of items — they are, after all, nearly the same model. That high agreement is usually treated as a disappointment. It is actually the asset: the more alike your checkpoints are, the more pairing buys you. Going from independent runs to 95% agreement takes the detectable gap from 12.8 points to 2.9 on an identical eval set, which is the same power you would otherwise have bought with roughly twenty times the items.

The practical form of this is unglamorous and worth being strict about: fix the item set, fix the sampling seed, fix the judge and its prompt, fix the decoding parameters, and change exactly one thing between arms. Every one of those that you let float is variance you are paying for in sample size.

04The score that matters is the one that went down

Everything so far has been about measuring the target metric accurately. But post-training rarely adds capability outright — it trades. You optimize for helpfulness and lose some terseness; you train on tool-use trajectories and the model starts reaching for tools on questions it should just answer; you fix a refusal behavior and loosen three others.

Which means the real risk in a post-training launch is not that the target metric fails to improve. It is that the target improves, you ship, and something you were not watching got quietly worse. And the selection problem above makes this considerably more likely, because the checkpoint most likely to be selected is the one that overfit hardest to the thing you were measuring — which is also, mechanically, the one most likely to have traded away something you were not.

What you runWhat it catchesWhat it misses
The target evalWhether the thing you optimized movedEverything else, by construction
A broad regression suiteDamage to capabilities you thought ofDamage to capabilities you did not — and with many small evals, multiple comparisons make some look broken by chance
Paired diffing on real trafficWhere old and new actually disagree, including things no eval encodesAnything your traffic does not exercise
Held-out behavioral probesSpecific known failure modes you have been bitten byNovel failure modes

The third row is the one I would add first if a team had none of them. Take a few thousand real prompts, run both checkpoints, and look only at the cases where they disagree. It requires no labels, no judge, and no benchmark design, and it surfaces the trades your eval suite has no category for. The disagreement set is also a free, self-updating regression suite: today’s disagreements are tomorrow’s eval items.

› Go deeper: your regression suite has its own multiple-comparisons problem

There is an irony in the standard advice to run a broad battery of small evals. If you run twenty 100-item evals against a checkpoint that changed nothing at all, the noise on each is ±4.6 points, and the expected worst result across twenty is several standard errors down. You will find a “regression” essentially every time, and teams then spend a day investigating a number that was never real.

The same statistics that make you over-credit your winner make you over-react to your worst loser. The fix is the same in both directions: pair the comparisons so the noise mostly cancels, decide in advance which evals are release-blocking rather than scanning for the red one after the fact, and treat any single eval’s movement inside its error bars as the non-event it is.

05Your eval is downstream of your reward

One more structural hazard specific to post-training. The eval and the training signal are usually close relatives — often literally the same judge model, the same rubric, sometimes the same prompt distribution. When that is true, the eval is not an independent check on the optimization. It is part of it.

The failure that follows is the one Goodhart predicts: the model learns what the judge rewards, the judge is delighted, and the capability underneath may not have moved at all. Preference-trained models learning to be longer, more hedged, and more confident-sounding is the canonical example, because those correlate with winning pairwise comparisons.

Which puts a hard requirement on at least one eval in your suite: it has to be something the training signal cannot see and cannot be talked into agreeing with. Execution-grounded outcomes — tests passing, code compiling, a computation reconciling against a known answer — are the strongest available option, because they are not opinions. If every eval you run is a model’s opinion, and your training signal was also a model’s opinion, you have built a closed loop and the number it produces is about the loop.

This also connects directly to the data flywheel problem: if your eval cannot distinguish genuinely-good from merely-persuasive, neither can the filter you use to select training data, and the contamination that follows is the same contamination.

06The protocol I would run

Compute your error bars before you look at any scores. Take √(p(1−p)/N) for every eval in the suite, write it next to the eval’s name, and keep it there permanently. Most arguments about whether a change is real end immediately once both numbers are on the same line.

Split selection from reporting. Use one held-out set to choose the checkpoint and a different one to report its score. The winner’s curse applies to the set you selected on, and only to that set — a clean reporting set gives an unbiased estimate. This is the single highest-value process change available, and it costs one more eval run.

Pair everything, and freeze everything you are not testing. Same items, same seeds, same judge, same decoding. Report the paired difference with its interval, not two separate levels.

Shrink your reported estimate toward the mean. If you must report the selected checkpoint’s score on the selection set, say plainly that it is biased upward and by roughly how much. “72.4 on the selection set, which is optimistic by a few points given we chose the best of twelve” is a sentence that will make you right more often than the bare number.

Pre-register the release-blocking evals. Decide which evals can block a launch before you see the results, so you are not scanning a wall of numbers for a reason to ship or not ship. This is the cheapest guard against both the winner’s curse and its mirror image.

Keep one eval the training signal cannot reach. Execution-grounded, not judged. It is the only thing standing between you and a closed loop.

Takeaways

  1. A 200-item eval measures to ±3.2 points. Checkpoints differ by one or two. Know your error bars before you argue about your numbers.
  2. Taking the max is selection, not measurement. Best-of-twelve on a small eval finds the genuinely best checkpoint 23% of the time and reports a score several points hot.
  3. Pair everything. Same items, same seeds, same judge. At realistic agreement it is a 4× power gain for zero extra compute — the best trade available in this entire post.
  4. Select on one set, report on another. One extra eval run buys you an honest number.
  5. Watch what went down. Post-training trades capabilities, and the checkpoint you selected is the one most likely to have traded something away.
  6. Keep one eval your reward model cannot see. Otherwise you are measuring the loop.

Related reading

Memory and weights: where knowledge belongs · Goodhart in the act · Evaluating long-horizon agents · Grading the wrong test · The verification gap

The lab is a Monte Carlo over a stated model — checkpoints drawn with a given true spread, measured with binomial noise — not telemetry from any training run. Every figure quoted here was computed independently before it was written, and the standard-error arithmetic is reproducible in two lines of any language you like.

Read next: Memory Is Post-Training With a Learning Rate of 1.0 · Goodhart in the Act

Cite this post
@misc{murugesan2026ptevals,
  author = {Murugesan, Sugeerth},
  title  = {You Didn't Pick the Best Checkpoint. You Picked the Luckiest One.},
  year   = {2026},
  month  = {oct},
  url    = {https://sugeerth.github.io/blog/post-training-evals/},
  note   = {Accessed: [date]}
}
SM
Sugeerth Murugesan Staff ML Engineer / Scientist · Intel / Intuit