Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q19IntermediateConcept

How do you make eval results statistically trustworthy?

30-second answerSay your answer out loud first, then reveal.

Confidence interval intuition

text
95% ≈ ±1.96 × SE
n examplesMeasured 80% pass95% CI (approx.)
5080%±11 points
20080%±5.5 points
1,00080%±2.5 points

"Prompt B scored 83% vs A's 80% on 100 examples" is likely noise.

Techniques

  1. Paired comparison: evaluate both versions on the same examples and look at the examples where they differ. That's far more sensitive than comparing two independent averages.
  2. Bootstrap confidence intervals for complex metrics.
  3. Multiple runs: for stochastic systems, run k times. Report mean ± spread, or pass^k (all runs pass) for reliability.
  4. Judge noise: include judge error in the uncertainty (Q17).
  5. Slices: a significant overall change may hide regressions in a small slice; but slicing into many groups risks false discoveries, so pre-define the important slices.
  6. Practical significance: a statistically significant 0.5% gain may not justify a 30% cost increase.
Interview signal. Spontaneously mentioning confidence intervals and paired comparisons is rare and impressive.

This is what real progress feels like.