1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you make eval results statistically trustworthy?
30-second answerSay your answer out loud first, then reveal.
Confidence interval intuition
95% ≈ ±1.96 × SE| n examples | Measured 80% pass | 95% CI (approx.) |
|---|---|---|
| 50 | 80% | ±11 points |
| 200 | 80% | ±5.5 points |
| 1,000 | 80% | ±2.5 points |
"Prompt B scored 83% vs A's 80% on 100 examples" is likely noise.
Techniques
- Paired comparison: evaluate both versions on the same examples and look at the examples where they differ. That's far more sensitive than comparing two independent averages.
- Bootstrap confidence intervals for complex metrics.
- Multiple runs: for stochastic systems, run k times. Report mean ± spread, or pass^k (all runs pass) for reliability.
- Judge noise: include judge error in the uncertainty (Q17).
- Slices: a significant overall change may hide regressions in a small slice; but slicing into many groups risks false discoveries, so pre-define the important slices.
- Practical significance: a statistically significant 0.5% gain may not justify a 30% cost increase.
Interview signal. Spontaneously mentioning confidence intervals and paired comparisons is rare and impressive.
Related
This is what real progress feels like.