1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is Goodhart's law in the context of evals, and how do you avoid overfitting to your eval set?
30-second answerSay your answer out loud first, then reveal.
How overfitting happens
- Adding few-shot examples copied from eval cases.
- Prompt rules that patch individual failures ("if the user asks about X, say Y").
- Optimising toward a judge's quirks (e.g. longer answers score higher).
- Automated prompt optimisers (e.g. DSPy-style) tuning directly on the test set.
Safeguards
| Safeguard | How |
|---|---|
| Train / dev / test splits | Iterate on dev; evaluate on test only at milestones |
| Fresh data | Monthly additions from production; rotate examples |
| Multiple metrics | Quality + conciseness + refusal rate + cost |
| Judge diversity and calibration | Re-validate judges; use different judge models occasionally |
| Human audits | Periodic blind human review of random production samples |
| Online validation | A/B tests on real user outcomes |
| Generalisation checks | Paraphrased versions of eval inputs should score similarly |
Interview line. "My eval set is a sample of reality, not reality. I protect a held-out set and keep refreshing from production."
Related
Little by little, you're building something great.