1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you validate that an LLM judge is trustworthy?
30-second answerSay your answer out loud first, then reveal.
Process
- Human labels: 100–300 outputs labelled by domain experts using the same criterion definition. Resolve disagreements to create gold labels; measure inter-human agreement as a ceiling.
- Split: dev set (to tune the judge prompt and examples) and test set (final measurement).
- Metrics:
• TPR (recall of passes): judge passes outputs humans passed
• TNR (recall of failures): judge catches outputs humans failed (often most important)
• Cohen's κ: agreement corrected for chance (above ~0.6 is often considered substantial) - Iterate: read the disagreements. Are the criteria ambiguous? Is context missing? Do the examples need refining?
- Correct aggregate estimates: if the judge's TPR and TNR are known, you can adjust the observed pass rate to estimate the true pass rate (and give a confidence interval).
- Monitor: re-validate on new samples monthly, or when the judge model or product changes.
Red flags
The judge passes nearly everything (leniency); agreement is high only because 95% of examples are passes (class imbalance, so look at TNR); different results across runs (use temperature 0, or majority voting).
Related
You understood something today that you didn't yesterday.