Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q17IntermediateConcept

How do you validate that an LLM judge is trustworthy?

30-second answerSay your answer out loud first, then reveal.

Process

  1. Human labels: 100–300 outputs labelled by domain experts using the same criterion definition. Resolve disagreements to create gold labels; measure inter-human agreement as a ceiling.
  2. Split: dev set (to tune the judge prompt and examples) and test set (final measurement).
  3. Metrics:
    • TPR (recall of passes): judge passes outputs humans passed
    • TNR (recall of failures): judge catches outputs humans failed (often most important)
    • Cohen's κ: agreement corrected for chance (above ~0.6 is often considered substantial)
  4. Iterate: read the disagreements. Are the criteria ambiguous? Is context missing? Do the examples need refining?
  5. Correct aggregate estimates: if the judge's TPR and TNR are known, you can adjust the observed pass rate to estimate the true pass rate (and give a confidence interval).
  6. Monitor: re-validate on new samples monthly, or when the judge model or product changes.

Red flags

The judge passes nearly everything (leniency); agreement is high only because 95% of examples are passes (class imbalance, so look at TNR); different results across runs (use temperature 0, or majority voting).

You understood something today that you didn't yesterday.