1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What are the pitfalls of using LLM-as-judge to evaluate agents, and how do you mitigate them?
30-second answerSay your answer out loud first, then reveal.
Known pitfalls
- Position bias: in pairwise comparisons the judge prefers whichever answer is first (or second).
- Verbosity bias: longer answers score higher.
- Self-preference: the judge favours outputs from its own model family.
- Leniency: judges tend to rate everything 4/5.
- No ground truth: a judge can't tell whether an order ID is correct unless you give it the data.
- Vague criteria: "rate helpfulness 1–10" gives noise.
Mitigations
- Prefer code checks for anything verifiable: state checks, tests, schemas, exact matches, required tool calls.
- Binary, specific criteria: "Did the agent confirm the order number before issuing the refund? yes/no" beats a 1–10 score.
- One criterion per judge call for complex rubrics.
- Give context: reference answers, retrieved docs, policies, the full trajectory.
- Calibrate: label 100+ examples by hand, measure judge-human agreement (precision/recall on failures), and iterate on the judge prompt. Re-check periodically.
- Pairwise with swapped order to cancel position bias.
- Ask for reasoning before the verdict, and log it for audit.
Follow-ups to expect
- How do you know your judge is good enough? It should agree with domain experts at an acceptable rate on a held-out labelled set, especially at catching failures.
Related
Slow is fine. Stopping is the only problem.