Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q33IntermediateConcept

What are the pitfalls of using LLM-as-judge to evaluate agents, and how do you mitigate them?

30-second answerSay your answer out loud first, then reveal.

Known pitfalls

  • Position bias: in pairwise comparisons the judge prefers whichever answer is first (or second).
  • Verbosity bias: longer answers score higher.
  • Self-preference: the judge favours outputs from its own model family.
  • Leniency: judges tend to rate everything 4/5.
  • No ground truth: a judge can't tell whether an order ID is correct unless you give it the data.
  • Vague criteria: "rate helpfulness 1–10" gives noise.

Mitigations

  1. Prefer code checks for anything verifiable: state checks, tests, schemas, exact matches, required tool calls.
  2. Binary, specific criteria: "Did the agent confirm the order number before issuing the refund? yes/no" beats a 1–10 score.
  3. One criterion per judge call for complex rubrics.
  4. Give context: reference answers, retrieved docs, policies, the full trajectory.
  5. Calibrate: label 100+ examples by hand, measure judge-human agreement (precision/recall on failures), and iterate on the judge prompt. Re-check periodically.
  6. Pairwise with swapped order to cancel position bias.
  7. Ask for reasoning before the verdict, and log it for audit.

Follow-ups to expect

  • How do you know your judge is good enough? It should agree with domain experts at an acceptable rate on a held-out labelled set, especially at catching failures.

Slow is fine. Stopping is the only problem.