Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q16IntermediateConcept

How do you design a good LLM-as-judge prompt?

30-second answerSay your answer out loud first, then reveal.

Template structure

text
You are evaluating whether a support answer is GROUNDED in the provided documents.

PASS if every factual claim in the answer is directly supported by the documents.
FAIL if any claim is not supported, contradicts the documents, or adds details
(numbers, dates, policy terms) not present in them.
Ignore style and tone.

Examples:
<example verdict="PASS">...</example>
<example verdict="FAIL">... (adds "within 7 days", docs say 10 days)</example>

<documents>...</documents>
<question>...</question>
<answer>...</answer>

Return JSON: {"reasoning": "<2-3 sentences>", "verdict": "PASS" | "FAIL"}

Design guidelines

  1. Binary verdicts are more reliable and actionable than 1–10 scores (what's the difference between a 6 and a 7?).
  2. Specific criteria from error analysis ("asked for order ID before checking status") beat generic ones ("helpfulness").
  3. Few-shot examples of borderline cases, labelled by experts.
  4. Reasoning before the verdict improves accuracy and makes debugging easier.
  5. A strong model as judge for nuanced criteria; small models for simple checks.
  6. Control biases: randomise order in pairwise comparisons; tell the judge not to favour length.
  7. Version the judge prompt like any other prompt, and re-validate it after changes.

Common mistakes

  • One judge prompt scoring "overall quality 1–10".
  • Not giving the judge the source documents, then asking it about faithfulness.

Little by little, you're building something great.