LLM as a judge
LLM as a judge is an evaluation method in which a capable language model reads a test case and grades it against written criteria, taking the place of a human expert.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
Five goldens can be checked by hand. Five thousand cannot. The judge is what lets an evaluation scale, and in RAGAS it is one object you build once and hand to every metric.
Replacing the expert with a model
The video goes back to its exam analogy. The teacher, the domain expert, marks every answer against the correct one. For thousands of questions that is slow and expensive: an expert might charge a hundred dollars an hour. So an LLM takes the teacher's place. It receives the question, the app's answer, the expected answer and the retrieved context, and one call grades them. The rules it grades by, different for each kind of question, are the metrics.
Then the video asks the class: should the judge be a small, cheap model or a large, smart one? A large one. Grading needs careful reasoning, the same reason you pick a stronger model for a hard task yourself.
The AI security course version of this app judged with llama-3.1-8b-instant, a small model picked to stay inside Groq's free token limits, and Groq has since retired it from the free plan. This course makes the same trade for the same reason: the judge runs on the smaller openai/gpt-oss-20b, which keeps evaluation fast and inside Groq's free daily token limit, while the app answers with qwen/qwen3.8-27b, a different model family, so the judge is not grading its own writing. Any model that returns structured output works as the judge; for a release decision, use the larger model the video recommends.
The llm_factory API
from openai import AsyncOpenAI
from ragas.llms import llm_factory
client = AsyncOpenAI(api_key=..., base_url=...) # any OpenAI-compatible server
judge = llm_factory("model-name", provider="openai", client=client) # the judge every metric usesA Groq client for the judge
Groq speaks the OpenAI API, so the openai package's client works with Groq's URL. It reads JUDGE_GROQ first, the video's name for a separate judge key, and falls back to GROQ_API_KEY.
groq = AsyncOpenAI(
api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
base_url="https://api.groq.com/openai/v1",
)Wrapping the model with llm_factory
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)provider="openai" names the API style, not the company: any server that speaks the OpenAI API, Groq included, uses it. These three lines are the video's build_judge function with the model name changed.
Embeddings for two of the metrics
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")Answer relevancy and answer correctness also compare meanings, which needs an embedding model. The video's repository loads a local sentence-transformers model for this; this course uses Gemini's hosted embeddings, so nothing is downloaded to your machine. gemini-embedding-2 merges a list of texts sent in one call into a single embedding, and answer relevancy embeds three questions at once, so the small subclass sends one text per call.
The judge.py file
Save the pieces as judge.py. Every lesson from here imports judge, and two import embeddings.
import os
from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
groq = AsyncOpenAI(
api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")Asking the judge the video's question
RAGAS metrics are built from small judge questions. The video's faithfulness slide shows one: Can this claim be fully inferred from the retrieved context, yes or no? A judge answers into a Pydantic class you name, so ask it that question about two claims from the video's bank example.
from pydantic import BaseModel
from judge import judge
class Verdict(BaseModel):
verdict: int # 1 = the context supports the claim, 0 = it does not
reason: str
context = "Minimum balance is ₹10,000 for urban branches. Non-maintenance fee is ₹350 + GST if balance falls below."
for claim in ["Min balance is ₹10,000 for urban branches", "Online fund transfer has no extra charge"]:
prompt = ("Can this claim be fully inferred from the retrieved context? Answer 1 for yes, 0 for no.\n"
f"Context: {context}\nClaim: {claim}")
result = judge.generate(prompt, Verdict)
print(result.verdict, "|", claim)
print(" ", result.reason)1 | Min balance is ₹10,000 for urban branches
The context explicitly states that the minimum balance for urban branches is ₹10,000, which directly supports the claim.
0 | Online fund transfer has no extra charge
The provided context only discusses minimum balance requirements and non‑maintenance fees; it does not mention any fees (or lack thereof) for online fund transfers, so the claim cannot be inferred.What the judge returned
- The first claim is in the context word for word, and the judge marks it 1.
- The second claim, the one the video's slide marks as hallucinated, is not in the context, and the judge marks it 0 with a reason.
- The result is a
Verdictobject, not text.judge.generate(prompt, Verdict)makes the model fill the class, which is how every RAGAS metric reads its judge's answers.
A human expert vs an LLM judge vs embeddings
| Human expert | LLM judge | Embedding similarity | |
|---|---|---|---|
| Quality | Best | Close to an expert with a strong model | Weak on complex answers |
| Cost | Highest | Per call | Cheapest |
| Scale | Tens of questions | Thousands | Thousands |
| Role in RAGAS | Writes the goldens | Runs the metrics | Part of two metrics |
The AI security course version makes the same comparison: embeddings are the cheap option for simple checks, and a strong LLM gives the better judgement when answers are complex.
When to use an LLM judge
- When the right answer can be worded many ways, so a string comparison fails good answers.
- When a quality has to be read, such as whether an answer is grounded, relevant or complete.
Related
- Previous: ID-based context precision and recall
- Next: RAGAS metrics
- Reference: LLMs reference
- Ask the judge about the claim
"Rural branch minimum is ₹2,500"with the same context and read its verdict. - Add a third field,
confidence: float, toVerdictand print it. - Change the model in
judge.pytoopenai/gpt-oss-120b, run the example again, and compare the reasons.
Slow is fine. Stopping is the only problem.