LLM as a judge
LLM as a judge is an evaluation method in which a language model reads a test case and returns a verdict on a fixed question, in place of a human grader.
Last updated: 09 Oct, 2026 · RAGAS 0.4
Goldens ended with test cases: for each query, the expected answer next to the actual one. Comparing five pairs by hand is tedious and a hundred is out of reach. The comparison needs a reader that scales.
Why a model does the grading
This part of the video starts at 1:50:13. The app in the video judged with llama-3.1-8b-instant, since retired on Groq; the judge below is openai/gpt-oss-20b.
The video puts an LLM where the teacher was. It is "a simple brain" that can take the query and the expected answer and tell whether the actual answer is near it. The clip weighs three ways of doing that comparison:
- A human expert. The best reader and the most expensive one.
- An embedding comparison. Turn the actual and the expected answer into vectors and measure how close they are. It is cheap and fine for a simple check. For complex queries the better reading comes from an LLM.
- An LLM. It reads like a person and costs a model call. Large models cost more than small ones, and all of them cost less than an expert's hours.
One problem remains, and the video names it: "there is lack of structure here". Asked only "is this answer good?", a model has nothing fixed to check, and you would have to explain again each time what to compare with what. A metric supplies the structure. It fixes which fields of the test case the judge receives, which question it is asked, and how its verdicts become a number. Frameworks such as RAGAS and DeepEval ship these metrics; the video then opens the RAGAS page for context recall, where the definition, the formula and the code are written down.
A stronger judge gives more reliable verdicts, which is why the clip reaches for a state-of-the-art model. The video's app still judged with a small 8B model, to stay inside free rate limits. The judge here is openai/gpt-oss-20b for the same reason: an evaluation run makes many calls, and the smaller model keeps them inside a free daily budget. For a release decision, use the strongest judge you can afford.
Setting up the judge with llm_factory
A Groq client and the judge
llm_factory wraps a chat model so that RAGAS metrics can ask it questions and get structured replies. Groq speaks the OpenAI API, so the client is the openai package's AsyncOpenAI pointed at Groq's URL.
import os
from openai import AsyncOpenAI
from ragas.llms import llm_factory
client = AsyncOpenAI(
base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"],
max_retries=6, # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)provider="openai"matches the client: an OpenAI SDK client, here pointed at Groq's OpenAI-compatible URL.AsyncOpenAIis the client to pass, because the metrics call the judge's async method,agenerate.temperature=0asks for the most repeatable replies the model can give.max_tokens=4096gives room for long structured replies. A reply cut off in the middle of its JSON makes the metric fail.max_retries=6on the client makes it wait and try again when Groq answers 429 because the per-minute token limit is used up.
The video's repository builds its judge the same way, in a build_judge function. Only the model name, temperature, max_tokens and max_retries differ here.
Embeddings for two of the metrics
Two of the five metrics also compare meanings, which takes an embedding model. The video's repository loads a local sentence-transformers model for this. The code here uses Gemini's hosted gemini-embedding-2 through RAGAS' GoogleEmbeddings, so nothing is downloaded. That model returns one vector per call even when it is given a list, so a small subclass sends each text in its own call. genai.Client() reads GEMINI_API_KEY from the environment.
from google import genai
from ragas.embeddings import GoogleEmbeddings
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 returns one vector per call, so each text is embedded on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")The metric lessons, starting with Faithfulness, reuse these two objects, judge and embeddings.
Asking the judge one question
RAGAS metrics are built from small questions. The faithfulness slide of the video shows one: "Can this claim be fully inferred from the retrieved context, yes or no?" The example asks the judge that question about four claims and one catalog entry, the SoundPods Pro. The judge answers into a Pydantic class, Verdict, so the reply is an object with fields and not free text. The last line checks the embeddings.
import asyncio
import os
import textwrap
from google import genai
from openai import AsyncOpenAI
from pydantic import BaseModel
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
client = AsyncOpenAI(
base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"],
max_retries=6, # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 returns one vector per call, so each text is embedded on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
class Verdict(BaseModel):
verdict: int # 1 = the context supports the claim, 0 = it does not
reason: str
CONTEXT = "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
CLAIMS = ["The SoundPods Pro play for 8 hours on one charge.",
"The SoundPods Pro are waterproof to 10 metres.",
"With the charging case the total battery life is 32 hours.",
"With the charging case the total battery life is 24 hours."]
async def main():
for claim in CLAIMS:
prompt = ("Can this claim be fully inferred from the retrieved context? Answer 1 for yes, 0 for no.\n"
f"Context: {CONTEXT}\nClaim: {claim}")
result = await judge.agenerate(prompt, Verdict) # the reply is a Verdict object
print(result.verdict, "|", claim)
print(textwrap.fill(result.reason, 100, initial_indent=" ", subsequent_indent=" "))
vector = await embeddings.aembed_text("return policy")
print("embedding length:", len(vector))
print("judge settings:", judge.model_args)
asyncio.run(main())judge settings: {'temperature': 0, 'top_p': 0.1, 'max_tokens': 4096}
1 | The SoundPods Pro play for 8 hours on one charge.
The context explicitly states that the SoundPods Pro provide 8 hours of playback per charge.
0 | The SoundPods Pro are waterproof to 10 metres.
The context states the earbuds have IPX4 water resistance, which only guarantees splash
resistance, not waterproofing to 10 metres. Therefore the claim cannot be inferred from the
provided context.
1 | With the charging case the total battery life is 32 hours.
The context states 8 hours of playback per charge plus 24 hours with the case, which sums to 32
hours.
1 | With the charging case the total battery life is 24 hours.
The context explicitly states that the earbuds provide 8 hours of playback per charge and 24
hours with the charging case, directly supporting the claim.
embedding length: 3072What the judge returned
- The judge's settings are visible.
judge.model_argsshows thetemperatureof 0 and themax_tokensof 4096 passed tollm_factory, next to atop_pof 0.1 that RAGAS sets itself. - The first two verdicts are the easy ones. The entry states 8 hours of playback per charge, so the first claim gets a 1. It gives a water resistance rating of IPX4 and nothing about 10 metres, so the second gets a 0. The remark that IPX4 means splash resistance is the judge's own knowledge; the entry does not say it.
- The last two claims contradict each other, and both get a 1. For 32 hours the judge adds the 8 and the 24. For 24 hours its reason quotes the 8 hours per charge and the 24 hours with the charging case, and calls the claim directly supported. The entry's wording allows both readings. The judge takes each question on its own, so it never notices that it agreed to two different totals.
- The embeddings work. One text came back as a vector of 3072 numbers.
Scoring the same sample three times
The 24-hour reading is the one the app itself chose for golden g003. A verdict is one model's reading, and a model does not have to read the same way twice, so the next check is to score that one sample several times. The sample is the answer the app gave to g003 in Goldens, with the SoundPods Pro entry as its context. The metric is faithfulness, the share of the answer's claims that the retrieved context supports; Faithfulness takes it apart. One asyncio.run wraps all three scores, and ascore is the async form of score.
import asyncio
import os
from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import Faithfulness
client = AsyncOpenAI(
base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"],
max_retries=6, # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)
metric = Faithfulness(llm=judge)
sample = {
"user_input": "How long is the battery life on the SoundPods Pro?",
"response": "The TechNest SoundPods Pro offer 8 hours of playback on a single charge, and the charging case extends this to a total of 24 hours.",
"retrieved_contexts": ["The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."],
}
async def main():
for run in range(1, 4):
result = await metric.ascore(**sample)
print(f"run {run}: faithfulness = {result.value}")
asyncio.run(main())run 1: faithfulness = 1.0 run 2: faithfulness = 1.0 run 3: faithfulness = 1.0
What three runs of one sample show
- All three runs gave 1.0. On this sample the judge is consistent.
- Consistent is not the same as correct. The answer says the case brings the total to 24 hours. The golden's reference says 32. The judge finds every claim supported all three times, as it accepted the 24-hour claim in the example before.
- Faithfulness never saw the reference. It reads the question, the answer and the context. A judge can only compare what the metric hands it, which is why one metric is never the whole evaluation.
How far to trust a judge
- A judged score is one judge's reading. Another model, or the same model on another day, can read the same case differently. Fix the judge model and its settings, and compare scores only between runs that used the same judge.
- Scores can move between runs. The video's app is run twice on its five goldens: the average faithfulness is 0.91 in the first results table and 0.97 in the second, while the other four averages stay the same. The answers were generated again between the two runs, so the generator or the judge can be behind the change. Either way, a score is a reading of one run.
- A judge favours its own writing. A model tends to rate text from itself or from its own model family higher, an effect called self-preference. It follows the model, not the company that hosts it. Here the generator is a Qwen model and the judge a gpt-oss model, two different families.
- A person stays in the loop. Read a sample of the verdicts yourself, starting with the lowest scores and with any high score that surprises you. This is the human in the review pipeline that LLM evaluation mentioned.
Human expert vs LLM judge vs embedding similarity
| Human expert | LLM judge | Embedding similarity | |
|---|---|---|---|
| Reads meaning | Yes | Yes, with errors of its own | Only how close two texts are |
| Cost per case | Highest | One or more model calls | One embedding call per text |
| Scale | Tens of cases | Thousands | Thousands |
| Gives a reason | Yes | Yes, on request | No |
| Repeatable | Varies between people | Close, not guaranteed | Yes |
| Role in an evaluation | Writes the goldens, audits the judge | Scores the test cases | A part of two of the RAGAS metrics |
Where you use LLM as a judge
- When a right answer can be worded many ways, so that comparing strings fails good answers.
- When a quality has to be read, such as whether an answer is supported by its sources or stays on the question.
- When the set is too large to read, with a person auditing a sample.
Related
- Previous: Goldens
- Next: RAGAS metrics
- Reference: RAGAS LLMs reference
- Delete the words
plus 24 hours with the casefromCONTEXTand run the first example again. Read what happens to the last two verdicts and to their reasons. - Change the judge's model to
"openai/gpt-oss-120b"in the three-run example and compare the scores with the ones above. - Change
range(1, 4)torange(1, 6)for five runs and see whether any score differs from the others.
You understood something today that you didn't yesterday.