DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

Custom judge model

A custom judge model is a class that inherits DeepEval's DeepEvalBaseLLM and answers a metric's prompts with any LLM you choose; DeepEval sends every judged metric's questions through it.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

Exact match and pattern match needed no model. Most DeepEval metrics read an answer the way a person would, and for that they ask an LLM. This lesson builds the judge the rest of the course uses: openai/gpt-oss-120b on Groq.

Why an LLM replaces the expert · from the Production RAG Live Marathon · 3:20:50 to 3:27:17

Replacing the expert with a model

The video asks whether a human expert should mark every answer. For five questions, maybe. For thousands, an expert charging a hundred dollars an hour is not scalable, even though the quality would be superb. So an LLM takes the teacher's place: it gets the question, the actual output, the expected output and the context, and one LLM call evaluates them. How it evaluates depends on the kind of question, the way an exam has a different marking scheme for an essay and a multiple-choice question. Those rules are the metrics: answer relevancy, faithfulness, contextual relevancy, contextual precision.

Then the video asks the class: should the judge be a small, cheap model or a large, smart one? A large one. Judging needs careful reasoning, the same reason you give a hard task to a stronger model; a paper checked by an eight-year-old is chaos, not an evaluation. This course follows that advice with openai/gpt-oss-120b, the larger of Groq's two gpt-oss models, while the bot writes its answers with qwen/qwen3.8-27b, a different model family.

Running a metric with the default judge

Every judged metric takes a model argument. Leave it out and DeepEval builds its default, an OpenAI model, which needs an OpenAI key the moment the metric is created.

Example
from deepeval.metrics import AnswerRelevancyMetric

metric = AnswerRelevancyMetric()

DeepEval has built-in classes for several providers, OpenAI, Gemini and Anthropic among them, and generic ones such as LocalModel and LiteLLMModel that can reach Groq's endpoint, shown in Integrations. This course writes its own class on DeepEval's base class instead, so the retries and the structured output are code you can read.

The DeepEvalBaseLLM API

python
from deepeval.models import DeepEvalBaseLLM

class MyJudge(DeepEvalBaseLLM):
    def load_model(self): ...                            # return the client or model object
    def get_model_name(self): ...                        # the name printed in reports
    def generate(self, prompt, schema=None): ...         # answer one prompt
    async def a_generate(self, prompt, schema=None): ... # the same, for async metrics

When a metric passes a schema, a Pydantic class, generate must return an instance of that class instead of text. That is how a metric gets a verdict or a list of claims it can read.

Two Groq clients

python
def __init__(self, model="openai/gpt-oss-120b"):
    self.model_name = model
    key = os.environ["GROQ_API_KEY"]
    # on a 429 (rate limit) the client waits and tries again, up to 8 times
    self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
    self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

Groq speaks the OpenAI API, so the openai package's clients work with Groq's URL. Metrics run asynchronously by default, which is why there are two. max_retries=8 matters on the free plan: one metric can send several prompts within seconds, and when Groq answers 429 for too many tokens in a minute, the client waits the time Groq asks for and tries again.

The name and the model object

python
def load_model(self):
    return self.client

def get_model_name(self):
    return self.model_name

Asking for JSON in the metric's shape

python
def request(self, prompt, schema):
    request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
    if schema is not None:
        # ask Groq for JSON in the shape of the metric's Pydantic schema
        json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
        request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
    return request

With a schema, the request carries Groq's structured output setting, built from the Pydantic class, so the model replies with JSON that has exactly those fields. temperature=0 keeps the judge's answers as steady as the model allows.

generate and a_generate

python
def generate(self, prompt, schema=None):
    reply = self.client.chat.completions.create(**self.request(prompt, schema))
    text = reply.choices[0].message.content
    return schema.model_validate_json(text) if schema else text

model_validate_json turns the JSON text into an instance of the schema. a_generate is the same method with await and the async client.

The judge.py file

Save the whole class as judge.py. Every lesson from here imports judge from it.

python
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))

The last line reads an optional JUDGE_MODEL variable. Groq's free plan gives each model its own daily token budget, so when openai/gpt-oss-120b runs out, export JUDGE_MODEL=openai/gpt-oss-20b switches the judge without editing the file. The smaller model sometimes returns JSON that does not fit the schema, and Groq answers 400 json_validate_failed; that error is not retried, so run the example again.

Asking the judge a plain question

ExampleAPI key
from judge import judge

print(judge.get_model_name())
print(judge.generate("In one sentence, what does an LLM judge do in an evaluation?"))

Asking for a verdict object

Metrics never ask for free text. They ask a narrow question and name the shape of the answer. Here is one such question about a claim the bot might make, checked against the return policy.

ExampleAPI key
from typing import Literal

from pydantic import BaseModel

from judge import judge


class Verdict(BaseModel):
    verdict: Literal["yes", "no"]  # yes = the context supports the claim
    reason: str


context = ("TechNest accepts returns within 30 days of the original purchase date. Customers are "
           "responsible for return shipping costs unless the item arrives defective or damaged. "
           "Refunds are processed within 5 to 7 business days of receiving the returned item.")
claims = ["Refunds are processed within 5 to 7 business days.", "TechNest pays the return shipping on every order."]
for claim in claims:
    prompt = ("Does the context support the claim? Answer yes or no, with a one-sentence reason.\n"
              f"Context: {context}\nClaim: {claim}")
    result = judge.generate(prompt, schema=Verdict)
    print(result.verdict, "|", claim)
    print("   ", result.reason)
print(type(result).__name__)

What the judge returned

  • The first claim is in the policy almost word for word, and the judge answers yes.
  • The second claim contradicts the policy, which makes customers pay return shipping unless the item is defective, and the judge answers no with that reason.
  • The last line is Verdict: the judge returned an object, not a string. This is the same path a DeepEval metric uses when it asks its own questions.

Passing the judge to a metric

python
from deepeval.metrics import AnswerRelevancyMetric
from judge import judge

metric = AnswerRelevancyMetric(model=judge)  # every judged metric takes model=

A built-in model vs a custom judge

Default (OpenAIModel)Built-in class (GeminiModel and others)Custom DeepEvalBaseLLM
Set up withOPENAI_API_KEYThe provider's key, and sometimes an extra packageA class you write
ProvidersOpenAIThe ones DeepEval ships classes forAny API or local model
Structured outputHandled for youHandled for youYour generate must return the schema
In this courseNot usedGemini, in IntegrationsGroqJudge

When to write your own judge class

  • When you want to see and control every judge call: how it retries, how it asks for JSON, which model it uses.
  • When every judge call should go through your own client: a gateway, a retry policy, or logging of each prompt.
Watch out. A generate that ignores schema and returns text works for a plain prompt but breaks metrics such as faithfulness, which read fields from the object they asked for. Test a new judge with a schema, as above, before running metrics on it.
Try it yourself
  • Add the claim "Returns are accepted within 60 days." to claims and read the judge's reason.
  • Add a field confidence: float to Verdict and print it for each claim.
  • Run the plain question with export JUDGE_MODEL=openai/gpt-oss-20b and check the model name it prints.

You understood something today that you didn't yesterday.