DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

G-Eval

G-Eval is a DeepEval metric that scores a test case against criteria you write in plain English: an LLM judge turns the criteria into evaluation steps, follows them, and returns a score from 0 to 1 with a reason.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

Exact match failed a correct answer worded differently, and Custom judge model built a Groq judge that can read meaning instead of characters. G-Eval is the first metric that puts that judge to work: you say what a correct answer is, and the judge grades against it.

A correctness judge with an LLM · from the Complete Agentic AI Course In 10 Hours · 9:43:00 to 9:48:36

Writing a correctness judge by hand

The video writes its LLM judge as a function. First it wraps the OpenAI client with the wrappers of LangSmith, the video's tracing tool, so every judge call is traced. Then it gives the model an instruction: it is an expert professor who grades students' answers to questions. The metric is called correctness. It receives the inputs, the outputs and the reference outputs, which are the ground truth, and it returns a boolean.

The user message holds the question, the real answer and the predicted answer that the chatbot wrote, and asks the model to respond with CORRECT or INCORRECT. The function calls gpt-4o-mini at temperature 0 with the instruction as the system message, reads the reply, and returns true when it says CORRECT and false otherwise. The rows it grades come from the dataset the video built shortly before this clip, such as the question What is LangChain? with the answer A framework for building LLM applications.

The video runs this in LangSmith with gpt-4o-mini; this page does it with DeepEval's G-Eval on the Groq judge. G-Eval writes the grading prompt for you and returns a score between 0 and 1 instead of true or false.

The GEval API

python
from deepeval.metrics import GEval
from deepeval.test_case import SingleTurnParams

metric = GEval(
    name="Correctness",                    # reports show it as "Correctness [GEval]"
    criteria="What a correct answer is, in plain English.",
    evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
    model=judge,                           # the Groq judge from judge.py
)
metric.measure(test_case)                  # sets metric.score (0 to 1) and metric.reason

The criteria in the video's words

criteria is the instruction the video gave its model, rewritten as a description of a good answer. G-Eval adds the test case fields and the request for a score.

python
criteria = ("You are an expert professor grading a student's answer to a question. "
            "Decide whether the predicted answer (actual output) is correct, given the real answer (expected output).")

The fields the judge reads

evaluation_params lists the test case fields that go into the judge's prompt. The video's prompt showed the question, the real answer and the predicted answer, so all three are listed. A field left out of this list never reaches the judge.

python
evaluation_params = [
    SingleTurnParams.INPUT,            # the question
    SingleTurnParams.ACTUAL_OUTPUT,    # the predicted answer
    SingleTurnParams.EXPECTED_OUTPUT,  # the real answer
]
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
judge.py
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))

Grading the video's LangChain question with G-Eval

Two predicted answers to the video's first question, written for this page: one correct in other words, one wrong. Run it next to judge.py from Custom judge model.

ExampleAPI keyFrom the video, run on Groq
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams

from judge import judge

correctness = GEval(
    name="Correctness",
    criteria=("You are an expert professor grading a student's answer to a question. "
              "Decide whether the predicted answer (actual output) is correct, given the real answer (expected output)."),
    evaluation_params=[SingleTurnParams.INPUT, SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
    model=judge,
)
for predicted in ["LangChain is an open-source framework for building apps on top of large language models.",
                  "LangChain is a vector database that stores embeddings."]:
    test_case = LLMTestCase(input="What is LangChain?", actual_output=predicted,
                            expected_output="A framework for building LLM applications.")
    correctness.measure(test_case)
    print(correctness.score, "|", predicted)
    print("   ", correctness.reason)

What the judge decided on the LangChain answers

  • The correct answer in other words scores 0.9. The judge accepts it as a framework for building LLM applications and takes a little off because it adds detail and is a full sentence instead of the short expected phrase.
  • The vector database answer scores 0.0, with a reason that names the mismatch.
  • The result is a number, not true or false. The video's function returned a boolean; G-Eval returns a score and leaves pass or fail to a threshold, which Threshold and strict mode sets.

The same kind of metric can grade the TechNest bot. The next answer to the return policy question states every fact of the expected answer, in other words.

Scoring a paraphrase with criteria only

ExampleAPI key
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams

from judge import judge

correctness = GEval(
    name="Correctness",
    criteria="Determine whether the actual output is correct based on the expected output.",
    evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
    model=judge,
)
test_case = LLMTestCase(
    input="What is TechNest's return policy?",
    actual_output=("You can send an item back within 30 days of buying it, in its original box with every "
                   "accessory. Return postage is on you unless the item came defective, and the refund "
                   "arrives 5 to 7 business days after we get it."),
    expected_output=("TechNest accepts returns within 30 days of purchase. Items must be in original "
                     "packaging with all accessories. Customers pay return shipping unless the item is "
                     "defective. Refunds are processed in 5 to 7 business days."),
)
correctness.measure(test_case)
for step in correctness.evaluation_steps:
    print("-", step)
print(correctness.score, correctness.reason)

Why the correct paraphrase failed

  • The judge wrote its own steps from the one-line criteria, and they ask for an element-by-element exact match.
  • The answer scores 0.3 although it states every fact of the expected answer. The reason faults the missing brand name TechNest, the different phrasing and the extra clauses: wording, not facts.
  • The steps change from run to run. They are written on the metric's first measure, so another run of this code can get other steps; one earlier run wrote similar exact-match steps and scored this answer 0.0.

Fixing the score with evaluation_steps

evaluation_steps replaces the steps the judge would write. G-Eval skips its first call and grades with your list, so the rule about wording is yours, not the judge's. The second answer below gets two facts wrong, to check the fixed metric still fails a bad answer.

ExampleAPI key
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams

from judge import judge

expected = ("TechNest accepts returns within 30 days of purchase. Items must be in original "
            "packaging with all accessories. Customers pay return shipping unless the item is "
            "defective. Refunds are processed in 5 to 7 business days.")
answers = [
    ("You can send an item back within 30 days of buying it, in its original box with every "
     "accessory. Return postage is on you unless the item came defective, and the refund "
     "arrives 5 to 7 business days after we get it."),
    "You can return any item within 60 days, and TechNest always pays the return shipping.",
]
correctness = GEval(
    name="Correctness",
    evaluation_steps=[
        "Check whether the facts in 'actual output' contradict any facts in 'expected output'.",
        "Penalize a missing fact from 'expected output'.",
        "Different wording is fine when the facts are the same.",
    ],
    evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
    model=judge,
)
for actual in answers:
    test_case = LLMTestCase(input="What is TechNest's return policy?", actual_output=actual, expected_output=expected)
    correctness.measure(test_case)
    print(correctness.score, "|", actual[:60] + "...")
    print("   ", correctness.reason)

What the written steps changed

  • The paraphrase scores 1.0. The reason lists the four facts and says the wording differences are acceptable, which is the third step.
  • The 60-day answer scores 0.0: it contradicts the window and who pays shipping, and it leaves out the packaging and refund facts.
  • One judge call per answer. With the steps given, G-Eval goes straight to scoring.

The docs say to pass criteria or evaluation_steps, not both. DeepEval 4.2.8 accepts both without an error and grades with the steps; the criteria then only shows up in logs.

criteria vs evaluation_steps

criteria onlyevaluation_steps
Who writes the grading stepsThe judge, on the first measureYou
Judge calls per test caseTwo the first time (steps, then score), one afterOne
Same steps on every runNo, a new metric can get new stepsYes
Good forExploring a new criterionA metric you rely on in tests

When to use G-Eval

  • When correctness depends on meaning, as with the paraphrased return policy.
  • When the rule is specific to your app and no built-in metric matches it: a tone, a format, a fact every answer must keep.
  • When one overall judgement is enough. For a rule with several fixed checks, DAG metric gives each path its own score.
Watch out. Every field in evaluation_params must be set on the test case. List SingleTurnParams.EXPECTED_OUTPUT and forget expected_output, and measure stops with MissingTestCaseParamsError: 'expected_output' cannot be None for the 'Correctness [GEval]' metric. The opposite mistake is silent: a field your steps mention but the list leaves out never reaches the judge.
Try it yourself
  • Add "LangChain is a Python library for building LLM apps." to the predicted answers in the video's example and read its score.
  • Remove SingleTurnParams.EXPECTED_OUTPUT from the fixed metric's evaluation_params and run it again: the judge no longer sees the policy, so read what its reason says about the 60-day answer.
  • Call correctness.measure(test_case) a second time in the criteria-only example and print the steps again: the metric keeps the steps from its first run.

Slow is fine. Stopping is the only problem.