G-Eval
G-Eval is a DeepEval metric that scores a test case against criteria you write in plain English: an LLM judge turns the criteria into evaluation steps, follows them, and returns a score from 0 to 1 with a reason.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
Exact match failed a correct answer worded differently, and Custom judge model built a Groq judge that can read meaning instead of characters. G-Eval is the first metric that puts that judge to work: you say what a correct answer is, and the judge grades against it.
Writing a correctness judge by hand
The video writes its LLM judge as a function. First it wraps the OpenAI client with the wrappers of LangSmith, the video's tracing tool, so every judge call is traced. Then it gives the model an instruction: it is an expert professor who grades students' answers to questions. The metric is called correctness. It receives the inputs, the outputs and the reference outputs, which are the ground truth, and it returns a boolean.
The user message holds the question, the real answer and the predicted answer that the chatbot wrote, and asks the model to respond with CORRECT or INCORRECT. The function calls gpt-4o-mini at temperature 0 with the instruction as the system message, reads the reply, and returns true when it says CORRECT and false otherwise. The rows it grades come from the dataset the video built shortly before this clip, such as the question What is LangChain? with the answer A framework for building LLM applications.
The video runs this in LangSmith with gpt-4o-mini; this page does it with DeepEval's G-Eval on the Groq judge. G-Eval writes the grading prompt for you and returns a score between 0 and 1 instead of true or false.
The GEval API
from deepeval.metrics import GEval
from deepeval.test_case import SingleTurnParams
metric = GEval(
name="Correctness", # reports show it as "Correctness [GEval]"
criteria="What a correct answer is, in plain English.",
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
model=judge, # the Groq judge from judge.py
)
metric.measure(test_case) # sets metric.score (0 to 1) and metric.reasonThe criteria in the video's words
criteria is the instruction the video gave its model, rewritten as a description of a good answer. G-Eval adds the test case fields and the request for a score.
criteria = ("You are an expert professor grading a student's answer to a question. "
"Decide whether the predicted answer (actual output) is correct, given the real answer (expected output).")The fields the judge reads
evaluation_params lists the test case fields that go into the judge's prompt. The video's prompt showed the question, the real answer and the predicted answer, so all three are listed. A field left out of this list never reaches the judge.
evaluation_params = [
SingleTurnParams.INPUT, # the question
SingleTurnParams.ACTUAL_OUTPUT, # the predicted answer
SingleTurnParams.EXPECTED_OUTPUT, # the real answer
]- written in Custom judge model
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
Grading the video's LangChain question with G-Eval
Two predicted answers to the video's first question, written for this page: one correct in other words, one wrong. Run it next to judge.py from Custom judge model.
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams
from judge import judge
correctness = GEval(
name="Correctness",
criteria=("You are an expert professor grading a student's answer to a question. "
"Decide whether the predicted answer (actual output) is correct, given the real answer (expected output)."),
evaluation_params=[SingleTurnParams.INPUT, SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
model=judge,
)
for predicted in ["LangChain is an open-source framework for building apps on top of large language models.",
"LangChain is a vector database that stores embeddings."]:
test_case = LLMTestCase(input="What is LangChain?", actual_output=predicted,
expected_output="A framework for building LLM applications.")
correctness.measure(test_case)
print(correctness.score, "|", predicted)
print(" ", correctness.reason)0.9 | LangChain is an open-source framework for building apps on top of large language models.
The actual output correctly identifies LangChain as a framework for building LLM applications, matching the core content of the expected answer, though it adds extra detail and uses a full sentence rather than the concise phrase expected.
0.0 | LangChain is a vector database that stores embeddings.
The actual output incorrectly describes LangChain as a vector database storing embeddings, which does not match the expected answer that it is a framework for building LLM applications; therefore it fails to answer the question correctly.What the judge decided on the LangChain answers
- The correct answer in other words scores 0.9. The judge accepts it as a framework for building LLM applications and takes a little off because it adds detail and is a full sentence instead of the short expected phrase.
- The vector database answer scores 0.0, with a reason that names the mismatch.
- The result is a number, not true or false. The video's function returned a boolean; G-Eval returns a score and leaves pass or fail to a threshold, which Threshold and strict mode sets.
The same kind of metric can grade the TechNest bot. The next answer to the return policy question states every fact of the expected answer, in other words.
Scoring a paraphrase with criteria only
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams
from judge import judge
correctness = GEval(
name="Correctness",
criteria="Determine whether the actual output is correct based on the expected output.",
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
model=judge,
)
test_case = LLMTestCase(
input="What is TechNest's return policy?",
actual_output=("You can send an item back within 30 days of buying it, in its original box with every "
"accessory. Return postage is on you unless the item came defective, and the refund "
"arrives 5 to 7 business days after we get it."),
expected_output=("TechNest accepts returns within 30 days of purchase. Items must be in original "
"packaging with all accessories. Customers pay return shipping unless the item is "
"defective. Refunds are processed in 5 to 7 business days."),
)
correctness.measure(test_case)
for step in correctness.evaluation_steps:
print("-", step)
print(correctness.score, correctness.reason)- Read the Expected Output and the Actual Output. - Compare them element‑by‑element (or line‑by‑line) for exact match. - If every element matches and there are no extra or missing parts, mark the Actual Output as correct; otherwise mark it incorrect. - Document any mismatches found. 0.3 The actual output does not exactly match the expected wording; it omits the brand name “TechNest,” uses different phrasing for packaging, shipping responsibility, and refund processing, and adds extra clauses, so the element‑by‑element comparison fails.
Why the correct paraphrase failed
- The judge wrote its own steps from the one-line criteria, and they ask for an element-by-element exact match.
- The answer scores 0.3 although it states every fact of the expected answer. The reason faults the missing brand name TechNest, the different phrasing and the extra clauses: wording, not facts.
- The steps change from run to run. They are written on the metric's first
measure, so another run of this code can get other steps; one earlier run wrote similar exact-match steps and scored this answer 0.0.
Fixing the score with evaluation_steps
evaluation_steps replaces the steps the judge would write. G-Eval skips its first call and grades with your list, so the rule about wording is yours, not the judge's. The second answer below gets two facts wrong, to check the fixed metric still fails a bad answer.
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams
from judge import judge
expected = ("TechNest accepts returns within 30 days of purchase. Items must be in original "
"packaging with all accessories. Customers pay return shipping unless the item is "
"defective. Refunds are processed in 5 to 7 business days.")
answers = [
("You can send an item back within 30 days of buying it, in its original box with every "
"accessory. Return postage is on you unless the item came defective, and the refund "
"arrives 5 to 7 business days after we get it."),
"You can return any item within 60 days, and TechNest always pays the return shipping.",
]
correctness = GEval(
name="Correctness",
evaluation_steps=[
"Check whether the facts in 'actual output' contradict any facts in 'expected output'.",
"Penalize a missing fact from 'expected output'.",
"Different wording is fine when the facts are the same.",
],
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
model=judge,
)
for actual in answers:
test_case = LLMTestCase(input="What is TechNest's return policy?", actual_output=actual, expected_output=expected)
correctness.measure(test_case)
print(correctness.score, "|", actual[:60] + "...")
print(" ", correctness.reason)1.0 | You can send an item back within 30 days of buying it, in it...
The actual output includes all required facts—30‑day return window, original packaging with accessories, customer pays shipping unless defective, and 5‑7 business day refund—matching the expected output with no contradictions; wording differences are acceptable, so it fully aligns.
0.0 | You can return any item within 60 days, and TechNest always ...
The actual output contradicts the expected facts by stating a 60‑day return window and that TechNest always pays shipping, whereas the expected says 30 days and customers pay unless defective. It also omits required details about original packaging and refund processing time, so it fails the evaluation criteria.What the written steps changed
- The paraphrase scores 1.0. The reason lists the four facts and says the wording differences are acceptable, which is the third step.
- The 60-day answer scores 0.0: it contradicts the window and who pays shipping, and it leaves out the packaging and refund facts.
- One judge call per answer. With the steps given, G-Eval goes straight to scoring.
The docs say to pass criteria or evaluation_steps, not both. DeepEval 4.2.8 accepts both without an error and grades with the steps; the criteria then only shows up in logs.
criteria vs evaluation_steps
criteria only | evaluation_steps | |
|---|---|---|
| Who writes the grading steps | The judge, on the first measure | You |
| Judge calls per test case | Two the first time (steps, then score), one after | One |
| Same steps on every run | No, a new metric can get new steps | Yes |
| Good for | Exploring a new criterion | A metric you rely on in tests |
When to use G-Eval
- When correctness depends on meaning, as with the paraphrased return policy.
- When the rule is specific to your app and no built-in metric matches it: a tone, a format, a fact every answer must keep.
- When one overall judgement is enough. For a rule with several fixed checks, DAG metric gives each path its own score.
evaluation_params must be set on the test case. List SingleTurnParams.EXPECTED_OUTPUT and forget expected_output, and measure stops with MissingTestCaseParamsError: 'expected_output' cannot be None for the 'Correctness [GEval]' metric. The opposite mistake is silent: a field your steps mention but the list leaves out never reaches the judge.Related
- Previous: Custom judge model
- Next: Verbose mode
- Reference: G-Eval
- Add
"LangChain is a Python library for building LLM apps."to the predicted answers in the video's example and read its score. - Remove
SingleTurnParams.EXPECTED_OUTPUTfrom the fixed metric'sevaluation_paramsand run it again: the judge no longer sees the policy, so read what its reason says about the 60-day answer. - Call
correctness.measure(test_case)a second time in the criteria-only example and print the steps again: the metric keeps the steps from its first run.
Slow is fine. Stopping is the only problem.