Verbose mode
Verbose mode is a metric setting, verbose_mode=True, that makes DeepEval print the steps a metric took to reach its score each time it measures a test case; for G-Eval that is the criteria, the evaluation steps, the rubric, the score and the reason.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
In G-Eval, a correct answer scored 0.3, and the cause only showed up after printing the steps the judge had written for itself. Verbose mode prints that working for you, for any metric, without extra code.
The verbose_mode API
metric = GEval(..., verbose_mode=True) # every metric takes verbose_mode
metric.measure(test_case) # prints the metric's steps, then returns the score
metric.verbose_logs # the same log as a string, kept on the metricA metric that writes its own steps
A criteria-only G-Eval metric is the case where the log helps most, because the judge writes the steps you would otherwise not see.
correctness = GEval(
name="Correctness",
criteria="Determine whether the actual output is correct based on the expected output.",
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
model=judge,
verbose_mode=True,
)A battery answer that leaves out the case
The answer gives the 8 hours per charge and says nothing about the 24 extra hours from the charging case.
test_case = LLMTestCase(
input="How long is the battery life on the SoundPods Pro?",
actual_output="The SoundPods Pro play for 8 hours on one charge.",
expected_output=("The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours "
"from the charging case, giving a total of 32 hours."),
)- written in Custom judge model
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
Printing G-Eval's steps for one answer
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams
from judge import judge
correctness = GEval(
name="Correctness",
criteria="Determine whether the actual output is correct based on the expected output.",
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
model=judge,
verbose_mode=True,
)
test_case = LLMTestCase(
input="How long is the battery life on the SoundPods Pro?",
actual_output="The SoundPods Pro play for 8 hours on one charge.",
expected_output=("The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours "
"from the charging case, giving a total of 32 hours."),
)
correctness.measure(test_case)
print("last line of verbose_logs:", correctness.verbose_logs.splitlines()[-1])**************************************************
Correctness [GEval] Verbose Logs
**************************************************
Criteria:
Determine whether the actual output is correct based on the expected output.
Evaluation Steps:
[
"Compare the Actual Output to the Expected Output line by line",
"If every line matches exactly (including order and formatting), consider the output correct; otherwise it is incorrect",
"Record any mismatches in content, order, or formatting as reasons for failure",
"Conclude with a final determination of correct or incorrect based on the comparison"
]
Rubric:
None
Score: 0.0
Reason: The actual output only mentions 8 hours on one charge, while the expected output includes additional 24 hours from the case and a total of 32 hours, so the content does not match line by line
======================================================================
last line of verbose_logs: Score: 0.0What the verbose log showed
- Criteria is the one line passed to the metric.
- Evaluation Steps are the judge's own, and once again they ask for a line-by-line exact match, including order and formatting: the same trap as in G-Eval, visible here without any extra print.
- Rubric: None means no score bands were given, so the judge scored on the default 0 to 10 range that DeepEval turns into 0 to 1.
- Score: 0.0, with a reason about the missing 24 hours from the case and the 32-hour total. The answer is incomplete, so a low score is fair, but the steps would have failed a complete answer too.
- The last line of
verbose_logsis the score. The string kept on the metric stops there; the reason is printed but not stored in it. Readmetric.reasonfor that.
verbose_mode vs verbose_logs vs reason
verbose_mode=True | metric.verbose_logs | metric.reason | |
|---|---|---|---|
| What it is | A setting | A string on the metric | A string on the metric |
| Printed for you | Yes, on every measure | No | No |
| Holds the steps | Yes | Yes | No |
| Holds the reason | Yes, at the end | No, it stops at the score | Yes |
G-Eval always writes a reason and has no include_reason argument; passing one raises TypeError. Metrics that write the reason in a separate judge call, such as DAG metric, take include_reason=False to skip it.
When to turn on verbose mode
- When a score looks wrong and you need to see which step or verdict produced it.
- When you are trying a new criteria-only G-Eval metric and want to read the steps the judge wrote before copying them into
evaluation_steps.
Related
- Previous: G-Eval
- Next: Threshold and strict mode
- Reference: Debug a metric judgement
- Replace
criteriawith the threeevaluation_stepsfrom the fixed metric in G-Eval and run again: the log showsCriteria: Noneand your steps. - Change the answer to
"8 hours per charge, plus 24 more from the case: 32 hours in total."and read the score and reason in the log: with exact-match steps, a complete answer in other words can still score low. - Print
correctness.reasonaftermeasureand check it is the reason the log printed.
This is what real progress feels like.