DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

Verbose mode

Verbose mode is a metric setting, verbose_mode=True, that makes DeepEval print the steps a metric took to reach its score each time it measures a test case; for G-Eval that is the criteria, the evaluation steps, the rubric, the score and the reason.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

In G-Eval, a correct answer scored 0.3, and the cause only showed up after printing the steps the judge had written for itself. Verbose mode prints that working for you, for any metric, without extra code.

The verbose_mode API

python
metric = GEval(..., verbose_mode=True)  # every metric takes verbose_mode
metric.measure(test_case)               # prints the metric's steps, then returns the score
metric.verbose_logs                     # the same log as a string, kept on the metric

A metric that writes its own steps

A criteria-only G-Eval metric is the case where the log helps most, because the judge writes the steps you would otherwise not see.

python
correctness = GEval(
    name="Correctness",
    criteria="Determine whether the actual output is correct based on the expected output.",
    evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
    model=judge,
    verbose_mode=True,
)

A battery answer that leaves out the case

The answer gives the 8 hours per charge and says nothing about the 24 extra hours from the charging case.

python
test_case = LLMTestCase(
    input="How long is the battery life on the SoundPods Pro?",
    actual_output="The SoundPods Pro play for 8 hours on one charge.",
    expected_output=("The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours "
                     "from the charging case, giving a total of 32 hours."),
)
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
judge.py
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))

Printing G-Eval's steps for one answer

ExampleAPI key
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams

from judge import judge

correctness = GEval(
    name="Correctness",
    criteria="Determine whether the actual output is correct based on the expected output.",
    evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
    model=judge,
    verbose_mode=True,
)
test_case = LLMTestCase(
    input="How long is the battery life on the SoundPods Pro?",
    actual_output="The SoundPods Pro play for 8 hours on one charge.",
    expected_output=("The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours "
                     "from the charging case, giving a total of 32 hours."),
)
correctness.measure(test_case)
print("last line of verbose_logs:", correctness.verbose_logs.splitlines()[-1])

What the verbose log showed

  • Criteria is the one line passed to the metric.
  • Evaluation Steps are the judge's own, and once again they ask for a line-by-line exact match, including order and formatting: the same trap as in G-Eval, visible here without any extra print.
  • Rubric: None means no score bands were given, so the judge scored on the default 0 to 10 range that DeepEval turns into 0 to 1.
  • Score: 0.0, with a reason about the missing 24 hours from the case and the 32-hour total. The answer is incomplete, so a low score is fair, but the steps would have failed a complete answer too.
  • The last line of verbose_logs is the score. The string kept on the metric stops there; the reason is printed but not stored in it. Read metric.reason for that.

verbose_mode vs verbose_logs vs reason

verbose_mode=Truemetric.verbose_logsmetric.reason
What it isA settingA string on the metricA string on the metric
Printed for youYes, on every measureNoNo
Holds the stepsYesYesNo
Holds the reasonYes, at the endNo, it stops at the scoreYes

G-Eval always writes a reason and has no include_reason argument; passing one raises TypeError. Metrics that write the reason in a separate judge call, such as DAG metric, take include_reason=False to skip it.

When to turn on verbose mode

  • When a score looks wrong and you need to see which step or verdict produced it.
  • When you are trying a new criteria-only G-Eval metric and want to read the steps the judge wrote before copying them into evaluation_steps.
Watch out. Verbose mode prints a block for every test case a metric measures. Leave it on in a run of fifty test cases and the scores are lost between the logs. Turn it on while you debug one case, then turn it off.
Try it yourself
  • Replace criteria with the three evaluation_steps from the fixed metric in G-Eval and run again: the log shows Criteria: None and your steps.
  • Change the answer to "8 hours per charge, plus 24 more from the case: 32 hours in total." and read the score and reason in the log: with exact-match steps, a complete answer in other words can still score low.
  • Print correctness.reason after measure and check it is the reason the log printed.
PreviousG-Eval

This is what real progress feels like.