DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

Threshold and strict mode

A threshold is the lowest score a DeepEval metric accepts as a pass: is_successful() compares the score with it, and strict_mode=True turns the score into 1 or 0 and sets the threshold to 1.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

The battery answer in Verbose mode scored 0.0 on the judge's own steps. A score alone does not say whether a test should fail. The threshold decides that, and you choose it per metric.

The threshold API

python
metric = GEval(..., threshold=0.7)    # pass when score >= 0.7
metric = GEval(..., threshold=None)   # score-only: no pass or fail
metric = GEval(..., strict_mode=True) # score is 1 or 0, threshold becomes 1
metric.measure(test_case)
metric.is_successful()                # True, False, or None in score-only mode

Every metric scores in the same direction: 1 is the best score and the threshold is a minimum. metric.success holds the same value is_successful() returns after a measure.

Moving the threshold on exact match

Exact match needs no judge, so it shows the three settings for free. The test case is the warranty email from First test run, where a full sentence scored 0 against the bare address.

Example
from deepeval.metrics import ExactMatchMetric
from deepeval.test_case import LLMTestCase

test_case = LLMTestCase(
    input="Which email do I write to for a warranty claim?",
    actual_output="Email support@technest.com with your order number.",
    expected_output="support@technest.com",
)
for threshold in [1.0, 0.0, None]:
    metric = ExactMatchMetric(threshold=threshold)
    metric.measure(test_case)
    print(f"threshold={threshold}  score={metric.score}  is_successful()={metric.is_successful()}")
  • The score is 0.0 every time. The threshold never changes the score, only the verdict on it.
  • threshold=0.0 passes a score of 0.0, because the check is score is at least the threshold. A threshold of 0 passes everything.
  • threshold=None returns None: score-only mode. The score is still computed and reported, but the metric has no opinion on pass or fail.

Moving the threshold on a G-Eval score

A judged metric costs a judge call per measure. The threshold is only read when is_successful() runs, so one measure can be checked against several thresholds. The steps and the battery answer are the ones from G-Eval and Verbose mode; save them as battery_case.py so the next two examples can import them.

python
from deepeval.test_case import LLMTestCase, SingleTurnParams

steps = [
    "Check whether the facts in 'actual output' contradict any facts in 'expected output'.",
    "Penalize a missing fact from 'expected output'.",
    "Different wording is fine when the facts are the same.",
]
params = [SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT]
test_case = LLMTestCase(
    input="How long is the battery life on the SoundPods Pro?",
    actual_output="The SoundPods Pro play for 8 hours on one charge.",
    expected_output=("The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours "
                     "from the charging case, giving a total of 32 hours."),
)
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
judge.py
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
ExampleAPI key
from deepeval.metrics import GEval

from battery_case import params, steps, test_case
from judge import judge

correctness = GEval(name="Correctness", evaluation_steps=steps, evaluation_params=params, model=judge)
correctness.measure(test_case)
for threshold in [0.3, 0.5, 0.8, None]:
    correctness.threshold = threshold
    print(f"threshold={threshold}  score={correctness.score}  passed={correctness.is_successful()}")

What the thresholds decided

  • The score is 0.5 for the answer that leaves out the case, from one judge call.
  • 0.3 and 0.5 pass, 0.8 fails. A score equal to the threshold passes. Changing threshold and calling is_successful() again made no new judge call.
  • None gives None, the same score-only result as with exact match.

Scoring the same answer with strict_mode

ExampleAPI key
from deepeval.metrics import GEval

from battery_case import params, steps, test_case
from judge import judge

strict = GEval(name="Correctness", evaluation_steps=steps, evaluation_params=params, model=judge, strict_mode=True)
print("threshold before measure:", strict.threshold)
strict.measure(test_case)
print(f"score={strict.score}  passed={strict.is_successful()}")
print(strict.reason)

What strict mode changed

  • The threshold is 1 before any measure. strict_mode=True sets it when the metric is built.
  • The score is 0, printed as a whole number, and the answer fails. Without strict mode the same answer scored 0.5: strict mode asks the judge for full marks or nothing.
  • The reason names the missing 24 hours from the case, the fact that cost the answer its point.

A threshold vs score-only vs strict mode

threshold=0.7threshold=Nonestrict_mode=True
Score0 to 10 to 11 or 0
is_successful()True or FalseNoneTrue only for a perfect answer
Can fail a testYesNoYes
Good forMost checksTracking a number without gating on itRules with no partial credit

When to set each one

  • Raise the threshold when a partly correct answer should fail, as with a battery answer that leaves out the case.
  • Use threshold=None for a metric you want to watch for a while before it can block anything. The docs note that every evaluate() or assert_test() call still needs at least one non-flaky metric with a threshold.
  • Use strict_mode when a half-right answer is as bad as a wrong one, such as a price or a refund window.
Watch out. strict_mode=True overrides the threshold you pass. GEval(..., threshold=0.5, strict_mode=True) has threshold 1, so a test you expected to pass at 0.5 fails on any answer that is not perfect.
Try it yourself
  • Add 0.5 to the thresholds in the exact match example: it fails, because the score is 0.0.
  • Pass threshold=0.5, strict_mode=True to the strict metric and check the printed threshold is still 1.
  • Change the battery answer to include the 24 hours from the case and the 32-hour total, and run the strict example again.
PreviousVerbose mode

Every expert started right here.