Threshold and strict mode
A threshold is the lowest score a DeepEval metric accepts as a pass: is_successful() compares the score with it, and strict_mode=True turns the score into 1 or 0 and sets the threshold to 1.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
The battery answer in Verbose mode scored 0.0 on the judge's own steps. A score alone does not say whether a test should fail. The threshold decides that, and you choose it per metric.
The threshold API
metric = GEval(..., threshold=0.7) # pass when score >= 0.7
metric = GEval(..., threshold=None) # score-only: no pass or fail
metric = GEval(..., strict_mode=True) # score is 1 or 0, threshold becomes 1
metric.measure(test_case)
metric.is_successful() # True, False, or None in score-only modeEvery metric scores in the same direction: 1 is the best score and the threshold is a minimum. metric.success holds the same value is_successful() returns after a measure.
Moving the threshold on exact match
Exact match needs no judge, so it shows the three settings for free. The test case is the warranty email from First test run, where a full sentence scored 0 against the bare address.
from deepeval.metrics import ExactMatchMetric
from deepeval.test_case import LLMTestCase
test_case = LLMTestCase(
input="Which email do I write to for a warranty claim?",
actual_output="Email support@technest.com with your order number.",
expected_output="support@technest.com",
)
for threshold in [1.0, 0.0, None]:
metric = ExactMatchMetric(threshold=threshold)
metric.measure(test_case)
print(f"threshold={threshold} score={metric.score} is_successful()={metric.is_successful()}")threshold=1.0 score=0.0 is_successful()=False threshold=0.0 score=0.0 is_successful()=True threshold=None score=0.0 is_successful()=None
- The score is 0.0 every time. The threshold never changes the score, only the verdict on it.
- threshold=0.0 passes a score of 0.0, because the check is score is at least the threshold. A threshold of 0 passes everything.
- threshold=None returns None: score-only mode. The score is still computed and reported, but the metric has no opinion on pass or fail.
Moving the threshold on a G-Eval score
A judged metric costs a judge call per measure. The threshold is only read when is_successful() runs, so one measure can be checked against several thresholds. The steps and the battery answer are the ones from G-Eval and Verbose mode; save them as battery_case.py so the next two examples can import them.
from deepeval.test_case import LLMTestCase, SingleTurnParams
steps = [
"Check whether the facts in 'actual output' contradict any facts in 'expected output'.",
"Penalize a missing fact from 'expected output'.",
"Different wording is fine when the facts are the same.",
]
params = [SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT]
test_case = LLMTestCase(
input="How long is the battery life on the SoundPods Pro?",
actual_output="The SoundPods Pro play for 8 hours on one charge.",
expected_output=("The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours "
"from the charging case, giving a total of 32 hours."),
)- written in Custom judge model
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
from deepeval.metrics import GEval
from battery_case import params, steps, test_case
from judge import judge
correctness = GEval(name="Correctness", evaluation_steps=steps, evaluation_params=params, model=judge)
correctness.measure(test_case)
for threshold in [0.3, 0.5, 0.8, None]:
correctness.threshold = threshold
print(f"threshold={threshold} score={correctness.score} passed={correctness.is_successful()}")threshold=0.3 score=0.5 passed=True threshold=0.5 score=0.5 passed=True threshold=0.8 score=0.5 passed=False threshold=None score=0.5 passed=None
What the thresholds decided
- The score is 0.5 for the answer that leaves out the case, from one judge call.
- 0.3 and 0.5 pass, 0.8 fails. A score equal to the threshold passes. Changing
thresholdand callingis_successful()again made no new judge call. - None gives None, the same score-only result as with exact match.
Scoring the same answer with strict_mode
from deepeval.metrics import GEval
from battery_case import params, steps, test_case
from judge import judge
strict = GEval(name="Correctness", evaluation_steps=steps, evaluation_params=params, model=judge, strict_mode=True)
print("threshold before measure:", strict.threshold)
strict.measure(test_case)
print(f"score={strict.score} passed={strict.is_successful()}")
print(strict.reason)threshold before measure: 1 score=0 passed=False Actual output only mentions 8‑hour battery life, omitting the additional 24‑hour case runtime stated in expected output.
What strict mode changed
- The threshold is 1 before any measure.
strict_mode=Truesets it when the metric is built. - The score is 0, printed as a whole number, and the answer fails. Without strict mode the same answer scored 0.5: strict mode asks the judge for full marks or nothing.
- The reason names the missing 24 hours from the case, the fact that cost the answer its point.
A threshold vs score-only vs strict mode
threshold=0.7 | threshold=None | strict_mode=True | |
|---|---|---|---|
| Score | 0 to 1 | 0 to 1 | 1 or 0 |
is_successful() | True or False | None | True only for a perfect answer |
| Can fail a test | Yes | No | Yes |
| Good for | Most checks | Tracking a number without gating on it | Rules with no partial credit |
When to set each one
- Raise the threshold when a partly correct answer should fail, as with a battery answer that leaves out the case.
- Use
threshold=Nonefor a metric you want to watch for a while before it can block anything. The docs note that everyevaluate()orassert_test()call still needs at least one non-flaky metric with a threshold. - Use
strict_modewhen a half-right answer is as bad as a wrong one, such as a price or a refund window.
strict_mode=True overrides the threshold you pass. GEval(..., threshold=0.5, strict_mode=True) has threshold 1, so a test you expected to pass at 0.5 fails on any answer that is not perfect.Related
- Previous: Verbose mode
- Next: DAG metric
- Reference: Metric thresholds
- Add
0.5to the thresholds in the exact match example: it fails, because the score is 0.0. - Pass
threshold=0.5, strict_mode=Trueto the strict metric and check the printed threshold is still 1. - Change the battery answer to include the 24 hours from the case and the 32-hour total, and run the strict example again.
Every expert started right here.