Custom metrics
A custom metric is a class that inherits DeepEval's BaseMetric and sets its own score, reason and pass or fail in measure, so any rule you can write in Python runs wherever a built-in metric runs.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
G-Eval and the DAG metric both ask a judge. Some rules need no judge at all: a length limit, a required word, a number in range. For those a few lines of Python are cheaper, faster and give the same score every run.
The video's concision rule
After the correctness judge, the video defines one more metric, called concision. The first metric had a system prompt, the user's question and an OpenAI client making the decision. This one is very simple: it checks whether the actual output is less than two times the length of the reference output.
The video writes the rule as a plain function and runs it in LangSmith; this page writes it as a DeepEval metric class.
The BaseMetric API
from deepeval.metrics import BaseMetric
class MyMetric(BaseMetric):
def __init__(self, threshold=0.5): ... # store the settings, at least threshold
def measure(self, test_case): ... # set self.score and self.success, return the score
async def a_measure(self, test_case): ... # the same, for async runs
def is_successful(self): ... # pass or fail
@property
def __name__(self): return "My Metric" # the name in reportsThe threshold in __init__
def __init__(self, threshold: float = 1.0):
self.threshold = thresholdThe concision score is 1.0 or 0.0, so a threshold of 1.0 means the answer must pass the rule.
The concision rule in measure
def measure(self, test_case: LLMTestCase) -> float:
actual, expected = len(test_case.actual_output), len(test_case.expected_output)
self.score = 1.0 if actual < 2 * expected else 0.0
self.reason = f"{actual} characters against a limit of {2 * expected - 1}."
self.success = self.score >= self.threshold
return self.scoreThis is the video's rule: the actual output must be shorter than twice the expected output. len counts characters, and the longest passing answer is one character under twice the expected length, which is what the reason prints. The docs say measure must set self.score and self.success; the reason is optional.
A metric with only measure
metric.measure works with that method alone. Run the same metric through assert_test from First test run:
from deepeval import assert_test
from deepeval.metrics import BaseMetric
from deepeval.test_case import LLMTestCase
class ConcisionMetric(BaseMetric):
def __init__(self, threshold: float = 1.0):
self.threshold = threshold
def measure(self, test_case: LLMTestCase) -> float:
self.score = 1.0 if len(test_case.actual_output) < 2 * len(test_case.expected_output) else 0.0
self.success = self.score >= self.threshold
return self.score
test_case = LLMTestCase(input="What is the price of the PixelPhone 15?",
actual_output="The PixelPhone 15 costs $899.",
expected_output="The TechNest PixelPhone 15 is priced at $899.")
metric = ConcisionMetric()
print("measure:", metric.measure(test_case))
assert_test(test_case, [metric])measure: 1.0
🎯 Evaluating test case #0 0% 0:00:00
Traceback (most recent call last):
File "main.py", line 21, in <module>
assert_test(test_case, [metric])
NotImplementedError: Async execution for ConcisionMetric not supported yet. Please set 'async_mode' to 'False'.measure returned 1.0, then assert_test failed with NotImplementedError. assert_test and evaluate() run metrics asynchronously by default, so they call a_measure, and the one inherited from BaseMetric only raises this error. The message suggests turning async off; this page writes a_measure instead.
a_measure for async runs
async def a_measure(self, test_case: LLMTestCase) -> float:
return self.measure(test_case)The rule makes no network call, so the async version can call measure, as the docs suggest for scoring with no async path.
is_successful and the name
def is_successful(self) -> bool:
if self.error is not None:
self.success = False
else:
try:
self.success = self.score >= self.threshold
except TypeError:
self.success = False
return self.successThis is the body the docs recommend copying: a metric that hit an error, or has no score yet, fails. The last piece is the name reports print:
@property
def __name__(self):
return "Concision"The concision.py file
from deepeval.metrics import BaseMetric
from deepeval.test_case import LLMTestCase
class ConcisionMetric(BaseMetric):
"""Passes when the answer is shorter than twice the expected answer."""
def __init__(self, threshold: float = 1.0):
self.threshold = threshold
def measure(self, test_case: LLMTestCase) -> float:
actual, expected = len(test_case.actual_output), len(test_case.expected_output)
self.score = 1.0 if actual < 2 * expected else 0.0
self.reason = f"{actual} characters against a limit of {2 * expected - 1}."
self.success = self.score >= self.threshold
return self.score
async def a_measure(self, test_case: LLMTestCase) -> float:
return self.measure(test_case)
def is_successful(self) -> bool:
if self.error is not None:
self.success = False
else:
try:
self.success = self.score >= self.threshold
except TypeError:
self.success = False
return self.success
@property
def __name__(self):
return "Concision"Scoring two PixelPhone answers for concision
The expected answer is 45 characters long. One answer is shorter; the other answers the price and then lists the phone's specs.
from deepeval.test_case import LLMTestCase
from concision import ConcisionMetric
expected = "The TechNest PixelPhone 15 is priced at $899."
answers = [
"The PixelPhone 15 costs $899.",
("The PixelPhone 15 costs $899. It has a 6.7-inch AMOLED screen, a 50MP triple camera, "
"8GB RAM and 256GB storage, and comes in Midnight Black and Arctic White."),
]
metric = ConcisionMetric()
for actual in answers:
metric.measure(LLMTestCase(input="What is the price of the PixelPhone 15?", actual_output=actual, expected_output=expected))
print(metric.__name__, metric.score, metric.is_successful(), "|", metric.reason)Concision 1.0 True | 29 characters against a limit of 89. Concision 0.0 False | 157 characters against a limit of 89.
What the concision metric decided
- The short answer is 29 characters against a limit of 89, so it scores 1.0 and passes.
- The answer with specs is 157 characters, over the limit, and scores 0.0. Every fact in it is true; the rule only measures length, which is why the video pairs it with the correctness metric.
- No key and no judge: the run makes no network call, and the scores are the same on every run.
BaseMetric vs GEval vs DAGMetric
Custom BaseMetric | GEval | DAGMetric | |
|---|---|---|---|
| Who scores | Your Python code | The judge | The judge picks verdicts, you set the scores |
| Needs a judge | Only if your code calls one | Yes | Yes |
| Same score every run | Yes, for plain code | Close, not always | Close, not always |
| Good for | Length, format, numbers, combining other metrics | Meaning and quality | Rules with gates |
When to write a custom metric
- When the rule is plain code, like the video's length limit or a check that a price is in the catalog.
- When you want several metrics' scores combined into one, for example the lowest of two judged scores; the docs call this a composite metric.
- When an existing scorer you trust, such as ROUGE, a word-overlap score for summaries, should report through DeepEval's tests and reports.
a_measure passes every quick check with metric.measure and then breaks the first assert_test or evaluate() run, as the error above showed. Write a_measure from the start, even when it only calls measure.Related
- Previous: DAG metric
- Next: RAG metrics
- Reference: Custom metrics
- Change
expectedto"$899."and run the two answers again: the limit drops to 9 characters and the short answer fails too. - Delete the
__name__property fromconcision.pyand run again: the name printed becomesBase Metric. - Add
a_measureto the broken example and run it:assert_testpasses with no error.
You understood something today that you didn't yesterday.