DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

Custom metrics

A custom metric is a class that inherits DeepEval's BaseMetric and sets its own score, reason and pass or fail in measure, so any rule you can write in Python runs wherever a built-in metric runs.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

G-Eval and the DAG metric both ask a judge. Some rules need no judge at all: a length limit, a required word, a number in range. For those a few lines of Python are cheaper, faster and give the same score every run.

A concision metric · from the Complete Agentic AI Course In 10 Hours · 9:48:37 to 9:49:41

The video's concision rule

After the correctness judge, the video defines one more metric, called concision. The first metric had a system prompt, the user's question and an OpenAI client making the decision. This one is very simple: it checks whether the actual output is less than two times the length of the reference output.

The video writes the rule as a plain function and runs it in LangSmith; this page writes it as a DeepEval metric class.

The BaseMetric API

python
from deepeval.metrics import BaseMetric

class MyMetric(BaseMetric):
    def __init__(self, threshold=0.5): ...            # store the settings, at least threshold
    def measure(self, test_case): ...                 # set self.score and self.success, return the score
    async def a_measure(self, test_case): ...         # the same, for async runs
    def is_successful(self): ...                      # pass or fail
    @property
    def __name__(self): return "My Metric"            # the name in reports

The threshold in __init__

python
def __init__(self, threshold: float = 1.0):
    self.threshold = threshold

The concision score is 1.0 or 0.0, so a threshold of 1.0 means the answer must pass the rule.

The concision rule in measure

python
def measure(self, test_case: LLMTestCase) -> float:
    actual, expected = len(test_case.actual_output), len(test_case.expected_output)
    self.score = 1.0 if actual < 2 * expected else 0.0
    self.reason = f"{actual} characters against a limit of {2 * expected - 1}."
    self.success = self.score >= self.threshold
    return self.score

This is the video's rule: the actual output must be shorter than twice the expected output. len counts characters, and the longest passing answer is one character under twice the expected length, which is what the reason prints. The docs say measure must set self.score and self.success; the reason is optional.

A metric with only measure

metric.measure works with that method alone. Run the same metric through assert_test from First test run:

Example
from deepeval import assert_test
from deepeval.metrics import BaseMetric
from deepeval.test_case import LLMTestCase


class ConcisionMetric(BaseMetric):
    def __init__(self, threshold: float = 1.0):
        self.threshold = threshold

    def measure(self, test_case: LLMTestCase) -> float:
        self.score = 1.0 if len(test_case.actual_output) < 2 * len(test_case.expected_output) else 0.0
        self.success = self.score >= self.threshold
        return self.score


test_case = LLMTestCase(input="What is the price of the PixelPhone 15?",
                        actual_output="The PixelPhone 15 costs $899.",
                        expected_output="The TechNest PixelPhone 15 is priced at $899.")
metric = ConcisionMetric()
print("measure:", metric.measure(test_case))
assert_test(test_case, [metric])

measure returned 1.0, then assert_test failed with NotImplementedError. assert_test and evaluate() run metrics asynchronously by default, so they call a_measure, and the one inherited from BaseMetric only raises this error. The message suggests turning async off; this page writes a_measure instead.

a_measure for async runs

python
async def a_measure(self, test_case: LLMTestCase) -> float:
    return self.measure(test_case)

The rule makes no network call, so the async version can call measure, as the docs suggest for scoring with no async path.

is_successful and the name

python
def is_successful(self) -> bool:
    if self.error is not None:
        self.success = False
    else:
        try:
            self.success = self.score >= self.threshold
        except TypeError:
            self.success = False
    return self.success

This is the body the docs recommend copying: a metric that hit an error, or has no score yet, fails. The last piece is the name reports print:

python
@property
def __name__(self):
    return "Concision"

The concision.py file

python
from deepeval.metrics import BaseMetric
from deepeval.test_case import LLMTestCase


class ConcisionMetric(BaseMetric):
    """Passes when the answer is shorter than twice the expected answer."""

    def __init__(self, threshold: float = 1.0):
        self.threshold = threshold

    def measure(self, test_case: LLMTestCase) -> float:
        actual, expected = len(test_case.actual_output), len(test_case.expected_output)
        self.score = 1.0 if actual < 2 * expected else 0.0
        self.reason = f"{actual} characters against a limit of {2 * expected - 1}."
        self.success = self.score >= self.threshold
        return self.score

    async def a_measure(self, test_case: LLMTestCase) -> float:
        return self.measure(test_case)

    def is_successful(self) -> bool:
        if self.error is not None:
            self.success = False
        else:
            try:
                self.success = self.score >= self.threshold
            except TypeError:
                self.success = False
        return self.success

    @property
    def __name__(self):
        return "Concision"

Scoring two PixelPhone answers for concision

The expected answer is 45 characters long. One answer is shorter; the other answers the price and then lists the phone's specs.

Example
from deepeval.test_case import LLMTestCase

from concision import ConcisionMetric

expected = "The TechNest PixelPhone 15 is priced at $899."
answers = [
    "The PixelPhone 15 costs $899.",
    ("The PixelPhone 15 costs $899. It has a 6.7-inch AMOLED screen, a 50MP triple camera, "
     "8GB RAM and 256GB storage, and comes in Midnight Black and Arctic White."),
]
metric = ConcisionMetric()
for actual in answers:
    metric.measure(LLMTestCase(input="What is the price of the PixelPhone 15?", actual_output=actual, expected_output=expected))
    print(metric.__name__, metric.score, metric.is_successful(), "|", metric.reason)

What the concision metric decided

  • The short answer is 29 characters against a limit of 89, so it scores 1.0 and passes.
  • The answer with specs is 157 characters, over the limit, and scores 0.0. Every fact in it is true; the rule only measures length, which is why the video pairs it with the correctness metric.
  • No key and no judge: the run makes no network call, and the scores are the same on every run.

BaseMetric vs GEval vs DAGMetric

Custom BaseMetricGEvalDAGMetric
Who scoresYour Python codeThe judgeThe judge picks verdicts, you set the scores
Needs a judgeOnly if your code calls oneYesYes
Same score every runYes, for plain codeClose, not alwaysClose, not always
Good forLength, format, numbers, combining other metricsMeaning and qualityRules with gates

When to write a custom metric

  • When the rule is plain code, like the video's length limit or a check that a price is in the catalog.
  • When you want several metrics' scores combined into one, for example the lowest of two judged scores; the docs call this a composite metric.
  • When an existing scorer you trust, such as ROUGE, a word-overlap score for summaries, should report through DeepEval's tests and reports.
Watch out. A custom metric without a_measure passes every quick check with metric.measure and then breaks the first assert_test or evaluate() run, as the error above showed. Write a_measure from the start, even when it only calls measure.
Try it yourself
  • Change expected to "$899." and run the two answers again: the limit drops to 9 characters and the short answer fails too.
  • Delete the __name__ property from concision.py and run again: the name printed becomes Base Metric.
  • Add a_measure to the broken example and run it: assert_test passes with no error.
PreviousDAG metric

You understood something today that you didn't yesterday.