DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

ExactMatchMetric and PatternMatchMetric

ExactMatchMetric is a DeepEval metric that scores 1 when the actual output equals the expected output and 0 otherwise, with no LLM involved; PatternMatchMetric does the same against a regular expression.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

A test case on its own scores nothing. A metric reads some of its fields and returns a number from 0 to 1, plus a reason. These two are the simplest DeepEval has: plain Python, no key, the same result on every run.

The metric API

python
from deepeval.metrics import ExactMatchMetric

metric = ExactMatchMetric(threshold=1.0)  # the score a test case needs to pass
metric.measure(test_case)                 # reads the test case, sets the fields below
metric.score, metric.reason, metric.is_successful()

Every DeepEval metric follows this shape: build it with its settings, call measure on one test case, then read score, reason and whether the score reached the threshold.

Three answers to one golden

The expected answer is the one TechNest wrote for the price question. The first answer copies it, the second says the same thing in other words, and the third has the wrong price.

python
expected = "The TechNest PixelPhone 15 is priced at $899."
answers = [
    "The TechNest PixelPhone 15 is priced at $899.",
    "The PixelPhone 15 costs $899.",
    "The PixelPhone 15 costs $799.",
]

One test case per answer

python
for actual in answers:
    test_case = LLMTestCase(input=question, actual_output=actual, expected_output=expected)
    metric.measure(test_case)

Scoring three answers with ExactMatchMetric

Example
from deepeval.metrics import ExactMatchMetric
from deepeval.test_case import LLMTestCase

question = "What is the price of the PixelPhone 15?"
expected = "The TechNest PixelPhone 15 is priced at $899."
answers = [
    "The TechNest PixelPhone 15 is priced at $899.",
    "The PixelPhone 15 costs $899.",
    "The PixelPhone 15 costs $799.",
]

metric = ExactMatchMetric(threshold=1.0)
for actual in answers:
    test_case = LLMTestCase(input=question, actual_output=actual, expected_output=expected)
    metric.measure(test_case)
    print(f"{metric.score}  passed={metric.is_successful()!s:5}  {actual}")
    print("     ", metric.reason)

What exact match decided

  • The copied answer scores 1.0 and passes.
  • The correct answer in other words scores 0.0, with the same reason as the wrong price. Exact match cannot tell a paraphrase from a mistake.
  • Spaces at either end do not count: the metric strips both outputs before comparing, so a trailing newline from a model does not fail a test.

Checking the price with PatternMatchMetric

Often the exact wording does not matter but one part does. PatternMatchMetric takes a regular expression instead of an expected output. The first pattern below looks for $899.

Example
from deepeval.metrics import PatternMatchMetric
from deepeval.test_case import LLMTestCase

test_case = LLMTestCase(
    input="What is the price of the PixelPhone 15?",
    actual_output="The PixelPhone 15 costs $899.",
)
for pattern in [r"\$899", r"(?s).*\$899\b.*"]:
    metric = PatternMatchMetric(pattern=pattern)
    metric.measure(test_case)
    print(f"{pattern:16} score={metric.score}  {metric.reason}")
  • \$899 scores 0.0 even though the answer contains $899. The metric uses a full match: the pattern has to describe the whole output, not a piece of it.
  • (?s).*\$899\b.* scores 1.0. .* on both sides allows any text around the price, (?s) lets it span several lines, and \b stops $8999 from passing.

ExactMatchMetric vs PatternMatchMetric vs a judged metric

ExactMatchMetricPatternMatchMetricJudged metric (from G-Eval)
Readsinput, actual_output, expected_output (compares the last two)actual_outputWhichever fields its criteria name
Passes a paraphrase?NoYes, if the key part matchesYes
Needs a keyNoNoYes, for the judge
Same score every runYesYesClose, not always identical

When to use a metric with no LLM

  • When the output has one correct form: a category label, an order id, a yes or no.
  • When one detail must be present whatever the wording: a price, a support email address, a refund window.
  • As a cheap first check in CI, before the judged metrics run.
Watch out. PatternMatchMetric matches the whole output, as the first pattern showed. A pattern written for re.search scores 0 on every real sentence and looks like the app is broken. Start patterns with (?s).* and end them with .* when you only care about one part.
Try it yourself
  • Add " The TechNest PixelPhone 15 is priced at $899. " with spaces at both ends to answers and check it still scores 1.0.
  • Pass ignore_case=True to PatternMatchMetric and try the pattern (?s).*pixelphone.*.
  • Change the price in the test case to $8999 and confirm the second pattern now scores 0.0.
PreviousLLMTestCase

Every expert started right here.