DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

First test run

assert_test is DeepEval's test assertion: it runs a list of metrics on one test case and fails the test when any score is below its threshold, so deepeval test run can run LLM checks like any pytest suite.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

Calling measure prints a score. A test suite needs more: a pass or fail per check, a report, and an exit code that can stop a deployment. DeepEval plugs into pytest for that, and the metrics from ExactMatchMetric and PatternMatchMetric are enough to see how.

The assert_test API

python
from deepeval import assert_test

def test_something():                        # pytest runs every function named test_*
    assert_test(test_case, [metric, ...])    # fails if any metric scores below its threshold

A test that should pass

The price check from the pattern match example, wrapped in a test function. The comment marks where your app's real answer goes; the later lesson deepeval test run calls the real bot there.

python
def test_price_is_mentioned():
    test_case = LLMTestCase(
        input="What is the price of the PixelPhone 15?",
        actual_output="The PixelPhone 15 costs $899.",  # your app's answer goes here
    )
    assert_test(test_case, [PatternMatchMetric(pattern=r"(?s).*\$899\b.*")])

A test that should fail

The expected output is only the email address, but the answer is a full sentence, so exact match scores 0.

python
def test_support_email():
    test_case = LLMTestCase(
        input="Which email do I write to for a warranty claim?",
        actual_output="Email support@technest.com with your order number.",
        expected_output="support@technest.com",
    )
    assert_test(test_case, [ExactMatchMetric()])

The test_first.py and pytest.ini files

Save both tests in test_first.py.

python
from deepeval import assert_test
from deepeval.metrics import ExactMatchMetric, PatternMatchMetric
from deepeval.test_case import LLMTestCase


def test_price_is_mentioned():
    test_case = LLMTestCase(
        input="What is the price of the PixelPhone 15?",
        actual_output="The PixelPhone 15 costs $899.",  # your app's answer goes here
    )
    assert_test(test_case, [PatternMatchMetric(pattern=r"(?s).*\$899\b.*")])


def test_support_email():
    test_case = LLMTestCase(
        input="Which email do I write to for a warranty claim?",
        actual_output="Email support@technest.com with your order number.",
        expected_output="support@technest.com",
    )
    assert_test(test_case, [ExactMatchMetric()])

Next to it, save a pytest.ini with one setting. Without it, pytest prints DeepEval's own source code under every failed assertion; the report DeepEval prints already says why a test failed.

ini
[pytest]
addopts = --tb=no

Running the tests with deepeval test run

Example
deepeval test run test_first.py

Reading the test run report

  • The pytest part comes first: FAILED test_first.py::test_support_email and 1 failed, 1 passed. That line, and the exit code behind it, is what CI reads.
  • The Test Results table lists each test case with its metric, score, status and reason: Pattern Match 1.0 PASSED, Exact Match 0.0 FAILED with The actual and expected outputs are different. evaluation model=n/a means neither metric used a judge.
  • The summary at the end gives the pass rate, 50%, and the time taken. The warning about hyperparameters and the Confident AI notes are DeepEval's own reminders; Comparing models shows hyperparameters.

deepeval test run vs pytest

assert_test raises an ordinary AssertionError, so plain pytest test_first.py runs the same file and reports the same pass and fail.

pytestdeepeval test run
Runs test_* functionsYesYes, through pytest
Results table with scores and reasonsNoYes
DeepEval flagsNo-n processes, -i ignore errors, -c use cache, -d failing
Saves the run for laterNoYes, in a local .deepeval folder

When to write evals as tests

  • When a change to the prompt or model should be blocked if it breaks an answer that used to pass.
  • When the checks belong next to your other tests and run with one command in CI, which Evals in CI sets up.
Watch out. pytest only collects functions whose names start with test_. A function called check_price is skipped without a word, and the run reports fewer tests than you wrote.
Try it yourself
  • Change expected_output in test_support_email to the full answer sentence and run the file again: both tests pass.
  • Run deepeval test run test_first.py -d failing and check that the table shows only the failing test case.
  • Rename test_support_email to check_support_email and count the tests in the report.

Little by little, you're building something great.