First test run
assert_test is DeepEval's test assertion: it runs a list of metrics on one test case and fails the test when any score is below its threshold, so deepeval test run can run LLM checks like any pytest suite.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
Calling measure prints a score. A test suite needs more: a pass or fail per check, a report, and an exit code that can stop a deployment. DeepEval plugs into pytest for that, and the metrics from ExactMatchMetric and PatternMatchMetric are enough to see how.
The assert_test API
from deepeval import assert_test
def test_something(): # pytest runs every function named test_*
assert_test(test_case, [metric, ...]) # fails if any metric scores below its thresholdA test that should pass
The price check from the pattern match example, wrapped in a test function. The comment marks where your app's real answer goes; the later lesson deepeval test run calls the real bot there.
def test_price_is_mentioned():
test_case = LLMTestCase(
input="What is the price of the PixelPhone 15?",
actual_output="The PixelPhone 15 costs $899.", # your app's answer goes here
)
assert_test(test_case, [PatternMatchMetric(pattern=r"(?s).*\$899\b.*")])A test that should fail
The expected output is only the email address, but the answer is a full sentence, so exact match scores 0.
def test_support_email():
test_case = LLMTestCase(
input="Which email do I write to for a warranty claim?",
actual_output="Email support@technest.com with your order number.",
expected_output="support@technest.com",
)
assert_test(test_case, [ExactMatchMetric()])The test_first.py and pytest.ini files
Save both tests in test_first.py.
from deepeval import assert_test
from deepeval.metrics import ExactMatchMetric, PatternMatchMetric
from deepeval.test_case import LLMTestCase
def test_price_is_mentioned():
test_case = LLMTestCase(
input="What is the price of the PixelPhone 15?",
actual_output="The PixelPhone 15 costs $899.", # your app's answer goes here
)
assert_test(test_case, [PatternMatchMetric(pattern=r"(?s).*\$899\b.*")])
def test_support_email():
test_case = LLMTestCase(
input="Which email do I write to for a warranty claim?",
actual_output="Email support@technest.com with your order number.",
expected_output="support@technest.com",
)
assert_test(test_case, [ExactMatchMetric()])Next to it, save a pytest.ini with one setting. Without it, pytest prints DeepEval's own source code under every failed assertion; the report DeepEval prints already says why a test failed.
[pytest]
addopts = --tb=noRunning the tests with deepeval test run
deepeval test run test_first.pyEvaluating 1 test case(s) in parallel 0% 0:00:00 FRunning teardown with pytest sessionfinish... =========================== slowest 10 durations =========================== 0.01s call test_first.py::test_price_is_mentioned (5 durations < 0.005s hidden. Use -vv to show these durations.) ========================= short test summary info ========================== FAILED test_first.py::test_support_email - AssertionError: Metrics: Exact Match (score: 0.0, threshold: 1.0, stric... 1 failed, 1 passed, 4 warnings in 0.02s Test Results ┏━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━━━━━━━━┓ ┃ ┃ ┃ ┃ ┃ Overall ┃ ┃ Test case ┃ Metric ┃ Score ┃ Status ┃ Success Rate ┃ ┡━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━━━━━━━━┩ │ test_price_is… │ │ │ │ 100.0% | │ │ │ │ │ │ passed=1 | │ │ │ │ │ │ failed=0 │ │ │ Pattern Match │ 1.0 │ PASSED │ │ │ │ │ (threshold=1.… │ │ │ │ │ │ evaluation │ │ │ │ │ │ model=n/a, │ │ │ │ │ │ reason=The │ │ │ │ │ │ actual output │ │ │ │ │ │ fully matches │ │ │ │ │ │ the pattern., │ │ │ │ │ │ error=None) │ │ │ │ │ │ │ │ │ │ test_support_… │ │ │ │ 0.0% | │ │ │ │ │ │ passed=0 | │ │ │ │ │ │ failed=1 │ │ │ Exact Match │ 0.0 │ FAILED │ │ │ │ │ (threshold=1.… │ │ │ │ │ │ evaluation │ │ │ │ │ │ model=n/a, │ │ │ │ │ │ reason=The │ │ │ │ │ │ actual and │ │ │ │ │ │ expected │ │ │ │ │ │ outputs are │ │ │ │ │ │ different., │ │ │ │ │ │ error=None) │ │ │ │ Note: Use │ │ │ │ │ │ Confident AI │ │ │ │ │ │ with DeepEval │ │ │ │ │ │ to analyze │ │ │ │ │ │ failed test │ │ │ │ │ │ cases for more │ │ │ │ │ │ details │ │ │ │ │ └────────────────┴───────────────┴────────────────┴────────┴───────────────┘ ⚠ WARNING: No hyperparameters logged. » Log hyperparameters to attribute prompts and models to your test runs. ============================================================================ ==== ✓ Evaluation completed 🎉! (time taken: 0.15s | token cost: None) » Test Results (2 total tests): » Pass Rate: 50.0% | Passed: 1 | Failed: 1 =========================================================================== ===== » Want to share evals with your team, or a place for your test cases to live? ❤️ 🏡 » Run 'deepeval view' to analyze and save testing results on Confident AI.
Reading the test run report
- The pytest part comes first:
FAILED test_first.py::test_support_emailand1 failed, 1 passed. That line, and the exit code behind it, is what CI reads. - The Test Results table lists each test case with its metric, score, status and reason: Pattern Match 1.0 PASSED, Exact Match 0.0 FAILED with The actual and expected outputs are different.
evaluation model=n/ameans neither metric used a judge. - The summary at the end gives the pass rate, 50%, and the time taken. The warning about hyperparameters and the Confident AI notes are DeepEval's own reminders; Comparing models shows hyperparameters.
deepeval test run vs pytest
assert_test raises an ordinary AssertionError, so plain pytest test_first.py runs the same file and reports the same pass and fail.
pytest | deepeval test run | |
|---|---|---|
Runs test_* functions | Yes | Yes, through pytest |
| Results table with scores and reasons | No | Yes |
| DeepEval flags | No | -n processes, -i ignore errors, -c use cache, -d failing |
| Saves the run for later | No | Yes, in a local .deepeval folder |
When to write evals as tests
- When a change to the prompt or model should be blocked if it breaks an answer that used to pass.
- When the checks belong next to your other tests and run with one command in CI, which Evals in CI sets up.
test_. A function called check_price is skipped without a word, and the run reports fewer tests than you wrote.Related
- Previous: ExactMatchMetric and PatternMatchMetric
- Next: Custom judge model
- Reference: Unit testing in CI/CD
- Change
expected_outputintest_support_emailto the full answer sentence and run the file again: both tests pass. - Run
deepeval test run test_first.py -d failingand check that the table shows only the failing test case. - Rename
test_support_emailtocheck_support_emailand count the tests in the report.
Little by little, you're building something great.