Evals in CI
Evals in CI means running DeepEval test files as a step of your build pipeline, so a pull request that makes the bot's answers worse fails before it is merged.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
deepeval test run runs the suite on your machine. A pipeline runs it on every change without anyone remembering to. The question the video answers first is which goldens to run, because judging every golden on every commit is slow and spends the judge's daily token budget.
Unit tests and the annual exam
The video returns to the exam analogy. A school runs short unit tests often, on two or three chapters, and one long annual exam on the whole syllabus. Holding the annual exam every month would exhaust the students and the teachers. An evaluation pipeline works the same way: out of thousands of goldens, pick the essential ones, about 50 in the video's example, that must pass before every deployment, and run those on every change. Run the goldens for a feature when that feature changes, and run the full set once a week or every two weeks.
Its diagram puts this into the CI/CD pipeline: a code or config change triggers the pipeline, a small daily change runs the unit evaluation, and a major release runs the comprehensive one. The video's app scores with RAGAS; the pipeline below runs DeepEval, with pytest marks to pick the subset.
The pytest mark API
pytest.param(golden, id="g001", marks=[pytest.mark.essential]) # tag one test case
deepeval test run test_ci.py -m essential # run only the tagged ones
deepeval test run test_ci.py # run every goldenTagging the essential goldens
Two of the five goldens become essential: the return policy, g001, and the PixelPhone price, g005. Each golden becomes one test case, named after the golden. The file reads goldens.json as plain dictionaries, because pytest.param needs each golden's name before any test runs.
ESSENTIAL = {"g001", "g005"} # the goldens every deployment must pass
GOLDENS = json.load(open("goldens.json", encoding="utf-8"))
CASES = [pytest.param(g, id=g["name"], marks=[pytest.mark.essential] if g["name"] in ESSENTIAL else [])
for g in GOLDENS]One test per golden
@pytest.mark.parametrize("golden", CASES)
def test_answer_is_relevant(golden):
response, contexts = answer(golden["input"])
test_case = LLMTestCase(input=golden["input"], actual_output=response, retrieval_context=contexts)
assert_test(test_case, [AnswerRelevancyMetric(model=judge, threshold=0.7)])- written in Custom judge model
- written in TechNest RAG app
- written in TechNest RAG app
- written in Datasets and goldens
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
[
{
"name": "g001",
"input": "What is TechNest's return policy?",
"expected_output": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
},
{
"name": "g002",
"input": "What are the RAM and storage specs of the ProBook X1?",
"expected_output": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
},
{
"name": "g003",
"input": "How long is the battery life on the SoundPods Pro?",
"expected_output": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
},
{
"name": "g004",
"input": "What are TechNest's shipping options and how long do returns take to process?",
"expected_output": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
},
{
"name": "g005",
"input": "What is the price of the PixelPhone 15?",
"expected_output": "The TechNest PixelPhone 15 is priced at $899."
}
]
The test_ci.py and pytest.ini files
Save the test file as test_ci.py, next to judge.py, technest.py, catalog.json and goldens.json.
import json
import pytest
from deepeval import assert_test
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase
from judge import judge
from technest import answer
ESSENTIAL = {"g001", "g005"} # the goldens every deployment must pass
GOLDENS = json.load(open("goldens.json", encoding="utf-8"))
CASES = [pytest.param(g, id=g["name"], marks=[pytest.mark.essential] if g["name"] in ESSENTIAL else [])
for g in GOLDENS]
@pytest.mark.parametrize("golden", CASES)
def test_answer_is_relevant(golden):
response, contexts = answer(golden["input"])
test_case = LLMTestCase(input=golden["input"], actual_output=response, retrieval_context=contexts)
assert_test(test_case, [AnswerRelevancyMetric(model=judge, threshold=0.7)])Replace pytest.ini with this version. It keeps the short tracebacks from First test run and registers the essential mark, so pytest does not warn about an unknown mark.
[pytest]
addopts = --tb=no
markers =
essential: goldens every deployment must passRunning only the essential goldens
-d failing prints only failing test cases in DeepEval's table, which keeps a green CI log short.
deepeval test run test_ci.py -m essential -d failingEvaluating 1 test case(s) in parallel 0% 0:00:02 FRunning teardown with pytest sessionfinish... =========================== slowest 10 durations =========================== 4.13s call test_ci.py::test_answer_is_relevant[g001] 3.06s call test_ci.py::test_answer_is_relevant[g005] (4 durations < 0.005s hidden. Use -vv to show these durations.) ========================= short test summary info ========================== FAILED test_ci.py::test_answer_is_relevant[g005] - AssertionError: Metrics: Answer Relevancy (score: 0.3333333333333333, t... 1 failed, 1 passed, 3 deselected, 4 warnings in 7.25s Test Results ┏━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━━━━━━━━┓ ┃ ┃ ┃ ┃ ┃ Overall ┃ ┃ Test case ┃ Metric ┃ Score ┃ Status ┃ Success Rate ┃ ┡━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━━━━━━━━┩ │ test_answer_i… │ │ │ │ 0.0% | │ │ │ │ │ │ passed=0 | │ │ │ │ │ │ failed=1 │ │ │ Answer │ 0.33 │ FAILED │ │ │ │ Relevancy │ (threshold=0.… │ │ │ │ │ │ evaluation │ │ │ │ │ │ model=openai/… │ │ │ │ │ │ reason=The │ │ │ │ │ │ score is 0.33 │ │ │ │ │ │ because the │ │ │ │ │ │ answer │ │ │ │ │ │ included │ │ │ │ │ │ unrelated │ │ │ │ │ │ statements │ │ │ │ │ │ about warranty │ │ │ │ │ │ and available │ │ │ │ │ │ colors, which │ │ │ │ │ │ do not address │ │ │ │ │ │ the price │ │ │ │ │ │ query, │ │ │ │ │ │ preventing a │ │ │ │ │ │ higher │ │ │ │ │ │ relevance │ │ │ │ │ │ rating., │ │ │ │ │ │ error=None) │ │ │ │ Note: Use │ │ │ │ │ │ Confident AI │ │ │ │ │ │ with DeepEval │ │ │ │ │ │ to analyze │ │ │ │ │ │ failed test │ │ │ │ │ │ cases for more │ │ │ │ │ │ details │ │ │ │ │ └────────────────┴───────────────┴────────────────┴────────┴───────────────┘ ⚠ WARNING: No hyperparameters logged. » Log hyperparameters to attribute prompts and models to your test runs. ============================================================================ ==== ✓ Evaluation completed 🎉! (time taken: 7.41s | token cost: None) » Test Results (2 total tests): » Pass Rate: 50.0% | Passed: 1 | Failed: 1 =========================================================================== ===== » Want to share evals with your team, or a place for your test cases to live? ❤️ 🏡 » Run 'deepeval view' to analyze and save testing results on Confident AI.
What the essential run did
- Two tests ran and three were deselected:
-m essentialkept g001 and g005 and skipped the rest. - g001 passed and g005 failed. The bot's price answer scored 0.33 against a threshold of 0.7, and the judge's reason says why: the answer added the warranty and the available colours, which do not answer "What is the price?".
- A pull request that made answers chattier would be stopped here. The price in the answer was right, but two thirds of it was padding, and a pull request that made answers chattier would be stopped here. The fix belongs in the bot's prompt, or in the golden's threshold if padding is acceptable, not in the CI step.
- The output is one run's. A judged run changes a little each time, so a different run can score g005 slightly differently.
The GitHub Actions workflow
Save this as .github/workflows/evals.yml. Add GROQ_API_KEY to the repository's secrets under Settings, Secrets and variables, Actions.
name: evals
on:
pull_request: # the unit test: essential goldens on every change
schedule:
- cron: "0 2 * * 1" # the annual exam: every golden, once a week
jobs:
deepeval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install "deepeval==4.2.8"
- name: Run the evals
env:
GROQ_API_KEY: ${{ secrets.GROQ_API_KEY }}
DEEPEVAL_RESULTS_FOLDER: eval-results
DEEPEVAL_TELEMETRY_OPT_OUT: "1"
run: |
if [ "${{ github.event_name }}" = "schedule" ]; then
deepeval test run test_ci.py
else
deepeval test run test_ci.py -m essential
fi
- uses: actions/upload-artifact@v4
if: always()
with:
name: eval-results
path: eval-results/- On a pull request it runs
-m essential: the unit test. - On the weekly schedule it runs every golden: the annual exam.
DEEPEVAL_RESULTS_FOLDERmakes DeepEval write each run as atest_run_<date>_<time>.jsonfile, with every test case's scores and reasons, and the last step keeps that folder as a build artifact you can download.- A failing metric fails the step, because
deepeval test runexits with pytest's non-zero code, and the pull request shows a red check.
Essential goldens vs the full set
| Essential goldens | Every golden | |
|---|---|---|
| In the exam analogy | A unit test | The annual exam |
| Runs | On every pull request | On a schedule, or before a major release |
| Size | A few to about 50 | The whole dataset |
| Command | deepeval test run test_ci.py -m essential | deepeval test run test_ci.py |
When to gate a pipeline on evals
- When a prompt, model or retrieval change can break answers that customers already rely on.
- When several people change the bot, so nobody has to remember which goldens to check by hand.
LLMTestCase(..., flaky=True), which turns its failure into a warning, or fix the golden, rather than raising every threshold.Related
- Previous: deepeval test run
- Next: Integrations
- Reference: Unit testing in CI/CD
- Add g003 to
ESSENTIALand run the essential command again: three tests run instead of two. - Run
deepeval test run test_ci.py -m "not essential"to run only the other three goldens. - Run
DEEPEVAL_RESULTS_FOLDER=eval-results deepeval test run test_ci.py -m essentialand open the JSON file it writes.
Little by little, you're building something great.