Evals in CI
Evals in CI is the practice of running a small golden set as an automated test on every change, so that a drop in quality fails the build before it reaches users.
Last updated: 09 Oct, 2026 · RAGAS 0.4
Reading evaluation results went through one run by eye. A team cannot do that after every commit. CI, continuous integration, is the service that runs a project's tests whenever the code changes; putting evaluations there means the golden set is checked every time, by a script, with a clear pass or fail.
Not every metric fits every commit. The five judge metrics cost many model calls per golden and their scores move a little from run to run. So there are two loops: a fast one with checks that need no judge on every commit, and the judged metrics on a schedule and before a release.
Checking answers without a judge
RAGAS ships metrics that are plain string operations. ExactMatch returns 1.0 when the answer equals the reference character for character. StringPresence returns 1.0 when a given string occurs in the answer. Neither takes a judge or embeddings, so they need no model call and give the same result every time.
The example applies both to two answers the video's app gave, saved as data. A CI job would load them from the results file of the latest run. Each golden gets a must_include list: the few strings a correct answer cannot do without.
from ragas.metrics.collections import ExactMatch, StringPresence
# two answers of the video's app, as shown on its Run Evaluation screen, with their goldens
saved = [
{"id": "g001",
"response": ("TechNest's return policy allows returns within 30 days of the original purchase date, as long as "
"items are in their original packaging with all accessories included. You're responsible for return "
"shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 "
"business days of receiving the returned item. Note that digital downloads and opened software are "
"non-refundable."),
"reference": ("TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all "
"accessories. Customers pay return shipping unless the item is defective. Refunds are processed in "
"5 to 7 business days."),
"must_include": ["30 days", "5 to 7 business days"]},
{"id": "g002",
"response": "The TechNest ProBook X1 features 16GB DDR5 RAM and a 512GB NVMe SSD.",
"reference": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD.",
"must_include": ["16GB DDR5", "512GB NVMe SSD"]},
]
exact, present = ExactMatch(), StringPresence()
for row in saved:
same = exact.score(reference=row["reference"], response=row["response"]).value
facts = [present.score(reference=fact, response=row["response"]).value for fact in row["must_include"]]
print(row["id"], "| exact match:", same, "| facts present:", facts)
wrong = "The ProBook X1 has 8GB RAM and a 512GB NVMe SSD."
print("wrong answer | facts present:",
[present.score(reference=fact, response=wrong).value for fact in saved[1]["must_include"]])g001 | exact match: 0.0 | facts present: [1.0, 1.0] g002 | exact match: 0.0 | facts present: [1.0, 1.0] wrong answer | facts present: [0.0, 1.0]
What the string checks catch
- Exact match is 0.0 for both answers, although both are right. "The TechNest ProBook X1 features" and "The ProBook X1 has" are different strings. Exact match suits answers with one fixed form, such as an id, a label or a number, and is too strict for sentences.
- The key facts are present in both answers: every entry is 1.0. This is the check to gate on.
- The wrong answer prints [0.0, 1.0]. It has the right SSD and the wrong RAM, and the missing "16GB DDR5" shows up as a 0.0.
- String presence is literal. "16 GB" with a space would also give 0.0, so pick strings that have one natural spelling, and keep the judge metrics for meaning.
Testing retrieval with ID-based context recall
Context recall needed a judge because it compared the reference's claims with chunk text. When each golden lists the ids of the catalog entries its answer needs, the same idea becomes a set operation: the share of the needed ids that the retriever returned. RAGAS documents this as ID-based context recall; it is short enough to write out.
The data: a small catalog and the goldens
The catalog holds six entries of the video's TechNest catalog, each as its title followed by a shortened version of its text. The goldens are the video's five questions. One of them, g004, needs two entries.
The retriever under test
def words(text):
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if len(w) > 3}
def retrieve(question, top_k):
"""The retriever under test: rank the entries by how many question words they share."""
overlap = {doc_id: len(words(question) & words(text)) for doc_id, text in CATALOG.items()}
return sorted(overlap, key=overlap.get, reverse=True)[:top_k]A tiny keyword retriever stands in for the app's embedding search so that the test runs with no model and no key. It ranks the entries by the number of longer words they share with the question. In a real project this function is your own retriever.
The metric and the test
def id_context_recall(retrieved_ids, reference_ids):
return len(set(retrieved_ids) & set(reference_ids)) / len(reference_ids)def test_retrieval(top_k, pass_mark=1.0):
recalls = {}
for golden in GOLDENS:
retrieved = retrieve(golden["user_input"], top_k)
recalls[golden["id"]] = id_context_recall(retrieved, golden["reference_context_ids"])
print(golden["id"], retrieved, "context recall", recalls[golden["id"]])
print("mean context recall:", sum(recalls.values()) / len(recalls))
for golden_id, recall in recalls.items():
assert recall >= pass_mark, f"{golden_id}: context recall {recall} is below {pass_mark}"
print("PASSED")The test prints the recall of every golden and their mean, then asserts on each golden separately. The file calls the function itself, so it runs with python test_evals.py. A failed assert raises an AssertionError that nothing catches, Python then exits with a non-zero exit code, and that is what makes a CI job fail. A test runner such as pytest finds tests by name instead, functions prefixed test in files named test_*.py, and treats a test function's arguments as fixtures, so under pytest each setting would be its own test function with no arguments. Save the example as test_evals.py:
import re
CATALOG = { # six entries of the TechNest catalog: the title, then a shortened text
"prod_001": "ProBook X1 Laptop. The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen "
"processor, 16GB DDR5 RAM, and a 512GB NVMe SSD.",
"prod_002": "PixelPhone 15. The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera "
"system, 8GB RAM, 256GB storage. Price: $899.",
"prod_003": "SoundPods Pro. The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation "
"(ANC), 8 hours of playback per charge plus 24 hours with the case.",
"policy_001": "Return Policy. TechNest accepts returns within 30 days of the original purchase date. Refunds are "
"processed within 5 to 7 business days of receiving the returned item.",
"policy_002": "Shipping Policy. TechNest offers free standard shipping on all orders over $50 within the "
"continental US. Standard shipping takes 3 to 5 business days.",
"policy_003": "Warranty Policy. All TechNest products include a minimum 1-year manufacturer warranty covering "
"defects in materials and workmanship.",
}
GOLDENS = [ # the video's five questions, each with the ids of the entries its answer needs
{"id": "g001", "user_input": "What is TechNest's return policy?", "reference_context_ids": ["policy_001"]},
{"id": "g002", "user_input": "What are the RAM and storage specs of the ProBook X1?",
"reference_context_ids": ["prod_001"]},
{"id": "g003", "user_input": "How long is the battery life on the SoundPods Pro?",
"reference_context_ids": ["prod_003"]},
{"id": "g004", "user_input": "What are TechNest's shipping options and how long do returns take to process?",
"reference_context_ids": ["policy_002", "policy_001"]},
{"id": "g005", "user_input": "What is the price of the PixelPhone 15?", "reference_context_ids": ["prod_002"]},
]
def words(text):
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if len(w) > 3}
def retrieve(question, top_k):
"""The retriever under test: rank the entries by how many question words they share."""
overlap = {doc_id: len(words(question) & words(text)) for doc_id, text in CATALOG.items()}
return sorted(overlap, key=overlap.get, reverse=True)[:top_k]
def id_context_recall(retrieved_ids, reference_ids):
return len(set(retrieved_ids) & set(reference_ids)) / len(reference_ids)
def test_retrieval(top_k, pass_mark=1.0):
recalls = {}
for golden in GOLDENS:
retrieved = retrieve(golden["user_input"], top_k)
recalls[golden["id"]] = id_context_recall(retrieved, golden["reference_context_ids"])
print(golden["id"], retrieved, "context recall", recalls[golden["id"]])
print("mean context recall:", sum(recalls.values()) / len(recalls))
for golden_id, recall in recalls.items():
assert recall >= pass_mark, f"{golden_id}: context recall {recall} is below {pass_mark}"
print("PASSED")
test_retrieval(top_k=2)
print()
test_retrieval(top_k=1) # someone lowers top_k to save tokensg001 ['policy_001', 'policy_002'] context recall 1.0
g002 ['prod_001', 'prod_002'] context recall 1.0
g003 ['prod_003', 'prod_001'] context recall 1.0
g004 ['policy_001', 'policy_002'] context recall 1.0
g005 ['prod_002', 'prod_001'] context recall 1.0
mean context recall: 1.0
PASSED
g001 ['policy_001'] context recall 1.0
g002 ['prod_001'] context recall 1.0
g003 ['prod_003'] context recall 1.0
g004 ['policy_001'] context recall 0.5
g005 ['prod_002'] context recall 1.0
mean context recall: 0.9
Traceback (most recent call last):
File "main.py", line 58, in <module>
test_retrieval(top_k=1) # someone lowers top_k to save tokens
File "main.py", line 52, in test_retrieval
assert recall >= pass_mark, f"{golden_id}: context recall {recall} is below {pass_mark}"
AssertionError: g004: context recall 0.5 is below 1.0What the failing run shows
- With top_k = 2 every golden reaches 1.0 and the test prints PASSED. For g004 both needed entries,
policy_001andpolicy_002, come back. - With top_k = 1 the test stops with an AssertionError that names the golden: g004 got one of its two entries, a context recall of 0.5.
- The mean is still 0.9. A gate on the average with a pass mark of 0.9 would have let this change through. Asserting per golden is what caught it.
- No model was called. The file needs no key and no network, so it can run on every commit.
Comparing a new run with the baseline
The judged metrics belong in the second loop. There the question is not "is every score above a line" but "what moved since the last good run". The example compares the two runs the video shows on its Results screens, treating the first as the baseline and the second as the new run.
import pandas as pd
metrics = ["faithfulness", "answer_relevancy", "context_precision", "context_recall", "answer_correctness"]
goldens = ["g001", "g002", "g003", "g004", "g005"]
# the two runs shown in the video, typed in from its Results screens
baseline = pd.DataFrame([[0.86, 1.00, 1.00, 1.00, 0.64], [1.00, 0.80, 1.00, 1.00, 0.97], [0.67, 0.85, 1.00, 1.00, 0.84],
[1.00, 0.73, 1.00, 0.55, 0.56], [1.00, 0.93, 1.00, 1.00, 1.00]], index=goldens, columns=metrics)
new_run = baseline.copy()
new_run.loc["g003", "faithfulness"] = 1.00
TOLERANCE = 0.05
change = new_run - baseline
moved = [(golden, metric, change.loc[golden, metric]) for golden in goldens for metric in metrics
if abs(change.loc[golden, metric]) > TOLERANCE]
print("cells that moved by more than", TOLERANCE)
for golden, metric, delta in moved:
print(f" {golden} {metric}: {baseline.loc[golden, metric]:.2f} -> {new_run.loc[golden, metric]:.2f} ({delta:+.2f})")
print("drops:", sum(delta < 0 for _, _, delta in moved), " rises:", sum(delta > 0 for _, _, delta in moved))
print("the same cells with the runs swapped:", [f"{-delta:+.2f}" for _, _, delta in moved])cells that moved by more than 0.05 g003 faithfulness: 0.67 -> 1.00 (+0.33) drops: 0 rises: 1 the same cells with the runs swapped: ['-0.33']
- One cell moved by more than the tolerance: the faithfulness of g003, from 0.67 to 1.00, a rise of 0.33.
- Nothing dropped, so this comparison would report no regression.
- Swapped, the same pair is a drop of 0.33. Had the second run been the baseline, a nightly job would have raised an alarm, with the same app and the same goldens.
- That swing is why judged scores make a poor hard gate on every commit. Report them against a baseline with a tolerance, repeat a run before acting on one cell, and keep the hard pass or fail for the checks that give the same result every time.
No-judge checks vs judge metrics
| No-judge checks | Judge metrics | |
|---|---|---|
| Examples | String presence, exact match, ID-based context recall | Faithfulness, answer relevancy, context precision, context recall, answer correctness |
| Model calls | None | Several per golden |
| Same result every run | Yes | No, verdicts can change |
| What they can see | Listed strings and ids | Meaning: paraphrases, unsupported claims, missing facts |
| Where they run | Every commit, as a hard pass or fail | Nightly and before a release, compared with a baseline |
The video says evaluations are not something to run on every commit of a development branch, because the judged run takes minutes and many calls. That holds for the judge metrics. The fast checks have no such cost, and the documentation in the video's repo asks for a run after every pipeline change and for all its metrics before a production release. The two loops give both.
Where you use evals in CI
- Pull requests that touch a prompt, the retriever or the chunking. The fast loop answers within the normal test run.
- Model upgrades. A provider retires a model and you switch ids: the nightly judged run against the old baseline shows what changed.
- Bug reports. Each bad answer from production becomes a golden with its
must_includestrings and its needed ids, so the same failure cannot come back unnoticed.
Related
- Previous: Reading evaluation results
- Next: Agent memory
- Reference: Traditional, non-LLM metrics in the RAGAS docs
- In
test_evals.py, change the second call totest_retrieval(top_k=1, pass_mark=0.5): both runs print PASSED, which shows how a loose pass mark waves the regression through. - In the string example, change the first fact of g002 to
"16 GB DDR5": the facts for g002 become [0.0, 1.0] although the answer is unchanged. - In the comparison, set
TOLERANCE = 0.4: no cell is listed and both counts are 0, because the one change of 0.33 is now inside the tolerance.
Every expert started right here.