RAGASragas 0.4.3 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

Evaluation results

Evaluation results are the per-golden, per-metric scores of a run, read as a table with a pass band for each score and saved as a file you can compare with the next run.

Last updated: 29 Sep, 2026 · RAGAS 0.4.3

The scored experiment printed numbers. A report turns them into the view the video's app shows in its Results tab: which golden is weak, on which metric, and how the averages look.

The Results tab: per-golden scores · from the Complete AI Security Course In 8 Hours · 160:40 to 163:46

Per-golden scores and details

The video opens the Results tab. First the per golden score table: golden one, the return-policy question written to test faithfulness, scores 0.86 on faithfulness while answer relevancy and context precision are good. One glance shows which golden lags and on which metric, which the raw RAG demo never showed. Below it are the per golden details: the scores, the RAG response, the retrieved chunks and the reference for each golden, and a button to download everything as JSON. It ends with the advice from the setup lesson: judge with a different provider from the generator.

The report pieces

python
rows = csv.DictReader(open("experiments/scored.csv"))  # the scoring run's saved table
badge(score)                    # good >= 0.75, fair >= 0.5, poor below: the app's bands
json.dump(report, open("results.json", "w"))

Pass bands from the video's app

python
def badge(score):
    return "good" if score >= 0.75 else "fair" if score >= 0.5 else "poor"

The video's reporter colours scores green from 0.75, yellow from 0.5 and red below. Words print where colours cannot.

The results file

python
report = {"per_golden": [{"id": r["id"], "scores": s} for r, s in zip(rows, scores)], "averages": averages}
json.dump(report, open("results.json", "w"), indent=2)

The scored table on disk

The scoring run in Scoring an experiment saved every row to experiments/scored.csv. This is the file it wrote: each golden's columns, the app's response and the two scores. The report reads it, so no judge call is repeated.

text
id,user_input,reference,response,faithfulness,context_recall
g001,What is TechNest's return policy?,TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days.,"TechNest accepts returns within 30 days of the original purchase date, provided the items are in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged, and refunds are processed within 5 to 7 business days of receiving the returned item.",1.0,1.0
g004,What are TechNest's shipping options and how long do returns take to process?,TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received.,"TechNest offers free standard shipping for US orders over $50, which takes 3 to 5 business days, while expedited and same-day options are available for $9.99 and $19.99 respectively. For returns, you have 30 days to send items back in their original packaging, and refunds are processed within 5 to 7 business days after TechNest receives the returned item.",1.0,1.0
g002,What are the RAM and storage specs of the ProBook X1?,The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD.,The TechNest ProBook X1 is equipped with 16GB of DDR5 RAM and a 512GB NVMe SSD.,1.0,1.0
g003,How long is the battery life on the SoundPods Pro?,"The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours.","The TechNest SoundPods Pro offer 8 hours of playback on a single charge. When used with the charging case, the total battery life extends to 24 hours.",1.0,1.0
g005,What is the price of the PixelPhone 15?,The TechNest PixelPhone 15 is priced at $899.,The TechNest PixelPhone 15 is priced at $899. It comes with a 1-year warranty and is available in Midnight Black and Arctic White.,1.0,1.0

Reporting the scored experiment

Save this as report.py in the same folder as scoring.py and run it after the scoring run. Its output reports the saved run above.

Example
import csv
import json


def badge(score):
    return "good" if score >= 0.75 else "fair" if score >= 0.5 else "poor"


metrics = ["faithfulness", "context_recall"]
rows = sorted(csv.DictReader(open("experiments/scored.csv", encoding="utf-8")), key=lambda r: r["id"])
scores = [{m: float(row[m]) for m in metrics} for row in rows]

print(f"{'golden':<8}" + "".join(f"{m:<22}" for m in metrics))
for row, s in zip(rows, scores):
    print(f"{row['id']:<8}" + "".join(f"{s[m]:.2f} {badge(s[m]):<17}" for m in metrics))
averages = {m: round(sum(s[m] for s in scores) / len(scores), 3) for m in metrics}
print(f"{'AVERAGE':<8}" + "".join(f"{averages[m]:.2f} {badge(averages[m]):<17}" for m in metrics))

report = {"per_golden": [{"id": r["id"], "scores": s} for r, s in zip(rows, scores)], "averages": averages}
json.dump(report, open("results.json", "w"), indent=2)
print("saved results.json")

Reading the report

  • Each row is one golden with a score and a band per metric. On this saved run every cell is 1.00 and good, so the report has nothing to flag; the capstone's failure run shows the same kind of table catching a broken retriever.
  • AVERAGE is the line to compare between runs.
  • results.json holds the same numbers, like the app's download button, so a later run can be diffed against it.

Averages vs per-golden scores

AveragesPer-golden scores
AnswersIs the app better than last run?Which question fails, on which metric?
HidesOne bad golden among good onesThe overall trend
Use forA dashboard or a CI thresholdDebugging

When to save results

  • After every evaluation run, so each change has a before and after.
  • Before a release, as the record of what quality shipped.
Watch out. Two runs with the same code give slightly different scores, because the judge is a model. Before calling a change an improvement, run the baseline twice to see how much the scores move by themselves.
Try it yourself
  • Print only the goldens with any score below 0.75.
  • Load results.json in a second script and print its averages.

You understood something today that you didn't yesterday.