Evaluation results
Evaluation results are the per-golden, per-metric scores of a run, read as a table with a pass band for each score and saved as a file you can compare with the next run.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
The scored experiment printed numbers. A report turns them into the view the video's app shows in its Results tab: which golden is weak, on which metric, and how the averages look.
Per-golden scores and details
The video opens the Results tab. First the per golden score table: golden one, the return-policy question written to test faithfulness, scores 0.86 on faithfulness while answer relevancy and context precision are good. One glance shows which golden lags and on which metric, which the raw RAG demo never showed. Below it are the per golden details: the scores, the RAG response, the retrieved chunks and the reference for each golden, and a button to download everything as JSON. It ends with the advice from the setup lesson: judge with a different provider from the generator.
The report pieces
rows = csv.DictReader(open("experiments/scored.csv")) # the scoring run's saved table
badge(score) # good >= 0.75, fair >= 0.5, poor below: the app's bands
json.dump(report, open("results.json", "w"))Pass bands from the video's app
def badge(score):
return "good" if score >= 0.75 else "fair" if score >= 0.5 else "poor"The video's reporter colours scores green from 0.75, yellow from 0.5 and red below. Words print where colours cannot.
The results file
report = {"per_golden": [{"id": r["id"], "scores": s} for r, s in zip(rows, scores)], "averages": averages}
json.dump(report, open("results.json", "w"), indent=2)The scored table on disk
The scoring run in Scoring an experiment saved every row to experiments/scored.csv. This is the file it wrote: each golden's columns, the app's response and the two scores. The report reads it, so no judge call is repeated.
id,user_input,reference,response,faithfulness,context_recall
g001,What is TechNest's return policy?,TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days.,"TechNest accepts returns within 30 days of the original purchase date, provided the items are in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged, and refunds are processed within 5 to 7 business days of receiving the returned item.",1.0,1.0
g004,What are TechNest's shipping options and how long do returns take to process?,TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received.,"TechNest offers free standard shipping for US orders over $50, which takes 3 to 5 business days, while expedited and same-day options are available for $9.99 and $19.99 respectively. For returns, you have 30 days to send items back in their original packaging, and refunds are processed within 5 to 7 business days after TechNest receives the returned item.",1.0,1.0
g002,What are the RAM and storage specs of the ProBook X1?,The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD.,The TechNest ProBook X1 is equipped with 16GB of DDR5 RAM and a 512GB NVMe SSD.,1.0,1.0
g003,How long is the battery life on the SoundPods Pro?,"The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours.","The TechNest SoundPods Pro offer 8 hours of playback on a single charge. When used with the charging case, the total battery life extends to 24 hours.",1.0,1.0
g005,What is the price of the PixelPhone 15?,The TechNest PixelPhone 15 is priced at $899.,The TechNest PixelPhone 15 is priced at $899. It comes with a 1-year warranty and is available in Midnight Black and Arctic White.,1.0,1.0Reporting the scored experiment
Save this as report.py in the same folder as scoring.py and run it after the scoring run. Its output reports the saved run above.
import csv
import json
def badge(score):
return "good" if score >= 0.75 else "fair" if score >= 0.5 else "poor"
metrics = ["faithfulness", "context_recall"]
rows = sorted(csv.DictReader(open("experiments/scored.csv", encoding="utf-8")), key=lambda r: r["id"])
scores = [{m: float(row[m]) for m in metrics} for row in rows]
print(f"{'golden':<8}" + "".join(f"{m:<22}" for m in metrics))
for row, s in zip(rows, scores):
print(f"{row['id']:<8}" + "".join(f"{s[m]:.2f} {badge(s[m]):<17}" for m in metrics))
averages = {m: round(sum(s[m] for s in scores) / len(scores), 3) for m in metrics}
print(f"{'AVERAGE':<8}" + "".join(f"{averages[m]:.2f} {badge(averages[m]):<17}" for m in metrics))
report = {"per_golden": [{"id": r["id"], "scores": s} for r, s in zip(rows, scores)], "averages": averages}
json.dump(report, open("results.json", "w"), indent=2)
print("saved results.json")golden faithfulness context_recall g001 1.00 good 1.00 good g002 1.00 good 1.00 good g003 1.00 good 1.00 good g004 1.00 good 1.00 good g005 1.00 good 1.00 good AVERAGE 1.00 good 1.00 good saved results.json
Reading the report
- Each row is one golden with a score and a band per metric. On this saved run every cell is 1.00 and good, so the report has nothing to flag; the capstone's failure run shows the same kind of table catching a broken retriever.
- AVERAGE is the line to compare between runs.
- results.json holds the same numbers, like the app's download button, so a later run can be diffed against it.
Averages vs per-golden scores
| Averages | Per-golden scores | |
|---|---|---|
| Answers | Is the app better than last run? | Which question fails, on which metric? |
| Hides | One bad golden among good ones | The overall trend |
| Use for | A dashboard or a CI threshold | Debugging |
When to save results
- After every evaluation run, so each change has a before and after.
- Before a release, as the record of what quality shipped.
Related
- Previous: Scoring an experiment
- Next: evaluate() and EvaluationDataset
- Reference: Experimentation
- Print only the goldens with any score below 0.75.
- Load
results.jsonin a second script and print its averages.
You understood something today that you didn't yesterday.