RAGASragas 0.4.3 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

Datasets

A RAGAS Dataset is a named table of rows stored by a backend, such as a CSV file, that holds the goldens an experiment runs over.

Last updated: 29 Sep, 2026 · RAGAS 0.4.3

So far the goldens came from goldens.json and lived in a Python list. A Dataset puts them where RAGAS experiments can read them, in a file your team can review and extend.

The Dataset API

python
from ragas import Dataset

dataset = Dataset(name="goldens", backend="local/csv", root_dir=".")  # datasets/goldens.csv
dataset.append({"user_input": ..., "reference": ...})                  # one row
dataset.save()                                                         # writes the file
Dataset.load(name="goldens", backend="local/csv", root_dir=".")        # reads it back

One row per golden

Each row keeps the golden's id, question and expected answer. The app's answers are not in it: they come from running the app, which is the next lesson.

python
for golden in goldens:
    dataset.append({"id": golden["id"], "user_input": golden["user_input"], "reference": golden["reference"]})
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
goldens.json
[
  {
    "id": "g001",
    "metric_focus": "faithfulness",
    "user_input": "What is TechNest's return policy?",
    "reference": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
  },
  {
    "id": "g002",
    "metric_focus": "answer_relevancy",
    "user_input": "What are the RAM and storage specs of the ProBook X1?",
    "reference": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
  },
  {
    "id": "g003",
    "metric_focus": "context_precision",
    "user_input": "How long is the battery life on the SoundPods Pro?",
    "reference": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
  },
  {
    "id": "g004",
    "metric_focus": "context_recall",
    "user_input": "What are TechNest's shipping options and how long do returns take to process?",
    "reference": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
  },
  {
    "id": "g005",
    "metric_focus": "answer_correctness",
    "user_input": "What is the price of the PixelPhone 15?",
    "reference": "The TechNest PixelPhone 15 is priced at $899."
  }
]

Saving the goldens as a CSV dataset

Run this in the folder with goldens.json. It writes datasets/goldens.csv there.

Example
import json

from ragas import Dataset

goldens = json.load(open("goldens.json", encoding="utf-8"))
dataset = Dataset(name="goldens", backend="local/csv", root_dir=".")
for golden in goldens:
    dataset.append({"id": golden["id"], "user_input": golden["user_input"], "reference": golden["reference"]})
dataset.save()

loaded = Dataset.load(name="goldens", backend="local/csv", root_dir=".")
print(len(loaded), "rows loaded")
print(loaded[4]["id"], "->", loaded[4]["reference"])
print()
print(open("datasets/goldens.csv", encoding="utf-8").read())

What was written

  • datasets/goldens.csv has one header line and one row per golden, with the three columns you appended.
  • Dataset.load reads the same rows back, so every later run starts from the same questions.
  • The folder name comes from the backend: local/csv keeps datasets under datasets/ inside root_dir.

A JSON list vs a RAGAS Dataset

goldens.json in a listRAGAS Dataset
Read byYour own code@experiment runs
Stored asWhatever you writeA backend: local/csv, local/jsonl and others
Edited byDevelopersAnyone with a spreadsheet, for CSV

When to keep goldens in a Dataset

  • When goldens grow past a handful and domain experts add rows in a spreadsheet.
  • When you want every experiment, this week and next month, to read the same file.
Watch out. dataset.save() rewrites the whole file from the rows in memory. Load, append and save; building a fresh Dataset with the same name and saving it replaces the rows a colleague added.
Try it yourself
  • Open datasets/goldens.csv in a spreadsheet, add a sixth row by hand, and load it again.
  • Save the same rows with backend="local/jsonl" and look at the file it writes.

This is what real progress feels like.