RAGASragas 0.4.3 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

Goldens

Goldens are a collection of carefully curated test cases, each holding a question and the expected answer, against which an AI application's outputs are evaluated.

Last updated: 29 Sep, 2026 · RAGAS 0.4.3

An evaluation needs a truth to compare against. The goldens are that truth: the questions your app must handle, with the answer a domain expert would give.

What a golden is · from the Production RAG Live Marathon · 208:52 to 211:59

A test case with an expected answer

The video reads the definition part by part. A golden is a collection of carefully curated test cases. Curated means someone chose the questions to cover the ways users ask, and wrote the expected answer, the ground truth, the way the app should respond. An airline's goldens would ask about cancelled flights and refunds; a LangChain docs bot's goldens would ask what LCEL and LangGraph are. Writing them needs domain expertise, because only someone who knows the domain knows the right answer.

A RAG golden always holds the question and the expected answer. It can hold more: the video's app adds a metric_focus field naming the metric the question was written to test.

The five TechNest goldens

Save the video's goldens as goldens.json, next to catalog.json.

json
[
  {
    "id": "g001",
    "metric_focus": "faithfulness",
    "user_input": "What is TechNest's return policy?",
    "reference": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
  },
  {
    "id": "g002",
    "metric_focus": "answer_relevancy",
    "user_input": "What are the RAM and storage specs of the ProBook X1?",
    "reference": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
  },
  {
    "id": "g003",
    "metric_focus": "context_precision",
    "user_input": "How long is the battery life on the SoundPods Pro?",
    "reference": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
  },
  {
    "id": "g004",
    "metric_focus": "context_recall",
    "user_input": "What are TechNest's shipping options and how long do returns take to process?",
    "reference": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
  },
  {
    "id": "g005",
    "metric_focus": "answer_correctness",
    "user_input": "What is the price of the PixelPhone 15?",
    "reference": "The TechNest PixelPhone 15 is priced at $899."
  }
]

Loading and listing the goldens

Example
import json

goldens = json.load(open("goldens.json", encoding="utf-8"))
print(len(goldens), "goldens")
for golden in goldens:
    print(golden["id"], golden["metric_focus"].ljust(18), golden["user_input"])

What each golden is built to test

  • g001, the return policy, has a four-part answer, so an answer that adds or changes a detail shows up in faithfulness.
  • g002 asks for two specs of one laptop, a good test of whether the answer sticks to the question.
  • g003 needs one product entry among eight similar ones, so the rank of that chunk matters.
  • g004 asks two things at once, shipping and returns, so the search must bring back two different policies.
  • g005 has one exact fact, a price, so a wrong number is easy to catch.

Hand-written vs generated goldens

Written by expertsGenerated from documents
QualityHighest: the expert knows the right answerGood, but needs review
CostExpert time is expensive and scarceA few LLM calls
ScaleTens of questionsHundreds of questions
ToolsA JSON or CSV fileRAGAS test set generation, DeepEval's synthesizer

The video recommends a mix: generate candidate questions, have an LLM review them, and have an expert fix the ones that are wrong. RAGAS has its own generator, covered in the docs under test set generation.

Where goldens come from in practice

  • Support tickets and chat logs: the questions customers ask in their own words.
  • Every bad answer found in production, added with the answer it should have been.
  • An expert session that writes the reference answers for the most important questions.
Watch out. Write the expected answer before you look at what the app says. A reference copied from the app's own output only records what the app already does, so every score comes out high and the goldens catch nothing.
Try it yourself
  • Add a sixth golden, g006, asking whether TechNest ships internationally, with its reference from the catalog's FAQ.
  • Print only the goldens whose metric_focus starts with context.

This is what real progress feels like.