Goldens
A golden is a test record that holds a query and the answer an expert expects for it, and a golden dataset is the set of such records an application is evaluated against.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
A custom evaluation starts with a question nobody else can answer for you: what should your application say? Benchmarks vs custom evaluation listed the golden dataset as the first thing to build. The video puts it in three words: "It is the truth."
What a golden holds
This part of the video starts at 1:32:37. On the whiteboard a golden is a small record with up to three parts:
- Queries. "Different types of queries which can be asked by the end users", chosen to cover every side of the chatbot.
- Expected answer. The truth for that query, the answer you would call perfect. Somebody has to write it, and the right somebody is an expert in the domain. Experts' time is scarce, so it has to be used with care.
- Expected context. The chunks or sources the pipeline should fetch to build that answer. This part is optional: "this is up to you as a developer".
The clip's own example comes from the document chat: the query is "what is attention", the app already gives an actual output, and the expected answer still has to be written by someone who knows the paper.
Frameworks name these parts differently. The whiteboard uses plain words, the TechNest app uses the RAGAS field names, and "golden" is the word DeepEval uses for such a record. The rest of this part uses the RAGAS names.
| On the whiteboard | RAGAS field | DeepEval field |
|---|---|---|
| Query | user_input | input |
| Expected answer | reference | expected_output |
| Expected context | reference_contexts | context |
| Actual answer | response | actual_output |
| Retrieved contexts | retrieved_contexts | retrieval_context |
The five goldens of the TechNest app
The video's repository keeps its goldens in goldens.json. Each one has an ID, a query and a reference:
{
"id": "g001",
"metric_focus": "faithfulness",
"user_input": "What is TechNest's return policy?",
"reference": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
}There are five, "for the sake of simplicity":
| ID | metric_focus | user_input |
|---|---|---|
| g001 | faithfulness | What is TechNest's return policy? |
| g002 | answer_relevancy | What are the RAM and storage specs of the ProBook X1? |
| g003 | context_precision | How long is the battery life on the SoundPods Pro? |
| g004 | context_recall | What are TechNest's shipping options and how long do returns take to process? |
| g005 | answer_correctness | What is the price of the PixelPhone 15? |
metric_focusis a label. It names the metric a golden was written to stress. Every golden is still scored on every metric; nothing in the scoring code reads this field.- There is no expected context. The file has no
reference_contexts. RAGAS' context recall then uses thereferenceas a stand-in for it. - g004 asks two things at once. Shipping options and return processing sit in two different catalog entries, so this golden tests whether the retriever brings back both.
A golden set is a test set
In machine learning terms, training data fits a model, validation data tunes it, and a test set is held out and only used to measure. A golden set plays the part of the test set. Nothing is fitted on it. It is run again after every change to the application, unchanged, which makes it a regression suite: a score that drops points at the change that caused it.
Running the app on the goldens
This part of the video starts at 1:44:31. Goldens alone cannot be scored. They say what should come; an evaluation also needs what did come.
The video uses an exam to explain it. A teacher marking the papers of ten students works from a master sheet: for the first question the answer is A, for the second it is B. The master sheet is the golden. The students' papers are the application's real results, and to get them the application has to sit the exam. So an evaluation run has two phases:
- Phase 1: run the application. Every golden query goes through the RAG pipeline. The run adds two things to each golden: the actual answer and the retrieved contexts. On the whiteboard these are written in a second colour under "Generated by RAG pipeline".
- Phase 2: evaluate. Compare the two sides, as the teacher compares a paper with the master sheet.
A golden plus the output of one run is a test case. With both sides in hand, the video names three comparisons: the actual answer against the expected answer, the retrieved chunks against the expected chunks, and the retrieved chunks against the query. The aim is an actual answer that "should be very near to this expected answer".
From a golden to a test case
In code, phase 1 is one call to the app per golden. The test case is the golden's fields plus the two the run produced.
golden = {"id": "g003", "user_input": "How long is the battery life on the SoundPods Pro?",
"reference": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours ..."}
response, contexts = answer(golden["user_input"]) # the run adds two fields
test_case = {**golden, "response": response, "retrieved_contexts": contexts}Phase 1 on three of the video's goldens
The example runs g001, g003 and g004 through the app built in LLM evaluation. Put this code in the same file, under the answer function, in place of the lines that asked one question. It writes the test cases to phase1.json.
import json
# Continues the file from the LLM evaluation lesson: CATALOG, embed, retrieve, generate and answer stay.
GOLDENS = [
{"id": "g001", "user_input": "What is TechNest's return policy?",
"reference": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."},
{"id": "g003", "user_input": "How long is the battery life on the SoundPods Pro?",
"reference": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."},
{"id": "g004", "user_input": "What are TechNest's shipping options and how long do returns take to process?",
"reference": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."},
]
test_cases = []
for golden in GOLDENS:
response, contexts = answer(golden["user_input"]) # phase 1: run the app on the query
test_cases.append({**golden, "response": response, "retrieved_contexts": contexts})
print(golden["id"], golden["user_input"])
print(textwrap.fill(" response: " + response, 100, subsequent_indent=" "))
print(textwrap.fill(" reference: " + golden["reference"], 100, subsequent_indent=" "))
for chunk in contexts:
print(" chunk: " + chunk[:62] + "...")
print()
with open("phase1.json", "w", encoding="utf-8") as fh:
json.dump(test_cases, fh, ensure_ascii=False, indent=1)
print("saved", len(test_cases), "test cases to phase1.json with the fields", list(test_cases[0]))g001 What is TechNest's return policy?
response: TechNest accepts returns within 30 days of the original purchase date, provided items
are in their original packaging with all accessories included. Customers are
responsible for return shipping costs unless the item arrives defective or damaged, and
refunds are processed within 5 to 7 business days of receiving the returned item.
Please note that digital downloads and opened software are non-refundable.
reference: TechNest accepts returns within 30 days of purchase. Items must be in original
packaging with all accessories. Customers pay return shipping unless the item is
defective. Refunds are processed in 5 to 7 business days.
chunk: TechNest accepts returns within 30 days of the original purcha...
chunk: All TechNest products include a minimum 1-year manufacturer wa...
chunk: TechNest offers free standard shipping on all orders over $50 ...
g003 How long is the battery life on the SoundPods Pro?
response: The TechNest SoundPods Pro offer 8 hours of playback on a single charge, and the
charging case extends this to a total of 24 hours.
reference: The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from
the charging case, giving a total of 32 hours.
chunk: The TechNest SoundPods Pro are true wireless earbuds with acti...
chunk: The TechNest ProBook X1 is a 14-inch laptop featuring an Intel...
chunk: The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone wit...
g004 What are TechNest's shipping options and how long do returns take to process?
response: TechNest offers free standard shipping for orders over $50 within the continental US,
which takes 3 to 5 business days, while expedited and same-day delivery are available
for $9.99 and $19.99 respectively. Returns must be initiated within 30 days of
purchase, and refunds are processed within 5 to 7 business days after the returned item
is received.
reference: TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and
expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30
days and refunds are processed in 5 to 7 business days after the item is received.
chunk: TechNest offers free standard shipping on all orders over $50 ...
chunk: TechNest accepts returns within 30 days of the original purcha...
chunk: All TechNest products include a minimum 1-year manufacturer wa...
saved 3 test cases to phase1.json with the fields ['id', 'user_input', 'reference', 'response', 'retrieved_contexts']What phase 1 produced
- Each golden became a test case. The three saved records carry five fields:
id,user_inputandreferencefrom the golden,responseandretrieved_contextsfrom the run. - g001 says more than its reference. The response adds that digital downloads and opened software are non-refundable. The catalog says so; the reference does not mention it. Whether that extra sentence is a fault depends on which metric you ask.
- g003 disagrees with its reference. The response gives a total of 24 hours with the case. The reference says 8 hours per charge and an additional 24 from the case, 32 in total. The retrieved chunk reads "8 hours of playback per charge plus 24 hours with the case", and the app read it differently from the person who wrote the golden.
- g004 needed two entries and got both. The shipping policy and the return policy are the first two chunks, and the response answers both halves of the question. It also lists same-day delivery, which the reference leaves out.
Reading three answers against three references by hand is possible. The video asks what happens with a bigger set: "imagine we have 100 goldens". LLM as a judge hands this comparison to a model.
Writing good goldens
- Cover the kinds of question users ask. Product facts, policies, and questions that need two entries at once, like g004.
- Include questions the knowledge base cannot answer. The expected answer is then that the app says it does not know.
- Include requests the app should refuse. A prompt injection attempt with the refusal as the expected answer checks the guardrails from AI guardrails with the same method.
- Write references as short, factual sentences. The metrics split a reference into separate claims, so one fact per sentence scores cleanly.
- Let a domain expert write or check them. A wrong reference makes every score built on it wrong. A second way to get goldens is to generate candidate questions and answers from your documents with an LLM, as RAGAS' test set generation does, and have the expert review them.
- Keep them in version control and keep adding. Start small, as the video does with five, and add every question that produced a bad answer in production.
A golden vs a test case
| Golden | Test case | |
|---|---|---|
| Holds | user_input, reference, optionally reference_contexts | The golden plus response and retrieved_contexts |
| Written by | A person who knows the right answer | The golden by a person, the rest by the application |
| Created | Once, before any run | Again on every run |
| Changes when | The knowledge or the policy changes | The prompt, model, retriever or catalog changes |
| Can be scored | No | Yes |
Where you use goldens
- As the regression suite of an LLM application. The same goldens run before every release.
- To compare two versions. Two retrievers or two models answer the same goldens, so the scores differ only because of the change.
- To record production failures. A bad answer reported by a user becomes a golden with the corrected answer as its reference.
Related
- Previous: Benchmarks vs custom evaluation
- Next: LLM as a judge
- Reference: RAGAS evaluation sample
- Add golden g005 from the table to
GOLDENS, with the reference"The TechNest PixelPhone 15 is priced at $899.", and run again: a fourth test case is saved. - Change the reference of g003 to any other sentence and run again. The response does not change, because the app never reads a reference. Only phase 2 does.
- In
answer, changeretrieve(question)toretrieve(question, top_k=1)and read the response for g004: with the shipping policy as its only chunk, the app can answer only the first half of the question.
Little by little, you're building something great.