AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Goldens

A golden is a test record that holds a query and the answer an expert expects for it, and a golden dataset is the set of such records an application is evaluated against.

Last updated: 09 Oct, 2026 · OpenAI SDK 3.3

A custom evaluation starts with a question nobody else can answer for you: what should your application say? Benchmarks vs custom evaluation listed the golden dataset as the first thing to build. The video puts it in three words: "It is the truth."

What a golden holds

What a golden holds · from the Complete AI Security Course in 8 Hours video · 1:32:37 to 1:35:16

This part of the video starts at 1:32:37. On the whiteboard a golden is a small record with up to three parts:

  • Queries. "Different types of queries which can be asked by the end users", chosen to cover every side of the chatbot.
  • Expected answer. The truth for that query, the answer you would call perfect. Somebody has to write it, and the right somebody is an expert in the domain. Experts' time is scarce, so it has to be used with care.
  • Expected context. The chunks or sources the pipeline should fetch to build that answer. This part is optional: "this is up to you as a developer".

The clip's own example comes from the document chat: the query is "what is attention", the app already gives an actual output, and the expected answer still has to be written by someone who knows the paper.

Frameworks name these parts differently. The whiteboard uses plain words, the TechNest app uses the RAGAS field names, and "golden" is the word DeepEval uses for such a record. The rest of this part uses the RAGAS names.

On the whiteboardRAGAS fieldDeepEval field
Queryuser_inputinput
Expected answerreferenceexpected_output
Expected contextreference_contextscontext
Actual answerresponseactual_output
Retrieved contextsretrieved_contextsretrieval_context

The five goldens of the TechNest app

The video's repository keeps its goldens in goldens.json. Each one has an ID, a query and a reference:

json
{
  "id": "g001",
  "metric_focus": "faithfulness",
  "user_input": "What is TechNest's return policy?",
  "reference": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
}

There are five, "for the sake of simplicity":

IDmetric_focususer_input
g001faithfulnessWhat is TechNest's return policy?
g002answer_relevancyWhat are the RAM and storage specs of the ProBook X1?
g003context_precisionHow long is the battery life on the SoundPods Pro?
g004context_recallWhat are TechNest's shipping options and how long do returns take to process?
g005answer_correctnessWhat is the price of the PixelPhone 15?
  • metric_focus is a label. It names the metric a golden was written to stress. Every golden is still scored on every metric; nothing in the scoring code reads this field.
  • There is no expected context. The file has no reference_contexts. RAGAS' context recall then uses the reference as a stand-in for it.
  • g004 asks two things at once. Shipping options and return processing sit in two different catalog entries, so this golden tests whether the retriever brings back both.

A golden set is a test set

In machine learning terms, training data fits a model, validation data tunes it, and a test set is held out and only used to measure. A golden set plays the part of the test set. Nothing is fitted on it. It is run again after every change to the application, unchanged, which makes it a regression suite: a score that drops points at the change that caused it.

Running the app on the goldens

Phase 1: running the RAG pipeline on the goldens · from the Complete AI Security Course in 8 Hours video · 1:44:31 to 1:49:26

This part of the video starts at 1:44:31. Goldens alone cannot be scored. They say what should come; an evaluation also needs what did come.

The video uses an exam to explain it. A teacher marking the papers of ten students works from a master sheet: for the first question the answer is A, for the second it is B. The master sheet is the golden. The students' papers are the application's real results, and to get them the application has to sit the exam. So an evaluation run has two phases:

  1. Phase 1: run the application. Every golden query goes through the RAG pipeline. The run adds two things to each golden: the actual answer and the retrieved contexts. On the whiteboard these are written in a second colour under "Generated by RAG pipeline".
  2. Phase 2: evaluate. Compare the two sides, as the teacher compares a paper with the master sheet.
Two columns that join into one test case. From the golden, written before any run: user_input (the query), reference (the expected answer) and, optionally, reference_contexts (the expected chunks). From the run of the RAG pipeline: response (the actual answer) and retrieved_contexts (the chunks the retriever returned). A judge compares the two sides.

A golden plus the output of one run is a test case. With both sides in hand, the video names three comparisons: the actual answer against the expected answer, the retrieved chunks against the expected chunks, and the retrieved chunks against the query. The aim is an actual answer that "should be very near to this expected answer".

From a golden to a test case

In code, phase 1 is one call to the app per golden. The test case is the golden's fields plus the two the run produced.

python
golden = {"id": "g003", "user_input": "How long is the battery life on the SoundPods Pro?",
          "reference": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours ..."}

response, contexts = answer(golden["user_input"])                # the run adds two fields
test_case = {**golden, "response": response, "retrieved_contexts": contexts}

Phase 1 on three of the video's goldens

The example runs g001, g003 and g004 through the app built in LLM evaluation. Put this code in the same file, under the answer function, in place of the lines that asked one question. It writes the test cases to phase1.json.

ExampleAPI keyFrom the video, run on Groq
import json

# Continues the file from the LLM evaluation lesson: CATALOG, embed, retrieve, generate and answer stay.
GOLDENS = [
    {"id": "g001", "user_input": "What is TechNest's return policy?",
     "reference": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."},
    {"id": "g003", "user_input": "How long is the battery life on the SoundPods Pro?",
     "reference": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."},
    {"id": "g004", "user_input": "What are TechNest's shipping options and how long do returns take to process?",
     "reference": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."},
]

test_cases = []
for golden in GOLDENS:
    response, contexts = answer(golden["user_input"])            # phase 1: run the app on the query
    test_cases.append({**golden, "response": response, "retrieved_contexts": contexts})
    print(golden["id"], golden["user_input"])
    print(textwrap.fill("  response:  " + response, 100, subsequent_indent="             "))
    print(textwrap.fill("  reference: " + golden["reference"], 100, subsequent_indent="             "))
    for chunk in contexts:
        print("  chunk:     " + chunk[:62] + "...")
    print()

with open("phase1.json", "w", encoding="utf-8") as fh:
    json.dump(test_cases, fh, ensure_ascii=False, indent=1)
print("saved", len(test_cases), "test cases to phase1.json with the fields", list(test_cases[0]))

What phase 1 produced

  • Each golden became a test case. The three saved records carry five fields: id, user_input and reference from the golden, response and retrieved_contexts from the run.
  • g001 says more than its reference. The response adds that digital downloads and opened software are non-refundable. The catalog says so; the reference does not mention it. Whether that extra sentence is a fault depends on which metric you ask.
  • g003 disagrees with its reference. The response gives a total of 24 hours with the case. The reference says 8 hours per charge and an additional 24 from the case, 32 in total. The retrieved chunk reads "8 hours of playback per charge plus 24 hours with the case", and the app read it differently from the person who wrote the golden.
  • g004 needed two entries and got both. The shipping policy and the return policy are the first two chunks, and the response answers both halves of the question. It also lists same-day delivery, which the reference leaves out.

Reading three answers against three references by hand is possible. The video asks what happens with a bigger set: "imagine we have 100 goldens". LLM as a judge hands this comparison to a model.

Writing good goldens

  • Cover the kinds of question users ask. Product facts, policies, and questions that need two entries at once, like g004.
  • Include questions the knowledge base cannot answer. The expected answer is then that the app says it does not know.
  • Include requests the app should refuse. A prompt injection attempt with the refusal as the expected answer checks the guardrails from AI guardrails with the same method.
  • Write references as short, factual sentences. The metrics split a reference into separate claims, so one fact per sentence scores cleanly.
  • Let a domain expert write or check them. A wrong reference makes every score built on it wrong. A second way to get goldens is to generate candidate questions and answers from your documents with an LLM, as RAGAS' test set generation does, and have the expert review them.
  • Keep them in version control and keep adding. Start small, as the video does with five, and add every question that produced a bad answer in production.

A golden vs a test case

GoldenTest case
Holdsuser_input, reference, optionally reference_contextsThe golden plus response and retrieved_contexts
Written byA person who knows the right answerThe golden by a person, the rest by the application
CreatedOnce, before any runAgain on every run
Changes whenThe knowledge or the policy changesThe prompt, model, retriever or catalog changes
Can be scoredNoYes

Where you use goldens

  • As the regression suite of an LLM application. The same goldens run before every release.
  • To compare two versions. Two retrievers or two models answer the same goldens, so the scores differ only because of the change.
  • To record production failures. A bad answer reported by a user becomes a golden with the corrected answer as its reference.
Watch out. Goldens go stale. When a price or a policy in the catalog changes, the old reference turns a correct new answer into a low score. Update the goldens in the same change that updates the knowledge base.
Try it yourself
  • Add golden g005 from the table to GOLDENS, with the reference "The TechNest PixelPhone 15 is priced at $899.", and run again: a fourth test case is saved.
  • Change the reference of g003 to any other sentence and run again. The response does not change, because the app never reads a reference. Only phase 2 does.
  • In answer, change retrieve(question) to retrieve(question, top_k=1) and read the response for g004: with the shipping policy as its only chunk, the app can answer only the first half of the question.

Little by little, you're building something great.