Datasets and goldens
A golden is a test case that holds a question and its expected answer but not the app's reply, and an EvaluationDataset is DeepEval's container that holds a list of goldens and the test cases built from them.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
The RAG metrics lessons built each test case inline, one at a time. An evaluation that runs after every change needs a fixed set of questions with known answers, kept apart from any one run of the bot. That set is the dataset.
What goes into a golden
The video defines goldens as a collection of carefully curated test cases that hold the expected or correct information, against which the app's outputs are evaluated. For the demo app, the goldens are its questions. Writing them is a crucial step before any evaluation pipeline: take your time, and cover the different kinds of question your bot should answer, because the results can only be as diverse as the questions.
The questions depend on the app. An airline chatbot needs goldens such as my flight got cancelled, how do I get a refund; a LangChain helper needs what is LCEL or what is the difference between LangChain and LangGraph. That is why goldens need domain expertise. Each golden is a container with the input and the expected output, the way the app should answer, which is the ground truth for that question. It can hold more fields when your use case needs them. The video's app keeps its goldens in its own format for RAGAS; DeepEval has a Golden class for the same idea.
The EvaluationDataset API
from deepeval.dataset import EvaluationDataset, Golden
golden = Golden(input="...", expected_output="...") # a question and its right answer
dataset = EvaluationDataset(goldens=[golden]) # a dataset from a list of goldens
dataset.add_goldens_from_json_file(file_path="goldens.json") # or add them from a file
dataset.goldens # the goldens, as a list
dataset.add_test_case(test_case) # a test case built from a golden
dataset.save_as(file_type="json", directory="saved") # write the goldens to diskThe goldens.json file
Save the five TechNest goldens as goldens.json, next to catalog.json and technest.py. They cover a policy question, two product specs, a question that needs two policies at once and a price.
[
{
"name": "g001",
"input": "What is TechNest's return policy?",
"expected_output": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
},
{
"name": "g002",
"input": "What are the RAM and storage specs of the ProBook X1?",
"expected_output": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
},
{
"name": "g003",
"input": "How long is the battery life on the SoundPods Pro?",
"expected_output": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
},
{
"name": "g004",
"input": "What are TechNest's shipping options and how long do returns take to process?",
"expected_output": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
},
{
"name": "g005",
"input": "What is the price of the PixelPhone 15?",
"expected_output": "The TechNest PixelPhone 15 is priced at $899."
}
]The keys name, input and expected_output are the names of Golden's own fields, so the loader needs no extra arguments; name_key_name="name" is already the default. A file with other key names still loads: pass them as input_key_name="question", expected_output_key_name=... or name_key_name=....
A golden written in code
The same golden can be built in Python. Only input is required; expected_output, context and name are optional.
from deepeval.dataset import Golden
golden = Golden(
name="g005",
input="What is the price of the PixelPhone 15?",
expected_output="The TechNest PixelPhone 15 is priced at $899.",
)Loading the file and saving a copy
dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="goldens.json")
path = dataset.save_as(file_type="json", directory="saved", file_name="technest_goldens")save_as writes the goldens to saved/technest_goldens.json and returns the path. It also accepts "csv" and "jsonl". Without file_name, the file is named after the current date and time.
Loading and saving the TechNest goldens
from deepeval.dataset import EvaluationDataset
dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="goldens.json")
print(len(dataset.goldens), "goldens")
for golden in dataset.goldens:
print(golden.name, "|", golden.input)
first = dataset.goldens[0]
print()
print(type(first).__name__)
print("expected:", first.expected_output)
print("actual:", first.actual_output)
path = dataset.save_as(file_type="json", directory="saved", file_name="technest_goldens")
print(path)5 goldens g001 | What is TechNest's return policy? g002 | What are the RAM and storage specs of the ProBook X1? g003 | How long is the battery life on the SoundPods Pro? g004 | What are TechNest's shipping options and how long do returns take to process? g005 | What is the price of the PixelPhone 15? Golden expected: TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days. actual: None Evaluation dataset saved at saved/technest_goldens.json! saved/technest_goldens.json
What the dataset holds
- Five goldens, g001 to g005, in the order of the file. Each keeps its
name, which the reports later use to label each row. - Each item is a
Goldenwith the expected answer from the file. actual: None: a golden has no reply from the bot yet. The bot writes one at evaluation time, every time you run it.- The saved file is at
saved/technest_goldens.json, and DeepEval prints where it put it. Open it and you will find everyGoldenfield, most of themnull.
Evaluating goldens directly
A golden looks like a test case, so the first thing many people try is to hand the goldens to evaluate(), the function that runs metrics on many test cases at once. It fails, with an error that does not say why.
from deepeval import evaluate
from deepeval.dataset import EvaluationDataset
from deepeval.metrics import ExactMatchMetric
dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="goldens.json")
evaluate(dataset.goldens, metrics=[ExactMatchMetric()])Traceback (most recent call last):
File "main.py", line 8, in <module>
evaluate(dataset.goldens, metrics=[ExactMatchMetric()])
AttributeError: 'Golden' object has no attribute 'mcp_servers'evaluate() expects LLMTestCase objects and reads fields a Golden does not have. And even if it ran, there would be nothing to score: no golden has an actual_output. Each golden has to go through the bot first.
- written in TechNest RAG app
- written in TechNest RAG app
View the code here
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
Turning goldens into test cases
Loop over the goldens, ask the bot each question with answer from technest.py, and build an LLMTestCase from the golden's fields plus what the bot returned. dataset.add_test_case keeps it in the dataset, next to its golden. Two goldens are enough to see the shape.
from deepeval.dataset import EvaluationDataset
from deepeval.test_case import LLMTestCase
from technest import answer
dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="goldens.json")
for golden in dataset.goldens[:2]:
response, contexts = answer(golden.input)
dataset.add_test_case(LLMTestCase(
name=golden.name,
input=golden.input,
actual_output=response,
expected_output=golden.expected_output,
retrieval_context=contexts,
))
for test_case in dataset.test_cases:
print(test_case.name, "|", test_case.input)
print(" actual: ", test_case.actual_output)
print(" expected:", test_case.expected_output)
print(" chunks: ", len(test_case.retrieval_context))g001 | What is TechNest's return policy? actual: TechNest accepts returns within 30 days of the original purchase date, provided items are in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged, and refunds are processed within 5 to 7 business days of receiving the returned item. expected: TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days. chunks: 3 g002 | What are the RAM and storage specs of the ProBook X1? actual: The TechNest ProBook X1 is equipped with 16GB of DDR5 RAM and a 512GB NVMe SSD. expected: The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD. chunks: 3
What the two test cases contain
- g001: the bot's answer covers the 30 days, the original packaging, who pays return shipping and the 5 to 7 business days, in longer sentences than the expected answer.
- g002: the bot gives 16GB of DDR5 RAM and a 512GB NVMe SSD, the same facts as the expected answer.
- chunks: 3: each test case also carries the three retrieved chunks, so faithfulness and the contextual metrics can run on it later.
- The goldens did not change. Running the bot again tomorrow, after a prompt change, builds new test cases from the same goldens.
Golden vs LLMTestCase
Golden | LLMTestCase | |
|---|---|---|
| Required fields | input | input, and in practice actual_output |
| Holds the bot's answer | No | Yes, actual_output |
| Holds the retrieved chunks | No | Yes, retrieval_context |
| Written by | A person who knows the domain, once | Your code, on every run |
| Scored by metrics | No | Yes |
| Lives in | A dataset file such as goldens.json | Memory, for the length of a run |
When to use a dataset
- When the same questions should be asked of every version of the bot: a new prompt, a new model, a new retriever.
- When non-programmers write the questions. A JSON or CSV file is easier to review than test cases in code.
- When the questions outgrow a handful. Five goldens are enough to learn with, not to judge an app.
goldens.json and evaluate those. The scores would describe the bot as it was when the file was written, not the code you are about to ship. DeepEval's end-to-end evaluation docs make the same point: treat goldens as inputs only, and produce actual_output fresh on each run.Related
- Previous: Contextual relevancy
- Next: Synthetic goldens
- Reference: Datasets
- Before the loop, add
dataset.add_golden(Golden(input="Do you ship to Canada?"))(importGoldenfromdeepeval.dataset) and printlen(dataset.goldens): it is now 6. - Change
dataset.goldens[:2]todataset.goldens[4:]and read the bot's answer to the PixelPhone 15 price question. - After the loop, call
dataset.save_as(file_type="json", directory="saved", file_name="with_cases", include_test_cases=True)and opensaved/with_cases.json: it now holds the test cases as well.
This is what real progress feels like.