DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

Synthetic goldens

The Synthesizer is DeepEval's golden generator that reads pieces of your knowledge base and asks an LLM to write new questions, with expected answers, grounded in them.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

goldens.json from Datasets and goldens has five hand-written questions. The catalog has fifteen entries, and most of them have no golden at all. Writing one by hand for each is slow; the Synthesizer writes first drafts that a person then checks.

The Synthesizer API

python
from deepeval.synthesizer import Synthesizer

synthesizer = Synthesizer(model=judge)           # the LLM that writes the goldens
goldens = synthesizer.generate_goldens_from_contexts(
    contexts=[["chunk", "chunk"], ["chunk"]],    # each inner list is one context
    max_goldens_per_context=1,                   # at most this many goldens per context
    include_expected_output=True,                # also write an expected answer
)

There are four ways to generate: from documents, from contexts, from scratch and from existing goldens. generate_goldens_from_contexts takes text you already have, such as catalog entries. generate_goldens_from_docs reads files instead, but it first splits and groups them with an embedding model, OpenAI's by default, and needs chromadb, langchain, langchain_community and langchain_text_splitters installed, so this lesson stays with contexts.

Picking contexts from the catalog

A context is a list of strings about one subject. Two catalog entries that no golden covers: the warranty policy and international shipping.

python
catalog = {item["id"]: item["content"] for item in json.load(open("catalog.json"))}
contexts = [[catalog["policy_003"]], [catalog["faq_002"]]]  # warranty, international shipping

A synthesizer on the Groq judge

python
synthesizer = Synthesizer(model=judge, max_concurrent=1)

Without model, the Synthesizer uses an OpenAI model. The same model also scores each draft question in the filtration step. max_concurrent=1 writes one golden at a time; the default allows many in parallel, which is more than a free Groq key accepts in a minute.

What happens to each golden

For every context, the model writes a question from the text. A filtration step scores the question for clarity and for making sense on its own, and asks for a new one when the score is too low. An evolution step then rewrites it to be harder or more realistic: a comparison, a reasoning question, a hypothetical. Last, the model writes the expected answer from the context. The evolutions used are kept in additional_metadata.

Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
judge.py
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
catalog.json
[
  {
    "id": "prod_001",
    "category": "product",
    "title": "ProBook X1 Laptop",
    "content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
  },
  {
    "id": "prod_002",
    "category": "product",
    "title": "PixelPhone 15",
    "content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
  },
  {
    "id": "prod_003",
    "category": "product",
    "title": "SoundPods Pro",
    "content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
  },
  {
    "id": "prod_004",
    "category": "product",
    "title": "UltraTab S2 Tablet",
    "content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
  },
  {
    "id": "prod_005",
    "category": "product",
    "title": "SmartWatch X",
    "content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
  },
  {
    "id": "prod_006",
    "category": "product",
    "title": "ProCam 4K Action Camera",
    "content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
  },
  {
    "id": "prod_007",
    "category": "product",
    "title": "BassBuds Max Headphones",
    "content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
  },
  {
    "id": "prod_008",
    "category": "product",
    "title": "SoundBar 360",
    "content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
  },
  {
    "id": "policy_001",
    "category": "policy",
    "title": "Return Policy",
    "content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
  },
  {
    "id": "policy_002",
    "category": "policy",
    "title": "Shipping Policy",
    "content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
  },
  {
    "id": "policy_003",
    "category": "policy",
    "title": "Warranty Policy",
    "content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
  },
  {
    "id": "policy_004",
    "category": "policy",
    "title": "Payment Policy",
    "content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
  },
  {
    "id": "faq_001",
    "category": "faq",
    "title": "Order Tracking",
    "content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
  },
  {
    "id": "faq_002",
    "category": "faq",
    "title": "International Shipping",
    "content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
  },
  {
    "id": "faq_003",
    "category": "faq",
    "title": "Bulk and Business Orders",
    "content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
  }
]

Generating goldens from two catalog entries

ExampleAPI key
import json

from deepeval.synthesizer import Synthesizer

from judge import judge

catalog = {item["id"]: item["content"] for item in json.load(open("catalog.json"))}
contexts = [[catalog["policy_003"]], [catalog["faq_002"]]]  # warranty, international shipping

synthesizer = Synthesizer(model=judge, max_concurrent=1)
goldens = synthesizer.generate_goldens_from_contexts(
    contexts=contexts,
    max_goldens_per_context=1,
    include_expected_output=True,
)

for golden in goldens:
    print("input:   ", golden.input)
    print("expected:", golden.expected_output)
    print("metadata:", golden.additional_metadata)
    print()

Reviewing what the synthesizer wrote

The synthesizer picks its question styles at random, so your goldens and quality scores will differ from these. Review yours the same way.

  • The warranty golden is a hypothetical: If a new TechNest device launches, does it retain at least a 1‑yr defect warranty? The context never mentions new launches, so the question tests something the catalog does not say, and its expected answer adds including new launches on its own. Its synthetic_input_quality is 0.3, below the default synthetic_input_quality_threshold, 0.5 in this version. The filtration step asks for better questions a few times and then keeps the best one it got, so a golden under the threshold can still come back.
  • The shipping golden scored 0.6, and its expected answer matches the context: rates from $14.99, 7 to 14 business days. But the question gives the answer away (What are TechNest's intl shipping rates and 7‑14 business day delivery times?), and no customer writes intl. A bot could pass it by repeating the question.
  • The metadata records the evolution each question went through, Hypothetical and Concretizing here, and the quality score from the filtration step. Sorting by that score is a quick way to decide which goldens to read first.
  • Neither golden is ready as it stands. Rewrite the first question or drop it, reword the second the way a customer would ask it, and only then add them to the dataset.

Saving the reviewed goldens

Once you have read and fixed the goldens, put them in a dataset and save them next to the hand-written ones. The name field is empty on synthetic goldens; set it so reports can label them.

python
from deepeval.dataset import EvaluationDataset

for number, golden in enumerate(goldens, start=6):
    golden.name = f"g{number:03d}"   # g006, g007

dataset = EvaluationDataset(goldens=goldens)
dataset.save_as(file_type="json", directory="saved", file_name="synthetic_goldens")

Synthetic vs hand-written goldens

Hand-written (goldens.json)Synthetic (Synthesizer)
Written bySomeone who knows the customersAn LLM, from your contexts
SpeedSlowSeconds per golden
Questions read likeReal customersWhatever the evolution step made of them
Expected answerChecked by a personWritten from the context; needs checking
CoversWhat you thought ofEvery context you pass, including ones you forgot

When to generate goldens

  • When you have no dataset yet and need a first set to start evaluating.
  • When the knowledge base grows: generate a few goldens for each new entry and review them, instead of writing all of them by hand.
  • Not as a replacement for real questions from users. Goldens need domain expertise, so a person reviews every synthetic one before it joins the dataset.
Watch out. The Synthesizer never writes actual_output. A synthetic golden is still a golden: the bot answers it at evaluation time, the same way as the hand-written ones.
Try it yourself
  • Pass styling_config=StylingConfig(scenario="Customers of an online electronics store chatting with its support bot", task="Answering customer questions about TechNest products and policies", input_format="A short, casual question a customer would type") to Synthesizer (import StylingConfig from deepeval.synthesizer.config) and compare how the questions read.
  • Pass evolution_config=EvolutionConfig(evolutions={Evolution.CONCRETIZING: 0.5, Evolution.CONSTRAINED: 0.5}) (import Evolution from deepeval.synthesizer and EvolutionConfig from deepeval.synthesizer.config): both evolutions stay with the context, so no hypothetical questions come back.
  • Set max_goldens_per_context=2 and keep only the warranty context: two different questions come back about the same entry.

Every expert started right here.