Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q5EasyConcept

What is a golden dataset? How do you build one and how big should it be?

30-second answerSay your answer out loud first, then reveal.

Composition (example for a support bot)

SliceShareWhy
Top intents (real distribution)50%Reflects actual usage
Hard / ambiguous cases15%Where systems fail
Unanswerable / out-of-scope10%Tests abstention
Adversarial (injection, abuse)10%Safety
Multilingual / Hinglish10%User base coverage
Recently changed policies5%Freshness

Each example includes: input (plus conversation history or context), expected output or rubric, tags (intent, difficulty, language), source, and labeller.

Sizing logic: with 100 examples, a measured 85% has a 95% confidence interval of roughly ±7 points. To detect a 3-point improvement reliably you need several hundred (Q19). Size so that decisions are statistically meaningful for the slices you care about.

Maintenance: version it, review labels when policies change, keep a held-out subset so prompts don't overfit (Q41), and add production failures weekly.

Every expert started right here.