1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is a golden dataset? How do you build one and how big should it be?
30-second answerSay your answer out loud first, then reveal.
Composition (example for a support bot)
| Slice | Share | Why |
|---|---|---|
| Top intents (real distribution) | 50% | Reflects actual usage |
| Hard / ambiguous cases | 15% | Where systems fail |
| Unanswerable / out-of-scope | 10% | Tests abstention |
| Adversarial (injection, abuse) | 10% | Safety |
| Multilingual / Hinglish | 10% | User base coverage |
| Recently changed policies | 5% | Freshness |
Each example includes: input (plus conversation history or context), expected output or rubric, tags (intent, difficulty, language), source, and labeller.
Sizing logic: with 100 examples, a measured 85% has a 95% confidence interval of roughly ±7 points. To detect a 3-point improvement reliably you need several hundred (Q19). Size so that decisions are statistically meaningful for the slices you care about.
Maintenance: version it, review labels when policies change, keep a held-out subset so prompts don't overfit (Q41), and add production failures weekly.
Related
Every expert started right here.