1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you build an evaluation dataset for RAG?
30-second answerSay your answer out loud first, then reveal.
Sources
- Real queries: search logs, chatbot logs, support tickets, FAQ pages. They match the real distribution.
- Subject-matter experts: write the questions users should be able to ask, with answers and sources.
- Synthetic generation: an LLM generates question–answer pairs from sampled chunks (RAGAS and others support this).
Pros: cheap coverage and automatic ground-truth context.
Cons: questions often copy the document's wording, which makes retrieval look too easy. Paraphrase them, and review samples manually.
Each example should have
- question, reference answer, relevant doc/chunk IDs, metadata (category, difficulty, type).
Coverage checklist
| Type | Example |
|---|---|
| Simple factoid | "What's the notice period?" |
| Multi-hop | "Who approves leave for the manager of the Pune office?" |
| Comparison | "How does plan A differ from plan B?" |
| Unanswerable | "What's the policy on Mars travel?" → should abstain |
| Ambiguous | "What's the limit?" → should clarify |
| Numerical / table | "What was Q3 APAC revenue?" |
| Recency | Questions about recently changed policies |
| Paraphrased / jargon | Slang, acronyms, typos, Hinglish |
Maintenance: version the dataset; update labels when documents change; add every production failure.
Related
Every expert started right here.