Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q25IntermediateConcept

How do you build an evaluation dataset for RAG?

30-second answerSay your answer out loud first, then reveal.

Sources

  1. Real queries: search logs, chatbot logs, support tickets, FAQ pages. They match the real distribution.
  2. Subject-matter experts: write the questions users should be able to ask, with answers and sources.
  3. Synthetic generation: an LLM generates question–answer pairs from sampled chunks (RAGAS and others support this).
    Pros: cheap coverage and automatic ground-truth context.
    Cons: questions often copy the document's wording, which makes retrieval look too easy. Paraphrase them, and review samples manually.

Each example should have

  • question, reference answer, relevant doc/chunk IDs, metadata (category, difficulty, type).

Coverage checklist

TypeExample
Simple factoid"What's the notice period?"
Multi-hop"Who approves leave for the manager of the Pune office?"
Comparison"How does plan A differ from plan B?"
Unanswerable"What's the policy on Mars travel?" → should abstain
Ambiguous"What's the limit?" → should clarify
Numerical / table"What was Q3 APAC revenue?"
RecencyQuestions about recently changed policies
Paraphrased / jargonSlang, acronyms, typos, Hinglish

Maintenance: version the dataset; update labels when documents change; add every production failure.

Every expert started right here.