1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you evaluate a RAG system?
30-second answerSay your answer out loud first, then reveal.

Retrieval metrics
Retrieval metrics need labelled relevant chunks or documents per question.
- Recall@k: fraction of relevant items found in the top k. The most important one: if it's not retrieved, the LLM can't use it.
- Precision@k: fraction of the top k that are relevant (noise level).
- MRR (Mean Reciprocal Rank): how high the first relevant result ranks.
- nDCG: ranking quality with graded relevance.
Generation metrics
- Faithfulness / groundedness: every claim in the answer is supported by the retrieved context.
- Answer relevance: does it address the question asked?
- Correctness: matches the reference answer (exact, semantic, or judged).
- Citation accuracy: citations point to supporting passages.
- Abstention quality: says "I don't know" for unanswerable questions.
Why separate them
Low quality with high retrieval recall means fix the prompt or generator. Low recall means fix chunking, embeddings or search. This diagnosis is the main point of the split.
Tools: RAGAS, DeepEval, TruLens, Promptfoo, LangSmith/Braintrust evaluators.
Related
Slow is fine. Stopping is the only problem.