Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q13EasyConcept

How do you evaluate a RAG system?

30-second answerSay your answer out loud first, then reveal.
A RAG pipeline from question to retriever, retrieved context, generator and answer, with retrieval eval checking the retrieved context and generation and end-to-end evals checking the answer.

Retrieval metrics

Retrieval metrics need labelled relevant chunks or documents per question.

  • Recall@k: fraction of relevant items found in the top k. The most important one: if it's not retrieved, the LLM can't use it.
  • Precision@k: fraction of the top k that are relevant (noise level).
  • MRR (Mean Reciprocal Rank): how high the first relevant result ranks.
  • nDCG: ranking quality with graded relevance.

Generation metrics

  • Faithfulness / groundedness: every claim in the answer is supported by the retrieved context.
  • Answer relevance: does it address the question asked?
  • Correctness: matches the reference answer (exact, semantic, or judged).
  • Citation accuracy: citations point to supporting passages.
  • Abstention quality: says "I don't know" for unanswerable questions.

Why separate them

Low quality with high retrieval recall means fix the prompt or generator. Low recall means fix chunking, embeddings or search. This diagnosis is the main point of the split.

Tools: RAGAS, DeepEval, TruLens, Promptfoo, LangSmith/Braintrust evaluators.

Slow is fine. Stopping is the only problem.