1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Explain the core RAGAS metrics: faithfulness, answer relevancy, context precision, and context recall.
30-second answerSay your answer out loud first, then reveal.
| Metric | Measures | Needs reference answer? | How computed (roughly) | Low score means |
|---|---|---|---|---|
| Faithfulness | Answer claims supported by context | No | Break answer into claims; LLM checks each against context; supported / total | Hallucination / parametric leakage |
| Answer relevancy | Answer addresses the question | No | Generate questions from the answer, compare similarity to the original question | Off-topic or evasive answers |
| Context precision | Relevant chunks ranked high | Yes (or relevance labels) | Precision weighted by rank of relevant chunks | Noisy retrieval; need reranking |
| Context recall | Context covers the needed information | Yes | Fraction of reference-answer statements attributable to context | Retrieval missing information |
Reading the scores together
- Low context recall + low correctness → retrieval problem.
- High context recall + low faithfulness → generation problem.
- High faithfulness + low correctness → context was wrong or outdated but the model faithfully repeated it.
Caveats (important in interviews)
- These are LLM-judged metrics: noisy, prompt-dependent, sensitive to the judge model. Calibrate against human labels.
- Faithfulness doesn't measure correctness. Faithfully repeating a wrong document scores high.
- Use them for comparison between versions more than as absolute truth.
Related
Little by little, you're building something great.