RAGAS metrics
A RAGAS metric is a predefined scoring rule, run by the judge as a chain of small questions or computed as plain code, that turns some fields of a sample into a score between 0 and 1.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
A single "quality" number hides where an app fails. RAGAS splits quality into metrics, each reading different fields of the sample, so a low score points at retrieval or at generation.
Metrics as the marking scheme
The video returns to the exam. Different kinds of question get different marking: a multiple-choice question is marked one way, an essay another. A RAG chatbot has its own metrics, and tool calling has others; the docs list them all. A student can be strong in English and weak in maths, and each subject's marks show it. In the same way, each metric shows one side of the app's quality. The video goes deep on three it calls essential for any RAG pipeline: faithfulness, answer relevancy and context recall. The demo app also scores context precision and answer correctness, and this course covers all five.
The ascore signature
metric = Faithfulness(llm=judge)
result = await metric.ascore(user_input=..., response=..., retrieved_contexts=...)
result.value # the score, between 0 and 1Every collections metric has an ascore for async code and a score for plain scripts. Their keyword arguments are the sample fields that metric reads, so the signature tells you what a metric needs before you run it.
Listing the fields each metric reads
import inspect
from ragas.metrics.collections import (AnswerCorrectness, AnswerRelevancy, ContextPrecision,
ContextRecall, Faithfulness)
for metric in [Faithfulness, AnswerRelevancy, ContextPrecision, ContextRecall, AnswerCorrectness]:
fields = [name for name in inspect.signature(metric.ascore).parameters if name != "self"]
print(metric.__name__.ljust(18), fields)Faithfulness ['user_input', 'response', 'retrieved_contexts'] AnswerRelevancy ['user_input', 'response'] ContextPrecision ['user_input', 'reference', 'retrieved_contexts'] ContextRecall ['user_input', 'retrieved_contexts', 'reference'] AnswerCorrectness ['user_input', 'response', 'reference']
What the signatures say
- Faithfulness reads the answer and the chunks, not the reference: it checks the answer against what was retrieved.
- AnswerRelevancy reads only the question and the answer. It needs no golden answer, so it can run on live traffic.
- ContextPrecision and ContextRecall read the chunks and the reference. They grade the search, not the answer.
- AnswerCorrectness reads the answer and the reference: it grades the answer against the truth.
Retrieval metrics vs generation metrics
| Metric | Grades | Needs a reference? | Needs embeddings? | A low score means |
|---|---|---|---|---|
| Faithfulness | Generation | No | No | The answer adds claims the chunks do not support |
| Answer relevancy | Generation | No | Yes | The answer drifts from the question |
| Context precision | Retrieval | Yes | No | Useful chunks are ranked below noise |
| Context recall | Retrieval | Yes | No | A chunk the answer needs was not retrieved |
| Answer correctness | Generation | Yes | Yes | The answer's facts differ from the golden's |
Choosing metrics for a question
- Worried the bot makes things up: faithfulness.
- Worried the search misses documents: context recall; worried it ranks them badly: context precision.
- Have expert answers and want to know if the bot is right: answer correctness.
Related
- Previous: LLM as a judge
- Next: Rate limits and cooldowns
- Reference: Available metrics
- Add
FactualCorrectnessandContextUtilizationto the list and read which fields they take. - For each of the five goldens'
metric_focus, say which fields that metric reads.
This is what real progress feels like.