RAGASragas 0.4.3 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

RAGAS metrics

A RAGAS metric is a predefined scoring rule, run by the judge as a chain of small questions or computed as plain code, that turns some fields of a sample into a score between 0 and 1.

Last updated: 29 Sep, 2026 · RAGAS 0.4.3

A single "quality" number hides where an app fails. RAGAS splits quality into metrics, each reading different fields of the sample, so a low score points at retrieval or at generation.

Different metrics, different sides of quality · from the Production RAG Live Marathon · 221:46 to 224:48

Metrics as the marking scheme

The video returns to the exam. Different kinds of question get different marking: a multiple-choice question is marked one way, an essay another. A RAG chatbot has its own metrics, and tool calling has others; the docs list them all. A student can be strong in English and weak in maths, and each subject's marks show it. In the same way, each metric shows one side of the app's quality. The video goes deep on three it calls essential for any RAG pipeline: faithfulness, answer relevancy and context recall. The demo app also scores context precision and answer correctness, and this course covers all five.

Four sample fields at the top, user_input, retrieved_contexts, response and reference, with lines down to the five metrics that read them: faithfulness reads user_input, response and retrieved_contexts; answer relevancy reads user_input and response; context precision and context recall read user_input, retrieved_contexts and reference; answer correctness reads user_input, response and reference.

The ascore signature

python
metric = Faithfulness(llm=judge)
result = await metric.ascore(user_input=..., response=..., retrieved_contexts=...)
result.value  # the score, between 0 and 1

Every collections metric has an ascore for async code and a score for plain scripts. Their keyword arguments are the sample fields that metric reads, so the signature tells you what a metric needs before you run it.

Listing the fields each metric reads

Example
import inspect

from ragas.metrics.collections import (AnswerCorrectness, AnswerRelevancy, ContextPrecision,
                                       ContextRecall, Faithfulness)

for metric in [Faithfulness, AnswerRelevancy, ContextPrecision, ContextRecall, AnswerCorrectness]:
    fields = [name for name in inspect.signature(metric.ascore).parameters if name != "self"]
    print(metric.__name__.ljust(18), fields)

What the signatures say

  • Faithfulness reads the answer and the chunks, not the reference: it checks the answer against what was retrieved.
  • AnswerRelevancy reads only the question and the answer. It needs no golden answer, so it can run on live traffic.
  • ContextPrecision and ContextRecall read the chunks and the reference. They grade the search, not the answer.
  • AnswerCorrectness reads the answer and the reference: it grades the answer against the truth.

Retrieval metrics vs generation metrics

MetricGradesNeeds a reference?Needs embeddings?A low score means
FaithfulnessGenerationNoNoThe answer adds claims the chunks do not support
Answer relevancyGenerationNoYesThe answer drifts from the question
Context precisionRetrievalYesNoUseful chunks are ranked below noise
Context recallRetrievalYesNoA chunk the answer needs was not retrieved
Answer correctnessGenerationYesYesThe answer's facts differ from the golden's

Choosing metrics for a question

  • Worried the bot makes things up: faithfulness.
  • Worried the search misses documents: context recall; worried it ranks them badly: context precision.
  • Have expert answers and want to know if the bot is right: answer correctness.
Watch out. A metric that needs a reference cannot score live traffic, where no expected answer exists. Plan which metrics run in development on goldens and which run on production answers.
Try it yourself
  • Add FactualCorrectness and ContextUtilization to the list and read which fields they take.
  • For each of the five goldens' metric_focus, say which fields that metric reads.

This is what real progress feels like.