RAGAS metrics
RAGAS metrics are named scoring rules from the RAGAS framework, each of which reads fixed fields of a test case and returns a score, usually between 0 and 1.
Last updated: 09 Oct, 2026 · RAGAS 0.4
LLM as a judge ended on a judge that answers one fixed question. A metric is that question plus the arithmetic that turns the verdicts into a number. The video's app scores every test case with five of them.
The five metrics of the video's app
Each metric looks at a different part of a RAG pipeline. Two judge the answer against what was retrieved or asked, two judge the retrieval, and one judges the answer against the golden. The definitions are those of the RAGAS documentation.
| Metric | The question it answers | How the score is formed |
|---|---|---|
| Faithfulness | Is the answer supported by the retrieved context? | Claims in the response that the context supports, divided by all claims in the response |
| Answer relevancy | Does the answer address the question that was asked? | Mean cosine similarity between the user input and questions generated from the response |
| Context precision | Are the useful chunks ranked above the useless ones? | Mean of precision@k over the ranks that hold a relevant chunk |
| Context recall | Did retrieval bring back what the expected answer needs? | Claims in the reference that the context supports, divided by all claims in the reference |
| Answer correctness | Does the answer match the expected answer? | A weighted mix of a factual F1 score and semantic similarity |
A metric returns a number and no pass mark. Where the line between pass and fail sits is your decision, made per metric from runs of your own application.
Which fields each metric reads
A test case has four fields: user_input, response, retrieved_contexts and reference. No metric reads all of them. The inputs of a metric are the parameters of its ascore method, so the installed library can list them.
import inspect
from ragas.metrics.collections import (AnswerCorrectness, AnswerRelevancy, ContextPrecision,
ContextRecall, Faithfulness)
for metric in [Faithfulness, AnswerRelevancy, ContextPrecision, ContextRecall, AnswerCorrectness]:
fields = list(inspect.signature(metric.ascore).parameters)[1:] # the inputs, without self
reference = "yes" if "reference" in fields else "no"
embeddings = "yes" if "embeddings" in inspect.signature(metric.__init__).parameters else "no"
print(f"{metric.__name__:18} reads {', '.join(fields):46} reference: {reference:4} embeddings: {embeddings}")Faithfulness reads user_input, response, retrieved_contexts reference: no embeddings: no AnswerRelevancy reads user_input, response reference: no embeddings: yes ContextPrecision reads user_input, reference, retrieved_contexts reference: yes embeddings: no ContextRecall reads user_input, retrieved_contexts, reference reference: yes embeddings: no AnswerCorrectness reads user_input, response, reference reference: yes embeddings: yes
What the inputs tell you
- Every metric reads
user_input. The question is the one field all five share. - Faithfulness and answer relevancy need no reference. They read only what the run produced, so they can score an answer nobody wrote a golden for. The other three read
referenceand only work on a golden set. - Two metrics take an embedding model. Answer relevancy and answer correctness accept
embeddings; the other three use the judge alone.
The RAG triad
The video draws a triangle that many RAG evaluations start from. Its corners are the three things a RAG run has: the query, the retrieved context and the response. Each side is one check. The name RAG triad comes from TruLens.
- Answer relevance (response to query): is the response relevant to the query? In RAGAS this is answer relevancy.
- Groundedness (response to context): is the response supported by the context? In RAGAS this is faithfulness.
- Context relevance (context to query): is the retrieved context relevant to the query? RAGAS has a
ContextRelevancemetric for it.
The triad needs no golden: all three corners come from the run. Context precision and context recall are not on the triangle. They also judge retrieval, but against the golden's reference, and context precision looks at the order of the chunks as well.
Counting the judge calls for one golden
This part of the video starts at 1:56:18. It sets the scene on the whiteboard: five goldens, each run through the RAG pipeline, each then scored on five metrics by the judge.
How many LLM calls is that? The installed library can answer exactly, without spending a single token, if the judge is replaced by a stand-in that counts.
A judge that counts its calls
To a metric, a judge is a subclass of InstructorBaseRagasLLM whose agenerate(prompt, response_model) method returns an instance of response_model. CountingJudge does that with a made-up reply and adds one to a counter. The scores it leads to mean nothing; the number of calls is what the real judge would make.
class CountingJudge(InstructorBaseRagasLLM):
"""Stands in for the judge: counts each call and returns a valid, made-up reply."""
calls = 0
async def agenerate(self, prompt, response_model):
self.calls += 1
return fill(response_model)
def generate(self, prompt, response_model):
raise NotImplementedErrorCounting calls and embedded texts per metric
fill builds the smallest valid reply for whichever reply class a metric asks for. count scores one sample with all five metrics and adds up the calls. The video's app hands the judge the top 2 of the 3 retrieved chunks, so the first count uses 2 chunks.
import inspect
import typing
from ragas.embeddings.base import BaseRagasEmbedding
from ragas.llms.base import InstructorBaseRagasLLM
from ragas.metrics.collections import (AnswerCorrectness, AnswerRelevancy, ContextPrecision,
ContextRecall, Faithfulness)
def fill(model):
"""The smallest valid reply for a judge's reply class: one item per list, 1 for a number."""
values = {}
for name, field in model.model_fields.items():
kind = field.annotation
if typing.get_origin(kind) is list:
item = typing.get_args(kind)[0]
values[name] = ["x"] if item is str else [fill(item)]
else:
values[name] = 1 if kind is int else "x"
return model(**values)
class CountingJudge(InstructorBaseRagasLLM):
"""Stands in for the judge: counts each call and returns a valid, made-up reply."""
calls = 0
async def agenerate(self, prompt, response_model):
self.calls += 1
return fill(response_model)
def generate(self, prompt, response_model):
raise NotImplementedError
class CountingEmbeddings(BaseRagasEmbedding):
"""Stands in for the embedding model: counts each text and returns a fixed vector."""
texts = 0
def embed_text(self, text, **kwargs):
self.texts += 1
return [1.0, 0.0]
async def aembed_text(self, text, **kwargs):
return self.embed_text(text)
def count(chunks, show=False):
sample = {"user_input": "q", "response": "a", "reference": "r", "retrieved_contexts": ["c"] * chunks}
total = 0
for metric_class in [Faithfulness, AnswerRelevancy, ContextPrecision, ContextRecall, AnswerCorrectness]:
judge, embeddings = CountingJudge(), CountingEmbeddings()
if "embeddings" in inspect.signature(metric_class.__init__).parameters:
metric = metric_class(llm=judge, embeddings=embeddings)
else:
metric = metric_class(llm=judge)
fields = list(inspect.signature(metric.ascore).parameters)
metric.score(**{name: sample[name] for name in fields})
total += judge.calls
if show:
print(f"{metric_class.__name__:18} judge calls: {judge.calls} texts embedded: {embeddings.texts}")
return total
per_golden = count(chunks=2, show=True) # the video's app hands 2 chunks to the judge
print("judge calls for one golden, 2 chunks:", per_golden)
print("judge calls for five goldens:", 5 * per_golden)
print("judge calls for one golden, 3 chunks:", count(chunks=3))Faithfulness judge calls: 2 texts embedded: 0 AnswerRelevancy judge calls: 3 texts embedded: 4 ContextPrecision judge calls: 2 texts embedded: 0 ContextRecall judge calls: 1 texts embedded: 0 AnswerCorrectness judge calls: 3 texts embedded: 2 judge calls for one golden, 2 chunks: 11 judge calls for five goldens: 55 judge calls for one golden, 3 chunks: 12
What the count shows
- No metric is one call. Faithfulness makes 2 judge calls, answer relevancy 3, context precision 2, context recall 1 and answer correctness 3.
- One golden takes 11 judge calls, five take 55. Five goldens times five metrics is 25 scores, and every score takes more than one call.
- Context precision grows with the chunks. With all 3 retrieved chunks the total is 12, one more than with 2, because that metric asks about each chunk separately.
- Two metrics also embed texts. Answer relevancy embeds 4 texts for a sample, the user input and the three generated questions, and answer correctness embeds 2, the response and the reference.
| Metric | Judge calls per sample | What the calls do | Texts embedded |
|---|---|---|---|
| Faithfulness | 2 | Split the response into claims, then check every claim against the context | 0 |
| Answer relevancy | 3 | Write one question the response would answer, three times | 4 |
| Context precision | 1 per chunk | Decide for each retrieved chunk whether it was useful | 0 |
| Context recall | 1 | Split the reference into claims and attribute each to the context, in one reply | 0 |
| Answer correctness | 3 | Split the response into claims, split the reference, then sort the claims | 2 |
Rate limits and cooldowns
This part of the video starts at 1:59:32. How full a context window should be has no documented figure; the metrics are separate calls because each has its own prompt and its own reply format.
A viewer asks in the clip why the five metrics are not scored in one LLM call. The video answers with a person handed ten projects at once: none of them gets done well. A judge asked one focused question, with one reply format, gives verdicts you can read and check one by one. Its work also decides whether a release goes out, so "we don't want this judge to be hallucinated".
This part of the video starts at 2:02:47. The Phase 2 panel on screen shows the app's two cooldown values, 25 and 35 seconds.
Many separate calls meet another limit. A hosted model allows only so many requests and tokens per minute, and a run that sends all its calls at once is answered with rate-limit errors. The video's app spaces its calls with cooldowns: "it is kind of a simple sleep call". It scores one sample at a time. This is the function that scores one metric, which the app calls an experiment, over all goldens:
async def score_experiment(metric, inputs, error_collector=None):
scores = []
for i, inp in enumerate(inputs):
score = await _score_one(metric, inp, error_collector=error_collector, sample_idx=i)
scores.append(score)
if i < len(inputs) - 1:
await asyncio.sleep(SAMPLE_COOLDOWN)
return scoresShown as in the video's repository, without its type hints and docstring, and not run here: it is one function of the Streamlit app. After it returns, the app sleeps EXPERIMENT_COOLDOWN, 35 seconds, before the next metric.
If a call still fails with a rate-limit error, the app waits 65 seconds and tries that sample once more. The waiting adds up:
GOLDENS, METRICS = 5, 5
SAMPLE_COOLDOWN, EXPERIMENT_COOLDOWN = 25, 35 # seconds, the values in the video's app
inside = METRICS * (GOLDENS - 1) * SAMPLE_COOLDOWN # after each golden but the last, in every metric
between = (METRICS - 1) * EXPERIMENT_COOLDOWN # after each metric but the last
total = inside + between
print("waiting between goldens:", inside, "s")
print("waiting between metrics:", between, "s")
print("total waiting:", total, "s, about", round(total / 60, 1), "minutes")waiting between goldens: 500 s waiting between metrics: 140 s total waiting: 640 s, about 10.7 minutes
For five goldens and five metrics the app sleeps 500 seconds between goldens and 140 seconds between metrics: 640 seconds, about 10.7 minutes, in which no call is made. The button in the video's app estimates the whole of phase 2 at 12 to 15 minutes.
Reading your own limits from a response
The right cooldown depends on your key and your model, and both change. Groq reports the current limits in the headers of every response, so one small call tells you what you have.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
raw = client.chat.completions.with_raw_response.create(
model="openai/gpt-oss-20b", temperature=0, max_tokens=300,
messages=[{"role": "user", "content": "Reply with the single word: ready"}])
reply = raw.parse()
print("reply:", reply.choices[0].message.content)
print("tokens used by this call:", reply.usage.total_tokens)
print("token limit per minute:", raw.headers.get("x-ratelimit-limit-tokens"))
print("tokens left this minute:", raw.headers.get("x-ratelimit-remaining-tokens"))reply: ready tokens used by this call: 110 token limit per minute: 8000 tokens left this minute: 4939
This key may send 8000 tokens a minute to openai/gpt-oss-20b. The call itself used 110 tokens, and 4939 were left in that minute because other calls had run shortly before it. A judge call carries a long prompt and a structured reply, so a few scores in a row can use up a minute's allowance. That is the pressure the cooldowns relieve; read your own numbers before you copy the 25 and 35 seconds of the video's app.
Reference-free vs reference-required metrics
| Reference-free | Reference-required | |
|---|---|---|
| Metrics here | Faithfulness, answer relevancy | Context precision, context recall, answer correctness |
| Needs a golden | No, only the run | Yes, the golden's reference |
| Can score live traffic | Yes | No |
| What it can catch | Unsupported claims, answers off the question | Missing facts, wrong facts, wrong chunks retrieved |
| What it cannot see | Whether the answer is the right one | Anything outside the golden set |
Where you use RAGAS metrics
- Finding the weak step of a RAG app. Low context recall points at retrieval; low faithfulness with good retrieval points at generation.
- Watching production. The two reference-free metrics can score a sample of real conversations, where no expected answer exists.
- Guarding releases. All five run on the golden set before a change ships, and each score is compared with the run before.
Related
- Previous: LLM as a judge
- Next: Faithfulness
- Reference: RAGAS available metrics
- Add
print(count(chunks=5, show=True))at the end of the counting example: context precision makes 5 judge calls and the total for one golden is 14. - In
count, changemetric_class(llm=judge, embeddings=embeddings)tometric_class(llm=judge, embeddings=embeddings, strictness=5) if metric_class is AnswerRelevancy else metric_class(llm=judge, embeddings=embeddings): answer relevancy then makes 5 judge calls and embeds 6 texts, and one golden takes 13. - Change
SAMPLE_COOLDOWNto 10 in the waiting example: the total drops to 340 seconds.
Slow is fine. Stopping is the only problem.