AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

RAGAS metrics

RAGAS metrics are named scoring rules from the RAGAS framework, each of which reads fixed fields of a test case and returns a score, usually between 0 and 1.

Last updated: 09 Oct, 2026 · RAGAS 0.4

LLM as a judge ended on a judge that answers one fixed question. A metric is that question plus the arithmetic that turns the verdicts into a number. The video's app scores every test case with five of them.

The five metrics of the video's app

Each metric looks at a different part of a RAG pipeline. Two judge the answer against what was retrieved or asked, two judge the retrieval, and one judges the answer against the golden. The definitions are those of the RAGAS documentation.

MetricThe question it answersHow the score is formed
FaithfulnessIs the answer supported by the retrieved context?Claims in the response that the context supports, divided by all claims in the response
Answer relevancyDoes the answer address the question that was asked?Mean cosine similarity between the user input and questions generated from the response
Context precisionAre the useful chunks ranked above the useless ones?Mean of precision@k over the ranks that hold a relevant chunk
Context recallDid retrieval bring back what the expected answer needs?Claims in the reference that the context supports, divided by all claims in the reference
Answer correctnessDoes the answer match the expected answer?A weighted mix of a factual F1 score and semantic similarity

A metric returns a number and no pass mark. Where the line between pass and fail sits is your decision, made per metric from runs of your own application.

Which fields each metric reads

A test case has four fields: user_input, response, retrieved_contexts and reference. No metric reads all of them. The inputs of a metric are the parameters of its ascore method, so the installed library can list them.

ExampleRun on RAGAS 0.4.3
import inspect

from ragas.metrics.collections import (AnswerCorrectness, AnswerRelevancy, ContextPrecision,
                                       ContextRecall, Faithfulness)

for metric in [Faithfulness, AnswerRelevancy, ContextPrecision, ContextRecall, AnswerCorrectness]:
    fields = list(inspect.signature(metric.ascore).parameters)[1:]          # the inputs, without self
    reference = "yes" if "reference" in fields else "no"
    embeddings = "yes" if "embeddings" in inspect.signature(metric.__init__).parameters else "no"
    print(f"{metric.__name__:18} reads {', '.join(fields):46} reference: {reference:4} embeddings: {embeddings}")

What the inputs tell you

  • Every metric reads user_input. The question is the one field all five share.
  • Faithfulness and answer relevancy need no reference. They read only what the run produced, so they can score an answer nobody wrote a golden for. The other three read reference and only work on a golden set.
  • Two metrics take an embedding model. Answer relevancy and answer correctness accept embeddings; the other three use the judge alone.
A table of the five RAGAS metrics and the fields each reads. Faithfulness reads user_input, response and retrieved_contexts. Answer relevancy reads user_input and response and needs embeddings. Context precision and context recall read user_input, retrieved_contexts and reference. Answer correctness reads user_input, response and reference and needs embeddings. The first two need no reference; the last three need the golden's reference.

The RAG triad

The video draws a triangle that many RAG evaluations start from. Its corners are the three things a RAG run has: the query, the retrieved context and the response. Each side is one check. The name RAG triad comes from TruLens.

A triangle with Query at the top, Context at the bottom right and Response at the bottom left. The side from Response to Query is answer relevance (RAGAS AnswerRelevancy), the side from Query to Context is context relevance (RAGAS ContextRelevance), and the side from Context to Response is groundedness (RAGAS Faithfulness).
  • Answer relevance (response to query): is the response relevant to the query? In RAGAS this is answer relevancy.
  • Groundedness (response to context): is the response supported by the context? In RAGAS this is faithfulness.
  • Context relevance (context to query): is the retrieved context relevant to the query? RAGAS has a ContextRelevance metric for it.

The triad needs no golden: all three corners come from the run. Context precision and context recall are not on the triangle. They also judge retrieval, but against the golden's reference, and context precision looks at the order of the chunks as well.

Counting the judge calls for one golden

Why an evaluation run makes many LLM calls · from the Complete AI Security Course in 8 Hours video · 1:56:18 to 1:57:38

This part of the video starts at 1:56:18. It sets the scene on the whiteboard: five goldens, each run through the RAG pipeline, each then scored on five metrics by the judge.

How many LLM calls is that? The installed library can answer exactly, without spending a single token, if the judge is replaced by a stand-in that counts.

A judge that counts its calls

To a metric, a judge is a subclass of InstructorBaseRagasLLM whose agenerate(prompt, response_model) method returns an instance of response_model. CountingJudge does that with a made-up reply and adds one to a counter. The scores it leads to mean nothing; the number of calls is what the real judge would make.

python
class CountingJudge(InstructorBaseRagasLLM):
    """Stands in for the judge: counts each call and returns a valid, made-up reply."""
    calls = 0

    async def agenerate(self, prompt, response_model):
        self.calls += 1
        return fill(response_model)

    def generate(self, prompt, response_model):
        raise NotImplementedError

Counting calls and embedded texts per metric

fill builds the smallest valid reply for whichever reply class a metric asks for. count scores one sample with all five metrics and adds up the calls. The video's app hands the judge the top 2 of the 3 retrieved chunks, so the first count uses 2 chunks.

ExampleRun on RAGAS 0.4.3
import inspect
import typing

from ragas.embeddings.base import BaseRagasEmbedding
from ragas.llms.base import InstructorBaseRagasLLM
from ragas.metrics.collections import (AnswerCorrectness, AnswerRelevancy, ContextPrecision,
                                       ContextRecall, Faithfulness)


def fill(model):
    """The smallest valid reply for a judge's reply class: one item per list, 1 for a number."""
    values = {}
    for name, field in model.model_fields.items():
        kind = field.annotation
        if typing.get_origin(kind) is list:
            item = typing.get_args(kind)[0]
            values[name] = ["x"] if item is str else [fill(item)]
        else:
            values[name] = 1 if kind is int else "x"
    return model(**values)


class CountingJudge(InstructorBaseRagasLLM):
    """Stands in for the judge: counts each call and returns a valid, made-up reply."""
    calls = 0

    async def agenerate(self, prompt, response_model):
        self.calls += 1
        return fill(response_model)

    def generate(self, prompt, response_model):
        raise NotImplementedError


class CountingEmbeddings(BaseRagasEmbedding):
    """Stands in for the embedding model: counts each text and returns a fixed vector."""
    texts = 0

    def embed_text(self, text, **kwargs):
        self.texts += 1
        return [1.0, 0.0]

    async def aembed_text(self, text, **kwargs):
        return self.embed_text(text)


def count(chunks, show=False):
    sample = {"user_input": "q", "response": "a", "reference": "r", "retrieved_contexts": ["c"] * chunks}
    total = 0
    for metric_class in [Faithfulness, AnswerRelevancy, ContextPrecision, ContextRecall, AnswerCorrectness]:
        judge, embeddings = CountingJudge(), CountingEmbeddings()
        if "embeddings" in inspect.signature(metric_class.__init__).parameters:
            metric = metric_class(llm=judge, embeddings=embeddings)
        else:
            metric = metric_class(llm=judge)
        fields = list(inspect.signature(metric.ascore).parameters)
        metric.score(**{name: sample[name] for name in fields})
        total += judge.calls
        if show:
            print(f"{metric_class.__name__:18} judge calls: {judge.calls}   texts embedded: {embeddings.texts}")
    return total


per_golden = count(chunks=2, show=True)                  # the video's app hands 2 chunks to the judge
print("judge calls for one golden, 2 chunks:", per_golden)
print("judge calls for five goldens:", 5 * per_golden)
print("judge calls for one golden, 3 chunks:", count(chunks=3))

What the count shows

  • No metric is one call. Faithfulness makes 2 judge calls, answer relevancy 3, context precision 2, context recall 1 and answer correctness 3.
  • One golden takes 11 judge calls, five take 55. Five goldens times five metrics is 25 scores, and every score takes more than one call.
  • Context precision grows with the chunks. With all 3 retrieved chunks the total is 12, one more than with 2, because that metric asks about each chunk separately.
  • Two metrics also embed texts. Answer relevancy embeds 4 texts for a sample, the user input and the three generated questions, and answer correctness embeds 2, the response and the reference.
Five goldens go through phase 1, where the app retrieves and generates with one generator call per golden, producing a response and retrieved contexts. In phase 2 the judge scores each golden: faithfulness takes 2 judge calls, answer relevancy 3, context precision 2 (one per chunk, with 2 chunks), context recall 1 and answer correctness 3, which is 11 judge calls per golden and 55 for five goldens, giving 25 scores.
MetricJudge calls per sampleWhat the calls doTexts embedded
Faithfulness2Split the response into claims, then check every claim against the context0
Answer relevancy3Write one question the response would answer, three times4
Context precision1 per chunkDecide for each retrieved chunk whether it was useful0
Context recall1Split the reference into claims and attribute each to the context, in one reply0
Answer correctness3Split the response into claims, split the reference, then sort the claims2

Rate limits and cooldowns

Why each metric gets its own call · from the Complete AI Security Course in 8 Hours video · 1:59:32 to 2:01:30

This part of the video starts at 1:59:32. How full a context window should be has no documented figure; the metrics are separate calls because each has its own prompt and its own reply format.

A viewer asks in the clip why the five metrics are not scored in one LLM call. The video answers with a person handed ten projects at once: none of them gets done well. A judge asked one focused question, with one reply format, gives verdicts you can read and check one by one. Its work also decides whether a release goes out, so "we don't want this judge to be hallucinated".

Cooldowns between judge calls · from the Complete AI Security Course in 8 Hours video · 2:02:47 to 2:03:21

This part of the video starts at 2:02:47. The Phase 2 panel on screen shows the app's two cooldown values, 25 and 35 seconds.

Many separate calls meet another limit. A hosted model allows only so many requests and tokens per minute, and a run that sends all its calls at once is answered with rate-limit errors. The video's app spaces its calls with cooldowns: "it is kind of a simple sleep call". It scores one sample at a time. This is the function that scores one metric, which the app calls an experiment, over all goldens:

python
async def score_experiment(metric, inputs, error_collector=None):
    scores = []
    for i, inp in enumerate(inputs):
        score = await _score_one(metric, inp, error_collector=error_collector, sample_idx=i)
        scores.append(score)
        if i < len(inputs) - 1:
            await asyncio.sleep(SAMPLE_COOLDOWN)
    return scores

Shown as in the video's repository, without its type hints and docstring, and not run here: it is one function of the Streamlit app. After it returns, the app sleeps EXPERIMENT_COOLDOWN, 35 seconds, before the next metric.

If a call still fails with a rate-limit error, the app waits 65 seconds and tries that sample once more. The waiting adds up:

ExampleRun on Python 3.12
GOLDENS, METRICS = 5, 5
SAMPLE_COOLDOWN, EXPERIMENT_COOLDOWN = 25, 35            # seconds, the values in the video's app

inside = METRICS * (GOLDENS - 1) * SAMPLE_COOLDOWN       # after each golden but the last, in every metric
between = (METRICS - 1) * EXPERIMENT_COOLDOWN            # after each metric but the last
total = inside + between
print("waiting between goldens:", inside, "s")
print("waiting between metrics:", between, "s")
print("total waiting:", total, "s, about", round(total / 60, 1), "minutes")

For five goldens and five metrics the app sleeps 500 seconds between goldens and 140 seconds between metrics: 640 seconds, about 10.7 minutes, in which no call is made. The button in the video's app estimates the whole of phase 2 at 12 to 15 minutes.

Reading your own limits from a response

The right cooldown depends on your key and your model, and both change. Groq reports the current limits in the headers of every response, so one small call tells you what you have.

ExampleAPI keyRun on Groq
import os

from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
raw = client.chat.completions.with_raw_response.create(
    model="openai/gpt-oss-20b", temperature=0, max_tokens=300,
    messages=[{"role": "user", "content": "Reply with the single word: ready"}])
reply = raw.parse()
print("reply:", reply.choices[0].message.content)
print("tokens used by this call:", reply.usage.total_tokens)
print("token limit per minute:", raw.headers.get("x-ratelimit-limit-tokens"))
print("tokens left this minute:", raw.headers.get("x-ratelimit-remaining-tokens"))

This key may send 8000 tokens a minute to openai/gpt-oss-20b. The call itself used 110 tokens, and 4939 were left in that minute because other calls had run shortly before it. A judge call carries a long prompt and a structured reply, so a few scores in a row can use up a minute's allowance. That is the pressure the cooldowns relieve; read your own numbers before you copy the 25 and 35 seconds of the video's app.

Reference-free vs reference-required metrics

Reference-freeReference-required
Metrics hereFaithfulness, answer relevancyContext precision, context recall, answer correctness
Needs a goldenNo, only the runYes, the golden's reference
Can score live trafficYesNo
What it can catchUnsupported claims, answers off the questionMissing facts, wrong facts, wrong chunks retrieved
What it cannot seeWhether the answer is the right oneAnything outside the golden set

Where you use RAGAS metrics

  • Finding the weak step of a RAG app. Low context recall points at retrieval; low faithfulness with good retrieval points at generation.
  • Watching production. The two reference-free metrics can score a sample of real conversations, where no expected answer exists.
  • Guarding releases. All five run on the golden set before a change ships, and each score is compared with the run before.
Watch out. Counting one LLM call per metric underestimates a run. With the video's settings it is 11 judge calls per golden, and context precision grows with every chunk you retrieve. Plan the token budget and the cooldowns from the real count.
Try it yourself
  • Add print(count(chunks=5, show=True)) at the end of the counting example: context precision makes 5 judge calls and the total for one golden is 14.
  • In count, change metric_class(llm=judge, embeddings=embeddings) to metric_class(llm=judge, embeddings=embeddings, strictness=5) if metric_class is AnswerRelevancy else metric_class(llm=judge, embeddings=embeddings): answer relevancy then makes 5 judge calls and embeds 6 texts, and one golden takes 13.
  • Change SAMPLE_COOLDOWN to 10 in the waiting example: the total drops to 340 seconds.

Slow is fine. Stopping is the only problem.