AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Context recall

Context recall is a RAGAS metric that measures the share of claims in the reference answer that the retrieved context supports.

Last updated: 09 Oct, 2026 · RAGAS 0.4

Context precision asks whether the chunks that came back are in a good order. Context recall asks the question before that one: did everything the answer needs come back at all? If a fact of the right answer is in no retrieved chunk, the generator can only leave it out or make it up.

Context recall on the reference claims · from the Complete AI Security Course in 8 Hours video · 2:24:57 to 2:27:44

This part of the video starts at 2:24:57. The metric compares two things, the golden's reference and the chunks the retriever returned, and the app's own answer plays no part in it.

Reading the video's reference claims

The inputs are the reference, the ideal answer stored in the golden, and the retrieved_contexts. The video's card describes the reference as a proxy for what the retriever must cover: nobody has labelled which chunks are the right ones, so the ideal answer stands in for them.

The judge breaks the reference into independent claims and decides for each one whether it can be attributed to the retrieved context. The video's example stays with the question "What is the minimum balance for my savings account?". The reference holds four claims, and the retriever returned four chunks.

Four claims from the reference answer on the left and four retrieved chunks on the right. The urban, fee and semi-urban claims are each supported by a chunk; the rural claim is marked not retrieved because no chunk carries it, and the KYC chunk supports no claim. Three supported claims out of four give a context recall of 0.75.

The urban balance is in chunk 1, the fee in chunk 2 and the semi-urban balance in chunk 4. No chunk mentions the rural balance, so that claim stays unsupported. The video reads this as the retrieval pipeline not doing its job for that fact. Three of four claims are supported, and the card shows 0.75 against a pass mark of 0.7.

Writing context recall as a formula

Context recall, as the RAGAS docs define it

Every claim weighs the same. In RAGAS 0.4 one judge call does the whole job: the chunks are joined into one text, and the reply lists each statement of the reference with a reason and an attributed value of 1 or 0. The reply does not say which chunk supported a claim; the chunk numbers on the card are there for the reader.

Counting the supported claims by hand

ExampleThe video's example, computed in plain Python
reference_claims = [
    ("Urban branches require ₹10,000 minimum balance", "Chunk 1"),
    ("Non-maintenance fee is ₹350 + taxes when balance falls below limit", "Chunk 2"),
    ("Rural branch minimum balance is ₹2,500", None),  # no retrieved chunk carries it
    ("Semi-urban branch minimum balance is ₹5,000", "Chunk 4"),
]

supported = sum(chunk is not None for claim, chunk in reference_claims)
score = supported / len(reference_claims)

for claim, chunk in reference_claims:
    print("supported" if chunk else "missing  ", "|", claim, "|", chunk or "not retrieved")
print(f"context recall = {supported} / {len(reference_claims)} = {score}")
print("pass mark 0.7:", "pass" if score >= 0.7 else "fail")

Three supported claims out of four give 0.75, which passes the card's mark of 0.7. The one missing claim is the whole story: a retriever that never fetched the rural figure.

Watching recall grow with top-k

The five chunks of the context precision card are the same knowledge base in ranked order, with the rural chunk at rank 5. Cutting that ranking at different values of k shows how many reference claims each cut covers. Each relevant chunk here carries one claim.

ExampleRun on matplotlib 3.11.2
import matplotlib.pyplot as plt

# the retriever's ranking, best first, and the four facts the reference needs
ranking = ["urban minimum balance", "non-maintenance fee", "KYC update",
           "semi-urban minimum balance", "rural minimum balance"]
needed = {"urban minimum balance", "non-maintenance fee", "rural minimum balance", "semi-urban minimum balance"}

ks = [1, 2, 3, 4, 5]
recalls = []
for k in ks:
    covered = needed & set(ranking[:k])
    recalls.append(len(covered) / len(needed))
    print(f"top_k = {k}: {len(covered)} of {len(needed)} reference claims covered, context recall {recalls[-1]:.2f}")

plt.figure(figsize=(7, 3.8))
bars = plt.bar(ks, recalls, color=["#d64541" if k == 4 else "#3a6fd8" for k in ks])
plt.bar_label(bars, fmt="%.2f")
plt.title("Context recall as top-k grows")
plt.xlabel("top_k: how many chunks the retriever returns (red: the card's case)")
plt.ylabel("Context recall")
plt.ylim(0, 1.15)
plt.show()
A bar chart of context recall for top-k from 1 to 5: 0.25, 0.50, 0.50, 0.75 and 1.00. The bar for top-k = 4, the card's case, is red.
  • top_k = 4 gives 0.75, the card's case: the rural chunk sits at rank 5 and is cut off.
  • top_k = 5 gives 1.00. The missing fact was in the knowledge base all along; k was too small.
  • top_k = 3 gives 0.50, the same as top_k = 2. The third chunk is the KYC noise, which adds nothing to recall.
  • Recall never falls as k grows. Precision is what pays for a larger k, because more noise comes in with it.

What a low context recall points to

What causes a low context recall · from the Complete AI Security Course in 8 Hours video · 2:28:06 to 2:30:01

This part of the video starts at 2:28:06. A chunk that was never retrieved cannot be fixed by reordering the retrieved ones: it needs a larger k, a better search (hybrid search, a better embedding model) or better chunking.

The card lists three causes of a low score:

  • Missing chunks. The retriever never fetched the document that holds a key fact of the reference.
  • k too small. The right chunk exists and ranks fourth, but only three chunks are kept.
  • Embedding gap. The right chunk exists but its embedding sits too far from the question's, so it ranks too low to surface.

A viewer asks whether an unsupported claim means the answer is hallucinating. It does not. The claims come from the reference, the ideal answer in the golden, which is taken as correct and is not under test. What is under test is the retrieved context: a low score says it does not cover every part of the ideal answer.

Scoring the reference with RAGAS

The video's app ran llama-3.1-8b-instant as its judge, since retired on Groq; the run below uses openai/gpt-oss-20b from LLM as a judge. The reference joins the card's four claims into one answer, and the chunks are the four the card shows as retrieved.

python
from ragas.metrics.collections import ContextRecall

result = ContextRecall(llm=judge).score(user_input=question, retrieved_contexts=chunks, reference=reference)
result.value  # supported reference claims / all reference claims

The wrapper from Faithfulness keeps the judge's single reply, so the example can print every statement with its attributed value and the reason.

ExampleAPI keyFrom the video, run on Groq
import os

from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import ContextRecall

client = AsyncOpenAI(
    base_url="https://api.groq.com/openai/v1",
    api_key=os.environ["GROQ_API_KEY"],
    max_retries=6,  # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)

replies = []  # each structured reply the judge sends back to the metric
ask = judge.agenerate


async def keep(prompt, response_model):
    reply = await ask(prompt, response_model)
    replies.append(reply)
    return reply


judge.agenerate = keep

question = "What is the minimum balance for my savings account?"
reference = ("Urban branches require a ₹10,000 minimum balance. The non-maintenance fee is ₹350 + taxes when the "
             "balance falls below the limit. Rural branch minimum balance is ₹2,500. "
             "Semi-urban branch minimum balance is ₹5,000.")
chunks = [
    "Minimum balance is ₹10,000 for urban branches.",
    "Non-maintenance fee is ₹350 + taxes if the balance falls below the minimum.",
    "KYC update is required every 8 years.",
    "Semi-urban branch minimum balance is ₹5,000.",
]

result = ContextRecall(llm=judge).score(user_input=question, retrieved_contexts=chunks, reference=reference)
for item in replies[-1].classifications:
    print(item.attributed, "|", item.statement)
    print("     ", item.reason)
print("context recall:", result.value)

What the judge attributed

  • The judge split the reference into four statements, one per sentence, the four claims of the card.
  • Three are attributed and one is not. The rural statement gets a 0, with the reason that the context does not mention any minimum balance for rural branches.
  • The score is 0.75, the card's value and the hand count.
  • The KYC chunk changed nothing. Recall asks only whether each reference claim is covered; a chunk that supports no claim neither helps nor hurts. Context precision is the metric that charges for it.

Context recall vs faithfulness

Context recallFaithfulness
Claims are taken fromThe reference (the golden's answer)The response (the app's answer)
Claims are checked againstThe retrieved contextThe retrieved context
Reads the app's answerNoYes
Needs a goldenYesNo
A low score meansThe retriever missed something the answer needsThe generator said something the context does not hold
Part of the app it gradesRetrieverGenerator

Where you use context recall

  • Choosing top-k. Raise k until recall stops growing, then use context precision to see what the extra chunks cost.
  • Questions that need more than one chunk. A golden such as the video's "What are TechNest's shipping options and how long do returns take to process?" needs a shipping chunk and a returns chunk; recall drops when only one comes back.
  • After changing chunking, the embedding model or the search method. Recall on the same goldens shows whether the facts still reach the generator.
Watch out. Context recall is only as complete as the reference. It checks only the facts the golden's answer mentions, so a short reference lets a weak retriever score 1.0. Write references that state every fact a good answer must contain.
Try it yourself
  • In the hand count, replace None with "Chunk 5", as if the rural chunk had been retrieved: the score becomes 4 / 4 = 1.0.
  • In the top-k example, move "rural minimum balance" to the front of ranking: recall reaches 1.00 at top_k = 5 as before, but top_k = 4 now gives 0.75 for a different reason, the semi-urban chunk is the one cut off.
  • In the RAGAS example, add "Rural branch minimum balance is ₹2,500." to chunks and run it again: the rural statement is attributed and the score rises to 1.0.

You understood something today that you didn't yesterday.