Context recall
Context recall is a RAGAS metric that measures the share of claims in the reference answer that the retrieved context supports.
Last updated: 09 Oct, 2026 · RAGAS 0.4
Context precision asks whether the chunks that came back are in a good order. Context recall asks the question before that one: did everything the answer needs come back at all? If a fact of the right answer is in no retrieved chunk, the generator can only leave it out or make it up.
This part of the video starts at 2:24:57. The metric compares two things, the golden's reference and the chunks the retriever returned, and the app's own answer plays no part in it.
Reading the video's reference claims
The inputs are the reference, the ideal answer stored in the golden, and the retrieved_contexts. The video's card describes the reference as a proxy for what the retriever must cover: nobody has labelled which chunks are the right ones, so the ideal answer stands in for them.
The judge breaks the reference into independent claims and decides for each one whether it can be attributed to the retrieved context. The video's example stays with the question "What is the minimum balance for my savings account?". The reference holds four claims, and the retriever returned four chunks.
The urban balance is in chunk 1, the fee in chunk 2 and the semi-urban balance in chunk 4. No chunk mentions the rural balance, so that claim stays unsupported. The video reads this as the retrieval pipeline not doing its job for that fact. Three of four claims are supported, and the card shows 0.75 against a pass mark of 0.7.
Writing context recall as a formula
Every claim weighs the same. In RAGAS 0.4 one judge call does the whole job: the chunks are joined into one text, and the reply lists each statement of the reference with a reason and an attributed value of 1 or 0. The reply does not say which chunk supported a claim; the chunk numbers on the card are there for the reader.
Counting the supported claims by hand
reference_claims = [
("Urban branches require ₹10,000 minimum balance", "Chunk 1"),
("Non-maintenance fee is ₹350 + taxes when balance falls below limit", "Chunk 2"),
("Rural branch minimum balance is ₹2,500", None), # no retrieved chunk carries it
("Semi-urban branch minimum balance is ₹5,000", "Chunk 4"),
]
supported = sum(chunk is not None for claim, chunk in reference_claims)
score = supported / len(reference_claims)
for claim, chunk in reference_claims:
print("supported" if chunk else "missing ", "|", claim, "|", chunk or "not retrieved")
print(f"context recall = {supported} / {len(reference_claims)} = {score}")
print("pass mark 0.7:", "pass" if score >= 0.7 else "fail")supported | Urban branches require ₹10,000 minimum balance | Chunk 1 supported | Non-maintenance fee is ₹350 + taxes when balance falls below limit | Chunk 2 missing | Rural branch minimum balance is ₹2,500 | not retrieved supported | Semi-urban branch minimum balance is ₹5,000 | Chunk 4 context recall = 3 / 4 = 0.75 pass mark 0.7: pass
Three supported claims out of four give 0.75, which passes the card's mark of 0.7. The one missing claim is the whole story: a retriever that never fetched the rural figure.
Watching recall grow with top-k
The five chunks of the context precision card are the same knowledge base in ranked order, with the rural chunk at rank 5. Cutting that ranking at different values of k shows how many reference claims each cut covers. Each relevant chunk here carries one claim.
import matplotlib.pyplot as plt
# the retriever's ranking, best first, and the four facts the reference needs
ranking = ["urban minimum balance", "non-maintenance fee", "KYC update",
"semi-urban minimum balance", "rural minimum balance"]
needed = {"urban minimum balance", "non-maintenance fee", "rural minimum balance", "semi-urban minimum balance"}
ks = [1, 2, 3, 4, 5]
recalls = []
for k in ks:
covered = needed & set(ranking[:k])
recalls.append(len(covered) / len(needed))
print(f"top_k = {k}: {len(covered)} of {len(needed)} reference claims covered, context recall {recalls[-1]:.2f}")
plt.figure(figsize=(7, 3.8))
bars = plt.bar(ks, recalls, color=["#d64541" if k == 4 else "#3a6fd8" for k in ks])
plt.bar_label(bars, fmt="%.2f")
plt.title("Context recall as top-k grows")
plt.xlabel("top_k: how many chunks the retriever returns (red: the card's case)")
plt.ylabel("Context recall")
plt.ylim(0, 1.15)
plt.show()top_k = 1: 1 of 4 reference claims covered, context recall 0.25 top_k = 2: 2 of 4 reference claims covered, context recall 0.50 top_k = 3: 2 of 4 reference claims covered, context recall 0.50 top_k = 4: 3 of 4 reference claims covered, context recall 0.75 top_k = 5: 4 of 4 reference claims covered, context recall 1.00
- top_k = 4 gives 0.75, the card's case: the rural chunk sits at rank 5 and is cut off.
- top_k = 5 gives 1.00. The missing fact was in the knowledge base all along; k was too small.
- top_k = 3 gives 0.50, the same as top_k = 2. The third chunk is the KYC noise, which adds nothing to recall.
- Recall never falls as k grows. Precision is what pays for a larger k, because more noise comes in with it.
What a low context recall points to
This part of the video starts at 2:28:06. A chunk that was never retrieved cannot be fixed by reordering the retrieved ones: it needs a larger k, a better search (hybrid search, a better embedding model) or better chunking.
The card lists three causes of a low score:
- Missing chunks. The retriever never fetched the document that holds a key fact of the reference.
- k too small. The right chunk exists and ranks fourth, but only three chunks are kept.
- Embedding gap. The right chunk exists but its embedding sits too far from the question's, so it ranks too low to surface.
A viewer asks whether an unsupported claim means the answer is hallucinating. It does not. The claims come from the reference, the ideal answer in the golden, which is taken as correct and is not under test. What is under test is the retrieved context: a low score says it does not cover every part of the ideal answer.
Scoring the reference with RAGAS
The video's app ran llama-3.1-8b-instant as its judge, since retired on Groq; the run below uses openai/gpt-oss-20b from LLM as a judge. The reference joins the card's four claims into one answer, and the chunks are the four the card shows as retrieved.
from ragas.metrics.collections import ContextRecall
result = ContextRecall(llm=judge).score(user_input=question, retrieved_contexts=chunks, reference=reference)
result.value # supported reference claims / all reference claimsThe wrapper from Faithfulness keeps the judge's single reply, so the example can print every statement with its attributed value and the reason.
import os
from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import ContextRecall
client = AsyncOpenAI(
base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"],
max_retries=6, # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)
replies = [] # each structured reply the judge sends back to the metric
ask = judge.agenerate
async def keep(prompt, response_model):
reply = await ask(prompt, response_model)
replies.append(reply)
return reply
judge.agenerate = keep
question = "What is the minimum balance for my savings account?"
reference = ("Urban branches require a ₹10,000 minimum balance. The non-maintenance fee is ₹350 + taxes when the "
"balance falls below the limit. Rural branch minimum balance is ₹2,500. "
"Semi-urban branch minimum balance is ₹5,000.")
chunks = [
"Minimum balance is ₹10,000 for urban branches.",
"Non-maintenance fee is ₹350 + taxes if the balance falls below the minimum.",
"KYC update is required every 8 years.",
"Semi-urban branch minimum balance is ₹5,000.",
]
result = ContextRecall(llm=judge).score(user_input=question, retrieved_contexts=chunks, reference=reference)
for item in replies[-1].classifications:
print(item.attributed, "|", item.statement)
print(" ", item.reason)
print("context recall:", result.value)1 | Urban branches require a ₹10,000 minimum balance.
The context explicitly states that the minimum balance for urban branches is ₹10,000.
1 | The non-maintenance fee is ₹350 + taxes when the balance falls below the limit.
The context specifies that a non-maintenance fee of ₹350 + taxes applies if the balance falls below the minimum.
0 | Rural branch minimum balance is ₹2,500.
The context does not mention any minimum balance for rural branches.
1 | Semi-urban branch minimum balance is ₹5,000.
The context states that the minimum balance for semi-urban branches is ₹5,000.
context recall: 0.75What the judge attributed
- The judge split the reference into four statements, one per sentence, the four claims of the card.
- Three are attributed and one is not. The rural statement gets a 0, with the reason that the context does not mention any minimum balance for rural branches.
- The score is 0.75, the card's value and the hand count.
- The KYC chunk changed nothing. Recall asks only whether each reference claim is covered; a chunk that supports no claim neither helps nor hurts. Context precision is the metric that charges for it.
Context recall vs faithfulness
| Context recall | Faithfulness | |
|---|---|---|
| Claims are taken from | The reference (the golden's answer) | The response (the app's answer) |
| Claims are checked against | The retrieved context | The retrieved context |
| Reads the app's answer | No | Yes |
| Needs a golden | Yes | No |
| A low score means | The retriever missed something the answer needs | The generator said something the context does not hold |
| Part of the app it grades | Retriever | Generator |
Where you use context recall
- Choosing top-k. Raise k until recall stops growing, then use context precision to see what the extra chunks cost.
- Questions that need more than one chunk. A golden such as the video's "What are TechNest's shipping options and how long do returns take to process?" needs a shipping chunk and a returns chunk; recall drops when only one comes back.
- After changing chunking, the embedding model or the search method. Recall on the same goldens shows whether the facts still reach the generator.
Related
- Previous: Context precision
- Next: Answer correctness
- Reference: Context recall in the RAGAS docs
- In the hand count, replace
Nonewith"Chunk 5", as if the rural chunk had been retrieved: the score becomes 4 / 4 = 1.0. - In the top-k example, move
"rural minimum balance"to the front ofranking: recall reaches 1.00 at top_k = 5 as before, but top_k = 4 now gives 0.75 for a different reason, the semi-urban chunk is the one cut off. - In the RAGAS example, add
"Rural branch minimum balance is ₹2,500."tochunksand run it again: the rural statement is attributed and the score rises to 1.0.
You understood something today that you didn't yesterday.