Contextual recall
ContextualRecallMetric is a DeepEval RAG metric that scores whether the retrieved chunks contain everything the expected answer says: the share of the expected answer's sentences that the judge can attribute to a chunk.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
Contextual precision grades the order of the chunks, and it cannot see a chunk that never came back. Contextual recall starts from the other end: it takes the expected answer and looks for each part of it in the chunks.
Claims from the reference, found in the chunks
The video lines up the metrics so far. Faithfulness works on the hallucination side of the pipeline; answer relevancy checks that the answer addresses the user's question. An answer can be faithful and still not relevant, so you need both. Context recall covers a third side, and it needs two pieces of data: the reference, the ground-truth answer from the goldens, and the retrieved context. Its goal is to check whether the retrieved context is sufficient.
The judge does it in two steps. It extracts the claims from the reference, the correct answer, then checks whether the retrieved chunks cover each claim. For a ProBook X1 question, the reference says 16GB DDR5 RAM and a 512GB NVMe SSD, so the chunks must say both. In the bank example the reference makes four claims, and no chunk contains the third, "rural branch minimum balance is ₹2,500". Three supported out of four is 0.75. The gap points at the ingestion pipeline or the semantic search, not at the answer.
The video scores this with RAGAS; DeepEval's metric works the same way: it reads expected_output and retrieval_context, the judge goes through the expected output sentence by sentence and says yes when a sentence can be attributed to a node of the context, and the score is attributable sentences over all sentences.
The ContextualRecallMetric API
from deepeval.metrics import ContextualRecallMetric
metric = ContextualRecallMetric(model=judge, threshold=0.5)
metric.measure(LLMTestCase(
input=...,
expected_output=..., # split into sentences, each one looked for in the chunks
retrieval_context=[...], # the order does not matter here
))
metric.score, metric.verdicts # one yes or no per sentence, with the node it came fromThe video's reference and the chunks without the rural balance
The reference has four sentences, one per claim on the video's slide. The chunks leave out the rural one.
expected = ("Urban branches require a ₹10,000 minimum balance. The non-maintenance fee is ₹350 + taxes "
"when the balance falls below the limit. Rural branch minimum balance is ₹2,500. "
"Semi-urban branch minimum balance is ₹5,000.")
chunks = [
"Minimum balance is ₹10,000 for urban branches.",
"Non-maintenance fee is ₹350 + taxes if the balance falls below the minimum.",
"KYC update is required every 8 years.",
"Semi-urban branch minimum balance is ₹5,000.",
]- written in Custom judge model
- written in TechNest RAG app
- written in TechNest RAG app
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
The video's missing rural balance, run on Groq
from deepeval.metrics import ContextualRecallMetric
from deepeval.test_case import LLMTestCase
from judge import judge
question = "What is the minimum balance for my savings account?"
expected = ("Urban branches require a ₹10,000 minimum balance. The non-maintenance fee is ₹350 + taxes "
"when the balance falls below the limit. Rural branch minimum balance is ₹2,500. "
"Semi-urban branch minimum balance is ₹5,000.")
chunks = [
"Minimum balance is ₹10,000 for urban branches.",
"Non-maintenance fee is ₹350 + taxes if the balance falls below the minimum.",
"KYC update is required every 8 years.",
"Semi-urban branch minimum balance is ₹5,000.",
]
metric = ContextualRecallMetric(model=judge)
metric.measure(LLMTestCase(input=question, expected_output=expected, retrieval_context=chunks))
print(f"recall: {metric.score:.2f}")
for verdict in metric.verdicts:
print(f" {verdict.verdict:3} {verdict.reason}")recall: 0.75 yes 1st node: Minimum balance is ₹10,000 … yes 2nd node: Non-maintenance fee is ₹350 + taxes … no No node mentions rural branch minimum balance yes 4th node: Semi-urban branch minimum balance is ₹5,000 …
The judge went through the four sentences and attributed three of them to a node, naming the node each time. The rural balance gets no: no node mentions it. Three out of four is 0.75, the same score as the video's slide. Each reason names a node by its position, which tells you where a fact came from.
A two-part TechNest question
The question asks about two policies, shipping and returns, so a complete answer needs two catalog entries.
question = "What are TechNest's shipping options and how long do returns take to process?"
expected = ("Standard shipping is free on orders over $50 and takes 3 to 5 business days. Expedited shipping "
"takes 1 to 2 business days for $9.99, and same-day delivery costs $19.99 in select cities. "
"Refunds are processed within 5 to 7 business days of receiving the returned item.")Three chunks, then one
for top_k in [3, 1]:
contexts = retrieve(question, top_k=top_k)
metric.measure(LLMTestCase(input=question, expected_output=expected, retrieval_context=contexts))Recall of the TechNest search at top_k=3 and top_k=1
from deepeval.metrics import ContextualRecallMetric
from deepeval.test_case import LLMTestCase
from judge import judge
from technest import retrieve
question = "What are TechNest's shipping options and how long do returns take to process?"
expected = ("Standard shipping is free on orders over $50 and takes 3 to 5 business days. Expedited shipping "
"takes 1 to 2 business days for $9.99, and same-day delivery costs $19.99 in select cities. "
"Refunds are processed within 5 to 7 business days of receiving the returned item.")
metric = ContextualRecallMetric(model=judge)
for top_k in [3, 1]:
contexts = retrieve(question, top_k=top_k)
metric.measure(LLMTestCase(input=question, expected_output=expected, retrieval_context=contexts))
print(f"top_k={top_k} recall={metric.score:.2f} first chunk: {contexts[0][:45]}...")
for verdict in metric.verdicts:
print(f" {verdict.verdict:3} {verdict.reason}")top_k=3 recall=1.00 first chunk: TechNest accepts returns within 30 days of th... yes 2nd node: "free standard shipping on all orders over $50" and "takes 3 to 5 business days" yes 2nd node: "Expedited shipping (1 to 2 business days) is available for $9.99" and "Same-day delivery ... for $19.99" yes 1st node: "Refunds are processed within 5 to 7 business days of receiving the returned item" top_k=1 recall=0.33 first chunk: TechNest accepts returns within 30 days of th... no No shipping info in 1st node. no No shipping info in 1st node. yes Matches 1st node: "Refunds are processed within 5 to 7 business days of receiving the returned item..."
What top_k=1 cut off
- top_k=3 scores 1.00. The two shipping sentences are found in the 2nd node, the shipping policy, and the refund sentence in the 1st, the return policy.
- top_k=1 scores 0.33. The first chunk is the return policy, so the refund sentence still gets
yes, and both shipping sentences getno: the shipping policy ranked second and was cut off. - The search ranked the right entries. Recall dropped because
top_kkept too few of them, the first lever the docs name for a low recall.
Contextual recall vs faithfulness
Both metrics split a text into parts and look for each part in the chunks. They start from opposite ends.
| Contextual recall | Faithfulness | |
|---|---|---|
| Parts come from | The expected answer | The bot's answer |
| Checked against | The retrieved chunks | The retrieved chunks |
| Grades | The retriever | The generator |
| Needs an expected answer? | Yes | No |
When to use contextual recall
- After changing chunk size,
top_kor the search, to check that nothing an answer needs stopped coming back. - When answers are incomplete and you need to know whether the missing fact was ever retrieved.
Related
- Previous: Contextual precision
- Next: Contextual relevancy
- Reference: Contextual recall
- Add
"Rural branch minimum balance is ₹2,500."to the bankchunksand check that recall reaches 1.00. - Run the TechNest question with
top_k=2: the return and shipping policies both come back. - Ask
"Do you ship to Canada?"with the expected answer"TechNest offers free standard shipping on orders over $50 within the continental US."and read why recall is 0.
You understood something today that you didn't yesterday.