DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

Contextual precision

ContextualPrecisionMetric is a DeepEval RAG metric that scores whether the retriever ranks the chunks that are useful for the expected answer above the ones that are not, as a weighted cumulative precision from 0 to 1.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

Faithfulness and Answer relevancy grade the generator. The next three metrics grade the retriever, starting with the order of the chunks it returns, because the model leans hardest on the ones at the top.

Context precision and the ranking penalty · from the Complete AI Security Course In 8 Hours · 2:18:57 to 2:24:42

Ranking the useful chunks first

The video's golden query is "What is the minimum balance for my savings account?", and retrieval returns ranked chunks, the top-k. You want the most relevant chunk first and the least relevant last, which is also what a reranker does, so a reranker lifts this metric. The judge, a Llama 3.1 model in the video, marks each chunk relevant or noise: the minimum balance and the non-maintenance fee are relevant, "KYC update required every 8 years" at rank 3 is noise, and the chunks at ranks 4 and 5 are relevant.

Then the position penalty. Noise at the top is the worst case; noise at the bottom is acceptable. At each rank the score is the share of relevant chunks so far: 1 and 1, then 0.67 at the noise, and the relevant chunk at rank 4 gets 0.75 instead of 1 because it comes after the noise, then 0.8. Only the relevant positions count: (1 + 1 + 0.75 + 0.8) / 4, about 0.89. A low context precision tells you to add reranking.

The video scores this with RAGAS; DeepEval's metric works the same way: the judge marks each node of retrieval_context relevant or not by comparing it with expected_output, and the score is that same average of precision at each relevant position.

The ContextualPrecisionMetric API

python
from deepeval.metrics import ContextualPrecisionMetric

metric = ContextualPrecisionMetric(model=judge, threshold=0.5)
metric.measure(LLMTestCase(
    input=...,
    expected_output=...,          # what each chunk is judged against
    retrieval_context=[...],      # best first: the order is what gets scored
))
metric.score, metric.verdicts     # one yes or no per chunk, in rank order

The judge never sees the bot's answer. A chunk counts as useful when it helps produce the expected answer.

The video's five ranked chunks

The chunks and the expected answer use the facts on the video's slide.

python
chunks = [
    "Minimum balance is ₹10,000 for urban branches.",
    "Non-maintenance fee is ₹350 + taxes if the balance falls below the minimum.",
    "KYC update is required every 8 years.",
    "Semi-urban branch minimum balance is ₹5,000.",
    "Rural branch minimum balance is ₹2,500.",
]

The same chunks with the noise first

python
noise_first = [chunks[2]] + chunks[:2] + chunks[3:]  # KYC chunk moved to rank 1
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
judge.py
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
technest.py
import json
import os
import re

from openai import OpenAI

groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"

SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""

with open("catalog.json", encoding="utf-8") as f:
    CATALOG = json.load(f)

SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
        "long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}


def words(text):
    """The words in a text that carry meaning, in lower case."""
    return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}


def retrieve(question, top_k=3):
    asked = words(question)
    ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
    return [item["content"] for item in ranked[:top_k]]


def generate(question, contexts):
    context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
    ]
    response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
    return response.choices[0].message.content.strip()


def answer(question, top_k=3):
    contexts = retrieve(question, top_k)
    return generate(question, contexts), contexts
catalog.json
[
  {
    "id": "prod_001",
    "category": "product",
    "title": "ProBook X1 Laptop",
    "content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
  },
  {
    "id": "prod_002",
    "category": "product",
    "title": "PixelPhone 15",
    "content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
  },
  {
    "id": "prod_003",
    "category": "product",
    "title": "SoundPods Pro",
    "content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
  },
  {
    "id": "prod_004",
    "category": "product",
    "title": "UltraTab S2 Tablet",
    "content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
  },
  {
    "id": "prod_005",
    "category": "product",
    "title": "SmartWatch X",
    "content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
  },
  {
    "id": "prod_006",
    "category": "product",
    "title": "ProCam 4K Action Camera",
    "content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
  },
  {
    "id": "prod_007",
    "category": "product",
    "title": "BassBuds Max Headphones",
    "content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
  },
  {
    "id": "prod_008",
    "category": "product",
    "title": "SoundBar 360",
    "content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
  },
  {
    "id": "policy_001",
    "category": "policy",
    "title": "Return Policy",
    "content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
  },
  {
    "id": "policy_002",
    "category": "policy",
    "title": "Shipping Policy",
    "content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
  },
  {
    "id": "policy_003",
    "category": "policy",
    "title": "Warranty Policy",
    "content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
  },
  {
    "id": "policy_004",
    "category": "policy",
    "title": "Payment Policy",
    "content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
  },
  {
    "id": "faq_001",
    "category": "faq",
    "title": "Order Tracking",
    "content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
  },
  {
    "id": "faq_002",
    "category": "faq",
    "title": "International Shipping",
    "content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
  },
  {
    "id": "faq_003",
    "category": "faq",
    "title": "Bulk and Business Orders",
    "content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
  }
]

The video's ranking, run on Groq

ExampleAPI keyFrom the video, run on Groq
from deepeval.metrics import ContextualPrecisionMetric
from deepeval.test_case import LLMTestCase
from judge import judge

question = "What is the minimum balance for my savings account?"
expected = ("Urban branches require a ₹10,000 minimum balance. The non-maintenance fee is ₹350 + taxes "
            "when the balance falls below the limit. Semi-urban branches require ₹5,000 and rural "
            "branches ₹2,500.")
chunks = [
    "Minimum balance is ₹10,000 for urban branches.",
    "Non-maintenance fee is ₹350 + taxes if the balance falls below the minimum.",
    "KYC update is required every 8 years.",
    "Semi-urban branch minimum balance is ₹5,000.",
    "Rural branch minimum balance is ₹2,500.",
]
noise_first = [chunks[2]] + chunks[:2] + chunks[3:]

metric = ContextualPrecisionMetric(model=judge)
for label, order in [("video's order", chunks), ("noise first", noise_first)]:
    metric.measure(LLMTestCase(input=question, expected_output=expected, retrieval_context=order))
    verdicts = " ".join(v.verdict for v in metric.verdicts)
    print(f"{metric.score:.4f}  {label:14} {verdicts}")

In the video's order the judge marks the KYC chunk no and the other four yes, and the score is 0.8875, the slide's 0.89: DeepEval computes the same (1 + 1 + 0.75 + 0.8) / 4. With the KYC chunk moved to rank 1, the verdicts are the same set in a new order and the score drops to 0.6792: every useful chunk now sits one place lower, behind the noise. The chunks did not change; only their ranking did.

Ranking the SoundPods Pro battery chunks

The TechNest search ranks by shared words. For the battery question the ProBook X1 entry shares two words with the question, battery and life, and the SoundPods entry shares two others, soundpods and pro. On a tie the catalog order wins, so the laptop comes first. The run scores the search's real order and the same three chunks with the SoundPods entry moved to the top.

ExampleAPI key
from deepeval.metrics import ContextualPrecisionMetric
from deepeval.test_case import LLMTestCase
from judge import judge
from technest import retrieve

question = "How long is the battery life on the SoundPods Pro?"
expected = "The SoundPods Pro give 8 hours of playback per charge, plus 24 hours with the charging case."
contexts = retrieve(question)
for rank, chunk in enumerate(contexts, 1):
    print(f"{rank}. {chunk[:55]}...")

metric = ContextualPrecisionMetric(model=judge)
reordered = [contexts[1], contexts[0], contexts[2]]
for label, order in [("search order", contexts), ("SoundPods first", reordered)]:
    metric.measure(LLMTestCase(input=question, expected_output=expected, retrieval_context=order))
    verdicts = " ".join(v.verdict for v in metric.verdicts)
    print(f"{metric.score:.2f}  {label:16} {verdicts}")
print(metric.reason)

What the two rankings scored

  • The search order scores 0.50. The verdicts are no yes no: the ProBook X1 at rank 1 is not useful for a SoundPods question, the SoundPods entry at rank 2 is, so the only useful chunk gets 1/2.
  • SoundPods first scores 1.00. The same three chunks with the useful one at rank 1. The two laptop and watch chunks below it cost nothing, because precision only scores the positions that hold useful chunks.
  • The fix is in the ranking, not in what was retrieved. The word-overlap search ties the two entries and keeps catalog order; a reranker or an embedding search would put the SoundPods entry first.

Contextual precision vs contextual recall

Contextual precisionContextual recall
AsksAre the useful chunks at the top?Did every part of the expected answer come back?
Judges eachChunk, in rank orderSentence of the expected answer
Changes when you reorder the chunks?YesNo
Typical fixA reranker, a better searchA higher top_k, better chunking

When to use contextual precision

  • When you add or tune a reranker, to measure whether it moved the useful chunks up.
  • When answers mix up products: the right chunk was retrieved, but a similar one ranked above it.
Watch out. Precision cannot see a chunk that never came back. If the search returns one useful chunk at rank 1 and misses another one entirely, precision is still 1.0. Contextual recall catches the missing chunk.
Try it yourself
  • Move the KYC chunk to the end of the bank list, delete the noise_first row from the loop, and predict the score before you run it.
  • Score retrieve(question, top_k=2) for the battery question: the ProBook X1 and the SoundPods Pro, in that order, and change reordered to [contexts[1], contexts[0]].
  • Change expected to "The ProBook X1 battery lasts 12 hours." and see which chunk becomes the useful one.

Little by little, you're building something great.