AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Context precision

Context precision is a RAGAS metric that measures whether the retriever ranks the relevant chunks above the irrelevant ones, as the mean of precision@k over the ranks that hold a relevant chunk.

Last updated: 09 Oct, 2026 · RAGAS 0.4

Faithfulness and Answer relevancy grade the generator. The next two metrics grade the retriever. A retriever returns its chunks in an order, the top-k, and whatever sits high in that list is what the model is most sure to receive. Two retrievers can return the same five chunks, one with the useful ones first and one with noise on top. Context precision is the score that tells them apart.

Context precision on five ranked chunks · from the Complete AI Security Course in 8 Hours video · 2:19:01 to 2:22:13

This part of the video starts at 2:19:01. The judge gives each chunk a yes or no verdict, so the score rewards relevant chunks placed above noise and does not rank two relevant chunks against each other.

Reading the video's five ranked chunks

The golden query is "What is the minimum balance for my savings account?". The retriever returns five chunks, ranked 1 to 5. A judge tags each one relevant or noise: the urban minimum balance and the non-maintenance fee are relevant, "KYC update required every 8 years" at rank 3 is noise, and the semi-urban and rural balances at ranks 4 and 5 are relevant.

Then comes the position penalty. Precision at rank k, written P@k, is the share of relevant chunks among the first k. It is 1/1 at rank 1 and 2/2 at rank 2. The noise at rank 3 pulls it to 2/3. Ranks 4 and 5 are relevant again, but they sit behind the noise, so they only reach 3/4 and 4/5. Noise at rank k lowers the value of every relevant chunk after it: the earlier the noise, the bigger the hit.

Five ranked chunks for the query about the minimum balance: ranks 1, 2, 4 and 5 are relevant and rank 3, the KYC chunk, is noise. Precision at each rank is 1.00, 1.00, 0.67, 0.75 and 0.80; the 0.67 of the noise rank is not counted, and the mean of the other four is 0.8875, shown as 0.89.

Only the ranks that hold a relevant chunk enter the average. The card adds 1.00, 1.00, 0.75 and 0.80, divides by 4 and shows 0.89, above its pass mark of 0.7.

Writing context precision as a formula

v_k is 1 when the chunk at rank k is relevant and 0 when it is not

This is the average precision of information retrieval, computed on yes or no verdicts. Two things follow from the formula. A verdict has no degrees, so the metric cannot say which of two relevant chunks deserved the higher rank. And the denominator counts the relevant chunks that were retrieved, so a relevant chunk that never came back is not this metric's business; that is what Context recall measures.

The class used below, ContextPrecision, needs the golden's reference: for each chunk the judge is asked whether that chunk was useful in arriving at the reference answer.

Computing precision@k by hand

ExampleThe video's example, computed in plain Python
from ragas.metrics.collections import ContextPrecision

chunks = ["Min balance ₹10,000 urban branches", "Non-maintenance fee ₹350 + taxes",
          "KYC update required every 8 years", "Semi-urban balance ₹5,000", "Rural balance ₹2,500"]
verdicts = [1, 1, 0, 1, 1]  # 1 = relevant, 0 = noise, as tagged on the card


def average_precision(verdicts):
    total, hits = 0.0, 0
    for k, verdict in enumerate(verdicts, 1):
        hits += verdict
        total += (hits / k) * verdict  # precision@k counts only at a relevant rank
    return total / sum(verdicts) if sum(verdicts) else 0.0


hits = 0
for k, (chunk, verdict) in enumerate(zip(chunks, verdicts), 1):
    hits += verdict
    note = "" if verdict else "  (noise: not counted)"
    print(f"rank {k}  P@{k} = {hits}/{k} = {hits / k:.2f}  {chunk}{note}")
print("context precision:", average_precision(verdicts))
print("the library's own helper:", ContextPrecision._calculate_average_precision(None, verdicts))
print("plain share of relevant chunks:", sum(verdicts) / len(verdicts))

What the five ranks add up to

  • P@1 to P@5 are 1.00, 1.00, 0.67, 0.75 and 0.80, the five values on the card's bars.
  • The score is 0.8875, the mean of the four counted values, which the card rounds to 0.89.
  • The library's own helper prints 0.8874999999778125. RAGAS adds a tiny 1e-10 to the denominator to avoid a division by zero, so its results sit a hair under the exact value. The helper is internal, so its name can change in a later release; it is called here only to check the arithmetic.
  • The plain share of relevant chunks is 0.8, four out of five. Context precision is higher because the relevant chunks hold the top two ranks.

Moving the noise chunk through the ranking

The same five chunks, with the one noise chunk placed at each rank in turn, show what the metric rewards.

ExampleRun on matplotlib 3.11.2
import matplotlib.pyplot as plt


def average_precision(verdicts):
    total, hits = 0.0, 0
    for k, verdict in enumerate(verdicts, 1):
        hits += verdict
        total += (hits / k) * verdict  # precision@k counts only at a relevant rank
    return total / sum(verdicts) if sum(verdicts) else 0.0


ranks = [1, 2, 3, 4, 5]
scores = []
for noise_rank in ranks:
    verdicts = [0 if k == noise_rank else 1 for k in ranks]
    scores.append(average_precision(verdicts))
    print(f"noise at rank {noise_rank}: verdicts {verdicts} -> {scores[-1]:.4f}")
print("plain share of relevant chunks, in every order:", 4 / 5)
for verdicts in ([1, 0], [1, 1], [0, 1]):
    print(f"two chunks {verdicts}: {average_precision(verdicts):.2f}")

plt.figure(figsize=(7, 3.8))
bars = plt.bar(ranks, scores, color=["#d64541" if rank == 3 else "#3a6fd8" for rank in ranks])
plt.bar_label(bars, fmt="%.2f", padding=4)
plt.axhline(0.8, color="gray", linestyle="--", zorder=0, label="plain share of relevant chunks: 4/5 = 0.80")
plt.title("Context precision when the one noise chunk moves")
plt.xlabel("Rank of the noise chunk (red: the card's order)")
plt.ylabel("Context precision")
plt.ylim(0, 1.15)
plt.legend(loc="lower right")
plt.show()
A bar chart of context precision for five chunks with the one noise chunk placed at ranks 1 to 5: 0.68, 0.80, 0.89, 0.95 and 1.00. The bar for rank 3, the card's order, is red. A dashed line marks the plain share of relevant chunks, 0.80, which is the same in every order.
  • Noise at rank 1 scores 0.6792, noise at rank 5 scores 1.0000. The set of chunks is the same; only the order changed.
  • The card's order, noise at rank 3, gives 0.8875, the red bar.
  • Noise after the last relevant chunk costs nothing. With the noise at rank 5 the score is a full 1.0000, so a perfect score does not mean that every retrieved chunk was useful.
  • The plain share stays 0.8 in every order. It cannot see ranking at all.
  • With two chunks, [1, 0] and [1, 1] both score 1.00. A short list hides noise at its end.

Scoring the five chunks with RAGAS

The video's app ran llama-3.1-8b-instant as its judge, since retired on Groq; the run below uses openai/gpt-oss-20b from LLM as a judge. The card shows the query and the chunks but no reference, so the reference below is written from the four chunks the card tags relevant.

python
from ragas.metrics.collections import ContextPrecision

result = ContextPrecision(llm=judge).score(user_input=question, reference=reference, retrieved_contexts=chunks)
result.value  # mean of precision@k over the relevant ranks

The metric calls the judge once per chunk. The wrapper from Faithfulness keeps the five replies, each a verdict (1 useful, 0 not) with the judge's reason, so the example can print them next to the chunks.

ExampleAPI keyFrom the video, run on Groq
import os

from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import ContextPrecision

client = AsyncOpenAI(
    base_url="https://api.groq.com/openai/v1",
    api_key=os.environ["GROQ_API_KEY"],
    max_retries=6,  # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)

replies = []  # each structured reply the judge sends back to the metric
ask = judge.agenerate


async def keep(prompt, response_model):
    reply = await ask(prompt, response_model)
    replies.append(reply)
    return reply


judge.agenerate = keep

question = "What is the minimum balance for my savings account?"
reference = ("Urban branches require a ₹10,000 minimum balance. The non-maintenance fee is ₹350 + taxes when the "
             "balance falls below the limit. Semi-urban branches require ₹5,000 and rural branches ₹2,500.")
chunks = [
    "Minimum balance is ₹10,000 for urban branches.",
    "Non-maintenance fee is ₹350 + taxes.",
    "KYC update is required every 8 years.",
    "Semi-urban branch minimum balance is ₹5,000.",
    "Rural branch minimum balance is ₹2,500.",
]

result = ContextPrecision(llm=judge).score(user_input=question, reference=reference, retrieved_contexts=chunks)
for rank, (chunk, reply) in enumerate(zip(chunks, replies), 1):
    print(rank, reply.verdict, "|", chunk)
    print("     ", reply.reason)
print("context precision:", round(result.value, 4))

Why this judge scored 0.7 and not 0.89

  • The judge marked two chunks as not useful, not one. Ranks 1, 4 and 5 got a 1. The KYC chunk at rank 3 got a 0, as on the card. The non-maintenance fee at rank 2 got a 0 as well.
  • Its reason for rank 2 is a narrow reading of the question: the fee chunk "does not provide any information about the minimum balance required for a savings account". The card tags the fee as relevant, and the reference mentions it, but this judge tied usefulness to the minimum balance alone.
  • The score follows from those verdicts. With useful chunks at ranks 1, 4 and 5, the counted values are 1/1, 2/4 and 3/5, and the example below confirms their mean.
ExampleBoth sets of verdicts through the same arithmetic
def average_precision(verdicts):
    total, hits = 0.0, 0
    for k, verdict in enumerate(verdicts, 1):
        hits += verdict
        total += (hits / k) * verdict  # precision@k counts only at a relevant rank
    return total / sum(verdicts) if sum(verdicts) else 0.0


print("the card's verdicts  [1, 1, 0, 1, 1]:", round(average_precision([1, 1, 0, 1, 1]), 4))
print("this judge's verdicts [1, 0, 0, 1, 1]:", round(average_precision([1, 0, 0, 1, 1]), 4))

The same function gives 0.8875 for the card's verdicts and 0.7 for this judge's. The arithmetic is fixed; the verdicts are one judge's reading. The card's 0.89 is the result for the card's tags, not a number every run reproduces. So a low context precision is worth opening before acting on it: sometimes the retriever ranked badly, and sometimes the judge and you disagree about what counts as relevant.

Context precision vs the plain share of relevant chunks

Context precisionPlain share of relevant chunks
What it measuresWhether relevant chunks are ranked above noiseHow many of the retrieved chunks are relevant
Uses the orderYesNo
The card's five chunks0.890.80
Same chunks, noise moved to rank 10.680.80
Noise after the last relevant chunkNo penaltyCounts against the score

Where you use context precision

  • When you add or tune a reranker. A reranker reorders the retrieved chunks, which is the thing this metric grades. The video's advice: if context precision is low, try a reranker.
  • When you shrink top-k. If the useful chunks are at the top, a smaller k loses nothing; a low score warns that cutting k would cut useful chunks.
  • When answers use the wrong product or policy although the right chunk was retrieved further down.
Watch out. A score of 1.00 over a short list says little. The video's app shows 1.00 for all five goldens, but its metric code hands the judge only the first two chunks of each retrieval, and with two chunks both [1, 0] and [1, 1] score 1.00. Score the same list the generator receives, and read context precision together with context recall.
Try it yourself
  • In the hand computation, change verdicts to [0, 0, 1, 1, 1]: two noise chunks on top give P@3 = 0.33, P@4 = 0.50 and P@5 = 0.60, and the score drops to 0.4778.
  • In the same example, change verdicts to [1, 1, 1, 1, 0]: every counted P@k is 1.00 and the score is 1.0, with a noise chunk still in the list.
  • In the RAGAS example, move the KYC chunk to the end of chunks and run it again: the printed verdicts move with it, and the score is the mean of P@k over the new positions of the 1s.

Little by little, you're building something great.