Context precision
Context precision is a RAGAS metric that measures whether the retriever ranks the relevant chunks above the irrelevant ones, as the mean of precision@k over the ranks that hold a relevant chunk.
Last updated: 09 Oct, 2026 · RAGAS 0.4
Faithfulness and Answer relevancy grade the generator. The next two metrics grade the retriever. A retriever returns its chunks in an order, the top-k, and whatever sits high in that list is what the model is most sure to receive. Two retrievers can return the same five chunks, one with the useful ones first and one with noise on top. Context precision is the score that tells them apart.
This part of the video starts at 2:19:01. The judge gives each chunk a yes or no verdict, so the score rewards relevant chunks placed above noise and does not rank two relevant chunks against each other.
Reading the video's five ranked chunks
The golden query is "What is the minimum balance for my savings account?". The retriever returns five chunks, ranked 1 to 5. A judge tags each one relevant or noise: the urban minimum balance and the non-maintenance fee are relevant, "KYC update required every 8 years" at rank 3 is noise, and the semi-urban and rural balances at ranks 4 and 5 are relevant.
Then comes the position penalty. Precision at rank k, written P@k, is the share of relevant chunks among the first k. It is 1/1 at rank 1 and 2/2 at rank 2. The noise at rank 3 pulls it to 2/3. Ranks 4 and 5 are relevant again, but they sit behind the noise, so they only reach 3/4 and 4/5. Noise at rank k lowers the value of every relevant chunk after it: the earlier the noise, the bigger the hit.
Only the ranks that hold a relevant chunk enter the average. The card adds 1.00, 1.00, 0.75 and 0.80, divides by 4 and shows 0.89, above its pass mark of 0.7.
Writing context precision as a formula
This is the average precision of information retrieval, computed on yes or no verdicts. Two things follow from the formula. A verdict has no degrees, so the metric cannot say which of two relevant chunks deserved the higher rank. And the denominator counts the relevant chunks that were retrieved, so a relevant chunk that never came back is not this metric's business; that is what Context recall measures.
The class used below, ContextPrecision, needs the golden's reference: for each chunk the judge is asked whether that chunk was useful in arriving at the reference answer.
Computing precision@k by hand
from ragas.metrics.collections import ContextPrecision
chunks = ["Min balance ₹10,000 urban branches", "Non-maintenance fee ₹350 + taxes",
"KYC update required every 8 years", "Semi-urban balance ₹5,000", "Rural balance ₹2,500"]
verdicts = [1, 1, 0, 1, 1] # 1 = relevant, 0 = noise, as tagged on the card
def average_precision(verdicts):
total, hits = 0.0, 0
for k, verdict in enumerate(verdicts, 1):
hits += verdict
total += (hits / k) * verdict # precision@k counts only at a relevant rank
return total / sum(verdicts) if sum(verdicts) else 0.0
hits = 0
for k, (chunk, verdict) in enumerate(zip(chunks, verdicts), 1):
hits += verdict
note = "" if verdict else " (noise: not counted)"
print(f"rank {k} P@{k} = {hits}/{k} = {hits / k:.2f} {chunk}{note}")
print("context precision:", average_precision(verdicts))
print("the library's own helper:", ContextPrecision._calculate_average_precision(None, verdicts))
print("plain share of relevant chunks:", sum(verdicts) / len(verdicts))rank 1 P@1 = 1/1 = 1.00 Min balance ₹10,000 urban branches rank 2 P@2 = 2/2 = 1.00 Non-maintenance fee ₹350 + taxes rank 3 P@3 = 2/3 = 0.67 KYC update required every 8 years (noise: not counted) rank 4 P@4 = 3/4 = 0.75 Semi-urban balance ₹5,000 rank 5 P@5 = 4/5 = 0.80 Rural balance ₹2,500 context precision: 0.8875 the library's own helper: 0.8874999999778125 plain share of relevant chunks: 0.8
What the five ranks add up to
- P@1 to P@5 are 1.00, 1.00, 0.67, 0.75 and 0.80, the five values on the card's bars.
- The score is 0.8875, the mean of the four counted values, which the card rounds to 0.89.
- The library's own helper prints 0.8874999999778125. RAGAS adds a tiny 1e-10 to the denominator to avoid a division by zero, so its results sit a hair under the exact value. The helper is internal, so its name can change in a later release; it is called here only to check the arithmetic.
- The plain share of relevant chunks is 0.8, four out of five. Context precision is higher because the relevant chunks hold the top two ranks.
Moving the noise chunk through the ranking
The same five chunks, with the one noise chunk placed at each rank in turn, show what the metric rewards.
import matplotlib.pyplot as plt
def average_precision(verdicts):
total, hits = 0.0, 0
for k, verdict in enumerate(verdicts, 1):
hits += verdict
total += (hits / k) * verdict # precision@k counts only at a relevant rank
return total / sum(verdicts) if sum(verdicts) else 0.0
ranks = [1, 2, 3, 4, 5]
scores = []
for noise_rank in ranks:
verdicts = [0 if k == noise_rank else 1 for k in ranks]
scores.append(average_precision(verdicts))
print(f"noise at rank {noise_rank}: verdicts {verdicts} -> {scores[-1]:.4f}")
print("plain share of relevant chunks, in every order:", 4 / 5)
for verdicts in ([1, 0], [1, 1], [0, 1]):
print(f"two chunks {verdicts}: {average_precision(verdicts):.2f}")
plt.figure(figsize=(7, 3.8))
bars = plt.bar(ranks, scores, color=["#d64541" if rank == 3 else "#3a6fd8" for rank in ranks])
plt.bar_label(bars, fmt="%.2f", padding=4)
plt.axhline(0.8, color="gray", linestyle="--", zorder=0, label="plain share of relevant chunks: 4/5 = 0.80")
plt.title("Context precision when the one noise chunk moves")
plt.xlabel("Rank of the noise chunk (red: the card's order)")
plt.ylabel("Context precision")
plt.ylim(0, 1.15)
plt.legend(loc="lower right")
plt.show()noise at rank 1: verdicts [0, 1, 1, 1, 1] -> 0.6792 noise at rank 2: verdicts [1, 0, 1, 1, 1] -> 0.8042 noise at rank 3: verdicts [1, 1, 0, 1, 1] -> 0.8875 noise at rank 4: verdicts [1, 1, 1, 0, 1] -> 0.9500 noise at rank 5: verdicts [1, 1, 1, 1, 0] -> 1.0000 plain share of relevant chunks, in every order: 0.8 two chunks [1, 0]: 1.00 two chunks [1, 1]: 1.00 two chunks [0, 1]: 0.50
- Noise at rank 1 scores 0.6792, noise at rank 5 scores 1.0000. The set of chunks is the same; only the order changed.
- The card's order, noise at rank 3, gives 0.8875, the red bar.
- Noise after the last relevant chunk costs nothing. With the noise at rank 5 the score is a full 1.0000, so a perfect score does not mean that every retrieved chunk was useful.
- The plain share stays 0.8 in every order. It cannot see ranking at all.
- With two chunks, [1, 0] and [1, 1] both score 1.00. A short list hides noise at its end.
Scoring the five chunks with RAGAS
The video's app ran llama-3.1-8b-instant as its judge, since retired on Groq; the run below uses openai/gpt-oss-20b from LLM as a judge. The card shows the query and the chunks but no reference, so the reference below is written from the four chunks the card tags relevant.
from ragas.metrics.collections import ContextPrecision
result = ContextPrecision(llm=judge).score(user_input=question, reference=reference, retrieved_contexts=chunks)
result.value # mean of precision@k over the relevant ranksThe metric calls the judge once per chunk. The wrapper from Faithfulness keeps the five replies, each a verdict (1 useful, 0 not) with the judge's reason, so the example can print them next to the chunks.
import os
from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import ContextPrecision
client = AsyncOpenAI(
base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"],
max_retries=6, # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)
replies = [] # each structured reply the judge sends back to the metric
ask = judge.agenerate
async def keep(prompt, response_model):
reply = await ask(prompt, response_model)
replies.append(reply)
return reply
judge.agenerate = keep
question = "What is the minimum balance for my savings account?"
reference = ("Urban branches require a ₹10,000 minimum balance. The non-maintenance fee is ₹350 + taxes when the "
"balance falls below the limit. Semi-urban branches require ₹5,000 and rural branches ₹2,500.")
chunks = [
"Minimum balance is ₹10,000 for urban branches.",
"Non-maintenance fee is ₹350 + taxes.",
"KYC update is required every 8 years.",
"Semi-urban branch minimum balance is ₹5,000.",
"Rural branch minimum balance is ₹2,500.",
]
result = ContextPrecision(llm=judge).score(user_input=question, reference=reference, retrieved_contexts=chunks)
for rank, (chunk, reply) in enumerate(zip(chunks, replies), 1):
print(rank, reply.verdict, "|", chunk)
print(" ", reply.reason)
print("context precision:", round(result.value, 4))1 1 | Minimum balance is ₹10,000 for urban branches.
The context supplied the key fact that urban branches require a ₹10,000 minimum balance, which is directly reflected in the answer. The answer also adds additional details about fees and other branch types, but the core information about the minimum balance comes from the context.
2 0 | Non-maintenance fee is ₹350 + taxes.
The context only states the non‑maintenance fee of ₹350 + taxes, which does not provide any information about the minimum balance required for a savings account. Therefore the context was not useful for determining the minimum balance mentioned in the answer.
3 0 | KYC update is required every 8 years.
The context about KYC updates does not provide any information about minimum balance requirements for savings accounts, so it was not useful in arriving at the answer.
4 1 | Semi-urban branch minimum balance is ₹5,000.
The context supplied the minimum balance for semi‑urban branches (₹5,000), which is included in the answer, making it useful.
5 1 | Rural branch minimum balance is ₹2,500.
The context supplied the rural branch minimum balance of ₹2,500, which is explicitly included in the answer. Thus the context was useful in arriving at the answer.
context precision: 0.7Why this judge scored 0.7 and not 0.89
- The judge marked two chunks as not useful, not one. Ranks 1, 4 and 5 got a 1. The KYC chunk at rank 3 got a 0, as on the card. The non-maintenance fee at rank 2 got a 0 as well.
- Its reason for rank 2 is a narrow reading of the question: the fee chunk "does not provide any information about the minimum balance required for a savings account". The card tags the fee as relevant, and the reference mentions it, but this judge tied usefulness to the minimum balance alone.
- The score follows from those verdicts. With useful chunks at ranks 1, 4 and 5, the counted values are 1/1, 2/4 and 3/5, and the example below confirms their mean.
def average_precision(verdicts):
total, hits = 0.0, 0
for k, verdict in enumerate(verdicts, 1):
hits += verdict
total += (hits / k) * verdict # precision@k counts only at a relevant rank
return total / sum(verdicts) if sum(verdicts) else 0.0
print("the card's verdicts [1, 1, 0, 1, 1]:", round(average_precision([1, 1, 0, 1, 1]), 4))
print("this judge's verdicts [1, 0, 0, 1, 1]:", round(average_precision([1, 0, 0, 1, 1]), 4))the card's verdicts [1, 1, 0, 1, 1]: 0.8875 this judge's verdicts [1, 0, 0, 1, 1]: 0.7
The same function gives 0.8875 for the card's verdicts and 0.7 for this judge's. The arithmetic is fixed; the verdicts are one judge's reading. The card's 0.89 is the result for the card's tags, not a number every run reproduces. So a low context precision is worth opening before acting on it: sometimes the retriever ranked badly, and sometimes the judge and you disagree about what counts as relevant.
Context precision vs the plain share of relevant chunks
| Context precision | Plain share of relevant chunks | |
|---|---|---|
| What it measures | Whether relevant chunks are ranked above noise | How many of the retrieved chunks are relevant |
| Uses the order | Yes | No |
| The card's five chunks | 0.89 | 0.80 |
| Same chunks, noise moved to rank 1 | 0.68 | 0.80 |
| Noise after the last relevant chunk | No penalty | Counts against the score |
Where you use context precision
- When you add or tune a reranker. A reranker reorders the retrieved chunks, which is the thing this metric grades. The video's advice: if context precision is low, try a reranker.
- When you shrink top-k. If the useful chunks are at the top, a smaller k loses nothing; a low score warns that cutting k would cut useful chunks.
- When answers use the wrong product or policy although the right chunk was retrieved further down.
Related
- Previous: Answer relevancy
- Next: Context recall
- Reference: Context precision in the RAGAS docs
- In the hand computation, change
verdictsto[0, 0, 1, 1, 1]: two noise chunks on top give P@3 = 0.33, P@4 = 0.50 and P@5 = 0.60, and the score drops to 0.4778. - In the same example, change
verdictsto[1, 1, 1, 1, 0]: every counted P@k is 1.00 and the score is 1.0, with a noise chunk still in the list. - In the RAGAS example, move the KYC chunk to the end of
chunksand run it again: the printed verdicts move with it, and the score is the mean of P@k over the new positions of the 1s.
Little by little, you're building something great.