Faithfulness
Faithfulness is a RAGAS metric that measures the share of claims in an answer that the retrieved context supports.
Last updated: 09 Oct, 2026 · RAGAS 0.4
RAGAS metrics listed the five scores the video's app prints for every golden. Faithfulness is the first one it computes. It answers one question: did the answer stay inside the chunks the retriever handed over, or did the model add something of its own? A claim with no chunk behind it is a hallucination, and this score counts them.
This part of the video starts at 2:05:53. On the whiteboard card it shows, three of the four claims are grounded and one is not, so the score is 3 / 4 = 0.75.
Reading the video's bank example
The example is a bank's help bot. The retriever returned two chunks from the knowledge base:
- Chunk 1: "Minimum balance is ₹10,000 for urban branches. Non-maintenance fee is ₹350 + GST if balance falls below."
- Chunk 2: "Semi-urban branch minimum balance is ₹5,000. Rural branch minimum balance is ₹2,500."
From them the model wrote this answer: "The minimum balance for urban branches is ₹10,000. The non-maintenance fee charged is ₹350 + GST. You can also transfer funds online at no extra charge. For rural branches the minimum is ₹2,500."
The judge works in two steps. First it breaks the answer into atomic claims, short statements that each carry one fact. Then it asks the card's judge question about every claim: "Can this claim be fully inferred from the retrieved context, yes or no?" A yes makes the claim grounded. A no makes it hallucinated.
The balance and the fee are in chunk 1 and the rural figure is in chunk 2. The sentence about online transfers is in neither chunk: the model added it. It may even be true for this bank, but nothing the retriever returned says so.
Writing faithfulness as a formula
Every claim weighs the same, and the score runs from 0 to 1. The metric reads the response and the retrieved_contexts, plus the user_input, which helps the judge split the answer into claims. It never reads the golden's reference, so it also works on live traffic where no reference exists.
Counting the grounded claims by hand
With the verdicts of the card typed in, the score is a count and a division. No model is needed for this part.
claims = [
("Min balance is ₹10,000 for urban branches", "Chunk 1"),
("Non-maintenance fee is ₹350 + GST", "Chunk 1"),
("Online fund transfer has no extra charge", None), # in no chunk
("Rural branch minimum is ₹2,500", "Chunk 2"),
]
grounded = sum(source is not None for claim, source in claims)
score = grounded / len(claims)
for claim, source in claims:
verdict = "grounded " if source else "hallucinated"
print(verdict, "|", claim, "|", source or "not in any chunk")
print(f"faithfulness = {grounded} / {len(claims)} = {score}")
for pass_mark in (0.8, 0.7):
print(f"pass mark {pass_mark}: {'pass' if score >= pass_mark else 'fail'}")grounded | Min balance is ₹10,000 for urban branches | Chunk 1 grounded | Non-maintenance fee is ₹350 + GST | Chunk 1 hallucinated | Online fund transfer has no extra charge | not in any chunk grounded | Rural branch minimum is ₹2,500 | Chunk 2 faithfulness = 3 / 4 = 0.75 pass mark 0.8: fail pass mark 0.7: pass
What the count shows
- Three of four claims are grounded, so faithfulness is 3 / 4 = 0.75.
- At a pass mark of 0.8 the answer fails; at 0.7 it passes. The card uses 0.8.
- The pass mark is yours to set. The video calls it a hyperparameter: you decide how strict the check is for your use case. RAGAS returns only the number.
Seeing what one unsupported claim costs
The same single slip costs a short answer more than a long one, which matters when you pick a pass mark.
for total in (2, 4, 7, 10):
print(f"{total - 1} of {total} claims supported: {(total - 1) / total:.2f}")1 of 2 claims supported: 0.50 3 of 4 claims supported: 0.75 6 of 7 claims supported: 0.86 9 of 10 claims supported: 0.90
One unsupported claim out of two halves the score, while one out of ten leaves 0.90. The Results screen of the video's app shows 0.86 for its return-policy golden, the value that six supported claims out of seven give.
Scoring the same answer with RAGAS
The video's app ran llama-3.1-8b-instant as its judge, since retired on Groq; the run below uses openai/gpt-oss-20b, set up as in LLM as a judge.
The judge
client = AsyncOpenAI(
base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"],
max_retries=6, # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)temperature=0 keeps the verdicts as steady as the model allows, and max_tokens=4096 leaves room for the structured reply of a reasoning model. max_retries=6 makes the client wait and try again when Groq's free tier answers that the per-minute token limit is used up.
The metric call
from ragas.metrics.collections import Faithfulness
result = Faithfulness(llm=judge).score(user_input=question, response=response, retrieved_contexts=chunks)
result.value # supported claims / all claimsKeeping the judge's replies
score() returns only the number. The claims and verdicts that the card shows are in the judge's replies, which the metric reads and then drops. Each of the five metrics calls judge.agenerate, so wrapping that one method keeps each reply in a list. The reply objects and their field names are internals of RAGAS 0.4 and can change in a later release.
replies = [] # each structured reply the judge sends back to the metric
ask = judge.agenerate
async def keep(prompt, response_model):
reply = await ask(prompt, response_model)
replies.append(reply)
return reply
judge.agenerate = keepFaithfulness calls the judge twice. The first reply holds the claims, the second holds the same claims with a verdict and a reason each, so replies[-1] is the one to print. The example needs the GROQ_API_KEY from Installing Python for AI security in the environment.
import os
from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import Faithfulness
client = AsyncOpenAI(
base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"],
max_retries=6, # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)
replies = [] # each structured reply the judge sends back to the metric
ask = judge.agenerate
async def keep(prompt, response_model):
reply = await ask(prompt, response_model)
replies.append(reply)
return reply
judge.agenerate = keep
question = "What is the minimum balance for my savings account?"
chunks = [
"Minimum balance is ₹10,000 for urban branches. Non-maintenance fee is ₹350 + GST if balance falls below.",
"Semi-urban branch minimum balance is ₹5,000. Rural branch minimum balance is ₹2,500.",
]
response = ("The minimum balance for urban branches is ₹10,000. The non-maintenance fee charged is ₹350 + GST. "
"You can also transfer funds online at no extra charge. For rural branches the minimum is ₹2,500.")
result = Faithfulness(llm=judge).score(user_input=question, response=response, retrieved_contexts=chunks)
for item in replies[-1].statements:
print(item.verdict, "|", item.statement)
if item.verdict == 0:
print(" reason:", item.reason)
print("faithfulness:", result.value)
print("PASS" if result.value >= 0.8 else "FAIL", "at the card's pass mark of 0.8")1 | The minimum balance for urban branches is ₹10,000.
1 | The non-maintenance fee charged is ₹350 plus GST.
0 | Customers can transfer funds online at no extra charge.
reason: The context does not mention online fund transfers or any related fees.
1 | The minimum balance for rural branches is ₹2,500.
faithfulness: 0.75
FAIL at the card's pass mark of 0.8What the judge returned
- The judge split the answer into four claims, the same four as on the card, in its own words: "₹350 + GST" became "₹350 plus GST", and "You can also transfer funds online" became "Customers can transfer funds online".
- Three claims got a 1 and one got a 0. The 0 is the online transfer claim, and the judge's reason is that the context does not mention online fund transfers or any related fees.
- The score is 0.75, the value counted by hand, and it fails the card's pass mark of 0.8.
- The claims are the judge's wording, not yours. Another judge model can split the same answer into more or fewer claims, which changes the denominator. Compare faithfulness scores between runs made with the same judge.
Faithfulness vs answer correctness
| Faithfulness | Answer correctness | |
|---|---|---|
| The answer is compared with | The retrieved chunks | The golden's reference answer |
| Needs a golden | No | Yes |
| An answer that copies a wrong chunk | Scores high: every claim is in the context | Scores low: the facts differ from the reference |
| A true fact that no chunk holds | Counts as unsupported | Counts as correct when the reference has it |
| What it tells you | Whether the generator stays inside its context | Whether the final answer is right |
Where you use faithfulness
- Hallucination checks on live traffic. It needs the answer and the chunks, not a golden, so it can score real conversations sampled from production.
- After a prompt or model change. A drop in faithfulness on the same goldens means the generator started adding facts of its own.
- Answers where an invented detail is costly. Fees, policies and medical or legal text deserve a high pass mark.
CONTEXT_LIMIT = 2 in its evals/metrics.py), so a correct claim taken from the third chunk would count as unsupported. Pass the judge the same context the generator saw.Related
- Previous: RAGAS metrics
- Next: Answer relevancy
- Reference: Faithfulness in the RAGAS docs
- In the hand count, add a fifth claim with
Noneas its source: the score becomes 3 / 5 = 0.6. - In the RAGAS example, delete the sentence "You can also transfer funds online at no extra charge." from
responseand run it again: only supported claims are left, and the score rises to 1.0. - Change
0.8to0.7in the last line of the RAGAS example: the same 0.75 now prints PASS.
This is what real progress feels like.