AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Faithfulness

Faithfulness is a RAGAS metric that measures the share of claims in an answer that the retrieved context supports.

Last updated: 09 Oct, 2026 · RAGAS 0.4

RAGAS metrics listed the five scores the video's app prints for every golden. Faithfulness is the first one it computes. It answers one question: did the answer stay inside the chunks the retriever handed over, or did the model add something of its own? A claim with no chunk behind it is a hallucination, and this score counts them.

Faithfulness on the bank example · from the Complete AI Security Course in 8 Hours video · 2:05:53 to 2:08:56

This part of the video starts at 2:05:53. On the whiteboard card it shows, three of the four claims are grounded and one is not, so the score is 3 / 4 = 0.75.

Reading the video's bank example

The example is a bank's help bot. The retriever returned two chunks from the knowledge base:

  • Chunk 1: "Minimum balance is ₹10,000 for urban branches. Non-maintenance fee is ₹350 + GST if balance falls below."
  • Chunk 2: "Semi-urban branch minimum balance is ₹5,000. Rural branch minimum balance is ₹2,500."

From them the model wrote this answer: "The minimum balance for urban branches is ₹10,000. The non-maintenance fee charged is ₹350 + GST. You can also transfer funds online at no extra charge. For rural branches the minimum is ₹2,500."

The judge works in two steps. First it breaks the answer into atomic claims, short statements that each carry one fact. Then it asks the card's judge question about every claim: "Can this claim be fully inferred from the retrieved context, yes or no?" A yes makes the claim grounded. A no makes it hallucinated.

The response is split into four claims. Two are tied to chunk 1 and one to chunk 2 and are marked grounded; the claim that online fund transfer has no extra charge is in no chunk and is marked hallucinated. Three grounded claims out of four give a faithfulness of 0.75, below the pass mark of 0.8 drawn on the bar.

The balance and the fee are in chunk 1 and the rural figure is in chunk 2. The sentence about online transfers is in neither chunk: the model added it. It may even be true for this bank, but nothing the retriever returned says so.

Writing faithfulness as a formula

Faithfulness, as the RAGAS docs define it

Every claim weighs the same, and the score runs from 0 to 1. The metric reads the response and the retrieved_contexts, plus the user_input, which helps the judge split the answer into claims. It never reads the golden's reference, so it also works on live traffic where no reference exists.

Counting the grounded claims by hand

With the verdicts of the card typed in, the score is a count and a division. No model is needed for this part.

ExampleThe video's example, computed in plain Python
claims = [
    ("Min balance is ₹10,000 for urban branches", "Chunk 1"),
    ("Non-maintenance fee is ₹350 + GST", "Chunk 1"),
    ("Online fund transfer has no extra charge", None),  # in no chunk
    ("Rural branch minimum is ₹2,500", "Chunk 2"),
]

grounded = sum(source is not None for claim, source in claims)
score = grounded / len(claims)

for claim, source in claims:
    verdict = "grounded    " if source else "hallucinated"
    print(verdict, "|", claim, "|", source or "not in any chunk")
print(f"faithfulness = {grounded} / {len(claims)} = {score}")
for pass_mark in (0.8, 0.7):
    print(f"pass mark {pass_mark}: {'pass' if score >= pass_mark else 'fail'}")

What the count shows

  • Three of four claims are grounded, so faithfulness is 3 / 4 = 0.75.
  • At a pass mark of 0.8 the answer fails; at 0.7 it passes. The card uses 0.8.
  • The pass mark is yours to set. The video calls it a hyperparameter: you decide how strict the check is for your use case. RAGAS returns only the number.

Seeing what one unsupported claim costs

The same single slip costs a short answer more than a long one, which matters when you pick a pass mark.

ExampleRun on Python 3.12
for total in (2, 4, 7, 10):
    print(f"{total - 1} of {total} claims supported: {(total - 1) / total:.2f}")

One unsupported claim out of two halves the score, while one out of ten leaves 0.90. The Results screen of the video's app shows 0.86 for its return-policy golden, the value that six supported claims out of seven give.

Scoring the same answer with RAGAS

The video's app ran llama-3.1-8b-instant as its judge, since retired on Groq; the run below uses openai/gpt-oss-20b, set up as in LLM as a judge.

The judge

python
client = AsyncOpenAI(
    base_url="https://api.groq.com/openai/v1",
    api_key=os.environ["GROQ_API_KEY"],
    max_retries=6,  # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)

temperature=0 keeps the verdicts as steady as the model allows, and max_tokens=4096 leaves room for the structured reply of a reasoning model. max_retries=6 makes the client wait and try again when Groq's free tier answers that the per-minute token limit is used up.

The metric call

python
from ragas.metrics.collections import Faithfulness

result = Faithfulness(llm=judge).score(user_input=question, response=response, retrieved_contexts=chunks)
result.value  # supported claims / all claims

Keeping the judge's replies

score() returns only the number. The claims and verdicts that the card shows are in the judge's replies, which the metric reads and then drops. Each of the five metrics calls judge.agenerate, so wrapping that one method keeps each reply in a list. The reply objects and their field names are internals of RAGAS 0.4 and can change in a later release.

python
replies = []  # each structured reply the judge sends back to the metric
ask = judge.agenerate


async def keep(prompt, response_model):
    reply = await ask(prompt, response_model)
    replies.append(reply)
    return reply


judge.agenerate = keep

Faithfulness calls the judge twice. The first reply holds the claims, the second holds the same claims with a verdict and a reason each, so replies[-1] is the one to print. The example needs the GROQ_API_KEY from Installing Python for AI security in the environment.

ExampleAPI keyFrom the video, run on Groq
import os

from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import Faithfulness

client = AsyncOpenAI(
    base_url="https://api.groq.com/openai/v1",
    api_key=os.environ["GROQ_API_KEY"],
    max_retries=6,  # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)

replies = []  # each structured reply the judge sends back to the metric
ask = judge.agenerate


async def keep(prompt, response_model):
    reply = await ask(prompt, response_model)
    replies.append(reply)
    return reply


judge.agenerate = keep

question = "What is the minimum balance for my savings account?"
chunks = [
    "Minimum balance is ₹10,000 for urban branches. Non-maintenance fee is ₹350 + GST if balance falls below.",
    "Semi-urban branch minimum balance is ₹5,000. Rural branch minimum balance is ₹2,500.",
]
response = ("The minimum balance for urban branches is ₹10,000. The non-maintenance fee charged is ₹350 + GST. "
            "You can also transfer funds online at no extra charge. For rural branches the minimum is ₹2,500.")

result = Faithfulness(llm=judge).score(user_input=question, response=response, retrieved_contexts=chunks)
for item in replies[-1].statements:
    print(item.verdict, "|", item.statement)
    if item.verdict == 0:
        print("    reason:", item.reason)
print("faithfulness:", result.value)
print("PASS" if result.value >= 0.8 else "FAIL", "at the card's pass mark of 0.8")

What the judge returned

  • The judge split the answer into four claims, the same four as on the card, in its own words: "₹350 + GST" became "₹350 plus GST", and "You can also transfer funds online" became "Customers can transfer funds online".
  • Three claims got a 1 and one got a 0. The 0 is the online transfer claim, and the judge's reason is that the context does not mention online fund transfers or any related fees.
  • The score is 0.75, the value counted by hand, and it fails the card's pass mark of 0.8.
  • The claims are the judge's wording, not yours. Another judge model can split the same answer into more or fewer claims, which changes the denominator. Compare faithfulness scores between runs made with the same judge.

Faithfulness vs answer correctness

FaithfulnessAnswer correctness
The answer is compared withThe retrieved chunksThe golden's reference answer
Needs a goldenNoYes
An answer that copies a wrong chunkScores high: every claim is in the contextScores low: the facts differ from the reference
A true fact that no chunk holdsCounts as unsupportedCounts as correct when the reference has it
What it tells youWhether the generator stays inside its contextWhether the final answer is right

Where you use faithfulness

  • Hallucination checks on live traffic. It needs the answer and the chunks, not a golden, so it can score real conversations sampled from production.
  • After a prompt or model change. A drop in faithfulness on the same goldens means the generator started adding facts of its own.
  • Answers where an invented detail is costly. Fees, policies and medical or legal text deserve a high pass mark.
Watch out. Faithfulness is judged against the chunks you hand the judge, not against the truth and not against what the generator read. The video's app gives the generator three chunks but passes only the first two to the metric (CONTEXT_LIMIT = 2 in its evals/metrics.py), so a correct claim taken from the third chunk would count as unsupported. Pass the judge the same context the generator saw.
Try it yourself
  • In the hand count, add a fifth claim with None as its source: the score becomes 3 / 5 = 0.6.
  • In the RAGAS example, delete the sentence "You can also transfer funds online at no extra charge." from response and run it again: only supported claims are left, and the score rises to 1.0.
  • Change 0.8 to 0.7 in the last line of the RAGAS example: the same 0.75 now prints PASS.
PreviousRAGAS metrics

This is what real progress feels like.