DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

Faithfulness

FaithfulnessMetric is a DeepEval RAG metric that scores how well an answer agrees with the retrieved chunks: the judge splits the answer into claims and counts the share that the chunks do not contradict.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

Of the five metrics in RAG metrics, faithfulness is the one that catches a bot promising what its documents never said. It grades the generator: the answer against the chunks it was given.

Faithfulness claim by claim · from the Production RAG Live Marathon · 3:44:52 to 3:49:59

Checking each claim against the chunks

The video keeps four things in view for any RAG evaluation: the question, the retrieved context, the generated answer, and the reference answer, the ground truth. Faithfulness takes two of them, the retrieved context and the LLM response. It does not use the reference.

The judge first breaks the response into claims, then asks of each one: can this claim be fully inferred from the retrieved context? In the video's bank example the answer makes four claims. The urban minimum balance of ₹10,000 and the ₹350 + GST non-maintenance fee come from chunk one, and the rural minimum of ₹2,500 from chunk two. "Transfer funds online at no extra charge" is in no chunk: the generator hallucinated it. Three grounded claims out of four is 0.75, and with a threshold of 0.8 this answer fails.

The video scores this with RAGAS. DeepEval's metric follows the same idea, claims from the answer checked against the context, with one difference in what counts against the answer. DeepEval's judge says no only when the context contradicts a claim. A claim the context never mentions gets borderline, and borderline counts as faithful unless you set penalize_ambiguous_claims=True.

The FaithfulnessMetric API

python
from deepeval.metrics import FaithfulnessMetric

metric = FaithfulnessMetric(
    model=judge,
    threshold=0.8,                    # the pass mark
    penalize_ambiguous_claims=False,  # True: an unmentioned claim counts against the answer
    truths_extraction_limit=None,     # a number: keep only that many facts from the chunks
)
metric.measure(LLMTestCase(input=..., actual_output=..., retrieval_context=[...]))
metric.score, metric.truths, metric.claims, metric.verdicts

One measure asks the judge three things, plus a fourth call for the reason (include_reason=False skips it). It extracts the truths, the facts stated in retrieval_context. It extracts the claims made in actual_output. Then it gives one verdict per claim, compared with the truths: yes, no or borderline. The score is the share of claims that pass. The docs describe the unmentioned case as an ambiguous claim; the installed 4.2.8 labels it borderline.

The bank chunks

The chunks and the response are the ones on the video's slide.

python
chunks = [
    "Minimum balance is ₹10,000 for urban branches. Non-maintenance fee is ₹350 + GST if balance falls below.",
    "Semi-urban branch minimum balance is ₹5,000. Rural branch minimum balance is ₹2,500.",
]

The answer with the online-transfer claim

python
response = ("The minimum balance for urban branches is ₹10,000. The non-maintenance fee charged is ₹350 + GST. "
            "You can also transfer funds online at no extra charge. For rural branches the minimum is ₹2,500.")

Scoring with and without the penalty

python
for penalize in [False, True]:
    metric = FaithfulnessMetric(model=judge, threshold=0.8, penalize_ambiguous_claims=penalize)
    metric.measure(test_case)
    for claim, verdict in zip(metric.claims, metric.verdicts):
        print(f"   {verdict.verdict:10} {claim}")

metric.claims and metric.verdicts line up one to one, so zip prints each claim next to the judge's decision.

Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
judge.py
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
technest.py
import json
import os
import re

from openai import OpenAI

groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"

SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""

with open("catalog.json", encoding="utf-8") as f:
    CATALOG = json.load(f)

SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
        "long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}


def words(text):
    """The words in a text that carry meaning, in lower case."""
    return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}


def retrieve(question, top_k=3):
    asked = words(question)
    ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
    return [item["content"] for item in ranked[:top_k]]


def generate(question, contexts):
    context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
    ]
    response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
    return response.choices[0].message.content.strip()


def answer(question, top_k=3):
    contexts = retrieve(question, top_k)
    return generate(question, contexts), contexts
catalog.json
[
  {
    "id": "prod_001",
    "category": "product",
    "title": "ProBook X1 Laptop",
    "content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
  },
  {
    "id": "prod_002",
    "category": "product",
    "title": "PixelPhone 15",
    "content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
  },
  {
    "id": "prod_003",
    "category": "product",
    "title": "SoundPods Pro",
    "content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
  },
  {
    "id": "prod_004",
    "category": "product",
    "title": "UltraTab S2 Tablet",
    "content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
  },
  {
    "id": "prod_005",
    "category": "product",
    "title": "SmartWatch X",
    "content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
  },
  {
    "id": "prod_006",
    "category": "product",
    "title": "ProCam 4K Action Camera",
    "content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
  },
  {
    "id": "prod_007",
    "category": "product",
    "title": "BassBuds Max Headphones",
    "content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
  },
  {
    "id": "prod_008",
    "category": "product",
    "title": "SoundBar 360",
    "content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
  },
  {
    "id": "policy_001",
    "category": "policy",
    "title": "Return Policy",
    "content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
  },
  {
    "id": "policy_002",
    "category": "policy",
    "title": "Shipping Policy",
    "content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
  },
  {
    "id": "policy_003",
    "category": "policy",
    "title": "Warranty Policy",
    "content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
  },
  {
    "id": "policy_004",
    "category": "policy",
    "title": "Payment Policy",
    "content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
  },
  {
    "id": "faq_001",
    "category": "faq",
    "title": "Order Tracking",
    "content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
  },
  {
    "id": "faq_002",
    "category": "faq",
    "title": "International Shipping",
    "content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
  },
  {
    "id": "faq_003",
    "category": "faq",
    "title": "Bulk and Business Orders",
    "content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
  }
]

The video's bank example

ExampleAPI keyFrom the video, run on Groq
from deepeval.metrics import FaithfulnessMetric
from deepeval.test_case import LLMTestCase
from judge import judge

question = "What is the minimum balance for my savings account?"
chunks = [
    "Minimum balance is ₹10,000 for urban branches. Non-maintenance fee is ₹350 + GST if balance falls below.",
    "Semi-urban branch minimum balance is ₹5,000. Rural branch minimum balance is ₹2,500.",
]
response = ("The minimum balance for urban branches is ₹10,000. The non-maintenance fee charged is ₹350 + GST. "
            "You can also transfer funds online at no extra charge. For rural branches the minimum is ₹2,500.")
test_case = LLMTestCase(input=question, actual_output=response, retrieval_context=chunks)

for penalize in [False, True]:
    metric = FaithfulnessMetric(model=judge, threshold=0.8, penalize_ambiguous_claims=penalize)
    metric.measure(test_case)
    print(f"penalize_ambiguous_claims={penalize}: {metric.score:.2f}  passed={metric.is_successful()}")
    for claim, verdict in zip(metric.claims, metric.verdicts):
        print(f"   {verdict.verdict:10} {claim}")

With the default setting the answer scores 1.00 and passes the 0.8 mark. The judge found the same four claims as the slide and marked the online-transfer claim borderline: the chunks never mention transfers, so nothing contradicts it, and a borderline claim counts as faithful. With penalize_ambiguous_claims=True the verdicts are the same, but borderline now counts against the answer, and the score is the video's 0.75, a fail. The flag decides whether an unmentioned promise is a problem.

Faithfulness of two TechNest answers

The TechNest version uses the real chunks the search returns for the return-policy question and two fixed answers: one that keeps to the policy, and one that keeps the 30 days but promises free return shipping.

ExampleAPI key
from deepeval.metrics import FaithfulnessMetric
from deepeval.test_case import LLMTestCase
from judge import judge
from technest import retrieve

question = "What is TechNest's return policy?"
contexts = retrieve(question)  # return, shipping and warranty policies
answers = {
    "faithful": ("You can return items within 30 days of purchase in their original packaging. "
                 "Return shipping is on you unless the item arrives defective or damaged."),
    "contradicts": "You can return items within 30 days of purchase, and TechNest pays the return shipping on every order.",
}

metric = FaithfulnessMetric(model=judge)
for label, text in answers.items():
    metric.measure(LLMTestCase(input=question, actual_output=text, retrieval_context=contexts))
    print(f"{metric.score:.2f}  {label}  ({len(metric.truths)} truths, {len(metric.claims)} claims)")
    for claim, verdict in zip(metric.claims, metric.verdicts):
        print(f"   {verdict.verdict:4} {claim}")
        if verdict.reason:
            print(f"        {verdict.reason}")

Why the contradicting answer scored 0.50

  • 14 truths came from the three chunks, the return, shipping and warranty policies, for both answers. The claims are checked against these facts, not against the raw text.
  • The faithful answer scores 1.00. Its two claims, 30 days in the original packaging and return shipping paid by the customer unless the item is defective, both get yes.
  • The contradicting answer scores 0.50. Its 30-day claim gets yes; "TechNest pays the return shipping on every order" gets no, and the reason states the policy it breaks. One of two claims passes.
  • Only the no verdict carries a reason here. The judge is asked for a reason only when a claim fails or is borderline.

Faithfulness in DeepEval vs RAGAS

DeepEval FaithfulnessMetricRAGAS faithfulness (the video)
Judge stepsTruths from the chunks, claims from the answer, one verdict per claimClaims from the answer, one verdict per claim against the chunks
A claim the chunks contradictFailsFails
A claim the chunks never mentionPasses as borderline, fails with penalize_ambiguous_claims=TrueFails
ScorePassing claims / all claimsSupported claims / all claims

When to use faithfulness

  • Support bots whose answers carry prices, policies or promises, where a detail that contradicts the documents costs money or trust.
  • After a prompt or model change, to check the generator still stays inside the chunks it was given.
  • On live traffic: it needs no expected answer, only the question, the answer and the retrieved chunks.
Watch out. With the default settings an invented detail that the chunks never mention does not lower the score, as the bank run shows. When a made-up promise is as bad as a wrong one, set penalize_ambiguous_claims=True.
Try it yourself
  • Delete the online-transfer sentence from the bank response and run it again: with the penalty on, the score reaches 1.00.
  • Change "₹2,500" in the bank response to "₹3,000" and check that the judge says no to that claim even without the penalty.
  • Pass truths_extraction_limit=3 to the TechNest metric and print metric.truths to see the judge squeeze the chunks into three truths.
PreviousRAG metrics

This is what real progress feels like.