RAG metrics
RAG metrics are DeepEval's five built-in metrics for retrieval-augmented apps that grade the retriever and the generator separately, each one reading a fixed set of LLMTestCase fields.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
Custom metrics showed how to write a metric of your own. For a RAG app you rarely need to: DeepEval ships five metrics for it, and each one compares a different pair of things the TechNest bot produces: the question, the chunks, the answer and the expected answer.
Four checks on a RAG pipeline
The video starts from a diagram in the LangSmith documentation. A search retriever, a document search or a web search, takes the question and returns documents. The first check is whether those documents are relevant to the input question: retrieval relevance. The LLM then writes the answer with the documents as its context, and the second check is whether that answer is grounded in the documents: groundedness. The third is correctness: does the answer match the ground truth answer, which means you need a ground truth. The fourth is answer relevance: does the answer address the question.
The video runs these checks in LangSmith; this page does them with DeepEval. Groundedness is DeepEval's FaithfulnessMetric, answer relevance is AnswerRelevancyMetric, and retrieval relevance is ContextualRelevancyMetric. Correctness has no RAG metric of its own: it is a G-Eval criterion, as in G-Eval. DeepEval adds two more retriever checks that need the expected answer: ContextualPrecisionMetric, are the useful chunks ranked first, and ContextualRecallMetric, did every needed fact come back.

The RAG metrics API
from deepeval.metrics import (
AnswerRelevancyMetric, # generator: is the answer about the question?
FaithfulnessMetric, # generator: is the answer backed by the chunks?
ContextualRelevancyMetric, # retriever: are the chunks about the question?
ContextualPrecisionMetric, # retriever: are the useful chunks ranked first?
ContextualRecallMetric, # retriever: do the chunks hold the expected answer?
)
metric = FaithfulnessMetric(model=judge, threshold=0.5) # all five take model= and threshold=
metric.measure(test_case)All five are judged by an LLM, so each gets model=judge, the Groq judge from Custom judge model. They differ in which fields of the test case they read.
Fields each RAG metric reads
| Metric | input | actual_output | expected_output | retrieval_context | Grades |
|---|---|---|---|---|---|
| Faithfulness | Yes | Yes | Yes | Generator | |
| Answer relevancy | Yes | Yes | Generator | ||
| Contextual relevancy | Yes | Yes | Retriever | ||
| Contextual precision | Yes | Yes | Yes, in rank order | Retriever | |
| Contextual recall | Yes | Yes | Yes | Retriever |
The docs pages of the three contextual metrics also list actual_output as required. The installed 4.2.8 does not check it for them: their required fields are the ones in the table, and none of the three reads the answer. A retriever metric grades the search, whatever the model wrote afterwards.
- written in Custom judge model
- written in TechNest RAG app
- written in TechNest RAG app
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
Running faithfulness without retrieval_context
A test case with only a question and an answer is enough for answer relevancy. Give it to faithfulness and the metric stops before it calls the judge, because the chunks are missing.
from deepeval.metrics import FaithfulnessMetric
from deepeval.test_case import LLMTestCase
from judge import judge
test_case = LLMTestCase(
input="What is TechNest's return policy?",
actual_output="TechNest accepts returns within 30 days of purchase.",
)
metric = FaithfulnessMetric(model=judge)
metric.measure(test_case)Traceback (most recent call last):
File "main.py", line 10, in <module>
metric.measure(test_case)
deepeval.errors.MissingTestCaseParamsError: 'retrieval_context' cannot be None for the 'Faithfulness' metricThe error names the missing field and the metric. The check runs at the start of measure, so it costs no judge tokens. The fix is to keep the chunks the app retrieved and pass them as retrieval_context.
One test case with all four fields
The bot's answer and chunks come from one real run of answer; the question and the expected answer are written in advance.
question = "What is TechNest's return policy?"
expected = ("TechNest accepts returns within 30 days of purchase. Items must be in original "
"packaging with all accessories. Customers pay return shipping unless the item is "
"defective. Refunds are processed in 5 to 7 business days.")
response, contexts = answer(question)
test_case = LLMTestCase(input=question, actual_output=response,
expected_output=expected, retrieval_context=contexts)The five metrics in a loop
metrics = [AnswerRelevancyMetric(model=judge), FaithfulnessMetric(model=judge),
ContextualRelevancyMetric(model=judge), ContextualPrecisionMetric(model=judge),
ContextualRecallMetric(model=judge)]
for metric in metrics:
metric.measure(test_case)
print(f"{metric.score:.2f} {metric.__name__}")__name__ is the metric's display name, the one reports print.
Scoring one bot answer with all five RAG metrics
from deepeval.metrics import (
AnswerRelevancyMetric,
ContextualPrecisionMetric,
ContextualRecallMetric,
ContextualRelevancyMetric,
FaithfulnessMetric,
)
from deepeval.test_case import LLMTestCase
from judge import judge
from technest import answer
question = "What is TechNest's return policy?"
expected = ("TechNest accepts returns within 30 days of purchase. Items must be in original "
"packaging with all accessories. Customers pay return shipping unless the item is "
"defective. Refunds are processed in 5 to 7 business days.")
response, contexts = answer(question)
test_case = LLMTestCase(input=question, actual_output=response,
expected_output=expected, retrieval_context=contexts)
metrics = [AnswerRelevancyMetric(model=judge), FaithfulnessMetric(model=judge),
ContextualRelevancyMetric(model=judge), ContextualPrecisionMetric(model=judge),
ContextualRecallMetric(model=judge)]
for metric in metrics:
metric.measure(test_case)
print(f"{metric.score:.2f} {metric.__name__}")1.00 Answer Relevancy 1.00 Faithfulness 0.36 Contextual Relevancy 1.00 Contextual Precision 1.00 Contextual Recall
What the five scores say
- Answer relevancy 1.00 and faithfulness 1.00: the generator did its job. Every statement in the bot's answer is about returns, and none of its claims goes against the chunks.
- Contextual relevancy 0.36 is the low one. The search also returned the shipping and warranty policies, and most of the statements in those two chunks have nothing to do with returns.
- Contextual precision 1.00: the one useful chunk, the return policy, is ranked first, so the noise below it costs nothing on this metric.
- Contextual recall 1.00: every sentence of the expected answer is in the return-policy chunk.
- Read together, the answer is good and the search is noisy but complete. Fewer chunks would help relevancy; Contextual relevancy runs that change.
Retriever metrics vs generator metrics
| Retriever metrics | Generator metrics | |
|---|---|---|
| Metrics | Contextual relevancy, precision, recall | Faithfulness, answer relevancy |
| Grade | The chunks retrieve returned | The answer generate wrote |
Need expected_output? | Precision and recall: yes. Relevancy: no | No |
| A low score points at | Chunk size, top_k, the search or embedding model, a reranker | The prompt or the model |
When to use each group
- Generator metrics when you change the prompt or the model and want to know whether answers stay grounded and on topic. Neither needs an expected answer, so both run on live traffic.
- Retriever metrics when you change the search:
top_k, chunking, the embedding model or a reranker. Contextual relevancy runs without an expected answer; precision and recall need goldens. - All five together when an answer is wrong and you do not know which half of the app to fix.
Related
- Previous: Custom metrics
- Next: Faithfulness
- Reference: RAG evaluation quickstart
- Add
expected_output="30 days"to the test case in the error example, swapFaithfulnessMetricforContextualPrecisionMetric, and read which field the error names now. - Delete
retrieval_context=contextsfrom the five-metric test case and run the loop: answer relevancy scores, then faithfulness stops withMissingTestCaseParamsError. - Change the question to
"Do you ship to Canada?"and compare the contextual relevancy score with the answer relevancy score.
Slow is fine. Stopping is the only problem.