Answer relevancy
AnswerRelevancyMetric is a DeepEval RAG metric that scores how much of an answer addresses the question: the judge splits the answer into statements and counts the share that are relevant to the input.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
Faithfulness checks that the answer stays inside the chunks. An answer can do that and still talk about something else. Answer relevancy grades the other side of the generator: the answer against the question.
Judging an answer by the question it answers
The video starts from the data each metric needs. Faithfulness did not use the user's question; answer relevancy does, together with the actual output and the retrieved context. Then come two steps. First the judge reverse engineers the answer: it writes a few questions that this answer would answer, the way a teacher reads a student's long answer about AI and writes down the questions it fits, such as "what is reranking?". Then it compares each generated question with the original question by similarity, averages the results, and the average is the score, with a threshold on top.
The video describes RAGAS's method; DeepEval's AnswerRelevancyMetric works differently: it reads only input and actual_output, splits the answer into statements, asks the judge whether each statement is relevant to the input, and scores relevant statements over all statements, with no generated questions and no embeddings.
The AnswerRelevancyMetric API
from deepeval.metrics import AnswerRelevancyMetric
metric = AnswerRelevancyMetric(model=judge, threshold=0.7)
metric.measure(LLMTestCase(input=..., actual_output=...)) # no chunks, no expected answer
metric.score # relevant statements / all statements
metric.statements # what the judge split the answer into
metric.verdicts # one yes, no or borderline per statementA borderline statement, one that is only partly on topic, counts as relevant in 4.2.8. Only no lowers the score.
A question about the SoundPods Pro battery
question = "How long is the battery life on the SoundPods Pro?"Two answers with the same battery facts
Both answers give the right battery life from the catalog. The second keeps going with an upsell: two other products and the free-shipping threshold.
answers = {
"on topic": "The SoundPods Pro play for 8 hours per charge, and the case adds another 24 hours.",
"upsell": ("The SoundPods Pro play for 8 hours per charge, and the case adds another 24 hours. "
"While you are here, the SmartWatch X costs $299 and tracks your heart rate. "
"The BassBuds Max make a great gift, and orders over $50 ship free."),
}Printing each statement with its verdict
for statement, verdict in zip(metric.statements, metric.verdicts):
print(f" {verdict.verdict:4} {statement}")- written in Custom judge model
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
Scoring an on-topic answer and an upsell
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase
from judge import judge
question = "How long is the battery life on the SoundPods Pro?"
answers = {
"on topic": "The SoundPods Pro play for 8 hours per charge, and the case adds another 24 hours.",
"upsell": ("The SoundPods Pro play for 8 hours per charge, and the case adds another 24 hours. "
"While you are here, the SmartWatch X costs $299 and tracks your heart rate. "
"The BassBuds Max make a great gift, and orders over $50 ship free."),
}
metric = AnswerRelevancyMetric(model=judge, threshold=0.7)
for label, text in answers.items():
metric.measure(LLMTestCase(input=question, actual_output=text))
print(f"{metric.score:.2f} passed={metric.is_successful()} {label}")
for statement, verdict in zip(metric.statements, metric.verdicts):
print(f" {verdict.verdict:4} {statement}")
print(metric.reason)1.00 passed=True on topic yes The SoundPods Pro play for 8 hours per charge. yes The case adds another 24 hours. 0.33 passed=False upsell yes The SoundPods Pro play for 8 hours per charge. yes The case adds another 24 hours. no The SmartWatch X costs $299. no The SmartWatch X tracks your heart rate. no The BassBuds Max make a great gift. no Orders over $50 ship free. The score is 0.33 because the answer drifted into unrelated content—pricing and features of SmartWatch X, details about BassBuds Max, and shipping policy—none of which address the SoundPods Pro battery‑life question.
Why the upsell answer scored low
- The on-topic answer scores 1.00. The judge split it into two statements, the 8 hours per charge and the 24 hours from the case, and both are about the battery.
- The upsell answer scores 0.33 and fails the 0.7 threshold. It has the same two relevant statements plus four that are not: the SmartWatch X price, its heart-rate tracking, the BassBuds gift idea and free shipping. Two out of six is 0.33.
- The facts did not save it. The upsell answer contains the full correct answer, and it still fails, because the metric scores the share of the answer that is on topic, not whether the answer is in there somewhere.
- The reason names the off-topic parts, which tells you what to cut from the prompt or the reply.
Answer relevancy vs faithfulness
| Answer relevancy | Faithfulness | |
|---|---|---|
| Compares | The answer with the question | The answer with the chunks |
| Reads | input, actual_output | input, actual_output, retrieval_context |
| Splits the answer into | Statements | Claims |
| A wrong number about the right product | Still on topic | Contradicted by the chunks |
When to use answer relevancy
- When users say the bot rambles: preamble, caveats and sales talk all show up as irrelevant statements.
- After a prompt change that adds instructions such as "suggest related products", to measure how much of each answer the extra text takes.
- On live traffic: it needs neither chunks nor an expected answer.
Related
- Previous: Faithfulness
- Next: Contextual precision
- Reference: Answer relevancy
- Add a third answer,
"The SoundPods Pro play for 40 hours per charge.", and check that it scores as on topic although the number is wrong. - Delete the BassBuds sentence from the upsell answer and predict the new score before you run it.
- Pass
strict_mode=Trueto the metric and read what the upsell answer scores now.
Every expert started right here.