DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

Answer relevancy

AnswerRelevancyMetric is a DeepEval RAG metric that scores how much of an answer addresses the question: the judge splits the answer into statements and counts the share that are relevant to the input.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

Faithfulness checks that the answer stays inside the chunks. An answer can do that and still talk about something else. Answer relevancy grades the other side of the generator: the answer against the question.

Answer relevancy · from the Production RAG Live Marathon · 3:52:50 to 3:58:06

Judging an answer by the question it answers

The video starts from the data each metric needs. Faithfulness did not use the user's question; answer relevancy does, together with the actual output and the retrieved context. Then come two steps. First the judge reverse engineers the answer: it writes a few questions that this answer would answer, the way a teacher reads a student's long answer about AI and writes down the questions it fits, such as "what is reranking?". Then it compares each generated question with the original question by similarity, averages the results, and the average is the score, with a threshold on top.

The video describes RAGAS's method; DeepEval's AnswerRelevancyMetric works differently: it reads only input and actual_output, splits the answer into statements, asks the judge whether each statement is relevant to the input, and scores relevant statements over all statements, with no generated questions and no embeddings.

The AnswerRelevancyMetric API

python
from deepeval.metrics import AnswerRelevancyMetric

metric = AnswerRelevancyMetric(model=judge, threshold=0.7)
metric.measure(LLMTestCase(input=..., actual_output=...))  # no chunks, no expected answer
metric.score       # relevant statements / all statements
metric.statements  # what the judge split the answer into
metric.verdicts    # one yes, no or borderline per statement

A borderline statement, one that is only partly on topic, counts as relevant in 4.2.8. Only no lowers the score.

A question about the SoundPods Pro battery

python
question = "How long is the battery life on the SoundPods Pro?"

Two answers with the same battery facts

Both answers give the right battery life from the catalog. The second keeps going with an upsell: two other products and the free-shipping threshold.

python
answers = {
    "on topic": "The SoundPods Pro play for 8 hours per charge, and the case adds another 24 hours.",
    "upsell": ("The SoundPods Pro play for 8 hours per charge, and the case adds another 24 hours. "
               "While you are here, the SmartWatch X costs $299 and tracks your heart rate. "
               "The BassBuds Max make a great gift, and orders over $50 ship free."),
}

Printing each statement with its verdict

python
for statement, verdict in zip(metric.statements, metric.verdicts):
    print(f"   {verdict.verdict:4} {statement}")
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
judge.py
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))

Scoring an on-topic answer and an upsell

ExampleAPI key
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase
from judge import judge

question = "How long is the battery life on the SoundPods Pro?"
answers = {
    "on topic": "The SoundPods Pro play for 8 hours per charge, and the case adds another 24 hours.",
    "upsell": ("The SoundPods Pro play for 8 hours per charge, and the case adds another 24 hours. "
               "While you are here, the SmartWatch X costs $299 and tracks your heart rate. "
               "The BassBuds Max make a great gift, and orders over $50 ship free."),
}

metric = AnswerRelevancyMetric(model=judge, threshold=0.7)
for label, text in answers.items():
    metric.measure(LLMTestCase(input=question, actual_output=text))
    print(f"{metric.score:.2f}  passed={metric.is_successful()}  {label}")
    for statement, verdict in zip(metric.statements, metric.verdicts):
        print(f"   {verdict.verdict:4} {statement}")
print(metric.reason)

Why the upsell answer scored low

  • The on-topic answer scores 1.00. The judge split it into two statements, the 8 hours per charge and the 24 hours from the case, and both are about the battery.
  • The upsell answer scores 0.33 and fails the 0.7 threshold. It has the same two relevant statements plus four that are not: the SmartWatch X price, its heart-rate tracking, the BassBuds gift idea and free shipping. Two out of six is 0.33.
  • The facts did not save it. The upsell answer contains the full correct answer, and it still fails, because the metric scores the share of the answer that is on topic, not whether the answer is in there somewhere.
  • The reason names the off-topic parts, which tells you what to cut from the prompt or the reply.

Answer relevancy vs faithfulness

Answer relevancyFaithfulness
ComparesThe answer with the questionThe answer with the chunks
Readsinput, actual_outputinput, actual_output, retrieval_context
Splits the answer intoStatementsClaims
A wrong number about the right productStill on topicContradicted by the chunks

When to use answer relevancy

  • When users say the bot rambles: preamble, caveats and sales talk all show up as irrelevant statements.
  • After a prompt change that adds instructions such as "suggest related products", to measure how much of each answer the extra text takes.
  • On live traffic: it needs neither chunks nor an expected answer.
Watch out. Answer relevancy does not check facts. An answer that gives the wrong battery life is still about the battery life, so it can pass. Run faithfulness next to it to catch that.
Try it yourself
  • Add a third answer, "The SoundPods Pro play for 40 hours per charge.", and check that it scores as on topic although the number is wrong.
  • Delete the BassBuds sentence from the upsell answer and predict the new score before you run it.
  • Pass strict_mode=True to the metric and read what the upsell answer scores now.
PreviousFaithfulness

Every expert started right here.