RAGASragas 0.4.3 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

Answer relevancy

Answer relevancy is a RAGAS metric that measures how directly an answer addresses the question, by having the judge write questions the answer would fit and comparing them with the real question using embeddings.

Last updated: 29 Sep, 2026 · RAGAS 0.4.3

An answer can be faithful to the chunks and still miss the point: correct facts about the wrong thing, or the right fact buried in filler. Answer relevancy measures that side of quality.

Answer relevancy: generate questions, then compare · from the Production RAG Live Marathon · 232:50 to 238:09

Reverse-engineering the question

The video starts from the data. Unlike faithfulness, this metric needs the user's question, plus the actual output. Then the judge works in two steps. First it reads the answer and invents N questions that the answer would answer, reverse engineering, like a teacher who reads a student's essay and writes the questions it must have been answering. Then it compares each invented question with the real one by cosine similarity of their embeddings, meaning, not keywords. The average is the score. The video's slide lists three causes of a low score: an off-topic answer, a padded answer full of filler, and an incomplete answer.

The video's slide also lists the retrieved context as an optional input. The RAGAS collections metric does not take it: AnswerRelevancy reads only the question and the answer.

The AnswerRelevancy API

python
from ragas.metrics.collections import AnswerRelevancy

relevancy = AnswerRelevancy(llm=judge, embeddings=embeddings, strictness=3)  # 3 invented questions
relevancy.score(user_input=..., response=...).value  # average cosine similarity, 0 to 1

The question from golden g002

python
question = "What are the RAM and storage specs of the ProBook X1?"

Four answers, one per cause

python
answers = {
    "direct": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD.",
    "padded": "Great question! TechNest loves helping customers. ...",
    "off-topic": "The ProBook X1 weighs 1.4kg and costs $1,299.",
    "evasive": "I am not sure about that, sorry.",
}

strictness is how many questions the judge invents per answer. It is set here so each score comes from three.

Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
judge.py
import os

from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory

groq = AsyncOpenAI(
    api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
    base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)


class OneTextPerCall(GoogleEmbeddings):
    """gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""

    def embed_texts(self, texts, **kwargs):
        return [self.embed_text(text) for text in texts]

    async def aembed_texts(self, texts, **kwargs):
        return [await self.aembed_text(text) for text in texts]


embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")

Scoring four answers to one question

ExampleAPI key
from judge import embeddings, judge
from ragas.metrics.collections import AnswerRelevancy

question = "What are the RAM and storage specs of the ProBook X1?"
answers = {
    "direct": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD.",
    "padded": ("Great question! TechNest loves helping customers find the right laptop, and we stock many "
               "models. Among other things, the ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."),
    "off-topic": "The ProBook X1 weighs 1.4kg and costs $1,299.",
    "evasive": "I am not sure about that, sorry.",
}

relevancy = AnswerRelevancy(llm=judge, embeddings=embeddings, strictness=3)
for label, text in answers.items():
    result = relevancy.score(user_input=question, response=text)
    print(f"{result.value:.2f}  {label}")

What the four scores show

  • direct, 1.00. The questions the judge invented from it are, in effect, the question that was asked.
  • padded, 1.00. The filler did not lower it on this run: the answer still ends with the exact specs, and the judge built its questions from that sentence. Padding hurts this metric when it replaces the answer, not when it only surrounds it.
  • off-topic, 0.86. Lower, but not low. Questions about the ProBook X1's weight and price sit close to a question about its RAM in embedding space, because they share the product. On this metric, a drop from 1.00 to the high 0.8s already means the answer drifted.
  • evasive, 0.00. When the judge marks every invented question as coming from a noncommittal answer, RAGAS multiplies the score by zero, whatever the similarity.

Answer relevancy vs faithfulness

Answer relevancyFaithfulness
ReadsQuestion and answerQuestion, answer and chunks
AsksIs the answer about the question?Is the answer backed by the chunks?
Judge calls per sampleOne per invented question, plus embeddingsTwo: claims, then verdicts
A faithful but off-topic answerScores lowScores high

When to use answer relevancy

  • On live traffic: it needs neither a reference nor the chunks.
  • When users complain the bot rambles or dodges, the padded and evasive cases above.
Watch out. Relevancy does not check facts. "The ProBook X1 has 64GB RAM" answers the question directly and scores high while being wrong. Faithfulness and answer correctness catch that.
Try it yourself
  • Add an incomplete answer, "The ProBook X1 has 16GB DDR5 RAM.", the slide's third cause, and see where it lands.
  • Set strictness=1 and score the padded answer twice; compare how much the score moves.
PreviousFaithfulness

You understood something today that you didn't yesterday.