Answer relevancy
Answer relevancy is a RAGAS metric that measures how directly an answer addresses the question, by having the judge write questions the answer would fit and comparing them with the real question using embeddings.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
An answer can be faithful to the chunks and still miss the point: correct facts about the wrong thing, or the right fact buried in filler. Answer relevancy measures that side of quality.
Reverse-engineering the question
The video starts from the data. Unlike faithfulness, this metric needs the user's question, plus the actual output. Then the judge works in two steps. First it reads the answer and invents N questions that the answer would answer, reverse engineering, like a teacher who reads a student's essay and writes the questions it must have been answering. Then it compares each invented question with the real one by cosine similarity of their embeddings, meaning, not keywords. The average is the score. The video's slide lists three causes of a low score: an off-topic answer, a padded answer full of filler, and an incomplete answer.
The video's slide also lists the retrieved context as an optional input. The RAGAS collections metric does not take it: AnswerRelevancy reads only the question and the answer.
The AnswerRelevancy API
from ragas.metrics.collections import AnswerRelevancy
relevancy = AnswerRelevancy(llm=judge, embeddings=embeddings, strictness=3) # 3 invented questions
relevancy.score(user_input=..., response=...).value # average cosine similarity, 0 to 1The question from golden g002
question = "What are the RAM and storage specs of the ProBook X1?"Four answers, one per cause
answers = {
"direct": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD.",
"padded": "Great question! TechNest loves helping customers. ...",
"off-topic": "The ProBook X1 weighs 1.4kg and costs $1,299.",
"evasive": "I am not sure about that, sorry.",
}strictness is how many questions the judge invents per answer. It is set here so each score comes from three.
- written in LLM as a judge
View the code here
import os
from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
groq = AsyncOpenAI(
api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
Scoring four answers to one question
from judge import embeddings, judge
from ragas.metrics.collections import AnswerRelevancy
question = "What are the RAM and storage specs of the ProBook X1?"
answers = {
"direct": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD.",
"padded": ("Great question! TechNest loves helping customers find the right laptop, and we stock many "
"models. Among other things, the ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."),
"off-topic": "The ProBook X1 weighs 1.4kg and costs $1,299.",
"evasive": "I am not sure about that, sorry.",
}
relevancy = AnswerRelevancy(llm=judge, embeddings=embeddings, strictness=3)
for label, text in answers.items():
result = relevancy.score(user_input=question, response=text)
print(f"{result.value:.2f} {label}")1.00 direct 1.00 padded 0.86 off-topic 0.00 evasive
What the four scores show
- direct, 1.00. The questions the judge invented from it are, in effect, the question that was asked.
- padded, 1.00. The filler did not lower it on this run: the answer still ends with the exact specs, and the judge built its questions from that sentence. Padding hurts this metric when it replaces the answer, not when it only surrounds it.
- off-topic, 0.86. Lower, but not low. Questions about the ProBook X1's weight and price sit close to a question about its RAM in embedding space, because they share the product. On this metric, a drop from 1.00 to the high 0.8s already means the answer drifted.
- evasive, 0.00. When the judge marks every invented question as coming from a noncommittal answer, RAGAS multiplies the score by zero, whatever the similarity.
Answer relevancy vs faithfulness
| Answer relevancy | Faithfulness | |
|---|---|---|
| Reads | Question and answer | Question, answer and chunks |
| Asks | Is the answer about the question? | Is the answer backed by the chunks? |
| Judge calls per sample | One per invented question, plus embeddings | Two: claims, then verdicts |
| A faithful but off-topic answer | Scores low | Scores high |
When to use answer relevancy
- On live traffic: it needs neither a reference nor the chunks.
- When users complain the bot rambles or dodges, the padded and evasive cases above.
Related
- Previous: Faithfulness
- Next: Context precision
- Reference: Answer relevancy
- Add an incomplete answer,
"The ProBook X1 has 16GB DDR5 RAM.", the slide's third cause, and see where it lands. - Set
strictness=1and score the padded answer twice; compare how much the score moves.
You understood something today that you didn't yesterday.