AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Answer relevancy

Answer relevancy is a RAGAS metric that measures how directly an answer addresses the question, as the mean cosine similarity between the question and questions a judge writes from the answer.

Last updated: 09 Oct, 2026 · RAGAS 0.4

An answer can be faithful to its chunks and still miss the point: right facts about the wrong thing, or a reply that dodges. Faithfulness cannot see that, because it never compares the answer with the question. Answer relevancy does.

Answer relevancy: the inputs and the two steps · from the Complete AI Security Course in 8 Hours video · 2:12:38 to 2:14:19

This part of the video starts at 2:12:38. It names the two inputs of the metric and the two steps that turn them into a score.

Working backwards from the answer

The metric needs two things: the question the user asked and the answer the app gave. The card in the video labels them input and actual_output; RAGAS calls them user_input and response. The retrieved chunks and the reference are not read.

The judge then works backwards, in two steps:

  1. Generate questions. The judge reads the answer and writes a question that this answer would be a good reply to. RAGAS asks for this N times, one question per call. N is the strictness argument, set to 3 below.
  2. Check similarity. Each generated question and the user's real question are turned into embeddings, and the cosine similarity of each pair is computed. The mean of those values is the score.

The idea is a reverse check. If the answer is on topic, a question written from it lands close to the question that was asked. If the answer talks about something else, the generated questions drift away and the similarity falls. The comparison is on meaning, through embeddings, not on shared keywords.

The judge reads the response and writes three questions it would answer, one per call, each with a noncommittal flag. The three generated questions and the user's question are embedded, the cosine similarity of each generated question with the user's question is computed, and the score is the mean of the three cosines, multiplied by 0 when every question is flagged noncommittal.

Writing answer relevancy as a formula

E_gi is the embedding of generated question i, E_o the embedding of the user's question, N the number of generated questions

Cosine similarity runs from −1 to 1, so the mean is not squeezed into 0 to 1 by any extra step. For texts on a shared topic it lands between 0 and 1 in practice, and the RAGAS docs say that range is usual but not guaranteed.

One more rule sits in the installed code. With each question the judge returns a noncommittal flag: 1 when the answer is evasive ("I am not sure"), 0 when it commits to something. If every generated question carries a 1, the score is multiplied by 0.

Computing the cosines and the mean by hand

Three toy vectors stand in for the embeddings of three generated questions, so each step of the formula can be checked on paper.

ExampleRun on NumPy 2.5.3
import numpy as np


def cosine(a, b):
    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))


question = np.array([1.0, 2.0, 0.0])      # a toy 3-number embedding of the user's question
generated = [np.array([1.0, 2.0, 0.0]),   # points the same way as the question
             np.array([2.0, 3.0, 1.0]),   # close to it
             np.array([0.0, 1.0, 3.0])]   # mostly about something else

cosines = [cosine(g, question) for g in generated]
for number, value in enumerate(cosines, 1):
    print(f"generated question {number}: cosine {value:.4f}")

for flags in ([0, 0, 0], [1, 1, 1]):
    score = np.mean(cosines) * int(not all(flags))
    print(f"noncommittal flags {flags}: answer relevancy {score:.4f}")
print("opposite direction:", round(cosine(-question, question), 4))

What the toy vectors show

  • Same direction gives 1.0000. The first generated question has the same vector as the user's question.
  • The further a vector turns away, the lower the cosine: 0.9562 for the close one, 0.2828 for the one about something else.
  • The score is their mean, 0.7463. One off-topic question pulls the whole score down.
  • All three flags at 1 give 0.0000, whatever the cosines were. That is how an evasive answer scores zero.
  • Vectors that point in opposite directions give −1.0, the lower end of the cosine range.

Scoring the video's answer with RAGAS

Answer relevancy on the TechNest return policy question · from the Complete AI Security Course in 8 Hours video · 2:15:48 to 2:16:44

This part of the video starts at 2:15:48. The comparison it describes is a cosine similarity computed directly between embedding vectors; nothing is searched.

The video's example for this metric is the first golden of its TechNest app: the question "What is TechNest's return policy?" and the answer the app generated for it, as shown on the Run Evaluation screen.

The video's app judges with llama-3.1-8b-instant, since retired on Groq, and embeds with a local sentence-transformers model, which has to be downloaded; the run below uses Google's hosted gemini-embedding-2 through RAGAS's own GoogleEmbeddings class, and the judge from LLM as a judge.

One text per embedding call

python
class OneTextPerCall(GoogleEmbeddings):
    """gemini-embedding-2 returns one vector per call, so each text is embedded on its own."""

    def embed_texts(self, texts, **kwargs):
        return [self.embed_text(text) for text in texts]

    async def aembed_texts(self, texts, **kwargs):
        return [await self.aembed_text(text) for text in texts]


embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")

gemini-embedding-2 answers a list of texts with a single vector. The metric embeds its three generated questions in one call and expects three vectors back, so this small subclass sends each text on its own.

The metric call

python
from ragas.metrics.collections import AnswerRelevancy

metric = AnswerRelevancy(llm=judge, embeddings=embeddings, strictness=3)  # 3 generated questions
metric.score(user_input=question, response=response).value

The example keeps the judge's replies with the same wrapper as in Faithfulness, prints each generated question with its cosine computed by hand on the real embedding vectors, and then prints RAGAS's score. It needs GROQ_API_KEY for the judge and GEMINI_API_KEY for the embeddings.

ExampleAPI keyFrom the video, run on Groq and Gemini
import os

import numpy as np
from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
from ragas.metrics.collections import AnswerRelevancy

client = AsyncOpenAI(
    base_url="https://api.groq.com/openai/v1",
    api_key=os.environ["GROQ_API_KEY"],
    max_retries=6,  # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)


class OneTextPerCall(GoogleEmbeddings):
    """gemini-embedding-2 returns one vector per call, so each text is embedded on its own."""

    def embed_texts(self, texts, **kwargs):
        return [self.embed_text(text) for text in texts]

    async def aembed_texts(self, texts, **kwargs):
        return [await self.aembed_text(text) for text in texts]


embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")

replies = []  # each structured reply the judge sends back to the metric
ask = judge.agenerate


async def keep(prompt, response_model):
    reply = await ask(prompt, response_model)
    replies.append(reply)
    return reply


judge.agenerate = keep

question = "What is TechNest's return policy?"
response = ("TechNest's return policy allows returns within 30 days of the original purchase date, as long as items "
            "are in their original packaging with all accessories included. You're responsible for return shipping "
            "costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days "
            "of receiving the returned item. Note that digital downloads and opened software are non-refundable.")

metric = AnswerRelevancy(llm=judge, embeddings=embeddings, strictness=3)
result = metric.score(user_input=question, response=response)

asked = np.array(embeddings.embed_text(question))
cosines = []
for reply in replies:
    written = np.array(embeddings.embed_text(reply.question))
    cosines.append(float(asked @ written / (np.linalg.norm(asked) * np.linalg.norm(written))))
    print(f"{cosines[-1]:.4f} | noncommittal {reply.noncommittal} | {reply.question}")
print("mean by hand:    ", round(sum(cosines) / len(cosines), 4))
print("answer relevancy:", round(result.value, 4))

What the three generated questions show

  • The judge wrote three questions, and two of them are word for word the same: "What are the conditions for returning items to TechNest?". At temperature 0 the three calls are near copies of each other, so a strictness of 3 did not give three independent guesses here.
  • All three sit close to the real question. Their cosines with "What is TechNest's return policy?" are 0.9605, 0.9450 and 0.9450.
  • The mean computed by hand, 0.9502, equals the score RAGAS returned. The metric is that mean.
  • Every noncommittal flag is 0. The answer commits to a policy, so the score is not set to zero.
  • The video's Results screen shows 1.00 for this golden. That run used another judge and another embedding model. Cosine values depend on the embedding model, so compare relevancy scores only between runs made with the same judge and the same embeddings.

What lowers answer relevancy

The card lists three causes of a low score: an off-topic answer, a padded response and an incomplete answer. The formula says how each one acts:

  • Off-topic answer. The questions the judge writes are about the other topic, so their cosines with the real question fall. This is the case the metric catches best.
  • Evasive answer. The noncommittal flags take the score to 0.
  • Padded or incomplete answer. These lower the score only when they change the questions the judge writes. Each generated question is a guess at the whole question behind the answer, not at one part of it, so filler around a direct answer, or one missing detail, can leave the score almost unchanged.

Answer relevancy vs faithfulness

Answer relevancyFaithfulness
The answer is compared withThe user's questionThe retrieved chunks
Reads the chunksNoYes
Needs embeddingsYesNo
CatchesAnswers that drift off the question or dodge itClaims that no chunk supports
MissesWrong facts stated directlyAnswers that are grounded but off the question

Where you use answer relevancy

  • On live traffic. It needs neither a golden nor the chunks, only the question and the answer.
  • When users say the bot rambles or dodges. Off-topic and evasive replies are what it is built to flag.
  • Next to faithfulness. Together they cover both sides of the generator: staying on the question and staying inside the context.
Watch out. Answer relevancy does not check facts. "The ProBook X1 has 64GB RAM" answers a question about the laptop's memory directly and scores high while being wrong. Faithfulness and answer correctness are the metrics that catch a wrong fact.
Try it yourself
  • In the toy example, replace the third vector with np.array([1.0, 2.0, 1.0]): its cosine rises to 0.9129 and the score to 0.9564.
  • In the toy example, change the second list of flags to [1, 1, 0]: one committed question is enough, and the score stays 0.7463.
  • In the RAGAS example, change strictness=3 to strictness=1: one generated question is printed, and the score equals its single cosine.
PreviousFaithfulness

Every expert started right here.