AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Answer correctness

Answer correctness is a RAGAS metric that compares an answer with the reference answer, blending a factual F1 score over their claims with the semantic similarity of the two texts.

Last updated: 09 Oct, 2026 · RAGAS 0.4

Faithfulness checks the answer against the chunks and Answer relevancy checks it against the question. Neither says whether the answer is right. Answer correctness is the one metric of the five that puts the app's answer next to the golden's reference and grades the match.

Answer correctness and the factual F1 · from the Complete AI Security Course in 8 Hours video · 2:30:42 to 2:34:41

This part of the video starts at 2:30:42. On the video's card, the two green claims are true positives and the two red claims are false positives: statements the response makes that the reference does not support.

Counting TP, FP and FN on the video's claims

The metric takes three inputs: the user_input, the response the app generated and the reference from the golden. The score has two components, and the first is factual.

The judge breaks the response into claims and the reference into claims, then sorts every claim into one of three groups:

  • TP, true positive: a claim of the response that the reference also makes.
  • FP, false positive: a claim of the response that the reference does not support.
  • FN, false negative: a claim of the reference that the response left out.

The card's response makes four claims. "Minimum balance is ₹10,000 for urban branches" and "Interest rate is 3.5% per annum" match the reference: two TP. "Penalty fee is ₹400 per month" and "Internet banking is available at no cost" are not in the reference: two FP. The reference also says "Non-maintenance fee is ₹350 + taxes" and "Passbook is issued free of charge", which the response never mentions: two FN. The missed claims can be found only because the reference is there to compare with.

The factual score: an F1 over the claims
ExampleThe video's example, computed in plain Python
response_claims = {
    "Minimum balance is ₹10,000 for urban branches",
    "Penalty fee is ₹400 per month",
    "Internet banking is available at no cost",
    "Interest rate is 3.5% per annum",
}
reference_claims = {
    "Minimum balance is ₹10,000 for urban branches",
    "Interest rate is 3.5% per annum",
    "Non-maintenance fee is ₹350 + taxes",
    "Passbook is issued free of charge",
}

tp = response_claims & reference_claims   # in both
fp = response_claims - reference_claims   # said by the response, not in the reference
fn = reference_claims - response_claims   # in the reference, missed by the response

for name, group in (("TP", tp), ("FP", fp), ("FN", fn)):
    print(name, len(group), sorted(group))
precision = len(tp) / (len(tp) + len(fp))
recall = len(tp) / (len(tp) + len(fn))
f1 = len(tp) / (len(tp) + 0.5 * (len(fp) + len(fn)))
print(f"precision {precision}  recall {recall}  F1 {f1}")

What the three counts give

  • TP, FP and FN are 2, 2 and 2, the three boxes on the card.
  • Precision is 0.5: half of what the response claims is backed by the reference.
  • Recall is 0.5: the response covers half of what the reference says.
  • F1 is 0.5, the single number that combines the two; it is the card's 2 / (2 + 0.5 × (2 + 2)) = 0.50.
  • Sets match exact strings. The hand count works because the claims are typed identically; the judge matches claims by meaning, which is why this step needs a model.

Blending in semantic similarity

Semantic similarity and the blend weights · from the Complete AI Security Course in 8 Hours video · 2:35:30 to 2:36:56

This part of the video starts at 2:35:30. The second component is the cosine similarity between the embedding of the response and the embedding of the reference, one number for the two whole texts.

The card shows 0.72 for it, a demo value set with a slider. The note beside it reads "Texts share domain structure even when factually wrong": two texts about the same subject share structure and vocabulary even when the facts in them differ, so similarity alone would reward a wrong answer that sounds right.

A weighted average; with weights that add up to 1 the denominator disappears
Two lanes merge into one score. In the first, the claims of the response are compared with the claims of the reference and counted as 2 true positives, 2 false positives and 2 false negatives, which gives a factual F1 of 0.50. In the second, the embeddings of the response and the reference give a cosine similarity of 0.72. With weights 0.75 and 0.25 the final score is 0.555.

The weights are yours to choose. The video calls them a hyperparameter: you decide whether exact facts or overall meaning matters more for your use case. The example reads the defaults from the installed class, then sets them explicitly.

ExampleThe video's example, run on RAGAS 0.4.3 and NumPy 2.5.3
import inspect

import matplotlib.pyplot as plt
import numpy as np
from ragas.metrics.collections import AnswerCorrectness

defaults = inspect.signature(AnswerCorrectness.__init__).parameters
print("installed defaults: weights", defaults["weights"].default, " beta", defaults["beta"].default)

f1, similarity = 0.50, 0.72  # the two components on the card
final = np.average([f1, similarity], weights=[0.75, 0.25])
print(f"0.75 × {f1} + 0.25 × {similarity} = {final:.3f}")
for weights in ([1.0, 0.0], [0.5, 0.5], [0.0, 1.0]):
    print(f"weights {weights}: {np.average([f1, similarity], weights=weights):.3f}")

w1 = np.linspace(0, 1, 21)  # weight of the factual part
plt.figure(figsize=(7, 3.8))
plt.plot(w1, w1 * f1 + (1 - w1) * similarity, color="#3a6fd8")
plt.scatter([0.75], [final], color="#d64541", zorder=3, label=f"default weights: {final:.3f}")
plt.title("Answer correctness for F1 = 0.50 and similarity = 0.72")
plt.xlabel("Weight of the factual F1 (the rest goes to semantic similarity)")
plt.ylabel("Answer correctness")
plt.ylim(0.45, 0.75)
plt.legend()
plt.show()
A straight line showing answer correctness falling from 0.72 at a factual weight of 0 to 0.50 at a factual weight of 1, with a red point at the default factual weight of 0.75, where the score is 0.555.
  • The installed defaults are weights [0.75, 0.25] and beta 1.0, the 75% and 25% of the card. beta tilts the factual score: above 1 it favours recall, below 1 precision.
  • 0.75 × 0.5 + 0.25 × 0.72 = 0.555. The card prints this with two decimals as 0.55.
  • Weights [1.0, 0.0] give 0.500, the factual F1 alone, and [0.0, 1.0] give 0.720, the similarity alone.
  • The line in the plot falls from 0.72 to 0.50 as weight moves to the factual part: here the facts are the weaker side of the answer, so trusting them more lowers the score.

Passing embeddings, or setting their weight to 0

With the default weights, similarity has a share, so the metric needs an embedding model. Leaving it out fails at once:

ExampleRun on RAGAS 0.4.3
import os

from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import AnswerCorrectness

client = AsyncOpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client)

AnswerCorrectness(llm=judge)  # the default weights give similarity a share, and no embeddings are passed

There are two ways out: pass embeddings=, as the full run below does, or use weights=[1.0, 0.0] for a factual-only score that needs no embedding model. Answer relevancy needs embeddings as well; the other three metrics do not.

Scoring the card's example with RAGAS

The card shows the claims, not the two texts they came from. The response and the reference below are written from those claims, one sentence per claim, and the question is a plain one that fits them. The video's app embeds with a local sentence-transformers model and judges with llama-3.1-8b-instant; the run uses hosted gemini-embedding-2 and openai/gpt-oss-20b, set up as in Answer relevancy.

python
from ragas.metrics.collections import AnswerCorrectness

metric = AnswerCorrectness(llm=judge, embeddings=embeddings, weights=[0.75, 0.25])  # [factual, semantic]
metric.score(user_input=question, response=response, reference=reference).value

The metric calls the judge three times: claims of the response, claims of the reference, then one call that sorts them into TP, FP and FN. The wrapper from Faithfulness keeps that last reply. Since the final score is 0.75 × F1 + 0.25 × similarity, the similarity part can be worked back out of it.

ExampleAPI keyFrom the video, run on Groq and Gemini
import os

from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
from ragas.metrics.collections import AnswerCorrectness

client = AsyncOpenAI(
    base_url="https://api.groq.com/openai/v1",
    api_key=os.environ["GROQ_API_KEY"],
    max_retries=6,  # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)


class OneTextPerCall(GoogleEmbeddings):
    """gemini-embedding-2 returns one vector per call, so each text is embedded on its own."""

    def embed_texts(self, texts, **kwargs):
        return [self.embed_text(text) for text in texts]

    async def aembed_texts(self, texts, **kwargs):
        return [await self.aembed_text(text) for text in texts]


embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")

replies = []  # each structured reply the judge sends back to the metric
ask = judge.agenerate


async def keep(prompt, response_model):
    reply = await ask(prompt, response_model)
    replies.append(reply)
    return reply


judge.agenerate = keep

question = "What are the terms of my savings account?"
reference = ("The minimum balance is ₹10,000 for urban branches. The interest rate is 3.5% per annum. "
             "The non-maintenance fee is ₹350 + taxes. The passbook is issued free of charge.")
response = ("The minimum balance is ₹10,000 for urban branches. The penalty fee is ₹400 per month. "
            "Internet banking is available at no cost. The interest rate is 3.5% per annum.")

metric = AnswerCorrectness(llm=judge, embeddings=embeddings, weights=[0.75, 0.25])
result = metric.score(user_input=question, response=response, reference=reference)

verdict = replies[-1]
for label in ("TP", "FP", "FN"):
    for item in getattr(verdict, label):
        print(label, "|", item.statement)
tp, fp, fn = len(verdict.TP), len(verdict.FP), len(verdict.FN)
f1 = tp / (tp + 0.5 * (fp + fn))
print("factual F1 from the judge's lists:", round(f1, 4))
print("answer correctness:", round(result.value, 4))
print("semantic similarity behind it:", round((result.value - 0.75 * f1) / 0.25, 4))

What the judge counted

  • The judge sorted the claims the way the card does: two TP, two FP and two FN, so the factual F1 from its lists is 0.5.
  • The final score is 0.6095, above the card's 0.555. The factual part is the same; the difference is the similarity.
  • The similarity behind the score works out to 0.938, where the card's slider shows 0.72. Two short texts about the same savings account sit very close in embedding space although half of their claims differ.
  • That is the card's warning in numbers. Similarity lifted a half-wrong answer from 0.5 to about 0.61. With weights=[1.0, 0.0] the score is the factual F1 alone.

Answer correctness vs answer relevancy

The video opens this metric with the question of how it differs from answer relevancy. The two read different things:

Answer correctnessAnswer relevancy
The answer is compared withThe golden's referenceThe user's question
Needs a goldenYesNo
Checks factsYes, claim by claimNo
A direct answer with a wrong numberScores lowScores high
A correct answer to a different questionScores low against this goldenScores low
Judge calls per sample in RAGAS 0.43One per generated question

Where you use answer correctness

  • Regression checks on a golden set. It is the closest of the five to "is the answer right", so it is the number to watch when a prompt, model or retriever changes.
  • Answers with exact values. Prices, limits and dates: weight the factual part high, up to weights=[1.0, 0.0].
  • Free-form answers. Where wording varies a lot and facts are few, give similarity a larger share.
Watch out. A false positive is a claim the reference does not have, which is wider than a false claim. A correct extra detail counts against the answer in the same way an invented one does. On the video's Results screen the return-policy golden scores 0.64 on answer correctness while its faithfulness is 0.86. One difference is visible on that screen: the app's answer ends with a sentence about digital downloads that its reference does not contain. Keep references as complete as a good answer, and read a low score claim by claim before blaming the model.
Try it yourself
  • In the hand count, add "Non-maintenance fee is ₹350 + taxes" to response_claims: TP becomes 3 and FN 1, and F1 rises to 0.6666666666666666.
  • In the blend example, change similarity to 0.95: the default weights now give 0.613, while weights [1.0, 0.0] still give 0.500.
  • In the example that fails, change the last line to AnswerCorrectness(llm=judge, weights=[1.0, 0.0]): the error is gone, because no similarity is computed.

Slow is fine. Stopping is the only problem.