Answer correctness
Answer correctness is a RAGAS metric that compares an answer with the reference answer, blending a factual F1 score over their claims with the semantic similarity of the two texts.
Last updated: 09 Oct, 2026 · RAGAS 0.4
Faithfulness checks the answer against the chunks and Answer relevancy checks it against the question. Neither says whether the answer is right. Answer correctness is the one metric of the five that puts the app's answer next to the golden's reference and grades the match.
This part of the video starts at 2:30:42. On the video's card, the two green claims are true positives and the two red claims are false positives: statements the response makes that the reference does not support.
Counting TP, FP and FN on the video's claims
The metric takes three inputs: the user_input, the response the app generated and the reference from the golden. The score has two components, and the first is factual.
The judge breaks the response into claims and the reference into claims, then sorts every claim into one of three groups:
- TP, true positive: a claim of the response that the reference also makes.
- FP, false positive: a claim of the response that the reference does not support.
- FN, false negative: a claim of the reference that the response left out.
The card's response makes four claims. "Minimum balance is ₹10,000 for urban branches" and "Interest rate is 3.5% per annum" match the reference: two TP. "Penalty fee is ₹400 per month" and "Internet banking is available at no cost" are not in the reference: two FP. The reference also says "Non-maintenance fee is ₹350 + taxes" and "Passbook is issued free of charge", which the response never mentions: two FN. The missed claims can be found only because the reference is there to compare with.
response_claims = {
"Minimum balance is ₹10,000 for urban branches",
"Penalty fee is ₹400 per month",
"Internet banking is available at no cost",
"Interest rate is 3.5% per annum",
}
reference_claims = {
"Minimum balance is ₹10,000 for urban branches",
"Interest rate is 3.5% per annum",
"Non-maintenance fee is ₹350 + taxes",
"Passbook is issued free of charge",
}
tp = response_claims & reference_claims # in both
fp = response_claims - reference_claims # said by the response, not in the reference
fn = reference_claims - response_claims # in the reference, missed by the response
for name, group in (("TP", tp), ("FP", fp), ("FN", fn)):
print(name, len(group), sorted(group))
precision = len(tp) / (len(tp) + len(fp))
recall = len(tp) / (len(tp) + len(fn))
f1 = len(tp) / (len(tp) + 0.5 * (len(fp) + len(fn)))
print(f"precision {precision} recall {recall} F1 {f1}")TP 2 ['Interest rate is 3.5% per annum', 'Minimum balance is ₹10,000 for urban branches'] FP 2 ['Internet banking is available at no cost', 'Penalty fee is ₹400 per month'] FN 2 ['Non-maintenance fee is ₹350 + taxes', 'Passbook is issued free of charge'] precision 0.5 recall 0.5 F1 0.5
What the three counts give
- TP, FP and FN are 2, 2 and 2, the three boxes on the card.
- Precision is 0.5: half of what the response claims is backed by the reference.
- Recall is 0.5: the response covers half of what the reference says.
- F1 is 0.5, the single number that combines the two; it is the card's 2 / (2 + 0.5 × (2 + 2)) = 0.50.
- Sets match exact strings. The hand count works because the claims are typed identically; the judge matches claims by meaning, which is why this step needs a model.
Blending in semantic similarity
This part of the video starts at 2:35:30. The second component is the cosine similarity between the embedding of the response and the embedding of the reference, one number for the two whole texts.
The card shows 0.72 for it, a demo value set with a slider. The note beside it reads "Texts share domain structure even when factually wrong": two texts about the same subject share structure and vocabulary even when the facts in them differ, so similarity alone would reward a wrong answer that sounds right.
The weights are yours to choose. The video calls them a hyperparameter: you decide whether exact facts or overall meaning matters more for your use case. The example reads the defaults from the installed class, then sets them explicitly.
import inspect
import matplotlib.pyplot as plt
import numpy as np
from ragas.metrics.collections import AnswerCorrectness
defaults = inspect.signature(AnswerCorrectness.__init__).parameters
print("installed defaults: weights", defaults["weights"].default, " beta", defaults["beta"].default)
f1, similarity = 0.50, 0.72 # the two components on the card
final = np.average([f1, similarity], weights=[0.75, 0.25])
print(f"0.75 × {f1} + 0.25 × {similarity} = {final:.3f}")
for weights in ([1.0, 0.0], [0.5, 0.5], [0.0, 1.0]):
print(f"weights {weights}: {np.average([f1, similarity], weights=weights):.3f}")
w1 = np.linspace(0, 1, 21) # weight of the factual part
plt.figure(figsize=(7, 3.8))
plt.plot(w1, w1 * f1 + (1 - w1) * similarity, color="#3a6fd8")
plt.scatter([0.75], [final], color="#d64541", zorder=3, label=f"default weights: {final:.3f}")
plt.title("Answer correctness for F1 = 0.50 and similarity = 0.72")
plt.xlabel("Weight of the factual F1 (the rest goes to semantic similarity)")
plt.ylabel("Answer correctness")
plt.ylim(0.45, 0.75)
plt.legend()
plt.show()installed defaults: weights [0.75, 0.25] beta 1.0 0.75 × 0.5 + 0.25 × 0.72 = 0.555 weights [1.0, 0.0]: 0.500 weights [0.5, 0.5]: 0.610 weights [0.0, 1.0]: 0.720
- The installed defaults are weights [0.75, 0.25] and beta 1.0, the 75% and 25% of the card.
betatilts the factual score: above 1 it favours recall, below 1 precision. - 0.75 × 0.5 + 0.25 × 0.72 = 0.555. The card prints this with two decimals as 0.55.
- Weights [1.0, 0.0] give 0.500, the factual F1 alone, and [0.0, 1.0] give 0.720, the similarity alone.
- The line in the plot falls from 0.72 to 0.50 as weight moves to the factual part: here the facts are the weaker side of the answer, so trusting them more lowers the score.
Passing embeddings, or setting their weight to 0
With the default weights, similarity has a share, so the metric needs an embedding model. Leaving it out fails at once:
import os
from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import AnswerCorrectness
client = AsyncOpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client)
AnswerCorrectness(llm=judge) # the default weights give similarity a share, and no embeddings are passedTraceback (most recent call last):
File "main.py", line 10, in <module>
AnswerCorrectness(llm=judge) # the default weights give similarity a share, and no embeddings are passed
ValueError: Embeddings are required for semantic similarity scoring. Either provide embeddings or set similarity weight to 0 (weights=[1.0, 0.0]) for pure factuality-only evaluation.There are two ways out: pass embeddings=, as the full run below does, or use weights=[1.0, 0.0] for a factual-only score that needs no embedding model. Answer relevancy needs embeddings as well; the other three metrics do not.
Scoring the card's example with RAGAS
The card shows the claims, not the two texts they came from. The response and the reference below are written from those claims, one sentence per claim, and the question is a plain one that fits them. The video's app embeds with a local sentence-transformers model and judges with llama-3.1-8b-instant; the run uses hosted gemini-embedding-2 and openai/gpt-oss-20b, set up as in Answer relevancy.
from ragas.metrics.collections import AnswerCorrectness
metric = AnswerCorrectness(llm=judge, embeddings=embeddings, weights=[0.75, 0.25]) # [factual, semantic]
metric.score(user_input=question, response=response, reference=reference).valueThe metric calls the judge three times: claims of the response, claims of the reference, then one call that sorts them into TP, FP and FN. The wrapper from Faithfulness keeps that last reply. Since the final score is 0.75 × F1 + 0.25 × similarity, the similarity part can be worked back out of it.
import os
from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
from ragas.metrics.collections import AnswerCorrectness
client = AsyncOpenAI(
base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"],
max_retries=6, # wait and retry when the per-minute token limit answers 429
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=client, temperature=0, max_tokens=4096)
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 returns one vector per call, so each text is embedded on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
replies = [] # each structured reply the judge sends back to the metric
ask = judge.agenerate
async def keep(prompt, response_model):
reply = await ask(prompt, response_model)
replies.append(reply)
return reply
judge.agenerate = keep
question = "What are the terms of my savings account?"
reference = ("The minimum balance is ₹10,000 for urban branches. The interest rate is 3.5% per annum. "
"The non-maintenance fee is ₹350 + taxes. The passbook is issued free of charge.")
response = ("The minimum balance is ₹10,000 for urban branches. The penalty fee is ₹400 per month. "
"Internet banking is available at no cost. The interest rate is 3.5% per annum.")
metric = AnswerCorrectness(llm=judge, embeddings=embeddings, weights=[0.75, 0.25])
result = metric.score(user_input=question, response=response, reference=reference)
verdict = replies[-1]
for label in ("TP", "FP", "FN"):
for item in getattr(verdict, label):
print(label, "|", item.statement)
tp, fp, fn = len(verdict.TP), len(verdict.FP), len(verdict.FN)
f1 = tp / (tp + 0.5 * (fp + fn))
print("factual F1 from the judge's lists:", round(f1, 4))
print("answer correctness:", round(result.value, 4))
print("semantic similarity behind it:", round((result.value - 0.75 * f1) / 0.25, 4))TP | Minimum balance for urban branches is ₹10,000. TP | Interest rate is 3.5% per annum. FP | Penalty fee per month is ₹400. FP | Internet banking is available at no cost. FN | The non-maintenance fee is ₹350 plus taxes. FN | The passbook is issued free of charge. factual F1 from the judge's lists: 0.5 answer correctness: 0.6095 semantic similarity behind it: 0.938
What the judge counted
- The judge sorted the claims the way the card does: two TP, two FP and two FN, so the factual F1 from its lists is 0.5.
- The final score is 0.6095, above the card's 0.555. The factual part is the same; the difference is the similarity.
- The similarity behind the score works out to 0.938, where the card's slider shows 0.72. Two short texts about the same savings account sit very close in embedding space although half of their claims differ.
- That is the card's warning in numbers. Similarity lifted a half-wrong answer from 0.5 to about 0.61. With
weights=[1.0, 0.0]the score is the factual F1 alone.
Answer correctness vs answer relevancy
The video opens this metric with the question of how it differs from answer relevancy. The two read different things:
| Answer correctness | Answer relevancy | |
|---|---|---|
| The answer is compared with | The golden's reference | The user's question |
| Needs a golden | Yes | No |
| Checks facts | Yes, claim by claim | No |
| A direct answer with a wrong number | Scores low | Scores high |
| A correct answer to a different question | Scores low against this golden | Scores low |
| Judge calls per sample in RAGAS 0.4 | 3 | One per generated question |
Where you use answer correctness
- Regression checks on a golden set. It is the closest of the five to "is the answer right", so it is the number to watch when a prompt, model or retriever changes.
- Answers with exact values. Prices, limits and dates: weight the factual part high, up to
weights=[1.0, 0.0]. - Free-form answers. Where wording varies a lot and facts are few, give similarity a larger share.
Related
- Previous: Context recall
- Next: Reading evaluation results
- Reference: Answer correctness in the RAGAS docs
- In the hand count, add
"Non-maintenance fee is ₹350 + taxes"toresponse_claims: TP becomes 3 and FN 1, and F1 rises to 0.6666666666666666. - In the blend example, change
similarityto0.95: the default weights now give 0.613, while weights [1.0, 0.0] still give 0.500. - In the example that fails, change the last line to
AnswerCorrectness(llm=judge, weights=[1.0, 0.0]): the error is gone, because no similarity is computed.
Slow is fine. Stopping is the only problem.