Answer correctness
Answer correctness is a RAGAS metric that grades an answer against the reference answer by combining a factual F1 score over their claims with the semantic similarity of the two texts.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
Faithfulness asks whether the answer stayed inside the chunks. Answer correctness asks whether it is right: does it say what the golden says? It is the closest thing RAGAS has to marking a paper against the answer key.
True positives, false positives and false negatives
The video first separates this metric from answer relevancy. It reads three inputs: the user input, the response the RAG app generated, and the reference, the expected answer. The score has two components: a factual one and a semantic one.
For the factual part, the judge extracts claims from both texts and sorts them. Claims in both are true positives; claims only in the response are false positives, including hallucinations; claims only in the reference are false negatives, facts the response missed. In the video's bank example the response gets two right (the ₹10,000 urban minimum and 3.5% interest), adds two wrong (a ₹400 monthly penalty and free internet banking), and misses two (the ₹350 fee and the free passbook). F1 = TP / (TP + 0.5 × (FP + FN)) = 2 / (2 + 0.5 × 4) = 0.50.
Blending in semantic similarity
The second component embeds the response and the reference and takes their cosine similarity. The final score is a weighted sum, w1 × factual F1 + w2 × similarity. The weights are yours to choose: raise w1 if the facts matter most. The video's slide uses 0.75 and 0.25, which with a similarity of 0.72 gives 0.75 × 0.50 + 0.25 × 0.72 = 0.55.
The AnswerCorrectness API
from ragas.metrics.collections import AnswerCorrectness
correctness = AnswerCorrectness(llm=judge, embeddings=embeddings, weights=[0.75, 0.25]) # [factual, semantic]
correctness.score(user_input=..., response=..., reference=...).valueThe reference: the answer key
reference = ("The minimum balance is ₹10,000 for urban branches. The savings account earns 3.5% interest "
"per annum. The non-maintenance fee is ₹350 + taxes. The passbook is issued free of charge.")The response with two right, two wrong
response = ("The minimum balance is ₹10,000 for urban branches. A penalty fee of ₹400 per month applies. "
"Internet banking is available at no cost. The interest rate is 3.5% per annum.")The weights are set to the slide's values rather than left to the default, so the run uses exactly the blend described above.
- written in LLM as a judge
- written in Goldens
View the code here
import os
from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
groq = AsyncOpenAI(
api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
[
{
"id": "g001",
"metric_focus": "faithfulness",
"user_input": "What is TechNest's return policy?",
"reference": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
},
{
"id": "g002",
"metric_focus": "answer_relevancy",
"user_input": "What are the RAM and storage specs of the ProBook X1?",
"reference": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
},
{
"id": "g003",
"metric_focus": "context_precision",
"user_input": "How long is the battery life on the SoundPods Pro?",
"reference": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
},
{
"id": "g004",
"metric_focus": "context_recall",
"user_input": "What are TechNest's shipping options and how long do returns take to process?",
"reference": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
},
{
"id": "g005",
"metric_focus": "answer_correctness",
"user_input": "What is the price of the PixelPhone 15?",
"reference": "The TechNest PixelPhone 15 is priced at $899."
}
]
The video's bank example
from judge import embeddings, judge
from ragas.metrics.collections import AnswerCorrectness
question = "What are the terms of my savings account?"
reference = ("The minimum balance is ₹10,000 for urban branches. The savings account earns 3.5% interest "
"per annum. The non-maintenance fee is ₹350 + taxes. The passbook is issued free of charge.")
response = ("The minimum balance is ₹10,000 for urban branches. A penalty fee of ₹400 per month applies. "
"Internet banking is available at no cost. The interest rate is 3.5% per annum.")
blended = AnswerCorrectness(llm=judge, embeddings=embeddings, weights=[0.75, 0.25])
factual_only = AnswerCorrectness(llm=judge, weights=[1.0, 0.0])
print("blended 0.75/0.25:", round(blended.score(user_input=question, response=response, reference=reference).value, 2))
print("factual F1 only: ", round(factual_only.score(user_input=question, response=response, reference=reference).value, 2))blended 0.75/0.25: 0.61 factual F1 only: 0.5
weights=[1.0, 0.0] turns the semantic part off, which also means no embeddings are needed, so the second line is the factual F1 alone: 0.5, the slide's 0.50 exactly. The blended score is 0.61 rather than the slide's 0.55 because Gemini's embeddings rate the two texts as more alike than the slide's 0.72; working back, 0.75 × 0.50 + 0.25 × s = 0.61 puts the similarity near 0.94.
Correctness of the TechNest price answer
import json
from judge import embeddings, judge
from ragas.metrics.collections import AnswerCorrectness
golden = json.load(open("goldens.json", encoding="utf-8"))[4]
correctness = AnswerCorrectness(llm=judge, embeddings=embeddings, weights=[0.75, 0.25])
for response in ["The PixelPhone 15 costs $899.", "The PixelPhone 15 costs $799."]:
result = correctness.score(user_input=golden["user_input"], response=response, reference=golden["reference"])
print(f"{result.value:.2f} {response}")0.98 The PixelPhone 15 costs $899. 0.22 The PixelPhone 15 costs $799.
What the two components did
- $899 matches the golden's claim, a true positive, and the texts are close in meaning, so both parts score high.
- $799 is a false positive and the golden's $899 becomes a false negative, so the factual part falls. The two sentences still read alike, so the semantic part stays high and keeps the blended score above zero: the video's point that texts can share structure while being factually wrong.
Answer correctness vs answer relevancy
| Answer correctness | Answer relevancy | |
|---|---|---|
| Compares the answer with | The reference | The question |
| Needs a golden? | Yes | No |
| A wrong price | Scores low | Can score high |
| Uses embeddings for | Similarity to the reference | Similarity of invented questions to the real one |
When to use answer correctness
- When you have expert-written references and want to know whether the bot gets the facts right.
- For single-fact goldens, such as prices and specs, where one wrong number must fail.
Related
- Previous: Context recall
- Next: Factual correctness
- Reference: Answer correctness
- Change the bank weights to
[0.5, 0.5]and see how far the blended score moves toward the similarity. - Score
"It is $899."for the PixelPhone and compare it with the full sentence.
Every expert started right here.