RAGASragas 0.4.3 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

Answer correctness

Answer correctness is a RAGAS metric that grades an answer against the reference answer by combining a factual F1 score over their claims with the semantic similarity of the two texts.

Last updated: 29 Sep, 2026 · RAGAS 0.4.3

Faithfulness asks whether the answer stayed inside the chunks. Answer correctness asks whether it is right: does it say what the golden says? It is the closest thing RAGAS has to marking a paper against the answer key.

Answer correctness: the factual component · from the Complete AI Security Course In 8 Hours · 150:42 to 154:47

True positives, false positives and false negatives

The video first separates this metric from answer relevancy. It reads three inputs: the user input, the response the RAG app generated, and the reference, the expected answer. The score has two components: a factual one and a semantic one.

For the factual part, the judge extracts claims from both texts and sorts them. Claims in both are true positives; claims only in the response are false positives, including hallucinations; claims only in the reference are false negatives, facts the response missed. In the video's bank example the response gets two right (the ₹10,000 urban minimum and 3.5% interest), adds two wrong (a ₹400 monthly penalty and free internet banking), and misses two (the ₹350 fee and the free passbook). F1 = TP / (TP + 0.5 × (FP + FN)) = 2 / (2 + 0.5 × 4) = 0.50.

The semantic component and the weights · from the Complete AI Security Course In 8 Hours · 154:49 to 156:56

Blending in semantic similarity

The second component embeds the response and the reference and takes their cosine similarity. The final score is a weighted sum, w1 × factual F1 + w2 × similarity. The weights are yours to choose: raise w1 if the facts matter most. The video's slide uses 0.75 and 0.25, which with a similarity of 0.72 gives 0.75 × 0.50 + 0.25 × 0.72 = 0.55.

The AnswerCorrectness API

python
from ragas.metrics.collections import AnswerCorrectness

correctness = AnswerCorrectness(llm=judge, embeddings=embeddings, weights=[0.75, 0.25])  # [factual, semantic]
correctness.score(user_input=..., response=..., reference=...).value

The reference: the answer key

python
reference = ("The minimum balance is ₹10,000 for urban branches. The savings account earns 3.5% interest "
             "per annum. The non-maintenance fee is ₹350 + taxes. The passbook is issued free of charge.")

The response with two right, two wrong

python
response = ("The minimum balance is ₹10,000 for urban branches. A penalty fee of ₹400 per month applies. "
            "Internet banking is available at no cost. The interest rate is 3.5% per annum.")

The weights are set to the slide's values rather than left to the default, so the run uses exactly the blend described above.

Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
judge.py
import os

from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory

groq = AsyncOpenAI(
    api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
    base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)


class OneTextPerCall(GoogleEmbeddings):
    """gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""

    def embed_texts(self, texts, **kwargs):
        return [self.embed_text(text) for text in texts]

    async def aembed_texts(self, texts, **kwargs):
        return [await self.aembed_text(text) for text in texts]


embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
goldens.json
[
  {
    "id": "g001",
    "metric_focus": "faithfulness",
    "user_input": "What is TechNest's return policy?",
    "reference": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
  },
  {
    "id": "g002",
    "metric_focus": "answer_relevancy",
    "user_input": "What are the RAM and storage specs of the ProBook X1?",
    "reference": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
  },
  {
    "id": "g003",
    "metric_focus": "context_precision",
    "user_input": "How long is the battery life on the SoundPods Pro?",
    "reference": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
  },
  {
    "id": "g004",
    "metric_focus": "context_recall",
    "user_input": "What are TechNest's shipping options and how long do returns take to process?",
    "reference": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
  },
  {
    "id": "g005",
    "metric_focus": "answer_correctness",
    "user_input": "What is the price of the PixelPhone 15?",
    "reference": "The TechNest PixelPhone 15 is priced at $899."
  }
]

The video's bank example

ExampleAPI keyFrom the video, run on Groq and Gemini
from judge import embeddings, judge
from ragas.metrics.collections import AnswerCorrectness

question = "What are the terms of my savings account?"
reference = ("The minimum balance is ₹10,000 for urban branches. The savings account earns 3.5% interest "
             "per annum. The non-maintenance fee is ₹350 + taxes. The passbook is issued free of charge.")
response = ("The minimum balance is ₹10,000 for urban branches. A penalty fee of ₹400 per month applies. "
            "Internet banking is available at no cost. The interest rate is 3.5% per annum.")

blended = AnswerCorrectness(llm=judge, embeddings=embeddings, weights=[0.75, 0.25])
factual_only = AnswerCorrectness(llm=judge, weights=[1.0, 0.0])
print("blended 0.75/0.25:", round(blended.score(user_input=question, response=response, reference=reference).value, 2))
print("factual F1 only:  ", round(factual_only.score(user_input=question, response=response, reference=reference).value, 2))

weights=[1.0, 0.0] turns the semantic part off, which also means no embeddings are needed, so the second line is the factual F1 alone: 0.5, the slide's 0.50 exactly. The blended score is 0.61 rather than the slide's 0.55 because Gemini's embeddings rate the two texts as more alike than the slide's 0.72; working back, 0.75 × 0.50 + 0.25 × s = 0.61 puts the similarity near 0.94.

Correctness of the TechNest price answer

ExampleAPI key
import json

from judge import embeddings, judge
from ragas.metrics.collections import AnswerCorrectness

golden = json.load(open("goldens.json", encoding="utf-8"))[4]
correctness = AnswerCorrectness(llm=judge, embeddings=embeddings, weights=[0.75, 0.25])
for response in ["The PixelPhone 15 costs $899.", "The PixelPhone 15 costs $799."]:
    result = correctness.score(user_input=golden["user_input"], response=response, reference=golden["reference"])
    print(f"{result.value:.2f}  {response}")

What the two components did

  • $899 matches the golden's claim, a true positive, and the texts are close in meaning, so both parts score high.
  • $799 is a false positive and the golden's $899 becomes a false negative, so the factual part falls. The two sentences still read alike, so the semantic part stays high and keeps the blended score above zero: the video's point that texts can share structure while being factually wrong.

Answer correctness vs answer relevancy

Answer correctnessAnswer relevancy
Compares the answer withThe referenceThe question
Needs a golden?YesNo
A wrong priceScores lowCan score high
Uses embeddings forSimilarity to the referenceSimilarity of invented questions to the real one

When to use answer correctness

  • When you have expert-written references and want to know whether the bot gets the facts right.
  • For single-fact goldens, such as prices and specs, where one wrong number must fail.
Watch out. With the semantic weight above zero, a confidently wrong answer that is worded like the reference still earns part of the score. For facts that must be exact, raise the factual weight or read the factual F1 on its own.
Try it yourself
  • Change the bank weights to [0.5, 0.5] and see how far the blended score moves toward the similarity.
  • Score "It is $899." for the PixelPhone and compare it with the full sentence.

Every expert started right here.