1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →
Factual correctness
Factual correctness is a RAGAS metric that compares the claims in an answer with the claims in a reference and reports their precision, recall or F1, with no embedding step.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
Answer correctness blends facts with wording. When you want the facts alone, and want to choose whether extra claims or missing claims matter more, factual correctness gives you the claim comparison on its own.
The FactualCorrectness API
from ragas.metrics.collections import FactualCorrectness
FactualCorrectness(llm=judge, mode="f1") # both extra and missing claims count
FactualCorrectness(llm=judge, mode="precision") # only extra claims lower the score
FactualCorrectness(llm=judge, mode="recall") # only missing claims lower the scoreIt reads only response and reference: no question, no chunks.
The reference from golden g002
reference = "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."A complete answer and a partial one
answers = ["The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD.",
"The ProBook X1 has 16GB DDR5 RAM."]Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
- written in LLM as a judge
View the code here
judge.py
import os
from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
groq = AsyncOpenAI(
api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
Three modes on two answers
from judge import judge
from ragas.metrics.collections import FactualCorrectness
reference = "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
answers = ["The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD.",
"The ProBook X1 has 16GB DDR5 RAM."]
for mode in ["precision", "recall", "f1"]:
metric = FactualCorrectness(llm=judge, mode=mode)
scores = [metric.score(response=text, reference=reference).value for text in answers]
print(f"{mode:<9} complete={scores[0]:.2f} partial={scores[1]:.2f}")Output
precision complete=1.00 partial=1.00 recall complete=1.00 partial=0.50 f1 complete=1.00 partial=0.67
How the modes treat a missing fact
- precision asks whether what the answer says is in the reference. The partial answer says nothing false, so it keeps a high precision.
- recall asks whether the reference's claims are in the answer. The partial answer leaves out the SSD, so its recall drops.
- f1 balances the two, so the partial answer lands between them.
Factual correctness vs answer correctness
| FactualCorrectness | AnswerCorrectness | |
|---|---|---|
| Reads | response, reference | user_input, response, reference |
| Embeddings | No | Yes, for the semantic part |
| Tunable by | mode, atomicity, coverage | weights |
| Best for | Facts only, with a choice of which error matters | One blended score |
When to use factual correctness
- When a short answer is fine but a wrong one is not: use precision.
- When the answer must list every item, such as all shipping options: use recall.
Watch out. The claim split depends on the judge. A long reference broken into many small claims makes recall harsher than the same reference broken into a few;
atomicity and coverage control how finely the judge splits.Related
- Previous: Answer correctness
- Next: Custom metrics
- Reference: Factual correctness
Try it yourself
- Score
"The ProBook X1 has 16GB DDR5 RAM, a 512GB NVMe SSD and a 4K screen.", which adds a claim, in all three modes. - Set
atomicity="high"andcoverage="high"in recall mode and compare the partial answer's score.
Little by little, you're building something great.