1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →
evaluate() and EvaluationDataset
evaluate() is the older RAGAS function that scores every sample of an EvaluationDataset on a list of metrics in one call and returns the averages, with a per-sample table.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
Most RAGAS examples written before version 0.4 use this one call. It still works in 0.4.3, with the older metric classes from ragas.metrics, so you will meet it in other people's code.
The evaluate API
from ragas import EvaluationDataset, evaluate
from ragas.metrics import Faithfulness, LLMContextRecall # the older metric classes
result = evaluate(dataset=EvaluationDataset(samples=[...]), metrics=[Faithfulness(), LLMContextRecall()], llm=judge)
result.to_pandas() # one row per sample, one column per metricTwo return-shipping samples
The video's notebook tests faithfulness with a question about return shipping. The first answer below is correct; the second is the notebook's wrong one, over the TechNest policy.
policy = ["Customers are responsible for return shipping costs unless the item arrives defective or damaged."]
right = "No. You pay return shipping unless the item arrives defective or damaged."
wrong = "Yes, we provide free prepaid return shipping labels for all returns."Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
- written in LLM as a judge
View the code here
judge.py
import os
from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
groq = AsyncOpenAI(
api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
Scoring two samples with evaluate()
from judge import judge
from ragas import EvaluationDataset, SingleTurnSample, evaluate
from ragas.metrics import Faithfulness, LLMContextRecall
question = "Do you offer free return shipping?"
policy = ["Customers are responsible for return shipping costs unless the item arrives defective or damaged."]
reference = "No. Customers pay return shipping unless the item is defective."
samples = [
SingleTurnSample(user_input=question, retrieved_contexts=policy, reference=reference,
response="No. You pay return shipping unless the item arrives defective or damaged."),
SingleTurnSample(user_input=question, retrieved_contexts=policy, reference=reference,
response="Yes, we provide free prepaid return shipping labels for all returns."),
]
result = evaluate(dataset=EvaluationDataset(samples=samples), metrics=[Faithfulness(), LLMContextRecall()],
llm=judge, show_progress=False)
print(result)
print(result.to_pandas()[["response", "faithfulness", "context_recall"]].to_string())Output
{'faithfulness': 0.5000, 'context_recall': 0.7500}
response faithfulness context_recall
0 No. You pay return shipping unless the item arrives defective or damaged. 1.0 1.0
1 Yes, we provide free prepaid return shipping labels for all returns. 0.0 0.5What evaluate() returned
- The first line is the average of each metric over both samples: faithfulness 0.5, context recall 0.75.
- Faithfulness is 1.0 for the right answer and 0.0 for the wrong one, whose only claim contradicts the policy.
- Context recall came out 1.0 and 0.5, though both rows share the same reference and chunk, so the answers should not matter to it. That is the judge answering the same question differently on two calls: the wobble the results lesson warns about, visible in one table.
evaluate() vs an experiment
| evaluate() | @experiment | |
|---|---|---|
| Metrics | Older classes from ragas.metrics | Collections metrics |
| Input | A list of finished samples | A dataset; the app runs inside |
| Output | Averages and to_pandas() | A saved CSV per run |
| Future | Metrics removed in RAGAS 1.0 | The current API |
When evaluate() is the right call
- Reading or updating code written before RAGAS 0.4.
- Scoring a list of finished samples with a single call, with a pandas table at the end.
Watch out. Importing metrics from
ragas.metrics prints a deprecation warning: they are removed in version 1.0. evaluate() only accepts these older classes; Collections and legacy metrics shows the error a collections metric gives.Related
- Previous: Evaluation results
- Next: Collections and legacy metrics
- Reference: evaluate() reference
Try it yourself
- Add a third sample whose answer is half right and read the new averages.
- Sort the pandas table by
faithfulnessand print the lowest row.
Slow is fine. Stopping is the only problem.