1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →
Collections and legacy metrics
RAGAS 0.4 ships two metrics APIs: the current collections metrics in ragas.metrics.collections, scored with keyword arguments, and the legacy classes in ragas.metrics, scored from a sample and used by evaluate().
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
Two imports in this course look almost the same and behave differently. Knowing which is which saves an afternoon with other people's examples.
Import paths for the two APIs
from ragas.metrics.collections import Faithfulness # current: .score(user_input=..., ...)
from ragas.metrics import Faithfulness # legacy: await .single_turn_ascore(sample), evaluate()The deprecation warning
Importing from the legacy path makes RAGAS say so itself.
Catching the legacy import warning
import warnings
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
from ragas.metrics import Faithfulness as LegacyFaithfulness
print(len(caught), "warning(s)")
print(str(caught[0].message)[:150])Output
2 warning(s) Importing Faithfulness from 'ragas.metrics' is deprecated and will be removed in v1.0. Please use 'ragas.metrics.collections' instead. Example: from r
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
- written in LLM as a judge
View the code here
judge.py
import os
from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
groq = AsyncOpenAI(
api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
Passing a collections metric to evaluate()
evaluate() checks its metrics against the legacy base class, so a collections metric is refused before any call is made.
from judge import judge
from ragas import EvaluationDataset, SingleTurnSample, evaluate
from ragas.metrics.collections import Faithfulness
sample = SingleTurnSample(user_input="q", response="a", retrieved_contexts=["c"])
evaluate(dataset=EvaluationDataset(samples=[sample]), metrics=[Faithfulness(llm=judge)])Output
Traceback (most recent call last):
File "main.py", line 6, in <module>
evaluate(dataset=EvaluationDataset(samples=[sample]), metrics=[Faithfulness(llm=judge)])
TypeError: All metrics must be initialised metric objects, e.g: metrics=[BleuScore(), AspectCritic()]What the two runs show
- The warning names the class, says the legacy path is removed in v1.0, and points to
ragas.metrics.collections. - The TypeError is
evaluate()refusing a metric from the other API. The two do not mix.
Collections vs legacy metrics
| ragas.metrics.collections (current) | ragas.metrics (legacy) | |
|---|---|---|
| How you score | metric.score(user_input=..., ...) or await metric.ascore(...) | await metric.single_turn_ascore(sample) |
| What you pass | Plain keyword arguments | A SingleTurnSample |
| The judge | Given when the metric is built: Faithfulness(llm=judge) | Given to evaluate(llm=judge) or the metric |
Works with evaluate() | No | Yes |
| Future | The API to write new code against | Removed in 1.0 |
Which API to use
- New code: collections metrics, scored directly or inside an experiment.
- Existing code built on
evaluate(): legacy metrics until you move it. - ID-based precision and recall: legacy only in 0.4.3, as the id-based lesson showed.
Watch out. Pin the RAGAS version. The metric APIs changed between 0.1, 0.2 and 0.4, and an example from a year ago may import a name that no longer exists. This course pins
ragas==0.4.3.Related
- Previous: evaluate() and EvaluationDataset
- Next: Evals in CI
- Reference: Faithfulness (both APIs)
Try it yourself
- Import both
Faithfulnessclasses in one file and print their__module__. - Import
LLMContextRecallfromragas.metricsinside the warning block and read its message.
This is what real progress feels like.