RAGASragas 0.4.3 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

Integrations

RAGAS integrations are the providers and backends each part of an evaluation can run on: the judge through llm_factory, the embeddings through an embeddings class, and datasets through a storage backend.

Last updated: 29 Sep, 2026 · RAGAS 0.4.3

This course ran on free Groq and Gemini keys, an in-memory retriever and CSV files. Each of those has a production counterpart, and swapping one is a single line, because every metric takes the judge and the embeddings as arguments.

What the course used and what replaces it

PieceIn this courseIn productionPackage
Judgellm_factory on Groq, openai/gpt-oss-20bOpenAI, Anthropic, Gemini, Azure, or any model through LiteLLMopenai, anthropic, litellm
EmbeddingsGoogleEmbeddings, gemini-embedding-2OpenAIEmbeddings, or HuggingFaceEmbeddings run locallyopenai, sentence-transformers
The app's searchA numpy array in memoryA vector store such as Chroma, Qdrant or pgvectorThe store's client
Datasets and experimentslocal/csvlocal/jsonl, Google Drive, or inmemoryragas[gdrive] for Drive
TracingNoneLangfuse or MLflowragas[tracing]

The llm_factory swap

python
judge = llm_factory("model-name", provider="openai", client=client)  # the only line that changes
Faithfulness(llm=judge)                                               # every metric stays as it is

Two judges on the same metric

The proof that a swap is one line: build a second judge on another model and hand it to the same metric. judge.py already exports the Groq client as groq.

python
for model in ["openai/gpt-oss-20b", "qwen/qwen3.8-27b"]:
    judge = llm_factory(model, provider="openai", client=groq)
    result = Faithfulness(llm=judge).score(user_input=question, response=response, retrieved_contexts=chunks)
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
judge.py
import os

from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory

groq = AsyncOpenAI(
    api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
    base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)


class OneTextPerCall(GoogleEmbeddings):
    """gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""

    def embed_texts(self, texts, **kwargs):
        return [self.embed_text(text) for text in texts]

    async def aembed_texts(self, texts, **kwargs):
        return [await self.aembed_text(text) for text in texts]


embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")

Scoring the bank example with two judges

ExampleAPI key
from judge import groq
from ragas.llms import llm_factory
from ragas.metrics.collections import Faithfulness

question = "What is the minimum balance for my savings account?"
chunks = [
    "Minimum balance is ₹10,000 for urban branches. Non-maintenance fee is ₹350 + GST if balance falls below.",
    "Semi-urban branch minimum balance is ₹5,000. Rural branch minimum balance is ₹2,500.",
]
response = ("The minimum balance for urban branches is ₹10,000. The non-maintenance fee charged is ₹350 + GST. "
            "You can also transfer funds online at no extra charge. For rural branches the minimum is ₹2,500.")

for model in ["openai/gpt-oss-20b", "qwen/qwen3.8-27b"]:
    judge = llm_factory(model, provider="openai", client=groq)
    result = Faithfulness(llm=judge).score(user_input=question, response=response, retrieved_contexts=chunks)
    print(f"{model:<22} faithfulness={result.value}")

What changed between the judges

  • Only the model string changed; the metric, the inputs and the call are identical.
  • Both judges scored 0.75. Each found the one unsupported claim among four. On a harder case two judges can disagree, and then the difference is the judge, not the app: one more reason to fix the judge before comparing runs.

Judges on other providers

Each block is the one line that changes, plus its client. They need that provider's key, and paid providers bill per call.

python
from openai import AsyncOpenAI
judge = llm_factory("gpt-4o-mini", client=AsyncOpenAI())  # OpenAI, reads OPENAI_API_KEY
python
gemini = AsyncOpenAI(api_key=os.environ["GOOGLE_API_KEY"],
                     base_url="https://generativelanguage.googleapis.com/v1beta/openai/")
judge = llm_factory("gemini-2.5-flash", provider="openai", client=gemini)  # Gemini through its OpenAI-compatible URL

The Gemini block uses Gemini's OpenAI-compatible URL, which the RAGAS source recommends over the native client for judges because of a known issue with the native route's safety settings. A judge from a different provider than the app's generator is the setup the video recommends.

Other evaluation tools the video names

ToolWhat it adds
DeepEvalMetrics for tool use and multi-agent apps, and a golden synthesizer
Arize PhoenixEvaluation and tracing in one tool, self-hostable
LangSmith, LangfuseTracing and evaluation dashboards for a deployed app

When to swap a piece

  • Swap the judge when free-tier limits slow your runs, or to judge with a provider different from your generator.
  • Swap the search for a vector store when the knowledge base no longer fits in memory; the evaluation code does not change.
Watch out. Change one piece at a time and re-run the baseline. A new judge changes every score, so a run with a new judge and a new retriever cannot tell you which change moved the numbers.
Try it yourself
  • Score the TechNest faithfulness example from the faithfulness lesson with the openai/gpt-oss-120b judge.
  • Save the goldens dataset with backend="local/jsonl" and run the experiment lesson against it.
PreviousEvals in CI

Little by little, you're building something great.