Integrations
RAGAS integrations are the providers and backends each part of an evaluation can run on: the judge through llm_factory, the embeddings through an embeddings class, and datasets through a storage backend.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
This course ran on free Groq and Gemini keys, an in-memory retriever and CSV files. Each of those has a production counterpart, and swapping one is a single line, because every metric takes the judge and the embeddings as arguments.
What the course used and what replaces it
| Piece | In this course | In production | Package |
|---|---|---|---|
| Judge | llm_factory on Groq, openai/gpt-oss-20b | OpenAI, Anthropic, Gemini, Azure, or any model through LiteLLM | openai, anthropic, litellm |
| Embeddings | GoogleEmbeddings, gemini-embedding-2 | OpenAIEmbeddings, or HuggingFaceEmbeddings run locally | openai, sentence-transformers |
| The app's search | A numpy array in memory | A vector store such as Chroma, Qdrant or pgvector | The store's client |
| Datasets and experiments | local/csv | local/jsonl, Google Drive, or inmemory | ragas[gdrive] for Drive |
| Tracing | None | Langfuse or MLflow | ragas[tracing] |
The llm_factory swap
judge = llm_factory("model-name", provider="openai", client=client) # the only line that changes
Faithfulness(llm=judge) # every metric stays as it isTwo judges on the same metric
The proof that a swap is one line: build a second judge on another model and hand it to the same metric. judge.py already exports the Groq client as groq.
for model in ["openai/gpt-oss-20b", "qwen/qwen3.8-27b"]:
judge = llm_factory(model, provider="openai", client=groq)
result = Faithfulness(llm=judge).score(user_input=question, response=response, retrieved_contexts=chunks)- written in LLM as a judge
View the code here
import os
from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
groq = AsyncOpenAI(
api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
Scoring the bank example with two judges
from judge import groq
from ragas.llms import llm_factory
from ragas.metrics.collections import Faithfulness
question = "What is the minimum balance for my savings account?"
chunks = [
"Minimum balance is ₹10,000 for urban branches. Non-maintenance fee is ₹350 + GST if balance falls below.",
"Semi-urban branch minimum balance is ₹5,000. Rural branch minimum balance is ₹2,500.",
]
response = ("The minimum balance for urban branches is ₹10,000. The non-maintenance fee charged is ₹350 + GST. "
"You can also transfer funds online at no extra charge. For rural branches the minimum is ₹2,500.")
for model in ["openai/gpt-oss-20b", "qwen/qwen3.8-27b"]:
judge = llm_factory(model, provider="openai", client=groq)
result = Faithfulness(llm=judge).score(user_input=question, response=response, retrieved_contexts=chunks)
print(f"{model:<22} faithfulness={result.value}")openai/gpt-oss-20b faithfulness=0.75 qwen/qwen3.8-27b faithfulness=0.75
What changed between the judges
- Only the model string changed; the metric, the inputs and the call are identical.
- Both judges scored 0.75. Each found the one unsupported claim among four. On a harder case two judges can disagree, and then the difference is the judge, not the app: one more reason to fix the judge before comparing runs.
Judges on other providers
Each block is the one line that changes, plus its client. They need that provider's key, and paid providers bill per call.
from openai import AsyncOpenAI
judge = llm_factory("gpt-4o-mini", client=AsyncOpenAI()) # OpenAI, reads OPENAI_API_KEYgemini = AsyncOpenAI(api_key=os.environ["GOOGLE_API_KEY"],
base_url="https://generativelanguage.googleapis.com/v1beta/openai/")
judge = llm_factory("gemini-2.5-flash", provider="openai", client=gemini) # Gemini through its OpenAI-compatible URLThe Gemini block uses Gemini's OpenAI-compatible URL, which the RAGAS source recommends over the native client for judges because of a known issue with the native route's safety settings. A judge from a different provider than the app's generator is the setup the video recommends.
Other evaluation tools the video names
| Tool | What it adds |
|---|---|
| DeepEval | Metrics for tool use and multi-agent apps, and a golden synthesizer |
| Arize Phoenix | Evaluation and tracing in one tool, self-hostable |
| LangSmith, Langfuse | Tracing and evaluation dashboards for a deployed app |
When to swap a piece
- Swap the judge when free-tier limits slow your runs, or to judge with a provider different from your generator.
- Swap the search for a vector store when the knowledge base no longer fits in memory; the evaluation code does not change.
Related
- Previous: Evals in CI
- Next: Project: TechNest support bot evaluation
- Reference: LLM adapters
- Score the TechNest faithfulness example from the faithfulness lesson with the
openai/gpt-oss-120bjudge. - Save the goldens dataset with
backend="local/jsonl"and run the experiment lesson against it.
Little by little, you're building something great.