Integrations
Integrations are the model classes, framework hooks and the Confident AI platform that DeepEval ships, which replace the pieces this course wrote by hand: the Groq judge class, the hand-made traces and the local result files.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
Everything in the course so far ran on one Groq key, with a judge class you wrote and results saved on your own disk. That is a complete setup, and it is also the smallest one. This lesson maps each piece to what DeepEval offers for it, and runs one of them for real: Gemini as a second judge.
The model classes API
from deepeval.models import GeminiModel, LiteLLMModel, OllamaModel, OpenAIModel
judge = GeminiModel(model="gemini-2.5-flash") # reads GOOGLE_API_KEY, needs google-genai
judge = LiteLLMModel(model="groq/openai/gpt-oss-120b", api_key=...) # needs litellm
metric = AnswerRelevancyMetric(model=judge) # used exactly like GroqJudgeEach class is a DeepEvalBaseLLM, the same base class GroqJudge inherits from Custom judge model, so any metric takes any of them as model=. The built-in ones already handle structured output and retries for their provider.
Adding Gemini as a second judgeOptional
This part needs a free Gemini key from Google AI Studio and one extra package. The free tier allows only a small number of requests per model per day, and one metric makes several, so run the example once rather than in a loop.
pip install "deepeval==4.2.8" "google-genai==2.28.0"export GOOGLE_API_KEY=AIza...- written in Custom judge model
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
Judging one answer with Groq and Gemini
The same answer relevancy metric, on the same padded answer, once with each judge. metric.evaluation_model is the name of the judge that scored it.
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.models import GeminiModel
from deepeval.test_case import LLMTestCase
from judge import judge
gemini = GeminiModel(model="gemini-2.5-flash", temperature=0)
test_case = LLMTestCase(
input="What is the price of the PixelPhone 15?",
actual_output="The PixelPhone 15 costs $899. We also sell the SmartWatch X for $299, with GPS and a 7-day battery.",
)
for model in [judge, gemini]:
metric = AnswerRelevancyMetric(model=model)
metric.measure(test_case)
print(f"{metric.evaluation_model:20} score={metric.score:.2f}")
print(" ", metric.reason)openai/gpt-oss-120b score=0.25
The score is 0.25 because the answer discusses the SmartWatch X’s price, features, and battery life, which are unrelated to the asked PixelPhone 15 price.
gemini-2.5-flash (Gemini) score=0.25
The score is 0.25 because the output contained multiple irrelevant statements discussing the price and features of a different product, the SmartWatch X, instead of focusing on the PixelPhone 15 as requested. This significantly reduced its relevancy to the input query about the PixelPhone 15's price.google-genai may also print a long warning about automatic function calling the first time Gemini is called; it does not affect the score and can be ignored.
What the two judges said
- Both judges scored the answer 0.25. Answer relevancy splits the answer into statements and keeps the share that address the question; only the price sentence does, and the three about the SmartWatch X do not.
- The names differ:
openai/gpt-oss-120bis whatGroqJudge.get_model_namereturns, andgemini-2.5-flash (Gemini)is how DeepEval's own class names itself. - Only the judge changed. The metric, the test case and the loop are the same code, which is what makes a second judge a cheap check on the first.
What you used and what to use in production
| What the course used | A production option | Package or setting |
|---|---|---|
GroqJudge, your own class | A built-in model class: OpenAIModel, GeminiModel, AnthropicModel, LiteLLMModel, OllamaModel for a local model | deepeval.models, plus the provider's SDK or litellm |
model=judge on every metric | A default judge for every metric, set once from the CLI | deepeval set-gemini --model=gemini-2.5-flash, deepeval set-litellm --model=... |
@observe on your own functions | Automatic traces from your framework | deepeval.integrations.langchain.CallbackHandler, and integrations for LlamaIndex, CrewAI, Pydantic AI, Google ADK, OpenAI Agents |
Goldens in goldens.json | Datasets edited by domain experts on Confident AI | dataset.pull(alias=...), after deepeval login |
Results in .deepeval and a results folder | Shared reports, regression comparison and tracing in production | Confident AI, with CONFIDENT_API_KEY |
The word-overlap retrieve | A vector store or search API | Your retriever; DeepEval only reads its output in retrieval_context |
A Groq judge through LiteLLM
If you would rather not maintain judge.py, LiteLLM reaches Groq too. It is a separate package, and the provider goes in front of the model name.
import os
from deepeval.models import LiteLLMModel # pip install litellm
judge = LiteLLMModel(model="groq/openai/gpt-oss-120b", api_key=os.environ["GROQ_API_KEY"], temperature=0)Traces from a LangChain agent
from deepeval.integrations.langchain import CallbackHandler # needs langchain installed
agent.invoke({"messages": [...]}, config={"callbacks": [CallbackHandler()]}) # spans recorded for youWhen to switch from the hand-built pieces
- Switch the judge class when your provider has a built-in one, so retries and structured output are maintained for you.
- Switch to framework hooks when the app is built on LangChain, LlamaIndex or another supported framework, so new steps are traced without new decorators.
- Add Confident AI when more than one person reads the results, or when you want each run compared with the last one without writing that comparison yourself.
Related
- Previous: Evals in CI
- Next: Project: TechNest support bot evaluation
- Reference: Gemini integration
- Change
gemini-2.5-flashtogemini-2.5-flash-liteand compare the reason it writes. - Remove the SmartWatch sentence from
actual_outputand run the example again: both judges should move to 1.00.
You understood something today that you didn't yesterday.