DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

Integrations

Integrations are the model classes, framework hooks and the Confident AI platform that DeepEval ships, which replace the pieces this course wrote by hand: the Groq judge class, the hand-made traces and the local result files.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

Everything in the course so far ran on one Groq key, with a judge class you wrote and results saved on your own disk. That is a complete setup, and it is also the smallest one. This lesson maps each piece to what DeepEval offers for it, and runs one of them for real: Gemini as a second judge.

The model classes API

python
from deepeval.models import GeminiModel, LiteLLMModel, OllamaModel, OpenAIModel

judge = GeminiModel(model="gemini-2.5-flash")       # reads GOOGLE_API_KEY, needs google-genai
judge = LiteLLMModel(model="groq/openai/gpt-oss-120b", api_key=...)  # needs litellm
metric = AnswerRelevancyMetric(model=judge)          # used exactly like GroqJudge

Each class is a DeepEvalBaseLLM, the same base class GroqJudge inherits from Custom judge model, so any metric takes any of them as model=. The built-in ones already handle structured output and retries for their provider.

Adding Gemini as a second judgeOptional

This part needs a free Gemini key from Google AI Studio and one extra package. The free tier allows only a small number of requests per model per day, and one metric makes several, so run the example once rather than in a loop.

pip install "deepeval==4.2.8" "google-genai==2.28.0"
export GOOGLE_API_KEY=AIza...
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
judge.py
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))

Judging one answer with Groq and Gemini

The same answer relevancy metric, on the same padded answer, once with each judge. metric.evaluation_model is the name of the judge that scored it.

ExampleAPI key
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.models import GeminiModel
from deepeval.test_case import LLMTestCase

from judge import judge

gemini = GeminiModel(model="gemini-2.5-flash", temperature=0)
test_case = LLMTestCase(
    input="What is the price of the PixelPhone 15?",
    actual_output="The PixelPhone 15 costs $899. We also sell the SmartWatch X for $299, with GPS and a 7-day battery.",
)
for model in [judge, gemini]:
    metric = AnswerRelevancyMetric(model=model)
    metric.measure(test_case)
    print(f"{metric.evaluation_model:20} score={metric.score:.2f}")
    print("   ", metric.reason)

google-genai may also print a long warning about automatic function calling the first time Gemini is called; it does not affect the score and can be ignored.

What the two judges said

  • Both judges scored the answer 0.25. Answer relevancy splits the answer into statements and keeps the share that address the question; only the price sentence does, and the three about the SmartWatch X do not.
  • The names differ: openai/gpt-oss-120b is what GroqJudge.get_model_name returns, and gemini-2.5-flash (Gemini) is how DeepEval's own class names itself.
  • Only the judge changed. The metric, the test case and the loop are the same code, which is what makes a second judge a cheap check on the first.

What you used and what to use in production

What the course usedA production optionPackage or setting
GroqJudge, your own classA built-in model class: OpenAIModel, GeminiModel, AnthropicModel, LiteLLMModel, OllamaModel for a local modeldeepeval.models, plus the provider's SDK or litellm
model=judge on every metricA default judge for every metric, set once from the CLIdeepeval set-gemini --model=gemini-2.5-flash, deepeval set-litellm --model=...
@observe on your own functionsAutomatic traces from your frameworkdeepeval.integrations.langchain.CallbackHandler, and integrations for LlamaIndex, CrewAI, Pydantic AI, Google ADK, OpenAI Agents
Goldens in goldens.jsonDatasets edited by domain experts on Confident AIdataset.pull(alias=...), after deepeval login
Results in .deepeval and a results folderShared reports, regression comparison and tracing in productionConfident AI, with CONFIDENT_API_KEY
The word-overlap retrieveA vector store or search APIYour retriever; DeepEval only reads its output in retrieval_context

A Groq judge through LiteLLM

If you would rather not maintain judge.py, LiteLLM reaches Groq too. It is a separate package, and the provider goes in front of the model name.

python
import os

from deepeval.models import LiteLLMModel  # pip install litellm

judge = LiteLLMModel(model="groq/openai/gpt-oss-120b", api_key=os.environ["GROQ_API_KEY"], temperature=0)

Traces from a LangChain agent

python
from deepeval.integrations.langchain import CallbackHandler  # needs langchain installed

agent.invoke({"messages": [...]}, config={"callbacks": [CallbackHandler()]})  # spans recorded for you

When to switch from the hand-built pieces

  • Switch the judge class when your provider has a built-in one, so retries and structured output are maintained for you.
  • Switch to framework hooks when the app is built on LangChain, LlamaIndex or another supported framework, so new steps are traced without new decorators.
  • Add Confident AI when more than one person reads the results, or when you want each run compared with the last one without writing that comparison yourself.
Watch out. Two judges agreed here, but on borderline answers they can score differently. A score from one judge model is not comparable with a score from another, so keep the judge fixed while you compare prompts or app models, and change it only between baselines.
Try it yourself
  • Change gemini-2.5-flash to gemini-2.5-flash-lite and compare the reason it writes.
  • Remove the SmartWatch sentence from actual_output and run the example again: both judges should move to 1.00.
PreviousEvals in CI

You understood something today that you didn't yesterday.