DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

evaluate()

evaluate() is DeepEval's function that runs a list of metrics on a list of test cases in one call and prints one report for the whole run.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

Datasets and goldens turned goldens into test cases by asking the bot each question. The next step is to score all of them at once, with the same metrics, and read one report instead of one score at a time.

Running the app on every golden · from the Production RAG Live Marathon · 3:09:08 to 3:13:33

Running the app on every golden

The video's demo app has five types of question; a real app has thousands. For each question, the RAG app runs for real and returns the actual output and the retrieved context. The expected output is already in the golden, and the input is the question itself. So the loop runs once per golden: five goldens, five runs of the whole pipeline.

Five questions are not enough to judge even the demo app, let alone a production app used by thousands of people at every level of knowledge. Everything the loop collects goes into a dataset that you can picture as a table: for What is TechNest's return policy?, the retrieved context, the RAG response and the reference answer, the correct one. At that point the data is ready for scoring. The video scores that table with RAGAS; on this page the same loop builds DeepEval test cases, and evaluate() scores them.

The evaluate() API

python
from deepeval import evaluate

result = evaluate(
    test_cases=[...],        # LLMTestCase objects, one per golden
    metrics=[...],           # every metric runs on every test case
    async_config=...,        # optional: how many test cases run at the same time
)

A test case per golden

The loop from the datasets lesson, as a function. It takes a list of goldens, so a run can pick which ones to score.

python
def build_test_cases(goldens):
    """Ask the bot every golden's question and wrap each reply in a test case."""
    test_cases = []
    for golden in goldens:
        response, contexts = answer(golden.input)
        test_cases.append(LLMTestCase(
            name=golden.name,
            input=golden.input,
            actual_output=response,
            expected_output=golden.expected_output,
            retrieval_context=contexts,
        ))
    return test_cases

Two metrics: relevancy and correctness

Answer relevancy asks whether the reply addresses the question; it reads only the input and the actual output. Correctness is a G-Eval metric, built as in G-Eval, that compares the reply with the expected answer. Its evaluation_steps are written out so the judge does not punish a correct answer for different wording. The last step also allows extra details, because the bot often adds context the expected answer leaves out.

python
relevancy = AnswerRelevancyMetric(model=judge)
correctness = GEval(
    name="Correctness",
    evaluation_steps=[
        "Check whether the facts in 'actual output' contradict any facts in 'expected output'.",
        "Penalize facts from 'expected output' that 'actual output' leaves out.",
        "Do not penalize extra details or different wording.",
    ],
    evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
    model=judge,
)
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
judge.py
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
technest.py
import json
import os
import re

from openai import OpenAI

groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"

SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""

with open("catalog.json", encoding="utf-8") as f:
    CATALOG = json.load(f)

SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
        "long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}


def words(text):
    """The words in a text that carry meaning, in lower case."""
    return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}


def retrieve(question, top_k=3):
    asked = words(question)
    ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
    return [item["content"] for item in ranked[:top_k]]


def generate(question, contexts):
    context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
    ]
    response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
    return response.choices[0].message.content.strip()


def answer(question, top_k=3):
    contexts = retrieve(question, top_k)
    return generate(question, contexts), contexts
catalog.json
[
  {
    "id": "prod_001",
    "category": "product",
    "title": "ProBook X1 Laptop",
    "content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
  },
  {
    "id": "prod_002",
    "category": "product",
    "title": "PixelPhone 15",
    "content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
  },
  {
    "id": "prod_003",
    "category": "product",
    "title": "SoundPods Pro",
    "content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
  },
  {
    "id": "prod_004",
    "category": "product",
    "title": "UltraTab S2 Tablet",
    "content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
  },
  {
    "id": "prod_005",
    "category": "product",
    "title": "SmartWatch X",
    "content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
  },
  {
    "id": "prod_006",
    "category": "product",
    "title": "ProCam 4K Action Camera",
    "content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
  },
  {
    "id": "prod_007",
    "category": "product",
    "title": "BassBuds Max Headphones",
    "content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
  },
  {
    "id": "prod_008",
    "category": "product",
    "title": "SoundBar 360",
    "content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
  },
  {
    "id": "policy_001",
    "category": "policy",
    "title": "Return Policy",
    "content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
  },
  {
    "id": "policy_002",
    "category": "policy",
    "title": "Shipping Policy",
    "content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
  },
  {
    "id": "policy_003",
    "category": "policy",
    "title": "Warranty Policy",
    "content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
  },
  {
    "id": "policy_004",
    "category": "policy",
    "title": "Payment Policy",
    "content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
  },
  {
    "id": "faq_001",
    "category": "faq",
    "title": "Order Tracking",
    "content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
  },
  {
    "id": "faq_002",
    "category": "faq",
    "title": "International Shipping",
    "content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
  },
  {
    "id": "faq_003",
    "category": "faq",
    "title": "Bulk and Business Orders",
    "content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
  }
]
goldens.json
[
  {
    "name": "g001",
    "input": "What is TechNest's return policy?",
    "expected_output": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
  },
  {
    "name": "g002",
    "input": "What are the RAM and storage specs of the ProBook X1?",
    "expected_output": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
  },
  {
    "name": "g003",
    "input": "How long is the battery life on the SoundPods Pro?",
    "expected_output": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
  },
  {
    "name": "g004",
    "input": "What are TechNest's shipping options and how long do returns take to process?",
    "expected_output": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
  },
  {
    "name": "g005",
    "input": "What is the price of the PixelPhone 15?",
    "expected_output": "The TechNest PixelPhone 15 is priced at $899."
  }
]

The technest_eval.py file

Save the dataset, the function and the two metrics as technest_eval.py, next to goldens.json, technest.py and judge.py. The next lessons import from it.

python
from deepeval.dataset import EvaluationDataset
from deepeval.metrics import AnswerRelevancyMetric, GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams

from judge import judge
from technest import answer

dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="goldens.json")


def build_test_cases(goldens):
    """Ask the bot every golden's question and wrap each reply in a test case."""
    test_cases = []
    for golden in goldens:
        response, contexts = answer(golden.input)
        test_cases.append(LLMTestCase(
            name=golden.name,
            input=golden.input,
            actual_output=response,
            expected_output=golden.expected_output,
            retrieval_context=contexts,
        ))
    return test_cases


relevancy = AnswerRelevancyMetric(model=judge)
correctness = GEval(
    name="Correctness",
    evaluation_steps=[
        "Check whether the facts in 'actual output' contradict any facts in 'expected output'.",
        "Penalize facts from 'expected output' that 'actual output' leaves out.",
        "Do not penalize extra details or different wording.",
    ],
    evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
    model=judge,
)

Calling evaluate

python
evaluate(
    test_cases=test_cases,
    metrics=[relevancy, correctness],
    async_config=AsyncConfig(max_concurrent=1),  # one test case at a time
)

By default evaluate() scores many test cases at the same time. AsyncConfig(max_concurrent=1) scores one test case at a time, which keeps the run inside the token limit of Groq's free plan; Rate limits and retries shows what happens without it.

Evaluating three TechNest goldens

Three goldens: the return policy (g001), the SoundPods Pro battery (g003) and the two-part question about shipping and return processing (g004).

ExampleAPI key
from deepeval import evaluate
from deepeval.evaluate import AsyncConfig

from technest_eval import build_test_cases, correctness, dataset, relevancy

goldens = [dataset.goldens[0], dataset.goldens[2], dataset.goldens[3]]  # g001, g003, g004
test_cases = build_test_cases(goldens)

evaluate(
    test_cases=test_cases,
    metrics=[relevancy, correctness],
    async_config=AsyncConfig(max_concurrent=1),
)

Reading the evaluate() report

  • The first two lines name each metric and the judge that runs it, openai/gpt-oss-120b.
  • g001 and g004 passed both metrics, and the report gives each a single line. Passing test cases are cut short by default, so their scores and reasons are not printed.
  • g003 failed, and its panel shows why. The bot says that with the case the total battery life extends to 24 hours. The catalog says 8 hours per charge plus 24 hours with the case, and the expected output adds them up to 32. Answer relevancy still gives 1.00, because the answer is about the battery life. Correctness gives 0.00, and its reason names the 24 against the 32.
  • The aggregate table averages each metric over the three test cases: relevancy 1.00 with every case passing, correctness 0.57 with two of three passing.
  • The summary gives the time taken and the pass rate, 66.67%: a test case passes only when all its metrics pass. The warning about hyperparameters comes back in Comparing models; the lines about deepeval view and Confident AI point to DeepEval's hosted platform, which this run does not use.

evaluate() vs measure()

metric.measure(test_case)evaluate(test_cases, metrics)
ScoresOne metric on one test caseEvery metric on every test case
PrintsNothing; read metric.scoreA panel per test case, an aggregate table and a summary
Runs test cases at the same timeNoYes, up to max_concurrent
ReturnsThe scoreAn EvaluationResult with every score and reason
Good forTrying one metricA whole dataset, after every change to the bot

When to use evaluate()

  • When you want the scores for a whole dataset from a script or a notebook, without writing pytest tests.
  • When you can run the app yourself on each golden, as here. For a run that builds test cases from traces of the app instead, see LLM tracing.
  • When you want the results as data: the return value, covered in EvaluationResult, holds every score and reason.
Watch out. evaluate() only sees the test cases you pass. If the loop that builds them fails halfway, for example because the bot's model returned an error, nothing is scored. Build the test cases first, check how many you have, then call evaluate().
Try it yourself
  • Score only g005 with goldens = [dataset.goldens[4]] and read the relevancy reason for the PixelPhone 15 price answer.
  • Remove relevancy from metrics and check that the aggregate table has one row.
  • Add dataset.goldens[1] (g002) to goldens and check that the summary counts four tests.

Little by little, you're building something great.