Comparing models
Comparing models is an evaluation that runs the same goldens and the same metrics against two versions of the app, with only the model changed, so the score difference belongs to the model.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
EvaluationResult turned one run into rows of scores. Two runs, one per model, give two sets of rows to put side by side. The goldens and metrics stay fixed; only the line that names the bot's model changes.
Two experiments on one dataset
The video runs its evaluation with three things: the function that calls the app for every input, the dataset of questions, and two evaluators, correctness and concision. Each run gets an experiment name, here one for gpt-4o-mini. The experiment page lists the scores per evaluator and, for each example, the input, the reference output and the app's output side by side.
Then the video changes one thing, the model the app calls, to gpt-4-turbo, and runs the same experiment under a new name. With both experiments on the same dataset, the scores answer the question of which model to keep, and the video picks gpt-4o-mini. The video runs this in LangSmith; this page does it with DeepEval, with the TechNest bot and two Groq models.
The hyperparameters API
evaluate(
test_cases=test_cases,
metrics=metrics,
hyperparameters={"model": "qwen/qwen3.8-27b", "temperature": 0}, # what this run used
)hyperparameters is a dictionary of the settings a run used, with string, number or prompt values. DeepEval stores it with the test run, so saved runs and Confident AI can tell them apart. Without it, evaluate() prints a warning that no hyperparameters were logged.
- written in Custom judge model
- written in TechNest RAG app
- written in TechNest RAG app
- written in Datasets and goldens
- written in evaluate()
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
[
{
"name": "g001",
"input": "What is TechNest's return policy?",
"expected_output": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
},
{
"name": "g002",
"input": "What are the RAM and storage specs of the ProBook X1?",
"expected_output": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
},
{
"name": "g003",
"input": "How long is the battery life on the SoundPods Pro?",
"expected_output": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
},
{
"name": "g004",
"input": "What are TechNest's shipping options and how long do returns take to process?",
"expected_output": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
},
{
"name": "g005",
"input": "What is the price of the PixelPhone 15?",
"expected_output": "The TechNest PixelPhone 15 is priced at $899."
}
]
from deepeval.dataset import EvaluationDataset
from deepeval.metrics import AnswerRelevancyMetric, GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams
from judge import judge
from technest import answer
dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="goldens.json")
def build_test_cases(goldens):
"""Ask the bot every golden's question and wrap each reply in a test case."""
test_cases = []
for golden in goldens:
response, contexts = answer(golden.input)
test_cases.append(LLMTestCase(
name=golden.name,
input=golden.input,
actual_output=response,
expected_output=golden.expected_output,
retrieval_context=contexts,
))
return test_cases
relevancy = AnswerRelevancyMetric(model=judge)
correctness = GEval(
name="Correctness",
evaluation_steps=[
"Check whether the facts in 'actual output' contradict any facts in 'expected output'.",
"Penalize facts from 'expected output' that 'actual output' leaves out.",
"Do not penalize extra details or different wording.",
],
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
model=judge,
)
Switching the bot's model
technest.py from TechNest RAG app reads the module variable CHAT_MODEL each time generate runs, so setting it before the loop changes the model for every answer that follows.
import technest
technest.CHAT_MODEL = "openai/gpt-oss-20b" # the next answer() call uses this modelOne run per model
for model in ["qwen/qwen3.8-27b", "openai/gpt-oss-20b"]:
technest.CHAT_MODEL = model
test_cases = build_test_cases(goldens)
result = evaluate(test_cases=test_cases, metrics=[relevancy, correctness],
hyperparameters={"model": model, "temperature": 0}, ...)The bot is asked again for each model, so each run scores that model's own answers. The scores go into a dictionary keyed by golden, metric and model, ready to print side by side.
Scoring two Groq models on the same goldens
The SoundPods Pro battery (g003) and the PixelPhone 15 price (g005), answered by qwen/qwen3.8-27b, the bot's model, and by openai/gpt-oss-20b. The judge stays openai/gpt-oss-120b for both runs.
from deepeval import evaluate
from deepeval.evaluate import AsyncConfig, DisplayConfig
import technest
from technest_eval import build_test_cases, correctness, dataset, relevancy
goldens = [dataset.goldens[2], dataset.goldens[4]] # g003, g005
scores = {}
for model in ["qwen/qwen3.8-27b", "openai/gpt-oss-20b"]:
technest.CHAT_MODEL = model # the bot now answers with this model
test_cases = build_test_cases(goldens)
result = evaluate(
test_cases=test_cases,
metrics=[relevancy, correctness],
hyperparameters={"model": model, "temperature": 0},
async_config=AsyncConfig(max_concurrent=1),
display_config=DisplayConfig(print_results=False, show_indicator=False),
)
for test in result.test_results:
for metric in test.metrics_data:
scores[(test.name, metric.name, model)] = metric.score
print(model, "|", test.name, "|", test.actual_output)
print()
print(f"{'golden':7}{'metric':22}{'qwen3.8-27b':>12}{'gpt-oss-20b':>12}")
for golden in goldens:
for name in ["Answer Relevancy", "Correctness [GEval]"]:
a = scores[(golden.name, name, "qwen/qwen3.8-27b")]
b = scores[(golden.name, name, "openai/gpt-oss-20b")]
print(f"{golden.name:7}{name:22}{a:12.2f}{b:12.2f}")⚠ WARNING: No prompts logged. » Log prompts to evaluate and optimize your prompt templates and models. ================================================================================ ✓ Evaluation completed 🎉! (time taken: 37.52s | token cost: None) » Test Results (2 total tests): » Pass Rate: 0.0% | Passed: 0 | Failed: 2 =============================================================================== = » Want to share evals with your team, or a place for your test cases to live? ❤️ 🏡 » Run 'deepeval view' to analyze and save testing results on Confident AI. qwen/qwen3.8-27b | g003 | The TechNest SoundPods Pro offer 8 hours of playback on a single charge. When used with the charging case, the total battery life extends to 24 hours. qwen/qwen3.8-27b | g005 | The TechNest PixelPhone 15 is priced at $899. It comes with a 1-year warranty and is available in Midnight Black and Arctic White. ⚠ WARNING: No prompts logged. » Log prompts to evaluate and optimize your prompt templates and models. ================================================================================ ✓ Evaluation completed 🎉! (time taken: 34.51s | token cost: None) » Test Results (2 total tests): » Pass Rate: 50.0% | Passed: 1 | Failed: 1 =============================================================================== = » Want to share evals with your team, or a place for your test cases to live? ❤️ 🏡 » Run 'deepeval view' to analyze and save testing results on Confident AI. openai/gpt-oss-20b | g003 | The SoundPods Pro offer 8 hours of playback on a single charge, and up to 24 hours total when you use the charging case. openai/gpt-oss-20b | g005 | The TechNest PixelPhone 15 is priced at $899. golden metric qwen3.8-27b gpt-oss-20b g003 Answer Relevancy 1.00 1.00 g003 Correctness [GEval] 0.00 0.10 g005 Answer Relevancy 0.33 1.00 g005 Correctness [GEval] 1.00 1.00
What the comparison shows
- The warning changed. With
hyperparameterslogged, each run warns No prompts logged instead; DeepEval tracks prompts as their ownPromptobjects, which the prompts docs page covers. - g003 fails correctness for both models. qwen3.8-27b says the total with the case extends to 24 hours, gpt-oss-20b says up to 24 hours total. Both read 8 hours per charge plus 24 hours with the case as 24 in all, so the model is not the cause; the prompt or the catalog wording is a better place to look.
- g005 splits the models on relevancy. qwen3.8-27b gives the price and adds the warranty and the two colours, and relevancy scores it 0.33. gpt-oss-20b gives only the price and scores 1.00. Correctness gives both 1.00, because both have the right price.
- The pass rates follow: 0% of test cases passed for qwen3.8-27b and 50% for gpt-oss-20b. On these two goldens the smaller gpt-oss model answers more to the point, which is a reason to test it on the whole dataset, not yet a reason to switch.
LangSmith experiments vs DeepEval runs
| LangSmith (the video) | DeepEval (this page) | |
|---|---|---|
| Runs the app | client.evaluate calls the target function | Your loop calls answer, then evaluate() scores |
| Names the run | experiment_prefix | hyperparameters, and optionally identifier |
| Scores with | Evaluator functions | DeepEval metrics on a judge model |
| Side by side view | The experiments page | Your own table, or Confident AI |
When to compare models
- Before switching the bot to a cheaper or faster model: the scores show what you give up, if anything.
- When a provider retires a model, as Groq has done with several, and you need a replacement that answers your goldens as well.
- When you change the prompt instead of the model. Keep the model fixed, change
SYSTEM_PROMPT, and run the same loop once per prompt.
Related
- Previous: EvaluationResult
- Next: LLM tracing
- Reference: Flags and configs: hyperparameters
- Before the loop, add
technest.SYSTEM_PROMPT += "\nWhen the context gives a number in parts, add the parts up and state the total.": qwen3.8-27b now answers g003 with 32 hours. Read its correctness score. - Replace every
openai/gpt-oss-20bin the example (the model list and theb =lookup) withopenai/gpt-oss-120b, the judge's own model, and compare. The judge then grades answers from its own model, which is worth knowing when you read the scores. - Add
dataset.goldens[0]togoldens: the table grows by two rows for g001.
This is what real progress feels like.