Project: TechNest support bot evaluation
The TechNest support bot evaluation is a complete DeepEval suite for a RAG bot: every golden is asked, each answer is scored for correctness, faithfulness and contextual recall by the Groq judge, and a change to the bot is compared with the last good run to catch a regression.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
The overview promised a score per question and per metric, and a check that catches a change for the worse. This project builds it from pieces the course already made, then changes one setting that looks harmless and lets the scores say what it cost.
The pieces and where they came from
| Piece | What it does here | Built in |
|---|---|---|
catalog.json, technest.py | The bot under test | TechNest RAG app |
judge.py | The Groq judge every metric uses | Custom judge model |
goldens.json | The five questions and expected answers | Datasets and goldens |
Correctness GEval | Are the answer's facts the expected ones? | G-Eval |
FaithfulnessMetric | Is every claim backed by a retrieved chunk? | Faithfulness |
ContextualRecallMetric | Did the search find what the expected answer needs? | Contextual recall |
evaluate() with AsyncConfig | One run over every golden, paced for the free tier | evaluate(), Rate limits and retries |
Pick one to watch it run, step by step.
The run function API
scores = run(top_k=3) # {"g001": {"Correctness [GEval]": 0.9, "Faithfulness": 1.0, ...}, ...}Three metrics on every answer
Correctness grades the answer against the golden, faithfulness grades it against the chunks, and contextual recall grades the search. Between them, a low score points at the part to fix.
METRICS = [correctness, FaithfulnessMetric(model=judge), ContextualRecallMetric(model=judge)]One test case per golden
for golden in GOLDENS:
response, contexts = answer(golden["input"], top_k=top_k)
test_cases.append(LLMTestCase(name=golden["name"], input=golden["input"], actual_output=response,
expected_output=golden["expected_output"], retrieval_context=contexts))name carries the golden's id into the result, so the scores come back labelled g001 to g005.
A quiet evaluate() call
result = evaluate(test_cases, METRICS, hyperparameters={"top_k": top_k},
async_config=AsyncConfig(max_concurrent=1),
display_config=DisplayConfig(print_results=False, show_indicator=False))
return {r.name: {m.name: m.score for m in r.metrics_data} for r in result.test_results}- written in Custom judge model
- written in TechNest RAG app
- written in TechNest RAG app
- written in Datasets and goldens
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
[
{
"name": "g001",
"input": "What is TechNest's return policy?",
"expected_output": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
},
{
"name": "g002",
"input": "What are the RAM and storage specs of the ProBook X1?",
"expected_output": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
},
{
"name": "g003",
"input": "How long is the battery life on the SoundPods Pro?",
"expected_output": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
},
{
"name": "g004",
"input": "What are TechNest's shipping options and how long do returns take to process?",
"expected_output": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
},
{
"name": "g005",
"input": "What is the price of the PixelPhone 15?",
"expected_output": "The TechNest PixelPhone 15 is priced at $899."
}
]
The evaluate_bot.py file
Save it as evaluate_bot.py, next to the other TechNest files.
import json
from deepeval import evaluate
from deepeval.evaluate import AsyncConfig, DisplayConfig
from deepeval.metrics import ContextualRecallMetric, FaithfulnessMetric, GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams
from judge import judge
from technest import answer
GOLDENS = json.load(open("goldens.json", encoding="utf-8"))
correctness = GEval(
name="Correctness",
evaluation_steps=[
"Check whether the facts in 'actual output' contradict any facts in 'expected output'.",
"Penalize a missing fact from 'expected output'.",
"Different wording is fine when the facts are the same.",
],
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
model=judge,
)
METRICS = [correctness, FaithfulnessMetric(model=judge), ContextualRecallMetric(model=judge)]
def run(top_k):
"""Ask the bot every golden with this top_k and return {golden: {metric: score}}."""
test_cases = []
for golden in GOLDENS:
response, contexts = answer(golden["input"], top_k=top_k)
test_cases.append(LLMTestCase(name=golden["name"], input=golden["input"], actual_output=response,
expected_output=golden["expected_output"], retrieval_context=contexts))
result = evaluate(test_cases, METRICS, hyperparameters={"top_k": top_k},
async_config=AsyncConfig(max_concurrent=1),
display_config=DisplayConfig(print_results=False, show_indicator=False))
return {r.name: {m.name: m.score for m in r.metrics_data} for r in result.test_results}Scoring the bot and a candidate change
Someone reads that fewer chunks means a shorter prompt and a cheaper call, and changes top_k from 3 to 1. The answers still read well. The suite scores both versions and lists every golden where a score fell by 0.3 or more.
from evaluate_bot import run
def show(label, scores):
print(label)
for name, by_metric in scores.items():
print(" ", name, " ".join(f"{metric} {score:.2f}" for metric, score in by_metric.items()))
baseline, candidate = run(top_k=3), run(top_k=1)
show("baseline, top_k=3", baseline)
show("candidate, top_k=1", candidate)
regressions = [f"{name} {metric}: {before:.2f} -> {candidate[name][metric]:.2f}"
for name in baseline for metric, before in baseline[name].items()
if before - candidate[name][metric] >= 0.3]
print("regressions:", regressions or "none")
print("verdict:", "block the change" if regressions else "safe to ship")⚠ WARNING: No prompts logged. » Log prompts to evaluate and optimize your prompt templates and models. ================================================================================ ✓ Evaluation completed 🎉! (time taken: 179.07s | token cost: None) » Test Results (5 total tests): » Pass Rate: 60.0% | Passed: 3 | Failed: 2 =============================================================================== = » Want to share evals with your team, or a place for your test cases to live? ❤️ 🏡 » Run 'deepeval view' to analyze and save testing results on Confident AI. ⚠ WARNING: No prompts logged. » Log prompts to evaluate and optimize your prompt templates and models. ================================================================================ ✓ Evaluation completed 🎉! (time taken: 192.29s | token cost: None) » Test Results (5 total tests): » Pass Rate: 60.0% | Passed: 3 | Failed: 2 =============================================================================== = » Want to share evals with your team, or a place for your test cases to live? ❤️ 🏡 » Run 'deepeval view' to analyze and save testing results on Confident AI. baseline, top_k=3 g001 Correctness [GEval] 1.00 Faithfulness 1.00 Contextual Recall 1.00 g002 Correctness [GEval] 1.00 Faithfulness 1.00 Contextual Recall 1.00 g003 Correctness [GEval] 0.00 Faithfulness 1.00 Contextual Recall 1.00 g004 Correctness [GEval] 0.30 Faithfulness 1.00 Contextual Recall 1.00 g005 Correctness [GEval] 1.00 Faithfulness 1.00 Contextual Recall 1.00 candidate, top_k=1 g001 Correctness [GEval] 1.00 Faithfulness 1.00 Contextual Recall 1.00 g002 Correctness [GEval] 1.00 Faithfulness 1.00 Contextual Recall 1.00 g003 Correctness [GEval] 0.20 Faithfulness 1.00 Contextual Recall 0.00 g004 Correctness [GEval] 0.30 Faithfulness 0.50 Contextual Recall 0.50 g005 Correctness [GEval] 1.00 Faithfulness 1.00 Contextual Recall 1.00 regressions: ['g003 Contextual Recall: 1.00 -> 0.00', 'g004 Faithfulness: 1.00 -> 0.50', 'g004 Contextual Recall: 1.00 -> 0.50'] verdict: block the change
What the two runs show
- The report lines come first.
print_results=Falsehides the per-test-case report, butevaluate()still prints its warning and the summary: both runs passed 3 of 5 goldens. - The baseline is not perfect. g003 scores 0.00 for correctness while faithfulness and recall are 1.00: the search found the SoundPods Pro entry and the answer is backed by it, but it gives 24 hours where the expected answer adds them up to 32. g004, the two-part golden, scores 0.30. Both are problems in the answer, and the retriever metrics say the search is not to blame.
- With
top_k=1the search breaks. g003's contextual recall falls from 1.00 to 0.00, and g004's from 1.00 to 0.50: one chunk cannot hold both the shipping and the return policy. - The answer drifts with it. g004's faithfulness drops to 0.50: half of its claims are no longer backed by the one chunk it got.
- Correctness alone would have missed most of it. It barely moved, since g003 and g004 were already failing it. The retriever metrics are what show the change made the search worse.
- The verdict is "block the change", with three regressions listed by golden and metric.
- A second run moved one regression. Run again, the same code blocked the change with three regressions too, but g004's faithfulness stayed at 1.00 and its correctness fell from 0.50 to 0.10 instead. The two recall drops repeated exactly, because the search is plain Python and gives the same chunks every time; the judged scores of the answers move between runs. That is why the comparison lists every regression, not one number.
Baseline vs candidate
Baseline, top_k=3 | Candidate, top_k=1 | |
|---|---|---|
| Chunks per answer | 3 | 1 |
| Prompt size | Larger | Smaller |
| What the suite decides | The reference run | Blocked: three regressions in g003 and g004 |
What the course left out
DeepEval's docs cover more than these lessons. Each of these is a documented feature that a TechNest suite could add next.
| Left out | What it is | Docs |
|---|---|---|
| Conversation simulator | Generates multi-turn conversations from goldens by playing the user | conversation-simulator |
| Argument correctness, step efficiency, plan quality and adherence | More agent metrics on tool arguments and plans | metrics-argument-correctness |
| Role adherence, conversation completeness, goal accuracy, topic adherence, turn-level RAG metrics | More multi-turn metrics | metrics-role-adherence |
| Hallucination, summarization, JSON correctness, non-advice, misuse, role violation | More single-turn metrics | metrics-hallucination |
| Arena G-Eval and arena test cases | Compare two apps' answers side by side | metrics-arena-g-eval |
| Conversational DAG | A decision tree over a whole conversation | metrics-conversational-dag |
| JevEval and the hybrid and system_one eval modes | Metrics judged by TypeSafe AI's Jev model; needs a TypeSafe key | metrics-jev-eval |
| Classifiers | Label answers (refusal, escalation, tone) and test the label | classifiers-introduction |
| MCP metrics | Score an app's use of MCP servers | metrics-mcp-use |
| Prompt optimization | GEPA, MIPROv2 and others improve a prompt against your metrics | prompt-optimization-introduction |
| Benchmarks | MMLU, HellaSwag, GSM8K and more, for models rather than apps | benchmarks-introduction |
| Multimodal and voice metrics | Images in test cases, and speech quality | evaluation-voice |
| Confident AI platform | Hosted datasets, reports, regression tracking and tracing | getting-started |
When to run this suite
- Before changing anything in
technest.py: the prompt, the model,top_kor the search. - On a schedule, as the full set in Evals in CI, with the essential goldens on every pull request.
Related
- Previous: Integrations
- Back to: DeepEval overview
- Reference: RAG evaluation quickstart
- Change the candidate to
run(top_k=5)and check whether any score drops. - Lower the regression gap from 0.3 to 0.1 and count how many lines it prints.
- Add
AnswerRelevancyMetric(model=judge)toMETRICSand look for g005's padded price answer in the scores.
Slow is fine. Stopping is the only problem.