LLMTestCase
LLMTestCase is DeepEval's class for one interaction with an LLM app: the input, the app's actual output, and optional fields such as the expected output and the retrieved chunks, which metrics read to score it.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
The TechNest bot returns an answer and the chunks it used. A test case puts those next to the question and the answer TechNest expects, so a metric has everything it needs in one object.
Four fields from two places
The video breaks the evaluation pipeline into its parts. The goldens are a container of input questions, each with its expected output. Each input goes to the RAG pipeline: the retriever fetches the context, and the LLM writes the actual output from it. One box collects all four: input, expected output, actual output and retrieved context. That box is the test case.
The actual output and the retrieved context come from the real run of the app, the student's answer in the exam analogy. The input and the expected output come from the goldens, the teacher's side. In DeepEval the box is an LLMTestCase, and the four fields are named input, expected_output, actual_output and retrieval_context.

The LLMTestCase API
from deepeval.test_case import LLMTestCase
test_case = LLMTestCase(
input="...", # what the user asked (required)
actual_output="...", # what the app answered
expected_output="...", # the answer TechNest expects
retrieval_context=["..."], # the chunks the app retrieved, best first
)The fields written in advance
The question and the expected answer are written before the app runs, by someone who knows the right answer.
question = "What is TechNest's return policy?"
expected = ("TechNest accepts returns within 30 days of purchase. Items must be in original "
"packaging with all accessories. Customers pay return shipping unless the item is "
"defective. Refunds are processed in 5 to 7 business days.")The fields from the bot's run
response, contexts = answer(question) # one real run of technest.py- written in TechNest RAG app
- written in TechNest RAG app
View the code here
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
Building a test case from one bot run
Run it in the folder that holds technest.py and catalog.json from TechNest RAG app.
from deepeval.test_case import LLMTestCase
from technest import answer
question = "What is TechNest's return policy?"
expected = ("TechNest accepts returns within 30 days of purchase. Items must be in original "
"packaging with all accessories. Customers pay return shipping unless the item is "
"defective. Refunds are processed in 5 to 7 business days.")
response, contexts = answer(question)
test_case = LLMTestCase(
input=question,
actual_output=response,
expected_output=expected,
retrieval_context=contexts,
)
print("input: ", test_case.input)
print("actual_output: ", test_case.actual_output)
print("expected_output: ", test_case.expected_output[:60], "...")
print("retrieval_context:", len(test_case.retrieval_context), "chunks")input: What is TechNest's return policy? actual_output: TechNest accepts returns within 30 days of the original purchase date, provided items are in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged, and refunds are processed within 5 to 7 business days of receiving the returned item. expected_output: TechNest accepts returns within 30 days of purchase. Items m ... retrieval_context: 3 chunks
What the test case holds
inputis the question exactly as the user typed it. It never includes the system prompt or the chunks; those belong to the app.actual_outputis the bot's real answer from this run, which is why it is the one field that changes when you run the example again.expected_outputis the answer TechNest expects, written from the catalog.retrieval_contextholds the three chunks the search returned, in rank order. Metrics that grade the search read this field.
retrieval_context vs context
LLMTestCase has a second list of strings, context, and the two are easy to mix up.
retrieval_context | context | |
|---|---|---|
| Holds | What the app's search returned on this run | The ideal facts for this question, written in advance |
| Changes between runs? | Yes, it comes from the app | No, it is static |
| Read by | Faithfulness and the contextual metrics | The hallucination metric |
Choosing what one test case covers
- The whole bot: question in, answer and chunks out, as above. This is where most evaluations start.
- One step: only the search (question in, chunks out), or only the generation. LLM tracing shows how DeepEval scores single steps.
- An agent: the same class with
tools_calledfilled in, which Tool correctness uses.
input is required to build a test case, so a missing answer is not caught when you create it. It is caught when a metric that reads the field runs: answer relevancy on a test case with no actual_output stops with MissingTestCaseParamsError: 'actual_output' cannot be None for the 'Answer Relevancy' metric.Related
- Previous: TechNest RAG app
- Next: ExactMatchMetric and PatternMatchMetric
- Reference: Single-turn test case
- Build a test case for
"What is the price of the PixelPhone 15?"with the expected answer"The TechNest PixelPhone 15 is priced at $899."and print itsactual_output. - In the original return-policy example, print
test_case.retrieval_context[0]and check that it is the return policy entry. - Add
context=["TechNest accepts returns within 30 days of purchase."]to the test case and printtest_case.context.
This is what real progress feels like.