LLM tracing
LLM tracing is DeepEval's way of recording each step of an app run as a span inside one trace, so a metric can score a single step, such as the LLM call, or the whole run from question to answer.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
Comparing two models told you which answers got better or worse. It did not tell you which step was to blame: the search that picked the chunks, or the model that wrote from them. A trace keeps every step apart, with its own input and output, so each step can get its own test case.
Spans, traces and the waterfall
The video recalls three words from observability: span, trace and waterfall. It opens the app's traces in Pydantic Logfire, recorded from the FastAPI backend. One line, one unit of execution, is a span, and it has a span name and a span ID. Many spans make up a trace. A step that runs smaller steps opens up like a trace inside a trace, and the waterfall on the right shows how long each span took.
DeepEval records the same tree with a decorator, @observe, and adds one thing the evaluation needs: any span, and the trace as a whole, can carry an LLMTestCase and metrics. In DeepEval's terms, one call of your app is one trace, and the steps inside it are spans nested under it.
The @observe API
from deepeval.tracing import observe, update_current_span, update_current_trace
@observe(type="retriever") # every call of this function becomes a span
def retrieve(question):
...
update_current_span(input=question, retrieval_context=contexts) # what this step did
update_current_trace(input=question, output=response) # the test case of the whole runThe outermost @observe function that runs starts the trace, and every decorated function it calls adds a span under it. type is a label: "llm", "retriever", "tool", "agent", or none for a plain step. The docs say it does not change any score; it names the role of the span in the tree. Both update functions take the fields of an LLMTestCase, with output standing for actual_output.
Tracing the retrieve step
@observe(type="retriever")
def retrieve(question):
contexts = technest.retrieve(question)
update_current_span(input=question, retrieval_context=contexts)
return contextsThe function wraps the word-overlap search from the TechNest bot and records what it was asked and which chunks it returned.
A metric on the generate span
@observe(type="llm", model=technest.CHAT_MODEL, metrics=[AnswerRelevancyMetric(model=judge)])
def generate(question, contexts):
response = technest.generate(question, contexts)
# the test case this span's metric scores
update_current_span(test_case=LLMTestCase(input=question, actual_output=response, retrieval_context=contexts))
return responsemetrics=[...] on the decorator attaches a metric to this span only, and update_current_span(test_case=...) gives it the test case to score. This is component-level evaluation: answer relevancy, from Answer relevancy, judges the LLM call alone.
The trace's own test case
@observe()
def answer(question):
contexts = retrieve(question)
response = generate(question, contexts)
# the test case for the whole run, scored by the metrics passed to evals_iterator
update_current_trace(input=question, output=response, retrieval_context=contexts)
return responseanswer is the outermost function, so each call is one trace with two spans under it. update_current_trace sets the trace's test case: the customer's question, the final answer and the chunks. A metric on the trace is an end-to-end evaluation, the app scored as a black box.
- written in Custom judge model
- written in TechNest RAG app
- written in TechNest RAG app
- written in Datasets and goldens
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
[
{
"name": "g001",
"input": "What is TechNest's return policy?",
"expected_output": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
},
{
"name": "g002",
"input": "What are the RAM and storage specs of the ProBook X1?",
"expected_output": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
},
{
"name": "g003",
"input": "How long is the battery life on the SoundPods Pro?",
"expected_output": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
},
{
"name": "g004",
"input": "What are TechNest's shipping options and how long do returns take to process?",
"expected_output": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
},
{
"name": "g005",
"input": "What is the price of the PixelPhone 15?",
"expected_output": "The TechNest PixelPhone 15 is priced at $899."
}
]
The traced_bot.py file
Save the three functions as traced_bot.py, next to technest.py, catalog.json and judge.py. It imports the bot instead of changing it.
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase
from deepeval.tracing import observe, update_current_span, update_current_trace
import technest
from judge import judge
@observe(type="retriever")
def retrieve(question):
contexts = technest.retrieve(question)
update_current_span(input=question, retrieval_context=contexts)
return contexts
@observe(type="llm", model=technest.CHAT_MODEL, metrics=[AnswerRelevancyMetric(model=judge)])
def generate(question, contexts):
response = technest.generate(question, contexts)
# the test case this span's metric scores
update_current_span(test_case=LLMTestCase(input=question, actual_output=response, retrieval_context=contexts))
return response
@observe()
def answer(question):
contexts = retrieve(question)
response = generate(question, contexts)
# the test case for the whole run, scored by the metrics passed to evals_iterator
update_current_trace(input=question, output=response, retrieval_context=contexts)
return responseCalling the traced bot outside an evaluation
from traced_bot import answer
print(answer("What is the price of the PixelPhone 15?"))[Confident AI Trace Log] No Confident AI API key found. Skipping trace posting. The TechNest PixelPhone 15 is priced at $899. It comes with a 1-year warranty and is available in Midnight Black and Arctic White.
The bot answers as before. No metric runs: the docs say an @observe function only evaluates inside an evaluation. The first line is DeepEval's trace log. With no Confident AI key, the trace stays on your machine and nothing is sent; CONFIDENT_TRACE_VERBOSE=0 in the environment hides the line.
Evaluating two goldens with evals_iterator
evals_iterator loops over a dataset's goldens, from Datasets and goldens, and turns each call of the traced app into a trace to score. Metrics passed to it score each trace's test case; the metric on generate scores the span. The run takes the first two goldens from goldens.json, which must sit in the same folder.
from deepeval.dataset import EvaluationDataset
from deepeval.evaluate import AsyncConfig
from deepeval.metrics import FaithfulnessMetric
from judge import judge
from traced_bot import answer
dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="goldens.json")
dataset.goldens = dataset.goldens[:2] # g001 and g002
for golden in dataset.evals_iterator(
metrics=[FaithfulnessMetric(model=judge)], # end-to-end: scores each trace
async_config=AsyncConfig(run_async=False), # one golden at a time
):
answer(golden.input)╭──────────────────────────────────────────────────────────────────────────────╮ │ 🚀 DeepEval Evaluation Results │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ ✅ g001 (Passed 1 metrics) │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ ✅ g002 (Passed 1 metrics) │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ │ │ ❌ generate │ │ ├── Input: What is TechNest's return policy? │ │ │ Actual Output: TechNest accepts returns within 30 days of the │ │ │ original purchase date, provided items are in │ │ │ their original packaging with all accessories │ │ │ included. Customers are responsible for return │ │ │ shipping costs unless the item arrives defective │ │ │ or damaged, and refunds are processed within 5 │ │ │ to 7 business days of receiving the returned │ │ │ item. │ │ └── Metrics │ │ Status ┃ Metric ┃ Score ┃ Threshold ┃ Reason │ │ ━━━━━━━━╇━━━━━━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━ │ │ PASS │ Answer Relevancy │ 1.00 │ 0.50 │ The score is 1.00 │ │ │ │ │ │ because the response │ │ │ │ │ │ fully ad... │ │ │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ │ │ ❌ generate │ │ ├── Input: What are the RAM and storage specs of the │ │ │ ProBook X1? │ │ │ Actual Output: The TechNest ProBook X1 is equipped with 16GB of │ │ │ DDR5 RAM and a 512GB NVMe SSD. │ │ └── Metrics │ │ Status ┃ Metric ┃ Score ┃ Threshold ┃ Reason │ │ ━━━━━━━━╇━━━━━━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━ │ │ PASS │ Answer Relevancy │ 1.00 │ 0.50 │ The score is 1.00 │ │ │ │ │ │ because the answer │ │ │ │ │ │ fully addr... │ │ │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ Aggregate Metrics │ │ │ │ Metric ┃ Average Score ┃ Pass Rate ┃ Total │ │ ━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━ │ │ Faithfulness │ 1.00 │ 100.00% | passed=2 | failed=0 │ 2 │ │ Answer Relevancy │ 1.00 │ 100.00% | passed=2 | failed=0 │ 2 │ ╰──────────────────────────────────────────────────────────────────────────────╯ ⚠ WARNING: No hyperparameters logged. » Log hyperparameters to attribute prompts and models to your test runs. ================================================================================ ✓ Evaluation completed 🎉! (time taken: 29.47s | token cost: None) » Test Results (2 total tests): » Pass Rate: 100.0% | Passed: 2 | Failed: 0 =============================================================================== = » Want to share evals with your team, or a place for your test cases to live? ❤️ 🏡 » Run 'deepeval view' to analyze and save testing results on Confident AI.
What the trace and span results show
- g001 and g002 are the two traces, named after their goldens. Each passed its one trace metric, faithfulness, so the report folds it into a single line.
- The two generate panels are the span results, one per trace. Their input is the golden's question and their output is the answer
generatewrote; answer relevancy scored 1.00 on both, PASS. - The red cross next to generate does not mean a failure. In DeepEval 4.2.8 a span result is drawn with ❌ even when every metric on it passes. Read the Status column and the aggregate table instead.
- Aggregate Metrics counts both scopes in one test run: Faithfulness 2 of 2 passed, Answer Relevancy 2 of 2 passed.
- The closing lines about
deepeval viewand hyperparameters are DeepEval's Confident AI notes. The run itself needed only the Groq key.
Component-level vs end-to-end evaluation
| Component-level | End-to-end | |
|---|---|---|
| Scores | One span, such as the LLM call or the search | The trace: question in, answer out |
| Metric goes on | @observe(metrics=[...]) | evals_iterator(metrics=[...]) |
| Test case from | update_current_span(test_case=...) | update_current_trace(...) |
| Tells you | Which step is weak | Whether the user got a good answer |
| In this run | Answer relevancy on generate | Faithfulness on each trace |
When to trace an app
- When an end-to-end score drops and you need to know whether retrieval or generation caused it.
- When one step can be graded on its own, such as the search with a contextual metric or a tool with tool correctness, which Tool correctness covers.
- When the app is an agent, whose steps a judge reads from the trace, as Task completion shows.
answer called technest.generate directly instead of the decorated generate, the span and its answer relevancy metric would silently disappear from the report, and only the trace metric would run.Related
- Previous: Comparing models
- Next: Tool correctness
- See also: Datasets and goldens for
goldens.json - Reference: LLM tracing
- Delete the
update_current_span(test_case=...)line ingenerateand run the evaluation again: the span's input becomes a dictionary of the function's arguments,questionandcontexts, and the metric scores that. - Run the outside-evaluation example with
CONFIDENT_TRACE_VERBOSE=0set in the environment and check that the trace log line is gone. - Change
[:2]to[2:3]in thedataset.goldensline to evaluate g003, the SoundPods Pro question whose search ranks the ProBook X1 first, and read both scores.
Every expert started right here.