Task completion
Task completion is a DeepEval agent metric that reads an agent's whole trace, works out the task the user gave and the outcome the agent reached, and asks the judge how well that outcome achieves the task.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
Tool correctness checked the agent's calls against a list you wrote by hand. Many tasks have no single right list, and a run can call the right tools and still leave the customer without what they asked for. Task completion needs no list: the judge reads what happened in the trace.
The TaskCompletionMetric API
from deepeval.metrics import TaskCompletionMetric
task_completion = TaskCompletionMetric(model=judge) # task=... pins the goal; by default it is read from the trace
for golden in dataset.evals_iterator(metrics=[task_completion]):
run_agent(golden.input) # a traced agent: every call is one trace to scoreThe docs say the metric analyses the full trace and needs tracing set up. Passing it to evals_iterator, as the docs' agent evaluation pages do, scores every trace in the loop as a whole.
The spans agent.py records
@observe(type="tool")
def track_order(order_id): ...
@observe(type="llm", model=AGENT_MODEL)
def call_model(messages): ...
@observe(type="agent", available_tools=["search_catalog", "track_order"])
def run_agent(question, max_steps=4): ...These decorators were in agent.py from Tool correctness; this is where they pay off. Each call of run_agent is a trace with an agent span at the top, an LLM span for every model reply and a tool span for every tool run, each holding its input and output. available_tools is extra information on the agent span. The judge reads this tree as JSON.
Setting the trace's input and output
if not message.tool_calls: # no tool asked for: this is the answer
update_current_trace(input=question, output=message.content, tools_called=tools_called)
return message.content, tools_calledWhen the agent answers, update_current_trace from LLM tracing fills the trace's test case with the question, the final reply and the tool calls.
Two tasks as goldens
dataset = EvaluationDataset(goldens=[
Golden(input="Where is my order TN-1001?"),
Golden(input="Please cancel my order TN-1002."),
])A golden for task completion needs only the input. The first task is one the agent's tools can do. The second is not: the agent can look orders up, but it has no tool that cancels one.
- written in Tool correctness
- written in Custom judge model
- written in TechNest RAG app
- written in TechNest RAG app
View the code here
import json
import os
from deepeval.test_case import ToolCall
from deepeval.tracing import observe, update_current_trace
from openai import OpenAI
from technest import retrieve
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
AGENT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are TechNest's support agent. Use search_catalog for questions about products and
store policies, and track_order for questions about an order. Answer only from what the tools return.
If no tool can do what the customer asks, say so. Reply in two or three plain sentences."""
ORDERS = {
"TN-1001": {"item": "SoundPods Pro", "status": "shipped", "arrives": "in 2 business days"},
"TN-1002": {"item": "ProBook X1", "status": "processing", "arrives": "in 5 to 7 business days"},
}
TOOLS = [
{"type": "function", "function": {
"name": "search_catalog",
"description": "Search TechNest's products and store policies.",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]},
}},
{"type": "function", "function": {
"name": "track_order",
"description": "Look up the status of an order by its id, such as TN-1001.",
"parameters": {"type": "object", "properties": {"order_id": {"type": "string"}}, "required": ["order_id"]},
}},
]
@observe(type="tool")
def search_catalog(query):
return "\n\n".join(retrieve(query))
@observe(type="tool")
def track_order(order_id):
order = ORDERS.get(order_id)
return json.dumps(order) if order else f"No order found with id {order_id}."
FUNCTIONS = {"search_catalog": search_catalog, "track_order": track_order}
@observe(type="llm", model=AGENT_MODEL)
def call_model(messages):
reply = groq.chat.completions.create(model=AGENT_MODEL, messages=messages, tools=TOOLS, temperature=0)
return reply.choices[0].message
@observe(type="agent", available_tools=["search_catalog", "track_order"])
def run_agent(question, max_steps=4):
messages = [{"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": question}]
tools_called = [] # every tool the agent ran, for the metrics
for _ in range(max_steps):
message = call_model(messages)
if not message.tool_calls: # no tool asked for: this is the answer
update_current_trace(input=question, output=message.content, tools_called=tools_called)
return message.content, tools_called
messages.append(message.model_dump(exclude_none=True))
for call in message.tool_calls: # run each tool the model asked for
args = json.loads(call.function.arguments)
output = FUNCTIONS[call.function.name](**args)
tools_called.append(ToolCall(name=call.function.name, input_parameters=args, output=output))
messages.append({"role": "tool", "tool_call_id": call.id, "content": output})
return "I could not finish this request.", tools_called
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
Scoring two tasks, one the agent cannot do
from deepeval.dataset import EvaluationDataset, Golden
from deepeval.evaluate import AsyncConfig
from deepeval.metrics import TaskCompletionMetric
from agent import run_agent
from judge import judge
dataset = EvaluationDataset(goldens=[
Golden(input="Where is my order TN-1001?"),
Golden(input="Please cancel my order TN-1002."),
])
task_completion = TaskCompletionMetric(model=judge)
for golden in dataset.evals_iterator(
metrics=[task_completion],
async_config=AsyncConfig(run_async=False), # one task at a time
):
run_agent(golden.input)╭──────────────────────────────────────────────────────────────────────────────╮ │ 🚀 DeepEval Evaluation Results │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ ✅ test_case_1 (Passed 1 metrics) │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ │ │ ❌ test_case_2 │ │ ├── Input: Please cancel my order TN-1002. │ │ │ Actual Output: Your order TN-1002 for the ProBook X1 is │ │ │ currently processing and expected to arrive in 5 │ │ │ to 7 business days. I'm unable to cancel the │ │ │ order directly, so please contact TechNest │ │ │ support to arrange the cancellation. │ │ └── Metrics │ │ Status ┃ Metric ┃ Score ┃ Threshold ┃ Reason │ │ ━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━ │ │ FAIL │ Task Completion │ 0.15 │ 0.50 │ The system only │ │ │ │ │ │ retrieved the order │ │ │ │ │ │ status and advised │ │ │ │ │ │ contacting support, │ │ │ │ │ │ but it did not │ │ │ │ │ │ perform the requested │ │ │ │ │ │ cancellation of order │ │ │ │ │ │ TN-1002, thus it │ │ │ │ │ │ largely fails to meet │ │ │ │ │ │ the task. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ Aggregate Metrics │ │ │ │ Metric ┃ Average Score ┃ Pass Rate ┃ Total │ │ ━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━ │ │ Task Completion │ 0.55 │ 50.00% | passed=1 | failed=1 │ 2 │ ╰──────────────────────────────────────────────────────────────────────────────╯ ⚠ WARNING: No hyperparameters logged. » Log hyperparameters to attribute prompts and models to your test runs. ================================================================================ ✓ Evaluation completed 🎉! (time taken: 38.25s | token cost: None) » Test Results (2 total tests): » Pass Rate: 50.0% | Passed: 1 | Failed: 1 =============================================================================== = » Want to share evals with your team, or a place for your test cases to live? ❤️ 🏡 » Run 'deepeval view' to analyze and save testing results on Confident AI.
Why the cancellation failed
- test_case_1, the order lookup, passed, so the report folds it into one line. DeepEval names a golden without a
nameby its position. - test_case_2, the cancellation, scored 0.15 and failed. The agent reported the order's status, a ProBook X1 still processing, and said it is unable to cancel the order. The judge's reason: the system only retrieved the order status and advised contacting support, but did not perform the cancellation.
- The agent behaved well and still failed. It did not invent a cancellation, and it pointed the customer to support. Task completion scores the outcome against the task, and the task was not done. Here the fix is in the product, a cancel tool, not in the prompt.
- Aggregate Metrics shows a 50% pass rate over the two tasks, one passed and one failed.
Task completion vs tool correctness
TaskCompletionMetric | ToolCorrectnessMetric | |
|---|---|---|
| Asks | Was the user's task achieved? | Were the expected tools called? |
| Reads | The whole trace | tools_called and expected_tools |
| Needs a reference | No | Yes, the expected tools |
| Scored by | The judge | Counting, no LLM |
| Cancel request | Fails: nothing was cancelled | Passes if you only expected track_order |
When to score task completion
- When tasks have many valid paths and writing
expected_toolsfor each would be guesswork. - When you need to find requests the agent cannot serve yet: a low score with an honest reply points at a missing tool.
- Together with tool correctness, which catches a right outcome reached with the wrong or extra tools.
Related
- Previous: Tool correctness
- Next: Multi-turn evaluation
- Reference: Task completion
- Add
Golden(input="Where is my order TN-9999?"), an order that does not exist.track_orderreturns No order found, the agent tells the customer it could not find the order, and the judge can pass that: reporting a missing order is the task done. - Pass
task="Cancel the customer's order."toTaskCompletionMetricand run the two goldens again: both are now judged against that one task. - Set
AGENT_MODEL = "openai/gpt-oss-20b"inagent.pyand run the cancel task again; compare what the reply says about cancelling. In one run the smaller model reported the order status and did not mention the cancellation at all.
You understood something today that you didn't yesterday.