Tool correctness
Tool correctness is a DeepEval agent metric that compares the tools an agent called with the tools it was expected to call and scores the share of expected tools it got right, with no LLM call unless you ask it to judge tool selection.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
Tracing gave every step of the RAG bot its own record. An agent goes further: it picks its own steps, deciding which tool to run and with which arguments. The first thing to check is whether it picked the right tools.
The Production RAG Live Marathon names DeepEval as its library for this check, the deterministic part of its evaluation pipeline where no LLM is needed, and its evaluation notebook has a smart-home example of it. This lesson runs that example with DeepEval's metric, then gives the TechNest bot two tools of its own.
The ToolCorrectnessMetric API
from deepeval.metrics import ToolCorrectnessMetric
from deepeval.test_case import LLMTestCase, ToolCall
test_case = LLMTestCase(
input="Where is my order TN-1001?",
actual_output="Your order has shipped.",
tools_called=[ToolCall(name="track_order", input_parameters={"order_id": "TN-1001"})], # what the agent ran
expected_tools=[ToolCall(name="track_order")], # what it should run
)
ToolCorrectnessMetric(model=judge).measure(test_case)A ToolCall records one call: its name, and optionally its input_parameters and output. By default only the names are compared.
Creating the metric without a judge
The score is plain counting, so it is tempting to leave out model. The constructor builds a judge model anyway, the OpenAI default, and it needs a key at that moment.
from deepeval.metrics import ToolCorrectnessMetric
metric = ToolCorrectnessMetric()Traceback (most recent call last):
File "main.py", line 3, in <module>
metric = ToolCorrectnessMetric()
deepeval.errors.DeepEvalError: OpenAI API key is not configured. Set OPENAI_API_KEY in your environment or pass `api_key` to OpenAIModel(...).Pass model=judge, the Groq judge from Custom judge model. The docs say the model is only asked anything when you also pass available_tools, to judge whether the chosen tools were the best ones; without that list no LLM is called.
The smart-home agent from the video's repository
The notebook gives a voice assistant three commands and records the tools it called: the right light tool for the living room, the light tool instead of music for the kitchen, and the thermostat plus an extra notification. It imports DeepEval's LLMTestCase and ToolCall to hold them, then scores them with a function of its own.
def tool_correctness_score(case):
called = {t.name for t in (case.tools_called or [])}
expected = {t.name for t in (case.expected_tools or [])}
if not expected:
return 1.0
union = len(called | expected)
return len(called & expected) / union if union > 0 else 0.0- written in Custom judge model
- written in TechNest RAG app
- written in TechNest RAG app
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
That is a Jaccard overlap: tools in both sets over tools in either set. The notebook's saved output shows 1.00, 0.00 and 0.50. Here are the same three test cases, scored by DeepEval's metric.
from deepeval.metrics import ToolCorrectnessMetric
from deepeval.test_case import LLMTestCase, ToolCall
from judge import judge
# ── Scenario: Smart Home Agent ───────────────────────────────────────────────
tool_cases = [
# Good: correct tool called
LLMTestCase(
input="Turn on the living room lights",
actual_output="I have turned on the living room lights for you.",
tools_called=[
ToolCall(name="control_smart_light", input_parameters={"room": "living room", "action": "on"}),
],
expected_tools=[ToolCall(name="control_smart_light")],
),
# Medium: wrong tool: user asked for music, agent triggered lights
LLMTestCase(
input="Play jazz music in the kitchen",
actual_output="Playing jazz in the kitchen.",
tools_called=[
ToolCall(name="control_smart_light", input_parameters={"room": "kitchen", "action": "on"}),
],
expected_tools=[ToolCall(name="play_music")],
),
# Bad: right primary tool but also called an extra unnecessary tool
LLMTestCase(
input="Set the thermostat to 22 degrees",
actual_output="Setting thermostat to 22 degrees Celsius.",
tools_called=[
ToolCall(name="set_thermostat", input_parameters={"temperature": 22}),
ToolCall(name="send_notification", input_parameters={"message": "Thermostat set to 22 C"}),
],
expected_tools=[ToolCall(name="set_thermostat")],
),
]
metric = ToolCorrectnessMetric(model=judge)
for case in tool_cases:
metric.measure(case)
print(f"{metric.score:.2f} | {case.input}")1.00 | Turn on the living room lights 0.00 | Play jazz music in the kitchen 1.00 | Set the thermostat to 22 degrees
DeepEval's scores vs the notebook's
- The living room lights score 1.00 in both: the one expected tool was called.
- The kitchen music scores 0.00 in both:
play_musicwas never called. - The thermostat splits them: the notebook says 0.50, DeepEval says 1.00. The installed DeepEval 4.2.8 divides the expected tools it found by the number of expected tools, so an extra call costs nothing. The docs page writes the formula as correctly used tools over total tools called, which would give 0.50 here, the notebook's number. The library decides the score, so read 1.00 as "every expected tool was called", not "no wrong tool was called".
Penalising the extra call with should_exact_match
To fail a run that calls anything beyond the list, ask for an exact match. Add these lines to the end of the same file.
strict = ToolCorrectnessMetric(model=judge, should_exact_match=True)
strict.measure(tool_cases[2]) # the thermostat case
print(f"{strict.score:.2f}")
print(strict.reason)0.00 [ Tool Calling Reason: Not an exact match: expected ['set_thermostat'], called ['set_thermostat', 'send_notification']. See details above. Tool Selection Reason: No available tools were provided to assess tool selection criteria ]
With should_exact_match=True the called list must equal the expected list, so the extra send_notification drops the score to 0.00, and the reason names both lists.
The TechNest agent
The smart-home cases were written by hand. A real agent produces tools_called itself, so the TechNest bot gets two tools and a loop that runs them.
The orders and the two tools
ORDERS = {
"TN-1001": {"item": "SoundPods Pro", "status": "shipped", "arrives": "in 2 business days"},
"TN-1002": {"item": "ProBook X1", "status": "processing", "arrives": "in 5 to 7 business days"},
}
@observe(type="tool")
def search_catalog(query):
return "\n\n".join(retrieve(query))
@observe(type="tool")
def track_order(order_id):
order = ORDERS.get(order_id)
return json.dumps(order) if order else f"No order found with id {order_id}."search_catalog is the word-overlap search from technest.py. track_order looks an id up in two fake orders. The @observe lines are the tracing from LLM tracing: here they only record spans, which Task completion scores.
Describing a tool to the model
{"type": "function", "function": {
"name": "track_order",
"description": "Look up the status of an order by its id, such as TN-1001.",
"parameters": {"type": "object", "properties": {"order_id": {"type": "string"}}, "required": ["order_id"]},
}},Each tool is described in the OpenAI tools format that Groq accepts: a name, a sentence the model reads to decide when to use it, and a JSON schema for its arguments. TOOLS holds one such entry per tool.
Asking the model until it answers
def run_agent(question, max_steps=4):
messages = [{"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": question}]
tools_called = [] # every tool the agent ran, for the metrics
for _ in range(max_steps):
message = call_model(messages)
if not message.tool_calls: # no tool asked for: this is the answer
update_current_trace(input=question, output=message.content, tools_called=tools_called)
return message.content, tools_called
...call_model sends the messages and the tool list to Groq. When the reply asks for no tool, it is the final answer, and the function returns it with the list of calls.
Recording each tool call
messages.append(message.model_dump(exclude_none=True))
for call in message.tool_calls: # run each tool the model asked for
args = json.loads(call.function.arguments)
output = FUNCTIONS[call.function.name](**args)
tools_called.append(ToolCall(name=call.function.name, input_parameters=args, output=output))
messages.append({"role": "tool", "tool_call_id": call.id, "content": output})Every call becomes a DeepEval ToolCall with the arguments the model chose and what the tool returned, and the tool's result goes back to the model for its next turn.
The agent.py file
Save the agent as agent.py, next to technest.py and catalog.json.
import json
import os
from deepeval.test_case import ToolCall
from deepeval.tracing import observe, update_current_trace
from openai import OpenAI
from technest import retrieve
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
AGENT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are TechNest's support agent. Use search_catalog for questions about products and
store policies, and track_order for questions about an order. Answer only from what the tools return.
If no tool can do what the customer asks, say so. Reply in two or three plain sentences."""
ORDERS = {
"TN-1001": {"item": "SoundPods Pro", "status": "shipped", "arrives": "in 2 business days"},
"TN-1002": {"item": "ProBook X1", "status": "processing", "arrives": "in 5 to 7 business days"},
}
TOOLS = [
{"type": "function", "function": {
"name": "search_catalog",
"description": "Search TechNest's products and store policies.",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]},
}},
{"type": "function", "function": {
"name": "track_order",
"description": "Look up the status of an order by its id, such as TN-1001.",
"parameters": {"type": "object", "properties": {"order_id": {"type": "string"}}, "required": ["order_id"]},
}},
]
@observe(type="tool")
def search_catalog(query):
return "\n\n".join(retrieve(query))
@observe(type="tool")
def track_order(order_id):
order = ORDERS.get(order_id)
return json.dumps(order) if order else f"No order found with id {order_id}."
FUNCTIONS = {"search_catalog": search_catalog, "track_order": track_order}
@observe(type="llm", model=AGENT_MODEL)
def call_model(messages):
reply = groq.chat.completions.create(model=AGENT_MODEL, messages=messages, tools=TOOLS, temperature=0)
return reply.choices[0].message
@observe(type="agent", available_tools=["search_catalog", "track_order"])
def run_agent(question, max_steps=4):
messages = [{"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": question}]
tools_called = [] # every tool the agent ran, for the metrics
for _ in range(max_steps):
message = call_model(messages)
if not message.tool_calls: # no tool asked for: this is the answer
update_current_trace(input=question, output=message.content, tools_called=tools_called)
return message.content, tools_called
messages.append(message.model_dump(exclude_none=True))
for call in message.tool_calls: # run each tool the model asked for
args = json.loads(call.function.arguments)
output = FUNCTIONS[call.function.name](**args)
tools_called.append(ToolCall(name=call.function.name, input_parameters=args, output=output))
messages.append({"role": "tool", "tool_call_id": call.id, "content": output})
return "I could not finish this request.", tools_calledThe agent runs on qwen/qwen3.8-27b, the bot's model, so it is not the judge's model.
Scoring the agent's tool calls
Three questions, each with the tools a person would expect. The second metric adds ToolCallParams.INPUT_PARAMETERS, so the arguments must match as well as the names.
from deepeval.metrics import ToolCorrectnessMetric
from deepeval.test_case import LLMTestCase, ToolCall, ToolCallParams
from agent import run_agent
from judge import judge
checks = [
("Where is my order TN-1001?",
[ToolCall(name="track_order", input_parameters={"order_id": "TN-1001"})]),
("How much RAM does the ProBook X1 have?",
[ToolCall(name="search_catalog", input_parameters={"query": "ProBook X1"})]),
("Has order TN-1002 shipped, and what is the battery life of the SoundPods Pro?",
[ToolCall(name="track_order", input_parameters={"order_id": "TN-1002"}),
ToolCall(name="search_catalog", input_parameters={"query": "SoundPods Pro battery"})]),
]
by_name = ToolCorrectnessMetric(model=judge)
with_arguments = ToolCorrectnessMetric(model=judge, evaluation_params=[ToolCallParams.INPUT_PARAMETERS])
for question, expected_tools in checks:
response, tools_called = run_agent(question)
test_case = LLMTestCase(input=question, actual_output=response,
tools_called=tools_called, expected_tools=expected_tools)
by_name.measure(test_case)
with_arguments.measure(test_case)
print(question)
print(" called:", [(tool.name, tool.input_parameters) for tool in tools_called])
print(f" names {by_name.score:.2f} | names and arguments {with_arguments.score:.2f}")[Confident AI Trace Log] No Confident AI API key found. Skipping trace posting.
Where is my order TN-1001?
called: [('track_order', {'order_id': 'TN-1001'})]
names 1.00 | names and arguments 1.00
[Confident AI Trace Log] No Confident AI API key found. Skipping trace posting.
How much RAM does the ProBook X1 have?
called: [('search_catalog', {'query': 'ProBook X1 RAM'})]
names 1.00 | names and arguments 0.00
[Confident AI Trace Log] No Confident AI API key found. Skipping trace posting.
Has order TN-1002 shipped, and what is the battery life of the SoundPods Pro?
called: [('track_order', {'order_id': 'TN-1002'}), ('search_catalog', {'query': 'SoundPods Pro battery life'})]
names 1.00 | names and arguments 0.50What the two scores show
- The trace log line before each question is printed by the
@observedecorators inagent.py: with no Confident AI key the trace is not sent anywhere. - TN-1001 scores 1.00 on both: the agent called
track_orderwith exactlyTN-1001. - The ProBook question scores 1.00 on names and 0.00 with arguments. The agent searched for
ProBook X1 RAM, the test expectedProBook X1, and DeepEval compares argument values for equality. Both searches would find the laptop; the metric cannot tell. - The two-part question made the agent call both tools. Names score 1.00. With arguments it scores 0.50:
TN-1002matches, but the agent searched forSoundPods Pro battery lifeand the test expectedSoundPods Pro battery, and the score averages the two expected tools.
Tool correctness vs argument correctness
ToolCorrectnessMetric | ArgumentCorrectnessMetric | |
|---|---|---|
| Checks | The right tools were called | The arguments passed to them make sense |
| Needs | tools_called and expected_tools | input and tools_called, no reference (the docs also list actual_output; 4.2.8 does not require it) |
| Scored by | Counting matches, no LLM | An LLM judge |
| Good for | Ids, fixed values, which tool ran | Free-text arguments such as a search query |
When to check tool correctness
- When a wrong tool has a cost: tracking the wrong order, or a light switching on instead of music.
- In CI, because the name check runs without any judge call and gives the same score every run for the same calls.
- With
INPUT_PARAMETERSonly for arguments that have one right value, such as an order id.
ToolCallParams.INPUT_PARAMETERS, every expected ToolCall needs input_parameters. In DeepEval 4.2.8 an expected tool without them stops the run with AttributeError: 'NoneType' object has no attribute 'keys' as soon as a called tool has the same name.Related
- Previous: LLM tracing
- Next: Task completion
- Reference: Tool correctness
- Delete the
send_notificationcall from the thermostat case and run the exact-match lines again: the score becomes 1.00. - In the TechNest checks, change the expected ProBook query to
"ProBook X1 RAM"and check that its arguments score becomes 1.00. - Set
AGENT_MODEL = "openai/gpt-oss-20b"inagent.pyand run the TechNest checks again; compare the queries it chooses.
Little by little, you're building something great.