Tool failures mid-run
A tool failure is an exception a tool raises mid-run; CrewAI sends the error text to the model, finishes the run, and records the failure on the result.
Last updated: 28 Sep, 2026 · CrewAI 1.15
The tool-calls lesson's tool always answered. A real order system goes down; here the tool raises instead, and the crew keeps going.
View the code here
import json
import os
import re
from crewai import BaseLLM
os.environ["CREWAI_DISABLE_TELEMETRY"] = "true"
os.environ["CREWAI_TRACING_ENABLED"] = "false"
os.environ["CREWAI_DISABLE_VERSION_CHECK"] = "true"
class ShopLLM(BaseLLM):
script: list = []
def supports_function_calling(self):
return True
def call(self, messages, tools=None, **kwargs):
if isinstance(messages, str):
messages = [{"role": "user", "content": messages}]
if self.script:
return self.script.pop(0)
names = [t["function"]["name"] for t in tools or []]
return self.decide(messages, names)
def decide(self, messages, tools):
last = messages[-1]
if last["role"] == "tool":
return last["content"]
text = last["content"]
orders = re.findall(r"\b[A-Z]\d+\b", text)
wanted = "refund_order" if "refund" in text.lower() else "lookup_order"
matches = [name for name in tools if name.endswith(wanted)]
if orders and matches:
args = json.dumps({"order_id": orders[0]})
return [{"id": f"call_{orders[0]}", "type": "function",
"function": {"name": matches[0], "arguments": args}}]
if orders:
return f"I have no way to look up {orders[0]} yet."
return "Hello. Which order is this about?"
The clerk and a broken tool
from crewai import Agent, Crew, Task
from shop_llm import ShopLLM
from tools import lookup_order
clerk = Agent(
role="Order clerk",
goal="Find the status of customers' orders",
backstory="You can look up any order in the shop's system.",
llm=ShopLLM(model="shop"),
tools=[lookup_order],
)
task = Task(description="Answer the customer: {question}",
expected_output="The order's status in one sentence.", agent=clerk)
crew = Crew(agents=[clerk], tasks=[task])from crewai.tools import tool
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
raise ConnectionError("the order database is not answering")A run that finishes on an error
result = crew.kickoff(inputs={"question": "Where is my order A17?"})
print(result.raw)Error executing tool: the order database is not answering
The crew finished, and its answer is the error. CrewAI caught the exception and sent Error executing tool: and the message to the model as the tool's result. Your model repeated it. A hosted model would write something politer, and the run would look like a success.
Checking the result for failures
result = crew.kickoff(inputs={"question": "Where is my order A17?"})
print(result.has_tool_failures)
for record in result.tool_failures:
print(record.tool_name, "|", record.failure.message)True lookup_order | the order database is not answering
CrewAI records the failure even though the run completed. The documentation's advice is to check has_tool_failures before treating raw as complete. A desk could send the ticket to a person instead of mailing the customer.
Stopping the run instead
task.tool_failure_policy = "raise"
try:
crew.kickoff(inputs={"question": "Where is my order A17?"})
except Exception as error:
print(type(error).__name__)
print(error)[CrewAIEventsBus] Warning: Event pairing mismatch. 'task_failed' closed 'agent_execution_started' (expected 'task_started') [CrewAIEventsBus] Warning: Event pairing mismatch. 'crew_kickoff_failed' closed 'task_started' (expected 'crew_kickoff_started') ToolExecutionFailedError Tool 'lookup_order' failed during 'Answer the customer: Where is my order A17?': the order database is not answering (code: ConnectionError)
tool_failure_policy takes three values: warn records the failure and continues, raise stops the run with ToolExecutionFailedError, and ignore continues and records nothing. Left unset the policy behaves like warn, which is what the first run did. The two warning lines come from CrewAI's event bus as the failed run unwinds. The policy can be set on a tool, a task, an agent or the crew, and the most specific setting wins.
Reading the failed runs
- The run completed: the error became the tool's result and the model passed it on.
- has_tool_failures is True even though raw looks like an answer, so check it before trusting raw.
- Setting the policy to raise stops the run with ToolExecutionFailedError for your code to catch.
The three policy values
| Value | On a tool error |
|---|---|
| warn | record the failure and continue (how an unset policy behaves) |
| raise | stop the run with ToolExecutionFailedError |
| ignore | continue and record nothing |
When to pick each policy
- warn for a desk that should still answer, then flag the ticket for a person.
- raise in a pipeline where a bad result must not flow on.
- ignore for a tool whose failure truly does not matter.
Related
- Previous: Tool calls: asking for a tool
- Next: max_iter: capping an agent's loop
- Set the policy to
"ignore"and printhas_tool_failures. - Set the policy on the agent instead of the task.
- Make the tool return a normal string for C40 and raise only for other ids.
Every expert started right here.