Tool errors, caught and retried
ToolErrorMiddleware is a middleware that catches a tool's exception and hands the model a message instead of crashing the run, while ToolRetryMiddleware calls the failing tool again first.
Last updated: 27 Sep, 2026 · LangChain 1.4
A tool that raises stops the whole run. The shop's order database times out now and then, so this version of lookup_order fails twice before it answers.
ShopModel, the stand-in from Several tool calls at once, asks exactly once, so every run shows only what the middleware did.ToolErrorMiddleware and ToolRetryMiddleware
from langchain.agents.middleware import ToolErrorMiddleware, ToolRetryMiddleware
# turn a tool exception into a message the model can read
errors = ToolErrorMiddleware(on_error) # on_error(exc, request) -> str or None
# call a failing tool again before giving up
retry = ToolRetryMiddleware(max_retries=3, initial_delay=0, on_failure="error")A tool that times out
Start with a tool that fails. OUTAGES holds two timeouts; each call pops one and raises until the list is empty.
from langchain.tools import tool
OUTAGES = ["timed out", "timed out"] # two failures, then success
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
if OUTAGES:
raise ConnectionError(f"order database {OUTAGES.pop()}")
return f"{order_id} shipped on 3 March."The plain agent
Add ShopModel below the tool.
import re
from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult
class ShopModel(BaseChatModel):
tools: list = []
@property
def _llm_type(self):
return "shop"
def bind_tools(self, tools, **kwargs):
return self.model_copy(update={"tools": tools}) # a copy holding the tools
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
message = self.decide(messages) # the reply comes from decide
return ChatResult(generations=[ChatGeneration(message=message)])
def decide(self, messages):
results = [] # the tool results at the end
for m in reversed(messages):
if not isinstance(m, ToolMessage):
break
results.insert(0, m.text)
if results: # results are back: answer with them
return AIMessage(" ".join(results))
text = messages[-1].text
orders = re.findall(r"\b[A-Z]\d+\b", text)
tool = "refund_order" if "refund" in text.lower() else "lookup_order"
if orders and tool in [t.name for t in self.tools]: # one call per order id
calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
for o in orders]
return AIMessage("", tool_calls=calls)
if orders: # that tool is not bound
return AIMessage(f"I have no way to look up {orders[0]} yet.")
return AIMessage("Hello. Which order is this about?")Without any handling
The first tool call raises, and the exception ends the run.
from langchain.agents import create_agent
agent = create_agent(ShopModel(), tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})Traceback (most recent call last):
File "main.py", line 4, in <module>
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})
ConnectionError: order database timed out
During task with name 'tools' and id 'aef9903f-af44-d04a-77dc-ca8420abe341'The exception came straight out of invoke. The model never heard about it, and the customer got no answer.
Errors as messages, and retries
on_error receives the exception and the tool request, and returns the text of a tool message; returning None lets the exception through. ToolRetryMiddleware calls the tool again, here with no wait between attempts; its default is one second, doubling each time. on_failure="error" tells it to re-raise the last exception once every attempt has failed, so something outside it can decide what the model sees; the default, "continue", would write its own failure message instead.
def on_error(exc, request):
return f"lookup failed: {exc}" # the tool message text; None re-raises
retry = ToolRetryMiddleware(max_retries=3, initial_delay=0, on_failure="error")
errors = ToolErrorMiddleware(on_error)The wrong order
Order the list [retry, errors], which reads naturally (retry, then catch), and print the tool message. Each example runs as a fresh program, so OUTAGES starts with two timeouts every time; it is module state, and in one long session you would refill it between runs.
agent = create_agent(ShopModel(), tools=[lookup_order], middleware=[retry, errors])
result = agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})
print(next(m for m in result["messages"] if m.type == "tool").text) # the tool's messagelookup failed: order database timed out
The tool was never retried. As The order middleware runs in showed, the first middleware is the outermost layer: errors sits inside retry, catches the first timeout and hands back a message, so retry sees a success.
The right order
Swap them to [errors, retry], with the same two outages waiting.
agent = create_agent(ShopModel(), tools=[lookup_order], middleware=[errors, retry])
result = agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})
print(next(m for m in result["messages"] if m.type == "tool").text) # the tool's messageA17 shipped on 3 March.
With errors first, retry is inside it and sees the raw exceptions. Two timeouts, a third attempt that worked, and a real answer. errors only gets involved if every attempt fails.
Why the order changed the result
- No middleware means a crash. The first timeout came straight out of
invoke, and the customer got nothing. - Order decides who wraps whom. The first middleware in the list is the outer layer, from the middleware-order lesson.
[retry, errors]retried nothing.errorswas inside, caught the first timeout, and handedretrya success to see.[errors, retry]worked.retrywas inside, saw the raw timeouts, tried again, and the third attempt returned the real status.
ToolErrorMiddleware vs ToolRetryMiddleware
| ToolErrorMiddleware | ToolRetryMiddleware | |
|---|---|---|
| When it acts | After the tool raises | When the tool raises, before giving up |
| What it does | Returns a message from on_error | Calls the tool again |
| On give-up | Passes the message to the model | Re-raises or reports, per on_failure |
| Put it | Outside, to catch the last failure | Inside, so it sees each raw error |
Where error handling fits
- A tool that calls a flaky network service which times out now and then.
- Turning a raised error into a message the model can explain to the user.
- Retrying a rate-limited API a few times before reporting it failed.
ToolErrorMiddleware before ToolRetryMiddleware in the list. The other way round, the error handler catches the first failure and the retry never runs.Related
- Previous: Changing the call: wrap_model_call
- Next: Call limits with ModelCallLimitMiddleware
- Reference: LangChain agent middleware
- Put three entries in
OUTAGESand run the working order again. - Return
Nonefromon_errorforConnectionErrorand see the exception come back. - Print the
statusof the tool message whenon_errorhandled the error.
You understood something today that you didn't yesterday.