Tool errors, caught and retried
ToolErrorMiddleware is a middleware that catches a tool's exception and hands the model a message instead of crashing the run, while ToolRetryMiddleware calls the failing tool again first.
Last updated: 27 Sep, 2026 · LangChain 1.4
A tool that raises stops the whole run. The shop's order database times out now and then, so this version of lookup_order fails twice before it answers.
ToolErrorMiddleware and ToolRetryMiddleware
from langchain.agents.middleware import ToolErrorMiddleware, ToolRetryMiddleware
# turn a tool exception into a message the model can read
errors = ToolErrorMiddleware(on_error) # on_error(exc, request) -> str or None
# call a failing tool again before giving up
retry = ToolRetryMiddleware(max_retries=3, initial_delay=0, on_failure="error")A tool that times out
Start with a tool that fails. OUTAGES holds two timeouts; each call pops one and raises until the list is empty.
from langchain.tools import tool
OUTAGES = ["timed out", "timed out"] # two failures, then success
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
if OUTAGES:
raise ConnectionError(f"order database {OUTAGES.pop()}")
return f"{order_id} shipped on 3 March."The plain agent
Build a plain agent with that tool and no middleware yet. ShopModel is the stand-in chat model from earlier lessons; it asks for lookup_order when it sees an order id.
from langchain.agents import create_agent
from shop_model import ShopModel
from tools import lookup_order
agent = create_agent(ShopModel(), tools=[lookup_order])Without any handling
The first tool call raises, and the exception ends the run.
agent = create_agent(ShopModel(), tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})The exception came straight out of invoke. The model never heard about it, and the customer got no answer.
Errors as messages, and retries
on_error receives the exception and the tool request, and returns the text of a tool message; returning None lets the exception through. ToolRetryMiddleware calls the tool again, here with no wait between attempts; its default is one second, doubling each time.
def on_error(exc, request):
return f"lookup failed: {exc}" # the tool message text; None re-raises
retry = ToolRetryMiddleware(max_retries=3, initial_delay=0, on_failure="error")
errors = ToolErrorMiddleware(on_error)The wrong order
Order the list [retry, errors], as the documentation's own example does, and print the tool message.
agent = create_agent(ShopModel(), tools=[lookup_order], middleware=[retry, errors])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][2].text)This is the order the documentation's example uses, and the tool was never retried. From lesson 17, the first middleware is the outermost layer: errors sits inside retry, catches the first timeout and hands back a message, so retry sees a success.
The right order
Swap them to [errors, retry].
agent = create_agent(ShopModel(), tools=[lookup_order], middleware=[errors, retry])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][2].text)With errors first, retry is inside it and sees the raw exceptions. Two timeouts, a third attempt that worked, and a real answer. errors only gets involved if every attempt fails.
Why the order changed the result
- No middleware means a crash. The first timeout came straight out of
invoke, and the customer got nothing. - Order decides who wraps whom. The first middleware in the list is the outer layer, from lesson 17.
[retry, errors]retried nothing.errorswas inside, caught the first timeout, and handedretrya success to see.[errors, retry]worked.retrywas inside, saw the raw timeouts, tried again, and the third attempt returned the real status.
ToolErrorMiddleware vs ToolRetryMiddleware
| ToolErrorMiddleware | ToolRetryMiddleware | |
|---|---|---|
| When it acts | After the tool raises | When the tool raises, before giving up |
| What it does | Returns a message from on_error | Calls the tool again |
| On give-up | Passes the message to the model | Re-raises or reports, per on_failure |
| Put it | Outside, to catch the last failure | Inside, so it sees each raw error |
Where error handling fits
- A tool that calls a flaky network service which times out now and then.
- Turning a raised error into a message the model can explain to the user.
- Retrying a rate-limited API a few times before reporting it failed.
ToolErrorMiddleware before ToolRetryMiddleware in the list. The other way round, the error handler catches the first failure and the retry never runs.Related
- Previous: Changing the call: wrap_model_call
- Next: Call limits with ModelCallLimitMiddleware
- Reference: LangChain agent middleware
- Put three entries in
OUTAGESand run the working order again. - Return
Nonefromon_errorforConnectionErrorand see the exception come back. - Print the
statusof the tool message whenon_errorhandled the error.
You understood something today that you didn't yesterday.