LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Tool errors, caught and retried

ToolErrorMiddleware is a middleware that catches a tool's exception and hands the model a message instead of crashing the run, while ToolRetryMiddleware calls the failing tool again first.

Last updated: 27 Sep, 2026 · LangChain 1.4

A tool that raises stops the whole run. The shop's order database times out now and then, so this version of lookup_order fails twice before it answers.

Why a stand-in here. The lesson is about what the middleware does with a failing tool. A real model, told that a lookup failed, may decide on its own to ask again, and that retry would hide which middleware did the work. ShopModel, the stand-in from Several tool calls at once, asks exactly once, so every run shows only what the middleware did.

ToolErrorMiddleware and ToolRetryMiddleware

python
from langchain.agents.middleware import ToolErrorMiddleware, ToolRetryMiddleware

# turn a tool exception into a message the model can read
errors = ToolErrorMiddleware(on_error)                # on_error(exc, request) -> str or None
# call a failing tool again before giving up
retry = ToolRetryMiddleware(max_retries=3, initial_delay=0, on_failure="error")

A tool that times out

Start with a tool that fails. OUTAGES holds two timeouts; each call pops one and raises until the list is empty.

python
from langchain.tools import tool

OUTAGES = ["timed out", "timed out"]   # two failures, then success

@tool
def lookup_order(order_id: str) -> str:
    """Look up an order's shipping status by its id, such as A17."""
    if OUTAGES:
        raise ConnectionError(f"order database {OUTAGES.pop()}")
    return f"{order_id} shipped on 3 March."

The plain agent

Add ShopModel below the tool.

python
import re

from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult


class ShopModel(BaseChatModel):
    tools: list = []

    @property
    def _llm_type(self):
        return "shop"

    def bind_tools(self, tools, **kwargs):
        return self.model_copy(update={"tools": tools})   # a copy holding the tools

    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        message = self.decide(messages)                   # the reply comes from decide
        return ChatResult(generations=[ChatGeneration(message=message)])

    def decide(self, messages):
        results = []                                # the tool results at the end
        for m in reversed(messages):
            if not isinstance(m, ToolMessage):
                break
            results.insert(0, m.text)
        if results:                                 # results are back: answer with them
            return AIMessage(" ".join(results))
        text = messages[-1].text
        orders = re.findall(r"\b[A-Z]\d+\b", text)
        tool = "refund_order" if "refund" in text.lower() else "lookup_order"
        if orders and tool in [t.name for t in self.tools]:   # one call per order id
            calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
                     for o in orders]
            return AIMessage("", tool_calls=calls)
        if orders:                                  # that tool is not bound
            return AIMessage(f"I have no way to look up {orders[0]} yet.")
        return AIMessage("Hello. Which order is this about?")

Without any handling

The first tool call raises, and the exception ends the run.

Example
from langchain.agents import create_agent

agent = create_agent(ShopModel(), tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})

The exception came straight out of invoke. The model never heard about it, and the customer got no answer.

Errors as messages, and retries

on_error receives the exception and the tool request, and returns the text of a tool message; returning None lets the exception through. ToolRetryMiddleware calls the tool again, here with no wait between attempts; its default is one second, doubling each time. on_failure="error" tells it to re-raise the last exception once every attempt has failed, so something outside it can decide what the model sees; the default, "continue", would write its own failure message instead.

python
def on_error(exc, request):
    return f"lookup failed: {exc}"      # the tool message text; None re-raises

retry = ToolRetryMiddleware(max_retries=3, initial_delay=0, on_failure="error")
errors = ToolErrorMiddleware(on_error)

The wrong order

Order the list [retry, errors], which reads naturally (retry, then catch), and print the tool message. Each example runs as a fresh program, so OUTAGES starts with two timeouts every time; it is module state, and in one long session you would refill it between runs.

Example
agent = create_agent(ShopModel(), tools=[lookup_order], middleware=[retry, errors])
result = agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})
print(next(m for m in result["messages"] if m.type == "tool").text)   # the tool's message

The tool was never retried. As The order middleware runs in showed, the first middleware is the outermost layer: errors sits inside retry, catches the first timeout and hands back a message, so retry sees a success.

The right order

Swap them to [errors, retry], with the same two outages waiting.

Example
agent = create_agent(ShopModel(), tools=[lookup_order], middleware=[errors, retry])
result = agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})
print(next(m for m in result["messages"] if m.type == "tool").text)   # the tool's message

With errors first, retry is inside it and sees the raw exceptions. Two timeouts, a third attempt that worked, and a real answer. errors only gets involved if every attempt fails.

Why the order changed the result

  • No middleware means a crash. The first timeout came straight out of invoke, and the customer got nothing.
  • Order decides who wraps whom. The first middleware in the list is the outer layer, from the middleware-order lesson.
  • [retry, errors] retried nothing. errors was inside, caught the first timeout, and handed retry a success to see.
  • [errors, retry] worked. retry was inside, saw the raw timeouts, tried again, and the third attempt returned the real status.

ToolErrorMiddleware vs ToolRetryMiddleware

ToolErrorMiddlewareToolRetryMiddleware
When it actsAfter the tool raisesWhen the tool raises, before giving up
What it doesReturns a message from on_errorCalls the tool again
On give-upPasses the message to the modelRe-raises or reports, per on_failure
Put itOutside, to catch the last failureInside, so it sees each raw error

Where error handling fits

  • A tool that calls a flaky network service which times out now and then.
  • Turning a raised error into a message the model can explain to the user.
  • Retrying a rate-limited API a few times before reporting it failed.
Watch out. Put ToolErrorMiddleware before ToolRetryMiddleware in the list. The other way round, the error handler catches the first failure and the retry never runs.
Try it yourself
  • Put three entries in OUTAGES and run the working order again.
  • Return None from on_error for ConnectionError and see the exception come back.
  • Print the status of the tool message when on_error handled the error.

You understood something today that you didn't yesterday.