Deep AgentsDeep Agents 0.7 · Python 3.11+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

Fault tolerance: call limits and tool retries

Fault tolerance in a deep agent is middleware that handles failure: a call limit stops a run that uses too many model calls, and a retry runs a failing tool again before the error reaches the model.

Last updated: 29 Sep, 2026 · Deep Agents 0.7

The video does not cover failures; this lesson follows the docs' fault-tolerance page. Two failures matter most for the trip planner: a run that loops and burns the free tier's tokens, and a flaky service that fails once and works on the second try.

Stopping a run with ModelCallLimitMiddleware

python
ModelCallLimitMiddleware(run_limit=2, exit_behavior="end")   # at most 2 model calls per invoke

Start trip.py with search_travel, the catalog tool from Tools: a travel search the agent can call. Everything below goes in the same file, under it.

python
from langchain.tools import tool

CATALOG = {
    "paris": {
        "flight": ["Return flight Delhi to Paris: 42,000 rupees"],
        "hotel": ["Seine Budget Inn, Latin Quarter: 5,200 rupees a night",
                  "Hotel Lumiere, Montmartre: 7,500 rupees a night",
                  "Le Grand Opera Hotel: 16,000 rupees a night"],
        "sight": ["Eiffel Tower summit: 3,100 rupees", "Louvre Museum: 2,000 rupees",
                  "Seine river cruise: 1,500 rupees", "Versailles day trip: 2,600 rupees",
                  "Montmartre walking tour: free"],
        "food": ["Cafe breakfast and bistro dinner: 3,000 rupees a day"],
    },
}


@tool
def search_travel(city: str, kind: str) -> str:
    """Search the travel catalog. kind is "flight", "hotel", "sight" or "food". Prices are in rupees."""
    entries = CATALOG.get(city.lower(), {}).get(kind)
    return "\n".join(entries) if entries else f"The catalog has no {kind} entries for {city}."
python
from deepagents import create_deep_agent
from langchain.chat_models import init_chat_model

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0, max_retries=6)
python
from langchain.agents.middleware import ModelCallLimitMiddleware

agent = create_deep_agent(
    model=model,
    tools=[search_travel],
    middleware=[ModelCallLimitMiddleware(run_limit=2, exit_behavior="end")],   # at most 2 model calls per run
    system_prompt="You are a travel planner. Look up each kind of price separately with search_travel.",
)

A request that needs more calls than allowed

Looked up one at a time, as this prompt asks, four lookups take five model calls. The limit allows two.

ExampleAPI keytrip.py, continued
result = agent.invoke({"messages": [{"role": "user", "content": "List the Paris flight, hotel, sight and food prices."}]})
for message in result["messages"]:
    print(f"{message.type:<5}", message.text or [(c["name"], c["args"]) for c in message.tool_calls])

What the limit did

  • Two model calls ran: each asked for one lookup, and both tool results came back.
  • The third call never happened. With exit_behavior="end" the run ended with a message saying the limit was reached, instead of an exception.
  • The answer is incomplete. A hard stop costs less than a runaway loop.

Retrying a flaky tool with ToolRetryMiddleware

python
ToolRetryMiddleware(max_retries=2, initial_delay=0.5, tools=["exchange_rate"])

An exchange-rate tool that fails once

This is a separate file, retry.py, with its own model. exchange_rate counts its calls and raises ConnectionError on the first one, the way a real service times out now and then.

python
from deepagents import create_deep_agent
from langchain.chat_models import init_chat_model

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0, max_retries=6)
python
from langchain.agents.middleware import ToolRetryMiddleware
from langchain.tools import tool

attempts = {"count": 0}


@tool
def exchange_rate(currency: str) -> str:
    """Get today's rate from rupees to another currency."""
    attempts["count"] += 1
    if attempts["count"] == 1:
        raise ConnectionError("rates service timed out")   # the first call fails
    return f"1 {currency} = 92 rupees"

The agent with the retry

python
agent = create_deep_agent(
    model=model,
    tools=[exchange_rate],
    middleware=[ToolRetryMiddleware(max_retries=2, initial_delay=0.5, tools=["exchange_rate"])],
    system_prompt="Answer in one sentence using exchange_rate.",
)

Retrying the rate lookup after a timeout

ExampleAPI keyretry.py, continued
result = agent.invoke({"messages": [{"role": "user", "content": "How many rupees is one euro today?"}]})
print("tool attempts:", attempts["count"])
print(result["messages"][-1].text)

What the retry did

  • The tool ran twice: the first attempt raised, the middleware waited about half a second and ran it again.
  • The model never saw the error. It received the second attempt's result and answered from it.

The fault-tolerance middleware compared

MiddlewareHandlesFrom
ModelCallLimitMiddlewareRunaway loopslangchain.agents.middleware
ToolRetryMiddlewareA tool that fails now and thenlangchain.agents.middleware
ModelRetryMiddlewareA model call that fails now and thenlangchain.agents.middleware
ModelFallbackMiddlewareA provider that is downlangchain.agents.middleware
max_retries on the modelRate limits such as Groq's per-minute capthe chat model

ModelRetryMiddleware() retries a failed model call with a growing wait. By default it retries every error except those the provider marks as not retryable; langchain-groq raises plain Groq errors, so this covers a failure you may meet in these lessons: now and then gpt-oss writes a malformed tool call, and Groq rejects the request with a 400 error that says tool_use_failed. Running the example again works; the middleware does that for you.

Where each fits

  • A call limit on any agent that runs unattended.
  • Tool retries around network calls: prices, weather, payments.
  • A fallback model when one provider's outage should not stop the app.
Watch out. Do not retry a tool that has side effects without care. Retrying a booking that timed out after it succeeded books twice; retry reads, and make writes safe to repeat.
Try it yourself
  • Set run_limit=6 and let the request finish.
  • Make exchange_rate fail three times in a row and read what the model receives.
  • Change exit_behavior to "error" and catch the exception it raises.

Slow is fine. Stopping is the only problem.