Fault tolerance: call limits and tool retries
Fault tolerance in a deep agent is middleware that handles failure: a call limit stops a run that uses too many model calls, and a retry runs a failing tool again before the error reaches the model.
Last updated: 29 Sep, 2026 · Deep Agents 0.7
The video does not cover failures; this lesson follows the docs' fault-tolerance page. Two failures matter most for the trip planner: a run that loops and burns the free tier's tokens, and a flaky service that fails once and works on the second try.
Stopping a run with ModelCallLimitMiddleware
ModelCallLimitMiddleware(run_limit=2, exit_behavior="end") # at most 2 model calls per invokeStart trip.py with search_travel, the catalog tool from Tools: a travel search the agent can call. Everything below goes in the same file, under it.
from langchain.tools import tool
CATALOG = {
"paris": {
"flight": ["Return flight Delhi to Paris: 42,000 rupees"],
"hotel": ["Seine Budget Inn, Latin Quarter: 5,200 rupees a night",
"Hotel Lumiere, Montmartre: 7,500 rupees a night",
"Le Grand Opera Hotel: 16,000 rupees a night"],
"sight": ["Eiffel Tower summit: 3,100 rupees", "Louvre Museum: 2,000 rupees",
"Seine river cruise: 1,500 rupees", "Versailles day trip: 2,600 rupees",
"Montmartre walking tour: free"],
"food": ["Cafe breakfast and bistro dinner: 3,000 rupees a day"],
},
}
@tool
def search_travel(city: str, kind: str) -> str:
"""Search the travel catalog. kind is "flight", "hotel", "sight" or "food". Prices are in rupees."""
entries = CATALOG.get(city.lower(), {}).get(kind)
return "\n".join(entries) if entries else f"The catalog has no {kind} entries for {city}."from deepagents import create_deep_agent
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0, max_retries=6)from langchain.agents.middleware import ModelCallLimitMiddleware
agent = create_deep_agent(
model=model,
tools=[search_travel],
middleware=[ModelCallLimitMiddleware(run_limit=2, exit_behavior="end")], # at most 2 model calls per run
system_prompt="You are a travel planner. Look up each kind of price separately with search_travel.",
)A request that needs more calls than allowed
Looked up one at a time, as this prompt asks, four lookups take five model calls. The limit allows two.
result = agent.invoke({"messages": [{"role": "user", "content": "List the Paris flight, hotel, sight and food prices."}]})
for message in result["messages"]:
print(f"{message.type:<5}", message.text or [(c["name"], c["args"]) for c in message.tool_calls])human List the Paris flight, hotel, sight and food prices.
ai [('search_travel', {'city': 'Paris', 'kind': 'flight'})]
tool Return flight Delhi to Paris: 42,000 rupees
ai [('search_travel', {'city': 'Paris', 'kind': 'hotel'})]
tool Seine Budget Inn, Latin Quarter: 5,200 rupees a night
Hotel Lumiere, Montmartre: 7,500 rupees a night
Le Grand Opera Hotel: 16,000 rupees a night
ai Model call limits exceeded: run limit (2/2)What the limit did
- Two model calls ran: each asked for one lookup, and both tool results came back.
- The third call never happened. With
exit_behavior="end"the run ended with a message saying the limit was reached, instead of an exception. - The answer is incomplete. A hard stop costs less than a runaway loop.
Retrying a flaky tool with ToolRetryMiddleware
ToolRetryMiddleware(max_retries=2, initial_delay=0.5, tools=["exchange_rate"])An exchange-rate tool that fails once
This is a separate file, retry.py, with its own model. exchange_rate counts its calls and raises ConnectionError on the first one, the way a real service times out now and then.
from deepagents import create_deep_agent
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0, max_retries=6)from langchain.agents.middleware import ToolRetryMiddleware
from langchain.tools import tool
attempts = {"count": 0}
@tool
def exchange_rate(currency: str) -> str:
"""Get today's rate from rupees to another currency."""
attempts["count"] += 1
if attempts["count"] == 1:
raise ConnectionError("rates service timed out") # the first call fails
return f"1 {currency} = 92 rupees"The agent with the retry
agent = create_deep_agent(
model=model,
tools=[exchange_rate],
middleware=[ToolRetryMiddleware(max_retries=2, initial_delay=0.5, tools=["exchange_rate"])],
system_prompt="Answer in one sentence using exchange_rate.",
)Retrying the rate lookup after a timeout
result = agent.invoke({"messages": [{"role": "user", "content": "How many rupees is one euro today?"}]})
print("tool attempts:", attempts["count"])
print(result["messages"][-1].text)tool attempts: 2 1 EUR = 92 rupees.
What the retry did
- The tool ran twice: the first attempt raised, the middleware waited about half a second and ran it again.
- The model never saw the error. It received the second attempt's result and answered from it.
The fault-tolerance middleware compared
| Middleware | Handles | From |
|---|---|---|
ModelCallLimitMiddleware | Runaway loops | langchain.agents.middleware |
ToolRetryMiddleware | A tool that fails now and then | langchain.agents.middleware |
ModelRetryMiddleware | A model call that fails now and then | langchain.agents.middleware |
ModelFallbackMiddleware | A provider that is down | langchain.agents.middleware |
max_retries on the model | Rate limits such as Groq's per-minute cap | the chat model |
ModelRetryMiddleware() retries a failed model call with a growing wait. By default it retries every error except those the provider marks as not retryable; langchain-groq raises plain Groq errors, so this covers a failure you may meet in these lessons: now and then gpt-oss writes a malformed tool call, and Groq rejects the request with a 400 error that says tool_use_failed. Running the example again works; the middleware does that for you.
Where each fits
- A call limit on any agent that runs unattended.
- Tool retries around network calls: prices, weather, payments.
- A fallback model when one provider's outage should not stop the app.
Related
- Previous: Permissions: rules for file writes
- Next: Custom middleware: logging every tool call
- Reference: Deep Agents fault tolerance
- Set
run_limit=6and let the request finish. - Make
exchange_ratefail three times in a row and read what the model receives. - Change
exit_behaviorto"error"and catch the exception it raises.
Slow is fine. Stopping is the only problem.