Retries and a fallback model
ModelRetryMiddleware is a middleware that calls the same model again after it fails, and ModelFallbackMiddleware moves the request on to another model when the first keeps failing.
Last updated: 27 Sep, 2026 · LangChain 1.4
When a provider goes down
A hosted model is a service, and services go down. The video starts from a real outage: OpenAI was down for four hours in November 2023, and apps that had GPT-4 hardcoded went dark. Behind a gateway, when one provider fails the request falls back to another, which the video calls a must for any production app. A fallback is that list of other models to try, in order, when the first one fails.
The video runs this with LiteLLM, a gateway that sits between your app and every provider; its completion function takes a fallbacks list. The first run uses gemini/gemini-1.5-flash as the primary with a Google key whose account had been suspended: it fails with a 403 PERMISSION_DENIED error, and the answer comes from gpt-4o-mini, the first fallback. The second run, below, uses a model name that does not exist, so the primary is certain to fail, with the same two backups. It needs an OpenAI key and pip install litellm:
from litellm import completion
# Force the primary to fail by using a fake model name
# Then watch the fallback chain rescue the call
response = completion(
model="openai/fake-nonexistent-model-9999", # 👈 will fail intentionally
messages=[{"role": "user", "content": "What is an LLM Gateway?"}],
fallbacks=[
"gpt-4o-mini", # 1st backup: real OpenAI model
"groq/llama-3.3-70b-versatile" # 2nd backup: Groq
]
)
print("✅ App still got a response, even though the primary failed!")
print(f"\n🤖 Model that actually answered: {response.model}")
print(f"\n📝 Response: {response.choices[0].message.content[:200]}...")✅ App still got a response, even though the primary failed! 🤖 Model that actually answered: gpt-4o-mini-2024-07-18 📝 Response: An LLM (Large Language Model) Gateway typically refers to an interface or platform that facilitates access to large language models like GPT-3, GPT-4, or other advanced natural language processing mod...
Shown as it ran in the video, not run here. It needs pip install litellm and an OpenAI key, and its second backup model has since been retired on Groq. The fallback middleware later in this lesson runs on the keys you have.
LiteLLM raised an error for the fake model, yet the app still got its answer from gpt-4o-mini, the first backup, and the calling code never had to handle the failure. Inside a LangChain agent the same safety net comes from middleware, which retries or switches models within the agent's loop.
Now the shop. A real provider cannot be made to fail on cue, so two stand-ins fail the way a provider does. FlakyModel fails twice and then works; DownModel never works.
ModelRetryMiddleware and ModelFallbackMiddleware
from langchain.agents.middleware import ModelFallbackMiddleware, ModelRetryMiddleware
retry = ModelRetryMiddleware(max_retries=2, initial_delay=0) # 3 attempts in all
backup = ModelFallbackMiddleware(other_model) # try other_model if the main raisesA flaky model and a dead one
The agent here runs on ShopModel, the stand-in chat model built in Several tool calls at once. Start the file with it.
import re
from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult
class ShopModel(BaseChatModel):
tools: list = []
@property
def _llm_type(self):
return "shop"
def bind_tools(self, tools, **kwargs):
return self.model_copy(update={"tools": tools}) # a copy holding the tools
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
message = self.decide(messages) # the reply comes from decide
return ChatResult(generations=[ChatGeneration(message=message)])
def decide(self, messages):
results = [] # the tool results at the end
for m in reversed(messages):
if not isinstance(m, ToolMessage):
break
results.insert(0, m.text)
if results: # results are back: answer with them
return AIMessage(" ".join(results))
text = messages[-1].text
orders = re.findall(r"\b[A-Z]\d+\b", text)
tool = "refund_order" if "refund" in text.lower() else "lookup_order"
if orders and tool in [t.name for t in self.tools]: # one call per order id
calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
for o in orders]
return AIMessage("", tool_calls=calls)
if orders: # that tool is not bound
return AIMessage(f"I have no way to look up {orders[0]} yet.")
return AIMessage("Hello. Which order is this about?")Both stand-ins extend ShopModel and override _generate, the method that produces a reply, to raise instead.
FAILURES = [ConnectionError("provider unavailable")] * 2 # two failures, then success
class FlakyModel(ShopModel):
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
if FAILURES:
raise FAILURES.pop()
return super()._generate(messages)
class DownModel(ShopModel):
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
raise ConnectionError("provider unavailable") # never worksThe tool and the agent
This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Add it below the models.
from langchain.tools import tool
ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
status = ORDERS.get(order_id)
return f"{order_id} {status}." if status else f"{order_id} is not an order we have."Without any handling
Nothing in the agent retries a failed model call on its own, so the first failure ends the run.
from langchain.agents import create_agent
agent = create_agent(FlakyModel(), tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})Traceback (most recent call last):
File "main.py", line 4, in <module>
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})
ConnectionError: provider unavailable
During task with name 'model' and id 'c9b0e398-d861-30d5-bce5-22e9ccc3182f'LangChain's hosted chat models retry a failed request a couple of times by default, inside their HTTP client, for network errors, rate limits and server errors, though not for a wrong key. FlakyModel raises straight from _generate, with no HTTP client underneath, so nothing retried it.
Trying again
ModelRetryMiddleware retries the same model. FlakyModel works on its third attempt.
from langchain.agents.middleware import ModelFallbackMiddleware, ModelRetryMiddleware
retry = ModelRetryMiddleware(max_retries=2, initial_delay=0)
agent = create_agent(FlakyModel(), tools=[lookup_order], middleware=[retry])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)A17 shipped on 3 March.
Two failures, a third attempt, a normal answer. max_retries=2 means three attempts in all. The defaults wait one second before the first retry and double the wait each time, with some randomness, up to a minute; initial_delay=0 keeps this example quick.
When the model never recovers, the retry gives up.
retry = ModelRetryMiddleware(max_retries=2, initial_delay=0)
agent = create_agent(DownModel(), tools=[lookup_order], middleware=[retry])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)Model call failed after 3 attempts with ConnectionError: provider unavailable
When every attempt fails, the default on_failure="continue" ends the run with an AI message describing the failure instead of raising.
Another model
ModelFallbackMiddleware takes one or more backup models, tried in order. When the main model raises, the request moves to the next one. With hosted models, the fallback is usually a different provider, so one outage does not take the shop's support desk down.
backup = ModelFallbackMiddleware(ShopModel())
agent = create_agent(DownModel(), tools=[lookup_order], middleware=[backup])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)A17 shipped on 3 March.
How each recovered from failure
- Nothing in the agent retries a model error.
FlakyModelhas no HTTP client underneath, so the first run crashed with the provider error. - Retry tries the same model.
FlakyModelrecovered on its third attempt. - Retry gives up cleanly. Against
DownModel, all three attempts failed and the run ended with a message, not a crash, because ofon_failure="continue". - Fallback switches models. When
DownModelraised, the request went toShopModel, which answered.
ModelRetryMiddleware vs ModelFallbackMiddleware
| ModelRetryMiddleware | ModelFallbackMiddleware | |
|---|---|---|
| What it does | Calls the same model again | Calls a different model, in order |
| Handles | A brief timeout or rate limit | A model or provider that is down |
| Attempts | max_retries + 1 on one model | Each model once, in the order given |
| Waiting | Backs off between attempts | Moves on to the next at once |
Where retries and fallback fit
- Riding out a brief provider timeout or rate limit with a retry.
- Keeping the support desk up during one provider's outage by falling back to another.
- Pairing both: retry the main model, then fall back if it stays down.
initial_delay low in tests, and keep the backoff in production so a struggling provider is not hammered.Related
- Previous: Call limits with ModelCallLimitMiddleware
- Next: Summarization with SummarizationMiddleware
- Reference: LangChain agent middleware
- Set
on_failure="error"on the retry and run it againstDownModel. - Give
ModelFallbackMiddlewaretwo models, aDownModelfirst, and check which one answers. - Put retry and fallback in one list and work out which runs first before you try it.
This is what real progress feels like.