Retries and a fallback model
ModelRetryMiddleware is a middleware that calls the same model again after it fails, and ModelFallbackMiddleware moves the request on to another model when the first keeps failing.
Last updated: 27 Sep, 2026 · LangChain 1.4
Hosted models fail sometimes: a timeout, a rate limit, an outage. Two stand-ins fail the way a provider does. FlakyModel fails twice and then works; DownModel never works.
ModelRetryMiddleware and ModelFallbackMiddleware
from langchain.agents.middleware import ModelFallbackMiddleware, ModelRetryMiddleware
retry = ModelRetryMiddleware(max_retries=2, initial_delay=0) # 3 attempts in all
backup = ModelFallbackMiddleware(other_model) # try other_model if the main raisesA flaky model and a dead one
Both stand-ins extend ShopModel and override _generate, the method that produces a reply, to raise instead.
from shop_model import ShopModel
FAILURES = [ConnectionError("provider unavailable")] * 2 # two failures, then success
class FlakyModel(ShopModel):
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
if FAILURES:
raise FAILURES.pop()
return super()._generate(messages)
class DownModel(ShopModel):
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
raise ConnectionError("provider unavailable") # never worksThe imports
Bring in the pieces for a plain agent on FlakyModel, no middleware yet.
from langchain.agents import create_agent
from langchain.agents.middleware import ModelFallbackMiddleware, ModelRetryMiddleware
from flaky_model import DownModel, FlakyModel
from shop_model import ShopModel
from tools import lookup_orderWithout any handling
The first failure comes from inside the model, so nothing catches it.
agent = create_agent(FlakyModel(), tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})LangChain's hosted chat models retry a failed request a couple of times by default, for network errors, rate limits and server errors, though not for a wrong key. This error came from inside the model, so nothing retried it.
Trying again
ModelRetryMiddleware retries the same model. FlakyModel works on its third attempt.
retry = ModelRetryMiddleware(max_retries=2, initial_delay=0)
agent = create_agent(FlakyModel(), tools=[lookup_order], middleware=[retry])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)Two failures, a third attempt, a normal answer. max_retries=2 means three attempts in all. The defaults wait one second before the first retry and double the wait each time, with some randomness, up to a minute; initial_delay=0 keeps this example quick.
When the model never recovers, the retry gives up.
retry = ModelRetryMiddleware(max_retries=2, initial_delay=0)
agent = create_agent(DownModel(), tools=[lookup_order], middleware=[retry])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)When every attempt fails, the default on_failure="continue" ends the run with an AI message describing the failure instead of raising.
Another model
ModelFallbackMiddleware takes one or more backup models. When the main model raises, the request moves to the next one.
backup = ModelFallbackMiddleware(ShopModel())
agent = create_agent(DownModel(), tools=[lookup_order], middleware=[backup])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)ModelFallbackMiddleware takes one or more models to try, in order, when the main one raises. With hosted models, the fallback is usually a different provider, so one outage does not take the shop's support desk down.
How each recovered from failure
- Errors from inside the model are not caught by the agent. The first run crashed with the provider error.
- Retry tries the same model.
FlakyModelrecovered on its third attempt. - Retry gives up cleanly. Against
DownModel, all three attempts failed and the run ended with a message, not a crash, because ofon_failure="continue". - Fallback switches models. When
DownModelraised, the request went toShopModel, which answered.
ModelRetryMiddleware vs ModelFallbackMiddleware
| ModelRetryMiddleware | ModelFallbackMiddleware | |
|---|---|---|
| What it does | Calls the same model again | Calls a different model, in order |
| Handles | A brief timeout or rate limit | A model or provider that is down |
| Attempts | max_retries + 1 on one model | Each model once, in the order given |
| Waiting | Backs off between attempts | Moves on to the next at once |
Where retries and fallback fit
- Riding out a brief provider timeout or rate limit with a retry.
- Keeping the support desk up during one provider's outage by falling back to another.
- Pairing both: retry the main model, then fall back if it stays down.
initial_delay low in tests, and keep the backoff in production so a struggling provider is not hammered.Related
- Previous: Call limits with ModelCallLimitMiddleware
- Next: Summarization with SummarizationMiddleware
- Reference: LangChain agent middleware
- Set
on_failure="error"on the retry and run it againstDownModel. - Give
ModelFallbackMiddlewaretwo models, aDownModelfirst, and check which one answers. - Put retry and fallback in one list and work out which runs first before you try it.
This is what real progress feels like.