LangChainLangChain 1.4 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Retries and a fallback model

ModelRetryMiddleware is a middleware that calls the same model again after it fails, and ModelFallbackMiddleware moves the request on to another model when the first keeps failing.

Last updated: 27 Sep, 2026 · LangChain 1.4

Hosted models fail sometimes: a timeout, a rate limit, an outage. Two stand-ins fail the way a provider does. FlakyModel fails twice and then works; DownModel never works.

ModelRetryMiddleware and ModelFallbackMiddleware

python
from langchain.agents.middleware import ModelFallbackMiddleware, ModelRetryMiddleware

retry = ModelRetryMiddleware(max_retries=2, initial_delay=0)   # 3 attempts in all
backup = ModelFallbackMiddleware(other_model)                  # try other_model if the main raises

A flaky model and a dead one

Both stand-ins extend ShopModel and override _generate, the method that produces a reply, to raise instead.

python
from shop_model import ShopModel

FAILURES = [ConnectionError("provider unavailable")] * 2   # two failures, then success

class FlakyModel(ShopModel):
    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        if FAILURES:
            raise FAILURES.pop()
        return super()._generate(messages)

class DownModel(ShopModel):
    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        raise ConnectionError("provider unavailable")     # never works

The imports

Bring in the pieces for a plain agent on FlakyModel, no middleware yet.

python
from langchain.agents import create_agent
from langchain.agents.middleware import ModelFallbackMiddleware, ModelRetryMiddleware
from flaky_model import DownModel, FlakyModel
from shop_model import ShopModel
from tools import lookup_order

Without any handling

The first failure comes from inside the model, so nothing catches it.

Example
agent = create_agent(FlakyModel(), tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})

LangChain's hosted chat models retry a failed request a couple of times by default, for network errors, rate limits and server errors, though not for a wrong key. This error came from inside the model, so nothing retried it.

Trying again

ModelRetryMiddleware retries the same model. FlakyModel works on its third attempt.

Example
retry = ModelRetryMiddleware(max_retries=2, initial_delay=0)
agent = create_agent(FlakyModel(), tools=[lookup_order], middleware=[retry])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)

Two failures, a third attempt, a normal answer. max_retries=2 means three attempts in all. The defaults wait one second before the first retry and double the wait each time, with some randomness, up to a minute; initial_delay=0 keeps this example quick.

When the model never recovers, the retry gives up.

Example
retry = ModelRetryMiddleware(max_retries=2, initial_delay=0)
agent = create_agent(DownModel(), tools=[lookup_order], middleware=[retry])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)

When every attempt fails, the default on_failure="continue" ends the run with an AI message describing the failure instead of raising.

Another model

ModelFallbackMiddleware takes one or more backup models. When the main model raises, the request moves to the next one.

Example
backup = ModelFallbackMiddleware(ShopModel())
agent = create_agent(DownModel(), tools=[lookup_order], middleware=[backup])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)

ModelFallbackMiddleware takes one or more models to try, in order, when the main one raises. With hosted models, the fallback is usually a different provider, so one outage does not take the shop's support desk down.

How each recovered from failure

  • Errors from inside the model are not caught by the agent. The first run crashed with the provider error.
  • Retry tries the same model. FlakyModel recovered on its third attempt.
  • Retry gives up cleanly. Against DownModel, all three attempts failed and the run ended with a message, not a crash, because of on_failure="continue".
  • Fallback switches models. When DownModel raised, the request went to ShopModel, which answered.

ModelRetryMiddleware vs ModelFallbackMiddleware

ModelRetryMiddlewareModelFallbackMiddleware
What it doesCalls the same model againCalls a different model, in order
HandlesA brief timeout or rate limitA model or provider that is down
Attemptsmax_retries + 1 on one modelEach model once, in the order given
WaitingBacks off between attemptsMoves on to the next at once

Where retries and fallback fit

  • Riding out a brief provider timeout or rate limit with a retry.
  • Keeping the support desk up during one provider's outage by falling back to another.
  • Pairing both: retry the main model, then fall back if it stays down.
Watch out. A retry with the default one-second backoff can add delay before it gives up. Set initial_delay low in tests, and keep the backoff in production so a struggling provider is not hammered.
Try it yourself
  • Set on_failure="error" on the retry and run it against DownModel.
  • Give ModelFallbackMiddleware two models, a DownModel first, and check which one answers.
  • Put retry and fallback in one list and work out which runs first before you try it.

This is what real progress feels like.