LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Retries and a fallback model

ModelRetryMiddleware is a middleware that calls the same model again after it fails, and ModelFallbackMiddleware moves the request on to another model when the first keeps failing.

Last updated: 27 Sep, 2026 · LangChain 1.4

Automatic fallbacks in an LLM gateway · from the Complete Agentic AI Course In 10 Hours · 649:14 to 651:59

When a provider goes down

A hosted model is a service, and services go down. The video starts from a real outage: OpenAI was down for four hours in November 2023, and apps that had GPT-4 hardcoded went dark. Behind a gateway, when one provider fails the request falls back to another, which the video calls a must for any production app. A fallback is that list of other models to try, in order, when the first one fails.

The video runs this with LiteLLM, a gateway that sits between your app and every provider; its completion function takes a fallbacks list. The first run uses gemini/gemini-1.5-flash as the primary with a Google key whose account had been suspended: it fails with a 403 PERMISSION_DENIED error, and the answer comes from gpt-4o-mini, the first fallback. The second run, below, uses a model name that does not exist, so the primary is certain to fail, with the same two backups. It needs an OpenAI key and pip install litellm:

Example
from litellm import completion

# Force the primary to fail by using a fake model name
# Then watch the fallback chain rescue the call
response = completion(
    model="openai/fake-nonexistent-model-9999",     # 👈 will fail intentionally
    messages=[{"role": "user", "content": "What is an LLM Gateway?"}],
    fallbacks=[
        "gpt-4o-mini",                              # 1st backup: real OpenAI model
        "groq/llama-3.3-70b-versatile"              # 2nd backup: Groq
    ]
)

print("✅ App still got a response, even though the primary failed!")
print(f"\n🤖 Model that actually answered: {response.model}")
print(f"\n📝 Response: {response.choices[0].message.content[:200]}...")

Shown as it ran in the video, not run here. It needs pip install litellm and an OpenAI key, and its second backup model has since been retired on Groq. The fallback middleware later in this lesson runs on the keys you have.

LiteLLM raised an error for the fake model, yet the app still got its answer from gpt-4o-mini, the first backup, and the calling code never had to handle the failure. Inside a LangChain agent the same safety net comes from middleware, which retries or switches models within the agent's loop.

Now the shop. A real provider cannot be made to fail on cue, so two stand-ins fail the way a provider does. FlakyModel fails twice and then works; DownModel never works.

ModelRetryMiddleware and ModelFallbackMiddleware

python
from langchain.agents.middleware import ModelFallbackMiddleware, ModelRetryMiddleware

retry = ModelRetryMiddleware(max_retries=2, initial_delay=0)   # 3 attempts in all
backup = ModelFallbackMiddleware(other_model)                  # try other_model if the main raises

A flaky model and a dead one

The agent here runs on ShopModel, the stand-in chat model built in Several tool calls at once. Start the file with it.

python
import re

from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult


class ShopModel(BaseChatModel):
    tools: list = []

    @property
    def _llm_type(self):
        return "shop"

    def bind_tools(self, tools, **kwargs):
        return self.model_copy(update={"tools": tools})   # a copy holding the tools

    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        message = self.decide(messages)                   # the reply comes from decide
        return ChatResult(generations=[ChatGeneration(message=message)])

    def decide(self, messages):
        results = []                                # the tool results at the end
        for m in reversed(messages):
            if not isinstance(m, ToolMessage):
                break
            results.insert(0, m.text)
        if results:                                 # results are back: answer with them
            return AIMessage(" ".join(results))
        text = messages[-1].text
        orders = re.findall(r"\b[A-Z]\d+\b", text)
        tool = "refund_order" if "refund" in text.lower() else "lookup_order"
        if orders and tool in [t.name for t in self.tools]:   # one call per order id
            calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
                     for o in orders]
            return AIMessage("", tool_calls=calls)
        if orders:                                  # that tool is not bound
            return AIMessage(f"I have no way to look up {orders[0]} yet.")
        return AIMessage("Hello. Which order is this about?")

Both stand-ins extend ShopModel and override _generate, the method that produces a reply, to raise instead.

python
FAILURES = [ConnectionError("provider unavailable")] * 2   # two failures, then success

class FlakyModel(ShopModel):
    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        if FAILURES:
            raise FAILURES.pop()
        return super()._generate(messages)

class DownModel(ShopModel):
    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        raise ConnectionError("provider unavailable")     # never works

The tool and the agent

This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Add it below the models.

python
from langchain.tools import tool

ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}


@tool
def lookup_order(order_id: str) -> str:
    """Look up an order's shipping status by its id, such as A17."""
    status = ORDERS.get(order_id)
    return f"{order_id} {status}." if status else f"{order_id} is not an order we have."

Without any handling

Nothing in the agent retries a failed model call on its own, so the first failure ends the run.

Example
from langchain.agents import create_agent

agent = create_agent(FlakyModel(), tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})

LangChain's hosted chat models retry a failed request a couple of times by default, inside their HTTP client, for network errors, rate limits and server errors, though not for a wrong key. FlakyModel raises straight from _generate, with no HTTP client underneath, so nothing retried it.

Trying again

ModelRetryMiddleware retries the same model. FlakyModel works on its third attempt.

Example
from langchain.agents.middleware import ModelFallbackMiddleware, ModelRetryMiddleware

retry = ModelRetryMiddleware(max_retries=2, initial_delay=0)
agent = create_agent(FlakyModel(), tools=[lookup_order], middleware=[retry])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)

Two failures, a third attempt, a normal answer. max_retries=2 means three attempts in all. The defaults wait one second before the first retry and double the wait each time, with some randomness, up to a minute; initial_delay=0 keeps this example quick.

When the model never recovers, the retry gives up.

Example
retry = ModelRetryMiddleware(max_retries=2, initial_delay=0)
agent = create_agent(DownModel(), tools=[lookup_order], middleware=[retry])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)

When every attempt fails, the default on_failure="continue" ends the run with an AI message describing the failure instead of raising.

Another model

ModelFallbackMiddleware takes one or more backup models, tried in order. When the main model raises, the request moves to the next one. With hosted models, the fallback is usually a different provider, so one outage does not take the shop's support desk down.

Example
backup = ModelFallbackMiddleware(ShopModel())
agent = create_agent(DownModel(), tools=[lookup_order], middleware=[backup])
print(agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"][-1].text)

How each recovered from failure

  • Nothing in the agent retries a model error. FlakyModel has no HTTP client underneath, so the first run crashed with the provider error.
  • Retry tries the same model. FlakyModel recovered on its third attempt.
  • Retry gives up cleanly. Against DownModel, all three attempts failed and the run ended with a message, not a crash, because of on_failure="continue".
  • Fallback switches models. When DownModel raised, the request went to ShopModel, which answered.

ModelRetryMiddleware vs ModelFallbackMiddleware

ModelRetryMiddlewareModelFallbackMiddleware
What it doesCalls the same model againCalls a different model, in order
HandlesA brief timeout or rate limitA model or provider that is down
Attemptsmax_retries + 1 on one modelEach model once, in the order given
WaitingBacks off between attemptsMoves on to the next at once

Where retries and fallback fit

  • Riding out a brief provider timeout or rate limit with a retry.
  • Keeping the support desk up during one provider's outage by falling back to another.
  • Pairing both: retry the main model, then fall back if it stays down.
Watch out. A retry with the default one-second backoff can add delay before it gives up. Set initial_delay low in tests, and keep the backoff in production so a struggling provider is not hammered.
Try it yourself
  • Set on_failure="error" on the retry and run it against DownModel.
  • Give ModelFallbackMiddleware two models, a DownModel first, and check which one answers.
  • Put retry and fallback in one list and work out which runs first before you try it.

This is what real progress feels like.