Pydantic AIPydantic AI 2.51 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
29 small wins to finish your pathNext lesson →

FallbackModel: when a provider is down

FallbackModel is a model that wraps a list of models and tries them in order, moving to the next when one fails with an API error. To the agent it looks like a single model.

Last updated: 28 Sep, 2026 · Pydantic AI 2.51

Providers have outages and rate limits. A fallback keeps the desk answering: if the first model is down, the request goes to the next, with no change to the agent around it.

Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
shop_model.py
import re

from pydantic_ai import ModelResponse, TextPart, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel


def sort_ticket(text):
    text = text.lower()
    if "charged" in text or "refund" in text:
        return "billing", 4
    if "parcel" in text or "arrived" in text:
        return "shipping", 3
    return "other", 1


def shop_reply(messages, info: AgentInfo) -> ModelResponse:
    prompts = [p.content for m in messages for p in m.parts if p.part_kind == "user-prompt"]
    ticket = prompts[-1]
    last = messages[-1].parts[-1]
    order = re.search(r"A-\d{4}", ticket)

    # 1. The ticket names an order and the agent has a tool: ask for it.
    if order and info.function_tools and last.part_kind == "user-prompt":
        tool = info.function_tools[0].name
        return ModelResponse(parts=[ToolCallPart(tool, {"order_id": order.group()})])

    # 2. A tool answered: write the reply from what it said.
    if last.part_kind == "tool-return" and info.allow_text_output:
        return ModelResponse(parts=[TextPart(f"Order {order.group()}: {last.content}.")])

    # 3. The agent wants a typed answer: fill in its output tool.
    category, priority = sort_ticket(ticket)
    if info.output_tools:
        args = {"category": category, "priority": priority}
        return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, args)])

    # 4. Otherwise, plain text.
    return ModelResponse(parts=[TextPart(f"Sorted as {category}.")])


shop_model = FunctionModel(shop_reply, model_name="shop")

A model that behaves like an outage

overloaded raises the same ModelHTTPError Pydantic AI raises for a real HTTP 503, so it stands in for a provider that is down.

python
from pydantic_ai import Agent
from pydantic_ai.exceptions import ModelHTTPError
from pydantic_ai.models.fallback import FallbackModel
from pydantic_ai.models.function import FunctionModel

from shop_model import shop_model


def overloaded(messages, info):
    raise ModelHTTPError(status_code=503, model_name="primary", body="overloaded")


primary = FunctionModel(overloaded, model_name="primary")

Falling through to the next model

Wrap the failing model and a working one in a FallbackModel and run.

Example
agent = Agent(FallbackModel(primary, shop_model))
result = agent.run_sync("My parcel never arrived")
print(result.output)
print(result.response.model_name)

FallbackModel(first, second, ...) is itself a model. primary failed, so the same request went to shop_model, and result.response.model_name says which model answered. In production the list would be real models from different providers, such as "groq:openai/gpt-oss-120b" then "google:gemini-2.5-flash".

When every model fails

If none of the models answers, the errors are collected together rather than lost.

Example
def backup_down(messages, info):
    raise ModelHTTPError(status_code=500, model_name="backup", body="internal error")


agent = Agent(FallbackModel(primary, FunctionModel(backup_down)))
try:
    agent.run_sync("My parcel never arrived")
except FallbackExceptionGroup as group:
    print(group)
    for error in group.exceptions:
        print("  ", error)

A FallbackExceptionGroup holds each model's error, so you can see why every attempt failed, not only the last.

What falls back and what does not

The failureWhat happens
A provider API error (ModelAPIError)The next model in the list is tried
A validation error on the outputA retry with the same model
ModelRetry from a toolA retry with the same model

By default only ModelAPIError, the errors from calling a provider, moves to the next model. A model that answers badly is not down, so it fixes its own mistake. fallback_on= takes other exception types if you need them.

When you reach for a fallback

  • A user-facing agent that must keep answering through a provider outage.
  • A cheaper or faster model first, a stronger one as the backup.
  • Spreading load across providers so one rate limit does not stop everything.
Watch out. A fallback only catches provider errors, not bad answers. A model that returns wrong or unvalidated output is handled by retries and validators, not by moving to the next model, so do not rely on a fallback to improve answer quality.
Try it yourself
  • Make overloaded raise ValueError instead and run the first agent.
  • Put shop_model first in the list.
  • Wrap a FunctionModel that raises ModelHTTPError(status_code=429, ...), a rate limit.

Every expert started right here.