FallbackModel: when a provider is down
FallbackModel is a model that wraps a list of models and tries them in order, moving to the next when one fails with an API error. To the agent it looks like a single model.
Last updated: 28 Sep, 2026 · Pydantic AI 2.51
Providers have outages and rate limits. A fallback keeps the desk answering: if the first model is down, the request goes to the next, with no change to the agent around it.
View the code here
import re
from pydantic_ai import ModelResponse, TextPart, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel
def sort_ticket(text):
text = text.lower()
if "charged" in text or "refund" in text:
return "billing", 4
if "parcel" in text or "arrived" in text:
return "shipping", 3
return "other", 1
def shop_reply(messages, info: AgentInfo) -> ModelResponse:
prompts = [p.content for m in messages for p in m.parts if p.part_kind == "user-prompt"]
ticket = prompts[-1]
last = messages[-1].parts[-1]
order = re.search(r"A-\d{4}", ticket)
# 1. The ticket names an order and the agent has a tool: ask for it.
if order and info.function_tools and last.part_kind == "user-prompt":
tool = info.function_tools[0].name
return ModelResponse(parts=[ToolCallPart(tool, {"order_id": order.group()})])
# 2. A tool answered: write the reply from what it said.
if last.part_kind == "tool-return" and info.allow_text_output:
return ModelResponse(parts=[TextPart(f"Order {order.group()}: {last.content}.")])
# 3. The agent wants a typed answer: fill in its output tool.
category, priority = sort_ticket(ticket)
if info.output_tools:
args = {"category": category, "priority": priority}
return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, args)])
# 4. Otherwise, plain text.
return ModelResponse(parts=[TextPart(f"Sorted as {category}.")])
shop_model = FunctionModel(shop_reply, model_name="shop")
A model that behaves like an outage
overloaded raises the same ModelHTTPError Pydantic AI raises for a real HTTP 503, so it stands in for a provider that is down.
from pydantic_ai import Agent
from pydantic_ai.exceptions import ModelHTTPError
from pydantic_ai.models.fallback import FallbackModel
from pydantic_ai.models.function import FunctionModel
from shop_model import shop_model
def overloaded(messages, info):
raise ModelHTTPError(status_code=503, model_name="primary", body="overloaded")
primary = FunctionModel(overloaded, model_name="primary")Falling through to the next model
Wrap the failing model and a working one in a FallbackModel and run.
agent = Agent(FallbackModel(primary, shop_model))
result = agent.run_sync("My parcel never arrived")
print(result.output)
print(result.response.model_name)Sorted as shipping. shop
FallbackModel(first, second, ...) is itself a model. primary failed, so the same request went to shop_model, and result.response.model_name says which model answered. In production the list would be real models from different providers, such as "groq:openai/gpt-oss-120b" then "google:gemini-2.5-flash".
When every model fails
If none of the models answers, the errors are collected together rather than lost.
def backup_down(messages, info):
raise ModelHTTPError(status_code=500, model_name="backup", body="internal error")
agent = Agent(FallbackModel(primary, FunctionModel(backup_down)))
try:
agent.run_sync("My parcel never arrived")
except FallbackExceptionGroup as group:
print(group)
for error in group.exceptions:
print(" ", error)All models from FallbackModel failed (2 sub-exceptions) status_code: 503, model_name: primary, body: overloaded status_code: 500, model_name: backup, body: internal error
A FallbackExceptionGroup holds each model's error, so you can see why every attempt failed, not only the last.
What falls back and what does not
| The failure | What happens |
|---|---|
A provider API error (ModelAPIError) | The next model in the list is tried |
| A validation error on the output | A retry with the same model |
ModelRetry from a tool | A retry with the same model |
By default only ModelAPIError, the errors from calling a provider, moves to the next model. A model that answers badly is not down, so it fixes its own mistake. fallback_on= takes other exception types if you need them.
When you reach for a fallback
- A user-facing agent that must keep answering through a provider outage.
- A cheaper or faster model first, a stronger one as the backup.
- Spreading load across providers so one rate limit does not stop everything.
Related
- Previous: Tool approval: approve, deny or edit a refund
- Next: Multi-agent delegation: agents that call agents
- Reference: FallbackModel
- Make
overloadedraiseValueErrorinstead and run the first agent. - Put
shop_modelfirst in the list. - Wrap a
FunctionModelthat raisesModelHTTPError(status_code=429, ...), a rate limit.
Every expert started right here.