Pydantic AIPydantic AI 2.51 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
29 small wins to finish your pathNext lesson →

Testing agents with pytest and TestModel

agent.override is a context manager that swaps the agent's model for a stand-in during a test, so the test runs the real agent code without ever calling a hosted model.

Last updated: 28 Sep, 2026 · Pydantic AI 2.51

You have run the desk on a scripted stand-in all along. A test needs the same swap, plus a guard so a forgotten override never spends money. ALLOW_MODEL_REQUESTS = False, from Real models: providers, keys and model names, is that guard.

The code under test

support.py names a real model and has an answer function the rest of the app calls. defer_model_check keeps it importable with no key:

python
from pydantic_ai import Agent

ORDERS = {"A-1001": "shipped on 12 March"}

agent = Agent("openai:gpt-4.1", defer_model_check=True, instructions="You answer support tickets.")


@agent.tool_plain
def lookup_order(order_id: str) -> str:
    """Look up the status of an order."""
    return ORDERS.get(order_id, "not found")


def answer(ticket: str) -> str:
    return agent.run_sync(ticket).output

Blocking real requests in tests

With ALLOW_MODEL_REQUESTS set to False, a real model refuses to send anything, even with a key present. Put it at the top of every test file so a test that forgot its override fails loudly:

Example
models.ALLOW_MODEL_REQUESTS = False
os.environ["OPENAI_API_KEY"] = "sk-test"

answer("Where is A-1001?")

Overriding the model for one block

agent.override(model=...) replaces the model for every run inside the with, even deep inside answer, which your test cannot pass a model to. capture_run_messages() collects the run's messages so a test can check which tools were called:

Example
with agent.override(model=TestModel()):
    with capture_run_messages() as messages:
        print(answer("Where is A-1001?"))

for message in messages:
    print(message.kind, [part.part_kind for part in message.parts])

What the captured messages show

  • TestModel made up the argument. The reply is {"lookup_order":"not found"} because the invented order id matched nothing, which is fine for a wiring check.
  • The four messages trace the run: the prompt, the tool call, the tool return, then the text answer.
  • TestModel proves the wiring, not your tool's behaviour: its arguments are invented, so it cannot check what the tool does with a real value.
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
shop_model.py
import re

from pydantic_ai import ModelResponse, TextPart, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel


def sort_ticket(text):
    text = text.lower()
    if "charged" in text or "refund" in text:
        return "billing", 4
    if "parcel" in text or "arrived" in text:
        return "shipping", 3
    return "other", 1


def shop_reply(messages, info: AgentInfo) -> ModelResponse:
    prompts = [p.content for m in messages for p in m.parts if p.part_kind == "user-prompt"]
    ticket = prompts[-1]
    last = messages[-1].parts[-1]
    order = re.search(r"A-\d{4}", ticket)

    if order and info.function_tools and last.part_kind == "user-prompt":
        tool = info.function_tools[0].name
        return ModelResponse(parts=[ToolCallPart(tool, {"order_id": order.group()})])

    if last.part_kind == "tool-return" and info.allow_text_output:
        return ModelResponse(parts=[TextPart(f"Order {order.group()}: {last.content}.")])

    category, priority = sort_ticket(ticket)
    if info.output_tools:
        args = {"category": category, "priority": priority}
        return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, args)])

    return ModelResponse(parts=[TextPart(f"Sorted as {category}.")])


shop_model = FunctionModel(shop_reply, model_name="shop")

Three tests with pytest

TestModel checks the wiring; the scripted shop_model gives real order ids to the tool, so it checks your tool's behaviour. Both go through override:

python
from pydantic_ai import capture_run_messages, models
from pydantic_ai.models.test import TestModel

from support import agent, answer
from shop_model import shop_model

models.ALLOW_MODEL_REQUESTS = False


def test_answer_runs_with_the_tool():
    with agent.override(model=TestModel()):
        with capture_run_messages() as messages:
            answer("Where is A-1001?")
    call = messages[1].parts[0]
    assert call.tool_name == "lookup_order"


def test_known_order_is_found():
    with agent.override(model=shop_model):
        assert answer("Where is A-1001?") == "Order A-1001: shipped on 12 March."


def test_unknown_order():
    with agent.override(model=shop_model):
        assert "not found" in answer("Where is A-4040?")
Example
pytest -q

Three tests, no network. Pydantic AI also hides its banner when it runs under pytest.

TestModel vs a scripted FunctionModel

Stand-inReads the promptGood for
TestModelNo; arguments are inventedChecking every tool is offered and the run completes
FunctionModelYes; you script it from the messagesChecking your tool's behaviour on real values

Where you use override and TestModel

  • A unit test that a tool is called with the right name.
  • A test that a known order returns the expected reply.
  • A CI job with ALLOW_MODEL_REQUESTS = False so no test can bill you.
Override covers only its block
override only redirects runs of that agent inside the with block. A run started outside the block, or on a different agent, still uses the real model. Keep the run you are testing inside the block, and set ALLOW_MODEL_REQUESTS = False so anything that escapes fails instead of paying.
Try it yourself
  • Add a test that passes TestModel(custom_output_text="Hello") and checks answer returns it.
  • Remove the override from one test and run pytest again.
  • Write a pytest fixture that wraps a test in agent.override(model=shop_model).

This is what real progress feels like.