Testing agents with pytest and TestModel
agent.override is a context manager that swaps the agent's model for a stand-in during a test, so the test runs the real agent code without ever calling a hosted model.
Last updated: 28 Sep, 2026 · Pydantic AI 2.51
You have run the desk on a scripted stand-in all along. A test needs the same swap, plus a guard so a forgotten override never spends money. ALLOW_MODEL_REQUESTS = False, from Real models: providers, keys and model names, is that guard.
The code under test
support.py names a real model and has an answer function the rest of the app calls. defer_model_check keeps it importable with no key:
from pydantic_ai import Agent
ORDERS = {"A-1001": "shipped on 12 March"}
agent = Agent("openai:gpt-4.1", defer_model_check=True, instructions="You answer support tickets.")
@agent.tool_plain
def lookup_order(order_id: str) -> str:
"""Look up the status of an order."""
return ORDERS.get(order_id, "not found")
def answer(ticket: str) -> str:
return agent.run_sync(ticket).output
Blocking real requests in tests
With ALLOW_MODEL_REQUESTS set to False, a real model refuses to send anything, even with a key present. Put it at the top of every test file so a test that forgot its override fails loudly:
models.ALLOW_MODEL_REQUESTS = False
os.environ["OPENAI_API_KEY"] = "sk-test"
answer("Where is A-1001?")Traceback (most recent call last):
File "main.py", line 4, in <module>
answer("Where is A-1001?")
RuntimeError: Model requests are not allowed, since ALLOW_MODEL_REQUESTS is FalseOverriding the model for one block
agent.override(model=...) replaces the model for every run inside the with, even deep inside answer, which your test cannot pass a model to. capture_run_messages() collects the run's messages so a test can check which tools were called:
with agent.override(model=TestModel()):
with capture_run_messages() as messages:
print(answer("Where is A-1001?"))
for message in messages:
print(message.kind, [part.part_kind for part in message.parts]){"lookup_order":"not found"}
request ['user-prompt']
response ['tool-call']
request ['tool-return']
response ['text']What the captured messages show
- TestModel made up the argument. The reply is
{"lookup_order":"not found"}because the invented order id matched nothing, which is fine for a wiring check. - The four messages trace the run: the prompt, the tool call, the tool return, then the text answer.
- TestModel proves the wiring, not your tool's behaviour: its arguments are invented, so it cannot check what the tool does with a real value.
View the code here
import re
from pydantic_ai import ModelResponse, TextPart, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel
def sort_ticket(text):
text = text.lower()
if "charged" in text or "refund" in text:
return "billing", 4
if "parcel" in text or "arrived" in text:
return "shipping", 3
return "other", 1
def shop_reply(messages, info: AgentInfo) -> ModelResponse:
prompts = [p.content for m in messages for p in m.parts if p.part_kind == "user-prompt"]
ticket = prompts[-1]
last = messages[-1].parts[-1]
order = re.search(r"A-\d{4}", ticket)
if order and info.function_tools and last.part_kind == "user-prompt":
tool = info.function_tools[0].name
return ModelResponse(parts=[ToolCallPart(tool, {"order_id": order.group()})])
if last.part_kind == "tool-return" and info.allow_text_output:
return ModelResponse(parts=[TextPart(f"Order {order.group()}: {last.content}.")])
category, priority = sort_ticket(ticket)
if info.output_tools:
args = {"category": category, "priority": priority}
return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, args)])
return ModelResponse(parts=[TextPart(f"Sorted as {category}.")])
shop_model = FunctionModel(shop_reply, model_name="shop")
Three tests with pytest
TestModel checks the wiring; the scripted shop_model gives real order ids to the tool, so it checks your tool's behaviour. Both go through override:
from pydantic_ai import capture_run_messages, models
from pydantic_ai.models.test import TestModel
from support import agent, answer
from shop_model import shop_model
models.ALLOW_MODEL_REQUESTS = False
def test_answer_runs_with_the_tool():
with agent.override(model=TestModel()):
with capture_run_messages() as messages:
answer("Where is A-1001?")
call = messages[1].parts[0]
assert call.tool_name == "lookup_order"
def test_known_order_is_found():
with agent.override(model=shop_model):
assert answer("Where is A-1001?") == "Order A-1001: shipped on 12 March."
def test_unknown_order():
with agent.override(model=shop_model):
assert "not found" in answer("Where is A-4040?")pytest -q... [100%] 3 passed in 0.32s
Three tests, no network. Pydantic AI also hides its banner when it runs under pytest.
TestModel vs a scripted FunctionModel
| Stand-in | Reads the prompt | Good for |
|---|---|---|
TestModel | No; arguments are invented | Checking every tool is offered and the run completes |
FunctionModel | Yes; you script it from the messages | Checking your tool's behaviour on real values |
Where you use override and TestModel
- A unit test that a tool is called with the right name.
- A test that a known order returns the expected reply.
- A CI job with
ALLOW_MODEL_REQUESTS = Falseso no test can bill you.
override only redirects runs of that agent inside the with block. A run started outside the block, or on a different agent, still uses the real model. Keep the run you are testing inside the block, and set ALLOW_MODEL_REQUESTS = False so anything that escapes fails instead of paying.Related
- Previous: Harness Coder: a ready-made coding agent
- Next: Evals: scoring the agent on many tickets
- See also: FunctionModel: write a stand-in model
- Reference: Testing
- Add a test that passes
TestModel(custom_output_text="Hello")and checksanswerreturns it. - Remove the
overridefrom one test and run pytest again. - Write a pytest fixture that wraps a test in
agent.override(model=shop_model).
This is what real progress feels like.