Support desk project: the finished agent
The support desk project is the finished agent that looks up orders, refunds small amounts, asks a person before a big refund, refuses a bad one, and ships with tests and an eval, using only what the course taught.
Last updated: 28 Sep, 2026 · Pydantic AI 2.51
Everything from Integrations: from stand-in to real backends back is now assembled into one build. The overview promised a desk that looks up orders and asks before big refunds; this is that desk, running end to end.
The agent, its deps and its tools
Desk holds the orders and a refund limit. Both tools raise ModelRetry for a mistake the model should fix, and refund_order asks for approval over the limit:
from dataclasses import dataclass
from pydantic_ai import Agent, ApprovalRequired, DeferredToolRequests, ModelRetry, RunContext
@dataclass
class Desk:
orders: dict[str, dict]
refund_limit: float = 50.0
agent = Agent(
"openai:gpt-4.1",
defer_model_check=True,
deps_type=Desk,
output_type=[str, DeferredToolRequests],
instructions="You answer support tickets for an online shop. Be short and kind.",
)
@agent.tool
def lookup_order(ctx: RunContext[Desk], order_id: str) -> str:
"""Look up an order.
Args:
order_id: The order id, like A-1001.
"""
order = ctx.deps.orders.get(order_id)
if order is None:
raise ModelRetry(f"There is no order {order_id}.")
return f"Order {order_id} is {order['status']}"
@agent.tool
def refund_order(ctx: RunContext[Desk], order_id: str, amount: float) -> str:
"""Refund part or all of an order.
Args:
order_id: The order id, like A-1001.
amount: The amount to refund, in euros.
"""
order = ctx.deps.orders.get(order_id)
if order is None:
raise ModelRetry(f"There is no order {order_id}.")
if amount > order["total"]:
raise ModelRetry(f"Order {order_id} only cost {order['total']:.2f} euros.")
if amount > ctx.deps.refund_limit and not ctx.tool_call_approved:
raise ApprovalRequired(metadata={"reason": f"{amount:.2f} euros is over the limit"})
order["refunded"] = amount
return f"Refunded {amount:.2f} euros on order {order_id}"
Deskholds the orders and the refund limit, from Dependencies: giving tools your data.output_type=[str, DeferredToolRequests]: a reply, or refunds waiting for a person, from Output functions and several output types and Tool approval: approve, deny or edit a refund.- Both tools raise
ModelRetrywith a message for the model, from Tool errors: ModelRetry and crashes. refund_orderrefuses more than the order cost, and asks for approval over the limit.
The scripted model for the project
The desk needs a stand-in that can look up orders and ask for refunds, so it reads amounts from the ticket too:
import re
from pydantic_ai import ModelResponse, TextPart, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel
def desk_reply(messages, info: AgentInfo) -> ModelResponse:
prompts = [p.content for m in messages for p in m.parts if p.part_kind == "user-prompt"]
ticket = prompts[-1].lower()
last = messages[-1].parts[-1]
order = re.search(r"a-\d{4}", ticket)
tools = {t.name for t in info.function_tools}
if last.part_kind == "user-prompt" and order:
order_id = order.group().upper()
if "refund" in ticket and "refund_order" in tools:
amount = float(re.search(r"(\d+(?:\.\d+)?) euros", ticket).group(1))
return ModelResponse(parts=[ToolCallPart("refund_order", {"order_id": order_id, "amount": amount})])
return ModelResponse(parts=[ToolCallPart("lookup_order", {"order_id": order_id})])
if last.part_kind == "retry-prompt":
return ModelResponse(parts=[TextPart(f"Sorry, I could not do that. {last.content}")])
if last.part_kind == "tool-return":
return ModelResponse(parts=[TextPart(last.content.rstrip(".") + ".")])
return ModelResponse(parts=[TextPart("Could you send me your order number, like A-1001?")])
desk_model = FunctionModel(desk_reply, model_name="desk")
View the code here
from pydantic_ai import DeferredToolRequests, DeferredToolResults, ToolDenied
from desk import Desk, agent
from desk_model import desk_model
desk = Desk(orders={
"A-1001": {"status": "shipped", "total": 40.0},
"A-1002": {"status": "waiting for stock", "total": 120.0},
})
def manager_approves(call):
"""A person decides. Here: refunds up to 100 euros are fine."""
return call.args["amount"] <= 100
def handle(ticket):
result = agent.run_sync(ticket, deps=desk, model=desk_model)
if isinstance(result.output, DeferredToolRequests):
decisions = DeferredToolResults()
for call in result.output.approvals:
ok = manager_approves(call)
print(f" approval needed: {call.args} -> {'yes' if ok else 'no'}")
decisions.approvals[call.tool_call_id] = True if ok else ToolDenied("A manager said no.")
result = agent.run_sync(message_history=result.all_messages(), deferred_tool_results=decisions,
deps=desk, model=desk_model)
return result.output
tickets = [
"Where is my order A-1001?",
"Please refund 30 euros on A-1001, it arrived broken",
"Refund 90 euros on A-1002, I cancelled",
"Refund 110 euros on A-1002 please",
"Where is A-9999?",
]
for ticket in tickets:
print(ticket)
print(" reply:", handle(ticket))
print(desk.orders)Running five tickets end to end
handle runs a ticket. If the run pauses for approval, it asks manager_approves, a stand-in for a person in your admin page, then runs again with the decision:
PYDANTIC_AI_NO_BANNER=1 python app.pyWhere is my order A-1001?
reply: Order A-1001 is shipped.
Please refund 30 euros on A-1001, it arrived broken
reply: Refunded 30.00 euros on order A-1001.
Refund 90 euros on A-1002, I cancelled
approval needed: {'order_id': 'A-1002', 'amount': 90.0} -> yes
reply: Refunded 90.00 euros on order A-1002.
Refund 110 euros on A-1002 please
approval needed: {'order_id': 'A-1002', 'amount': 110.0} -> no
reply: A manager said no.
Where is A-9999?
reply: Sorry, I could not do that. There is no order A-9999.
{'A-1001': {'status': 'shipped', 'total': 40.0, 'refunded': 30.0}, 'A-1002': {'status': 'waiting for stock', 'total': 120.0, 'refunded': 90.0}}What each ticket did
- A status question: one tool call, one reply.
- 30 euros: under the limit, refunded without asking.
- 90 euros on A-1002: over 50, so the run paused, a manager said yes, and the refund went through.
- 110 euros: the manager said no, the tool never ran, and the customer got the refusal.
- An unknown order:
ModelRetry, and the model passed the reason on.
The last line shows the orders: only the approved and small refunds were recorded. The denied 110 euro refund is the failure mode the overview promised, working.
Pick one to watch it run, step by step.
A second failure: a tool that keeps failing
A tool that raises ModelRetry gives the model a chance to fix its call. If the model keeps making the same failing call, the run does not loop forever: it stops with UnexpectedModelBehavior once the tool's retry limit is reached:
from pydantic_ai import UnexpectedModelBehavior, ModelResponse, ToolCallPart
from pydantic_ai.models.function import FunctionModel
from desk import Desk, agent
desk = Desk(orders={})
def keeps_trying(messages, info):
# the model keeps looking up an order that is not there
return ModelResponse(parts=[ToolCallPart("lookup_order", {"order_id": "A-9999"})])
try:
agent.run_sync("Where is A-9999?", deps=desk, model=FunctionModel(keeps_trying))
except UnexpectedModelBehavior as exc:
print("run stopped:", exc)run stopped: Tool 'lookup_order' exceeded max retries count of 1. Consider raising the retry limit, or see the docs on tool retries: https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-retries
The message names the tool, the limit it hit, and where to read more. In your app, catch this and show the customer a fallback, rather than letting the run raise.
Testing the refund rules
The tests pin the rules that matter: small refunds go through, big ones wait, a refund above the order cost is refused, an unknown order is handled:
from pydantic_ai import DeferredToolRequests, DeferredToolResults, models
from desk import Desk, agent
from desk_model import desk_model
models.ALLOW_MODEL_REQUESTS = False
def make_desk():
return Desk(orders={"A-1001": {"status": "shipped", "total": 40.0},
"A-1002": {"status": "waiting for stock", "total": 120.0}})
def test_small_refund_needs_no_approval():
desk = make_desk()
with agent.override(model=desk_model):
result = agent.run_sync("Refund 20 euros on A-1001", deps=desk)
assert result.output == "Refunded 20.00 euros on order A-1001."
assert desk.orders["A-1001"]["refunded"] == 20.0
def test_large_refund_waits_for_a_person():
desk = make_desk()
with agent.override(model=desk_model):
result = agent.run_sync("Refund 90 euros on A-1002", deps=desk)
assert isinstance(result.output, DeferredToolRequests)
assert "refunded" not in desk.orders["A-1002"]
call = result.output.approvals[0]
with agent.override(model=desk_model):
final = agent.run_sync(message_history=result.all_messages(), deps=desk,
deferred_tool_results=DeferredToolResults(approvals={call.tool_call_id: True}))
assert desk.orders["A-1002"]["refunded"] == 90.0
assert "Refunded 90.00" in final.output
def test_refund_above_order_total_is_refused():
desk = make_desk()
with agent.override(model=desk_model):
result = agent.run_sync("Refund 45 euros on A-1001", deps=desk)
assert "refunded" not in desk.orders["A-1001"]
assert result.output == "Sorry, I could not do that. Order A-1001 only cost 40.00 euros."
def test_unknown_order():
with agent.override(model=desk_model):
result = agent.run_sync("Where is A-7777?", deps=make_desk())
assert result.output == "Sorry, I could not do that. There is no order A-7777."pytest -q.... [100%] 4 passed in 0.26s
Scoring the desk with an eval
The eval shows how the desk handles tickets it was not written for. The stand-in has no rule for 'Can I get 25 euros back', so it looked the order up instead, which the eval catches:
from pydantic_evals import Case, Dataset
from pydantic_evals.evaluators import Contains, Evaluator, EvaluatorContext
from desk import Desk, agent
from desk_model import desk_model
class ShortReply(Evaluator):
def evaluate(self, ctx: EvaluatorContext) -> bool:
return len(ctx.output) <= 80
async def answer(ticket: str) -> str:
desk = Desk(orders={"A-1001": {"status": "shipped", "total": 40.0}})
result = await agent.run(ticket, deps=desk, model=desk_model)
return str(result.output)
dataset = Dataset(
name="desk",
cases=[
Case(name="status", inputs="Where is A-1001?", evaluators=[Contains("shipped")]),
Case(name="small refund", inputs="Refund 10 euros on A-1001", evaluators=[Contains("Refunded 10.00")]),
Case(name="no order id", inputs="Where is my parcel?", evaluators=[Contains("order number")]),
Case(name="lower case id", inputs="where is a-1001", evaluators=[Contains("shipped")]),
Case(name="unknown order", inputs="Where is A-7777?", evaluators=[Contains("no order A-7777")]),
Case(name="money back", inputs="Can I get 25 euros back on A-1001?", evaluators=[Contains("Refunded 25.00")]),
],
evaluators=[ShortReply()],
)
report = dataset.evaluate_sync(answer, progress=False)
for case in report.cases:
failed = [name for name, result in case.assertions.items() if not result.value]
print(f"{case.name:14} {case.output}")
if failed:
print(" failed:", failed)
print(f"passed: {report.averages().assertions:.1%}")PYDANTIC_AI_NO_BANNER=1 python eval_desk.pystatus Order A-1001 is shipped.
small refund Refunded 10.00 euros on order A-1001.
no order id Could you send me your order number, like A-1001?
lower case id Order A-1001 is shipped.
unknown order Sorry, I could not do that. There is no order A-7777.
money back Order A-1001 is shipped.
failed: ['Contains']
passed: 91.7%That failing case is one to keep in the dataset while you change the model or the prompt. To use a real model, follow Integrations: from stand-in to real backends: drop model=desk_model and set the provider key. The tests keep the stand-in through override.
What we left out
The course taught the desk's spine. These pages of the docs it named but did not cover, each worth reading when you need it:
| Topic | Why it was skipped | Docs |
|---|---|---|
| pydantic-graph | State machines for control flow beyond one agent loop | graph |
| Durable execution | Runs that survive a crash, with Temporal, DBOS, Prefect or Restate | durable_execution |
| Long-term memory | A store with semantic search; the core has none, the Harness Memory capability is the closest | harness/memory |
| Direct model requests | Calling a model once without an agent | direct |
| Native and built-in tools | Provider-run tools such as web search | native-tools |
| Thinking | Reading a model's reasoning tokens | thinking |
| Multimodal input | Images, audio and documents in a prompt | input |
| Realtime | Speech-to-speech voice runs | realtime |
| AG-UI and Vercel UI | Streaming an agent to a web front end | ag-ui |
Where you take the desk next
- Serve
handlefrom a web endpoint, storing paused runs by conversation id. - Add a real model behind a FallbackModel, and keep the tests on the stand-in.
- Add usage limits to every run and a test that checks the cap is hit.
manager_approves here is a stub. In a real app the decision comes from a person through your admin page, and the paused run must be stored between the two calls, not held in memory. Persist it, as the integrations lesson describes, or a restart loses the pending refund.Related
- Previous: Integrations: from stand-in to real backends
- See also: Tool approval: approve, deny or edit a refund, Evals: scoring the agent on many tickets
- Reference: Pydantic AI docs
- Serve
handlefrom a FastAPI endpoint and store paused runs by conversation id. - Add a
FallbackModelof two real models todesk.py. - Add
usage_limitsto every run inapp.pyand a test that checks it.
Slow is fine. Stopping is the only problem.