Pydantic AIPydantic AI 2.51 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
29 small wins to finish your path

Support desk project: the finished agent

The support desk project is the finished agent that looks up orders, refunds small amounts, asks a person before a big refund, refuses a bad one, and ships with tests and an eval, using only what the course taught.

Last updated: 28 Sep, 2026 · Pydantic AI 2.51

Everything from Integrations: from stand-in to real backends back is now assembled into one build. The overview promised a desk that looks up orders and asks before big refunds; this is that desk, running end to end.

The agent, its deps and its tools

Desk holds the orders and a refund limit. Both tools raise ModelRetry for a mistake the model should fix, and refund_order asks for approval over the limit:

python
from dataclasses import dataclass

from pydantic_ai import Agent, ApprovalRequired, DeferredToolRequests, ModelRetry, RunContext


@dataclass
class Desk:
    orders: dict[str, dict]
    refund_limit: float = 50.0


agent = Agent(
    "openai:gpt-4.1",
    defer_model_check=True,
    deps_type=Desk,
    output_type=[str, DeferredToolRequests],
    instructions="You answer support tickets for an online shop. Be short and kind.",
)


@agent.tool
def lookup_order(ctx: RunContext[Desk], order_id: str) -> str:
    """Look up an order.

    Args:
        order_id: The order id, like A-1001.
    """
    order = ctx.deps.orders.get(order_id)
    if order is None:
        raise ModelRetry(f"There is no order {order_id}.")
    return f"Order {order_id} is {order['status']}"


@agent.tool
def refund_order(ctx: RunContext[Desk], order_id: str, amount: float) -> str:
    """Refund part or all of an order.

    Args:
        order_id: The order id, like A-1001.
        amount: The amount to refund, in euros.
    """
    order = ctx.deps.orders.get(order_id)
    if order is None:
        raise ModelRetry(f"There is no order {order_id}.")
    if amount > order["total"]:
        raise ModelRetry(f"Order {order_id} only cost {order['total']:.2f} euros.")
    if amount > ctx.deps.refund_limit and not ctx.tool_call_approved:
        raise ApprovalRequired(metadata={"reason": f"{amount:.2f} euros is over the limit"})
    order["refunded"] = amount
    return f"Refunded {amount:.2f} euros on order {order_id}"

The scripted model for the project

The desk needs a stand-in that can look up orders and ask for refunds, so it reads amounts from the ticket too:

python
import re

from pydantic_ai import ModelResponse, TextPart, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel


def desk_reply(messages, info: AgentInfo) -> ModelResponse:
    prompts = [p.content for m in messages for p in m.parts if p.part_kind == "user-prompt"]
    ticket = prompts[-1].lower()
    last = messages[-1].parts[-1]
    order = re.search(r"a-\d{4}", ticket)
    tools = {t.name for t in info.function_tools}

    if last.part_kind == "user-prompt" and order:
        order_id = order.group().upper()
        if "refund" in ticket and "refund_order" in tools:
            amount = float(re.search(r"(\d+(?:\.\d+)?) euros", ticket).group(1))
            return ModelResponse(parts=[ToolCallPart("refund_order", {"order_id": order_id, "amount": amount})])
        return ModelResponse(parts=[ToolCallPart("lookup_order", {"order_id": order_id})])

    if last.part_kind == "retry-prompt":
        return ModelResponse(parts=[TextPart(f"Sorry, I could not do that. {last.content}")])

    if last.part_kind == "tool-return":
        return ModelResponse(parts=[TextPart(last.content.rstrip(".") + ".")])

    return ModelResponse(parts=[TextPart("Could you send me your order number, like A-1001?")])


desk_model = FunctionModel(desk_reply, model_name="desk")
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
app.py
from pydantic_ai import DeferredToolRequests, DeferredToolResults, ToolDenied

from desk import Desk, agent
from desk_model import desk_model

desk = Desk(orders={
    "A-1001": {"status": "shipped", "total": 40.0},
    "A-1002": {"status": "waiting for stock", "total": 120.0},
})


def manager_approves(call):
    """A person decides. Here: refunds up to 100 euros are fine."""
    return call.args["amount"] <= 100


def handle(ticket):
    result = agent.run_sync(ticket, deps=desk, model=desk_model)
    if isinstance(result.output, DeferredToolRequests):
        decisions = DeferredToolResults()
        for call in result.output.approvals:
            ok = manager_approves(call)
            print(f"  approval needed: {call.args} -> {'yes' if ok else 'no'}")
            decisions.approvals[call.tool_call_id] = True if ok else ToolDenied("A manager said no.")
        result = agent.run_sync(message_history=result.all_messages(), deferred_tool_results=decisions,
                                deps=desk, model=desk_model)
    return result.output


tickets = [
    "Where is my order A-1001?",
    "Please refund 30 euros on A-1001, it arrived broken",
    "Refund 90 euros on A-1002, I cancelled",
    "Refund 110 euros on A-1002 please",
    "Where is A-9999?",
]
for ticket in tickets:
    print(ticket)
    print("  reply:", handle(ticket))
print(desk.orders)

Running five tickets end to end

handle runs a ticket. If the run pauses for approval, it asks manager_approves, a stand-in for a person in your admin page, then runs again with the decision:

Example
PYDANTIC_AI_NO_BANNER=1 python app.py

What each ticket did

  • A status question: one tool call, one reply.
  • 30 euros: under the limit, refunded without asking.
  • 90 euros on A-1002: over 50, so the run paused, a manager said yes, and the refund went through.
  • 110 euros: the manager said no, the tool never ran, and the customer got the refusal.
  • An unknown order: ModelRetry, and the model passed the reason on.

The last line shows the orders: only the approved and small refunds were recorded. The denied 110 euro refund is the failure mode the overview promised, working.

Five tickets through the desk
ticketoutputrecordsasksyes or nohandle(ticket)runs, pauses, rerunsmanager_approvesup to 100 euros: yesAgentdesk_model, deps=Desklookup_orderModelRetry if unknownrefund_orderapproval over 50desk.orderswhat was refunded
Hover or tap a piece to see what it is and which lesson built it.
Trace a ticket

Pick one to watch it run, step by step.

A second failure: a tool that keeps failing

A tool that raises ModelRetry gives the model a chance to fix its call. If the model keeps making the same failing call, the run does not loop forever: it stops with UnexpectedModelBehavior once the tool's retry limit is reached:

Example
from pydantic_ai import UnexpectedModelBehavior, ModelResponse, ToolCallPart
from pydantic_ai.models.function import FunctionModel

from desk import Desk, agent

desk = Desk(orders={})


def keeps_trying(messages, info):
    # the model keeps looking up an order that is not there
    return ModelResponse(parts=[ToolCallPart("lookup_order", {"order_id": "A-9999"})])


try:
    agent.run_sync("Where is A-9999?", deps=desk, model=FunctionModel(keeps_trying))
except UnexpectedModelBehavior as exc:
    print("run stopped:", exc)

The message names the tool, the limit it hit, and where to read more. In your app, catch this and show the customer a fallback, rather than letting the run raise.

Testing the refund rules

The tests pin the rules that matter: small refunds go through, big ones wait, a refund above the order cost is refused, an unknown order is handled:

python
from pydantic_ai import DeferredToolRequests, DeferredToolResults, models

from desk import Desk, agent
from desk_model import desk_model

models.ALLOW_MODEL_REQUESTS = False


def make_desk():
    return Desk(orders={"A-1001": {"status": "shipped", "total": 40.0},
                        "A-1002": {"status": "waiting for stock", "total": 120.0}})


def test_small_refund_needs_no_approval():
    desk = make_desk()
    with agent.override(model=desk_model):
        result = agent.run_sync("Refund 20 euros on A-1001", deps=desk)
    assert result.output == "Refunded 20.00 euros on order A-1001."
    assert desk.orders["A-1001"]["refunded"] == 20.0


def test_large_refund_waits_for_a_person():
    desk = make_desk()
    with agent.override(model=desk_model):
        result = agent.run_sync("Refund 90 euros on A-1002", deps=desk)
    assert isinstance(result.output, DeferredToolRequests)
    assert "refunded" not in desk.orders["A-1002"]

    call = result.output.approvals[0]
    with agent.override(model=desk_model):
        final = agent.run_sync(message_history=result.all_messages(), deps=desk,
                               deferred_tool_results=DeferredToolResults(approvals={call.tool_call_id: True}))
    assert desk.orders["A-1002"]["refunded"] == 90.0
    assert "Refunded 90.00" in final.output


def test_refund_above_order_total_is_refused():
    desk = make_desk()
    with agent.override(model=desk_model):
        result = agent.run_sync("Refund 45 euros on A-1001", deps=desk)
    assert "refunded" not in desk.orders["A-1001"]
    assert result.output == "Sorry, I could not do that. Order A-1001 only cost 40.00 euros."


def test_unknown_order():
    with agent.override(model=desk_model):
        result = agent.run_sync("Where is A-7777?", deps=make_desk())
    assert result.output == "Sorry, I could not do that. There is no order A-7777."
Example
pytest -q

Scoring the desk with an eval

The eval shows how the desk handles tickets it was not written for. The stand-in has no rule for 'Can I get 25 euros back', so it looked the order up instead, which the eval catches:

python
from pydantic_evals import Case, Dataset
from pydantic_evals.evaluators import Contains, Evaluator, EvaluatorContext

from desk import Desk, agent
from desk_model import desk_model


class ShortReply(Evaluator):
    def evaluate(self, ctx: EvaluatorContext) -> bool:
        return len(ctx.output) <= 80


async def answer(ticket: str) -> str:
    desk = Desk(orders={"A-1001": {"status": "shipped", "total": 40.0}})
    result = await agent.run(ticket, deps=desk, model=desk_model)
    return str(result.output)


dataset = Dataset(
    name="desk",
    cases=[
        Case(name="status", inputs="Where is A-1001?", evaluators=[Contains("shipped")]),
        Case(name="small refund", inputs="Refund 10 euros on A-1001", evaluators=[Contains("Refunded 10.00")]),
        Case(name="no order id", inputs="Where is my parcel?", evaluators=[Contains("order number")]),
        Case(name="lower case id", inputs="where is a-1001", evaluators=[Contains("shipped")]),
        Case(name="unknown order", inputs="Where is A-7777?", evaluators=[Contains("no order A-7777")]),
        Case(name="money back", inputs="Can I get 25 euros back on A-1001?", evaluators=[Contains("Refunded 25.00")]),
    ],
    evaluators=[ShortReply()],
)

report = dataset.evaluate_sync(answer, progress=False)
for case in report.cases:
    failed = [name for name, result in case.assertions.items() if not result.value]
    print(f"{case.name:14} {case.output}")
    if failed:
        print("    failed:", failed)
print(f"passed: {report.averages().assertions:.1%}")
Example
PYDANTIC_AI_NO_BANNER=1 python eval_desk.py

That failing case is one to keep in the dataset while you change the model or the prompt. To use a real model, follow Integrations: from stand-in to real backends: drop model=desk_model and set the provider key. The tests keep the stand-in through override.

What we left out

The course taught the desk's spine. These pages of the docs it named but did not cover, each worth reading when you need it:

TopicWhy it was skippedDocs
pydantic-graphState machines for control flow beyond one agent loopgraph
Durable executionRuns that survive a crash, with Temporal, DBOS, Prefect or Restatedurable_execution
Long-term memoryA store with semantic search; the core has none, the Harness Memory capability is the closestharness/memory
Direct model requestsCalling a model once without an agentdirect
Native and built-in toolsProvider-run tools such as web searchnative-tools
ThinkingReading a model's reasoning tokensthinking
Multimodal inputImages, audio and documents in a promptinput
RealtimeSpeech-to-speech voice runsrealtime
AG-UI and Vercel UIStreaming an agent to a web front endag-ui

Where you take the desk next

  • Serve handle from a web endpoint, storing paused runs by conversation id.
  • Add a real model behind a FallbackModel, and keep the tests on the stand-in.
  • Add usage limits to every run and a test that checks the cap is hit.
The approval must outlive the request
The desk records a refund only after approval, but manager_approves here is a stub. In a real app the decision comes from a person through your admin page, and the paused run must be stored between the two calls, not held in memory. Persist it, as the integrations lesson describes, or a restart loses the pending refund.
Try it yourself
  • Serve handle from a FastAPI endpoint and store paused runs by conversation id.
  • Add a FallbackModel of two real models to desk.py.
  • Add usage_limits to every run in app.py and a test that checks it.

Slow is fine. Stopping is the only problem.