Pydantic AIPydantic AI 2.51 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
29 small wins to finish your pathNext lesson →

Multi-agent delegation: agents that call agents

Delegation is a multi-agent pattern that lets one agent call another inside a tool, so the second does part of the first agent's job and returns its answer.

Last updated: 28 Sep, 2026 · Pydantic AI 2.51

The FallbackModel: when a provider is down lesson kept one agent working when a provider failed. A real desk needs more than one skill: one agent to read a ticket and decide what to do, another that is good at wording a reply. Pydantic AI lets an agent run another agent from inside a tool, which is called delegation.

The delegation call

A tool on the first agent runs the second agent and returns its output. Passing usage=ctx.usage counts both agents' requests against one budget:

python
@desk.tool
async def draft_reply(ctx: RunContext, ticket: str) -> str:
    result = await writer.run(ticket, usage=ctx.usage)  # writer runs inside the desk run
    return result.output                                # its reply becomes the tool result

Two agents with different jobs

writer words the reply; desk handles the ticket. Each has its own model, instructions and tools, so the writer could later run on a cheaper model:

python
from pydantic_ai import Agent, ModelResponse, RunContext, TextPart, ToolCallPart
from pydantic_ai.models.function import FunctionModel


def write(messages, info):
    return ModelResponse(parts=[TextPart("Sorry about the double charge. The extra payment will be back within 5 days.")])


writer = Agent(FunctionModel(write), instructions="Write a short, kind reply to the customer.")


def lead_model(messages, info):
    last = messages[-1].parts[-1]
    if last.part_kind == "user-prompt":
        return ModelResponse(parts=[ToolCallPart("draft_reply", {"ticket": last.content})])
    return ModelResponse(parts=[TextPart(f"Reply ready: {last.content}")])


desk = Agent(FunctionModel(lead_model), instructions="Handle the ticket. Use draft_reply for the wording.")

Delegating to the writer inside a tool

The desk's model calls a draft_reply tool, and that tool runs the writer and hands back its wording:

Example
@desk.tool
async def draft_reply(ctx: RunContext, ticket: str) -> str:
    """Ask the writer for the reply to send."""
    result = await writer.run(ticket, usage=ctx.usage)
    return result.output


result = desk.run_sync("I was charged twice")
print(result.output)
print(result.usage)

What the usage count shows

  • Three requests. Two are the desk's own model calls and one is the writer's, because usage=ctx.usage added the writer's request to the desk's tally.
  • One tool call. The desk called draft_reply once, and the reply it printed is the writer's real text, not a transfer record.
  • The writer's messages stay hidden. The desk only sees the tool's return value, so the writer's own turns are not in the desk's history.

Counting usage without ctx.usage

Drop usage=ctx.usage and the writer's request is no longer added to the desk's total, so a usage limit on the desk would never see the writer's calls:

Example
@desk.tool
async def draft_reply(ctx: RunContext, ticket: str) -> str:
    """Ask the writer for the reply to send."""
    result = await writer.run(ticket)
    return result.output


print(desk.run_sync("I was charged twice").usage)

Delegation vs hand-off

PatternWho runs the next agentHistory
DelegationThe first agent, inside a tool, during its runThe child's turns stay inside the tool; the parent sees only the return value
Hand-offYour own code, after reading the first agent's outputYou choose the next agent and pass whatever history you want

The triage step from Output functions and several output types, where your if picks what runs next, is a hand-off. Delegation keeps the choice inside the model's tool call.

Where you use delegation

  • A lead agent that routes work to specialists, such as a billing agent and a shipping agent.
  • A drafting step where a wording agent writes the customer-facing text.
  • A cheap-model worker called by an expensive lead agent to save cost.
Pass the usage through
Without usage=ctx.usage, the child agent's requests and tokens are not counted against the parent's budget. A UsageLimits cap on the parent then cannot stop a child that loops, and your cost report is wrong. Pass ctx.usage whenever you delegate.
Try it yourself
  • Give writer a deps_type and pass deps=ctx.deps from the tool.
  • Run the desk with usage_limits=UsageLimits(request_limit=2), with and without usage=ctx.usage, and read the error.
  • Write the hand-off version: run a triage agent, then run the writer only for tickets it marks as replies.

Little by little, you're building something great.