Multi-agent delegation: agents that call agents
Delegation is a multi-agent pattern that lets one agent call another inside a tool, so the second does part of the first agent's job and returns its answer.
Last updated: 28 Sep, 2026 · Pydantic AI 2.51
The FallbackModel: when a provider is down lesson kept one agent working when a provider failed. A real desk needs more than one skill: one agent to read a ticket and decide what to do, another that is good at wording a reply. Pydantic AI lets an agent run another agent from inside a tool, which is called delegation.
The delegation call
A tool on the first agent runs the second agent and returns its output. Passing usage=ctx.usage counts both agents' requests against one budget:
@desk.tool
async def draft_reply(ctx: RunContext, ticket: str) -> str:
result = await writer.run(ticket, usage=ctx.usage) # writer runs inside the desk run
return result.output # its reply becomes the tool resultTwo agents with different jobs
writer words the reply; desk handles the ticket. Each has its own model, instructions and tools, so the writer could later run on a cheaper model:
from pydantic_ai import Agent, ModelResponse, RunContext, TextPart, ToolCallPart
from pydantic_ai.models.function import FunctionModel
def write(messages, info):
return ModelResponse(parts=[TextPart("Sorry about the double charge. The extra payment will be back within 5 days.")])
writer = Agent(FunctionModel(write), instructions="Write a short, kind reply to the customer.")
def lead_model(messages, info):
last = messages[-1].parts[-1]
if last.part_kind == "user-prompt":
return ModelResponse(parts=[ToolCallPart("draft_reply", {"ticket": last.content})])
return ModelResponse(parts=[TextPart(f"Reply ready: {last.content}")])
desk = Agent(FunctionModel(lead_model), instructions="Handle the ticket. Use draft_reply for the wording.")Delegating to the writer inside a tool
The desk's model calls a draft_reply tool, and that tool runs the writer and hands back its wording:
@desk.tool
async def draft_reply(ctx: RunContext, ticket: str) -> str:
"""Ask the writer for the reply to send."""
result = await writer.run(ticket, usage=ctx.usage)
return result.output
result = desk.run_sync("I was charged twice")
print(result.output)
print(result.usage)Reply ready: Sorry about the double charge. The extra payment will be back within 5 days. RunUsage(input_tokens=177, output_tokens=48, requests=3, tool_calls=1)
What the usage count shows
- Three requests. Two are the desk's own model calls and one is the writer's, because
usage=ctx.usageadded the writer's request to the desk's tally. - One tool call. The desk called
draft_replyonce, and the reply it printed is the writer's real text, not a transfer record. - The writer's messages stay hidden. The desk only sees the tool's return value, so the writer's own turns are not in the desk's history.
Counting usage without ctx.usage
Drop usage=ctx.usage and the writer's request is no longer added to the desk's total, so a usage limit on the desk would never see the writer's calls:
@desk.tool
async def draft_reply(ctx: RunContext, ticket: str) -> str:
"""Ask the writer for the reply to send."""
result = await writer.run(ticket)
return result.output
print(desk.run_sync("I was charged twice").usage)RunUsage(input_tokens=123, output_tokens=33, requests=2, tool_calls=1)
Delegation vs hand-off
| Pattern | Who runs the next agent | History |
|---|---|---|
| Delegation | The first agent, inside a tool, during its run | The child's turns stay inside the tool; the parent sees only the return value |
| Hand-off | Your own code, after reading the first agent's output | You choose the next agent and pass whatever history you want |
The triage step from Output functions and several output types, where your if picks what runs next, is a hand-off. Delegation keeps the choice inside the model's tool call.
Where you use delegation
- A lead agent that routes work to specialists, such as a billing agent and a shipping agent.
- A drafting step where a wording agent writes the customer-facing text.
- A cheap-model worker called by an expensive lead agent to save cost.
usage=ctx.usage, the child agent's requests and tokens are not counted against the parent's budget. A UsageLimits cap on the parent then cannot stop a child that loops, and your cost report is wrong. Pass ctx.usage whenever you delegate.Related
- Previous: FallbackModel: when a provider is down
- Next: MCP servers: tools from another program
- See also: Usage limits: stopping a run that loops
- Reference: Multi-agent applications
- Give
writeradeps_typeand passdeps=ctx.depsfrom the tool. - Run the desk with
usage_limits=UsageLimits(request_limit=2), with and withoutusage=ctx.usage, and read the error. - Write the hand-off version: run a triage agent, then run the writer only for tickets it marks as replies.
Little by little, you're building something great.