Multi-agent: routing between two agents
An AgentWorkflow is a group of agents that hand a request to each other, so a triage agent sends each question to the specialist that can answer it.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
One agent with many tools grows hard to steer. Splitting the work gives each agent a small job: a triage agent routes, a search agent answers customer questions, a policy agent answers staff questions from a separate index.
A handoff and a specialist
Each agent has a name, a description, and the agents it may hand off to. The triage agent holds no tools; it only routes with a built-in handoff call.
from llama_index.core.agent.workflow import AgentWorkflow, FunctionAgent
triage = FunctionAgent(name="triage_agent", description="Sends a question to the right specialist.",
llm=MockFunctionCallingLLM(response_generator=routes), tools=[],
can_handoff_to=["search_agent", "policy_agent"], system_prompt="Route the question.")Routing to the right agent
The triage generator reads the question and emits a handoff call naming the target agent. A question about approvals goes to the policy agent, everything else to the search agent.
def routes(messages, **kwargs):
question = last_user(messages).lower()
to = "policy_agent" if ("approve" in question or "approval" in question) else "search_agent"
return ChatMessage(role=MessageRole.ASSISTANT, blocks=[ToolCallBlock(
tool_call_id="h1", tool_name="handoff",
tool_kwargs={"to_agent": to, "reason": "specialist question"})])A specialist that answers, not repeats the handoff
After a handoff, the specialist sees the handoff note as a tool message. Its generator skips that note and calls its own tool, so the final reply is the real answer, not the transfer record.
def calls_tool(tool_name):
def generate(messages, **kwargs):
answered = [m for m in messages if m.role == MessageRole.TOOL and "is now handling" not in (m.content or "")]
if answered:
return ChatMessage(role=MessageRole.ASSISTANT, content=answered[-1].content)
return ChatMessage(role=MessageRole.ASSISTANT, blocks=[ToolCallBlock(
tool_call_id="c1", tool_name=tool_name, tool_kwargs={"input": last_user(messages)})])
return generateTwo questions routed to two agents
The whole program. The triage agent routes each question, and the specialist that received it answers from its own index.
import asyncio
from uuid import uuid4
from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.agent.workflow import AgentWorkflow, FunctionAgent
from llama_index.core.base.llms.types import ChatMessage, MessageRole, ToolCallBlock
from llama_index.core.llms import MockFunctionCallingLLM
from llama_index.core.tools import QueryEngineTool
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from extractive_llm import ExtractiveLLM
Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
customer = VectorStoreIndex.from_documents(SimpleDirectoryReader("help").load_data())
staff = VectorStoreIndex.from_documents(SimpleDirectoryReader("staff").load_data())
help_tool = QueryEngineTool.from_defaults(
customer.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2),
name="help_centre", description="Customer refund, delivery and lamp questions.")
policy_tool = QueryEngineTool.from_defaults(
staff.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2),
name="staff_policy", description="Internal refund-approval rules for staff.")
HANDOFF = "is now handling"
def last_user(messages):
return next((m.content for m in reversed(messages) if m.role == MessageRole.USER), "")
def calls_tool(tool_name):
def generate(messages, **kwargs):
answered = [m for m in messages if m.role == MessageRole.TOOL and HANDOFF not in (m.content or "")]
if answered:
return ChatMessage(role=MessageRole.ASSISTANT, content=answered[-1].content)
return ChatMessage(role=MessageRole.ASSISTANT, blocks=[ToolCallBlock(
tool_call_id="c" + uuid4().hex[:6], tool_name=tool_name,
tool_kwargs={"input": last_user(messages)})])
return generate
def routes(messages, **kwargs):
question = last_user(messages).lower()
to = "policy_agent" if ("approve" in question or "approval" in question) else "search_agent"
return ChatMessage(role=MessageRole.ASSISTANT, blocks=[ToolCallBlock(
tool_call_id="h" + uuid4().hex[:6], tool_name="handoff",
tool_kwargs={"to_agent": to, "reason": "specialist question"})])
triage = FunctionAgent(name="triage_agent", description="Sends a question to the right specialist.",
llm=MockFunctionCallingLLM(response_generator=routes), tools=[],
can_handoff_to=["search_agent", "policy_agent"], system_prompt="Route the question.")
search = FunctionAgent(name="search_agent", description="Answers customer help-centre questions.",
llm=MockFunctionCallingLLM(response_generator=calls_tool("help_centre")), tools=[help_tool],
system_prompt="Answer from the help centre.")
policy = FunctionAgent(name="policy_agent", description="Answers staff refund-approval questions.",
llm=MockFunctionCallingLLM(response_generator=calls_tool("staff_policy")), tools=[policy_tool],
system_prompt="Answer from the staff policy.")
workflow = AgentWorkflow(agents=[triage, search, policy], root_agent="triage_agent")
async def main():
for question in ["Can I refund a sale item?", "Who approves a refund over 200?"]:
result = await workflow.run(user_msg=question)
print(question)
print(" ", result)
asyncio.run(main())Where each answer came from
- The sale-item question was routed to the search agent, which answered from the customer help files.
- The approval question was routed to the policy agent, which answered from the staff-only file.
- The final reply is the specialist's answer, not the handoff note, because the specialist skipped that note and ran its own tool.
One agent vs an AgentWorkflow
| Shape | Each agent's job | Good for |
|---|---|---|
| Single agent, many tools | Everything | A few tools with clear names |
| AgentWorkflow, handoffs | One small job each | Separate knowledge or permissions per agent |
When to split into several agents
- Customer and staff knowledge that must stay in separate indexes.
- A triage step that routes to the right team before answering.
- Specialists you want to test and change on their own.
Related
- Previous: Workflows: a custom RAG pipeline with @step
- Next: Observability: tracing what a run did
- See also: Metadata filters: who may see which documents
- Reference: Multi-agent workflows
- Ask a delivery question and confirm it routes to the search agent.
- Add the word "approval" to a delivery question and watch the route change.
- Add a third agent for lamp questions and route to it on the word "lamp".
This is what real progress feels like.