LlamaIndexllama-index-core 0.14 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Multi-agent: routing between two agents

An AgentWorkflow is a group of agents that hand a request to each other, so a triage agent sends each question to the specialist that can answer it.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

One agent with many tools grows hard to steer. Splitting the work gives each agent a small job: a triage agent routes, a search agent answers customer questions, a policy agent answers staff questions from a separate index.

A handoff and a specialist

Each agent has a name, a description, and the agents it may hand off to. The triage agent holds no tools; it only routes with a built-in handoff call.

python
from llama_index.core.agent.workflow import AgentWorkflow, FunctionAgent

triage = FunctionAgent(name="triage_agent", description="Sends a question to the right specialist.",
    llm=MockFunctionCallingLLM(response_generator=routes), tools=[],
    can_handoff_to=["search_agent", "policy_agent"], system_prompt="Route the question.")

Routing to the right agent

The triage generator reads the question and emits a handoff call naming the target agent. A question about approvals goes to the policy agent, everything else to the search agent.

python
def routes(messages, **kwargs):
    question = last_user(messages).lower()
    to = "policy_agent" if ("approve" in question or "approval" in question) else "search_agent"
    return ChatMessage(role=MessageRole.ASSISTANT, blocks=[ToolCallBlock(
        tool_call_id="h1", tool_name="handoff",
        tool_kwargs={"to_agent": to, "reason": "specialist question"})])

A specialist that answers, not repeats the handoff

After a handoff, the specialist sees the handoff note as a tool message. Its generator skips that note and calls its own tool, so the final reply is the real answer, not the transfer record.

python
def calls_tool(tool_name):
    def generate(messages, **kwargs):
        answered = [m for m in messages if m.role == MessageRole.TOOL and "is now handling" not in (m.content or "")]
        if answered:
            return ChatMessage(role=MessageRole.ASSISTANT, content=answered[-1].content)
        return ChatMessage(role=MessageRole.ASSISTANT, blocks=[ToolCallBlock(
            tool_call_id="c1", tool_name=tool_name, tool_kwargs={"input": last_user(messages)})])
    return generate
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
help/lamps.md
# Lamps

The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.

All lamps come with a two year guarantee against electrical faults.

Bulbs are not covered by the refund policy once they have been used.

The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
help/refunds.md
# Refunds

You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.

Items bought in a sale can be refunded too, but the delivery charge is not returned.

To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.

Personalised items cannot be refunded unless they arrive damaged.
help/delivery.md
# Delivery

Standard delivery takes 3 to 5 working days and is free on orders over 40.

Express delivery arrives the next working day if you order before 2pm. It costs 6.

We deliver to the mainland only. Parcels to islands take 2 extra working days.

If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
extractive_llm.py
import re

from llama_index.core.llms import CompletionResponse, CustomLLM, LLMMetadata
from llama_index.core.llms.callbacks import llm_completion_callback


def stems(text):
    """Words longer than three letters, cut to five letters, so refund and refunds match."""
    return {w[:5] for w in re.findall(r"[a-z0-9-]+", text.lower()) if len(w) > 3}


class ExtractiveLLM(CustomLLM):
    """Answers with the context sentence that shares most words with the question."""

    @property
    def metadata(self):
        return LLMMetadata(model_name="extractive")

    @llm_completion_callback()
    def complete(self, prompt, formatted=False, **kwargs):
        context = prompt.split("---------------------")[1]
        question = prompt.split("Query:")[1].split("Answer:")[0]
        asked = stems(question)
        sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context)]
        sentences = [s for s in sentences if s and not s.startswith("#") and ": " not in s[:20]]
        best = max(sentences, key=lambda s: len(asked & stems(s)), default="")
        if len(asked & stems(best)) < 2:
            return CompletionResponse(text="I could not find that in the documents.")
        return CompletionResponse(text=best)

    @llm_completion_callback()
    def stream_complete(self, prompt, formatted=False, **kwargs):
        yield self.complete(prompt)
staff/refund-approvals.md
# Refund approvals

Refunds over 200 need a team lead's approval before they are paid.

A refund on a personalised item always needs a team lead to check the damage photos first.

Two questions routed to two agents

The whole program. The triage agent routes each question, and the specialist that received it answers from its own index.

Example
import asyncio
from uuid import uuid4

from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.agent.workflow import AgentWorkflow, FunctionAgent
from llama_index.core.base.llms.types import ChatMessage, MessageRole, ToolCallBlock
from llama_index.core.llms import MockFunctionCallingLLM
from llama_index.core.tools import QueryEngineTool
from llama_index.embeddings.huggingface import HuggingFaceEmbedding

from extractive_llm import ExtractiveLLM

Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")

customer = VectorStoreIndex.from_documents(SimpleDirectoryReader("help").load_data())
staff = VectorStoreIndex.from_documents(SimpleDirectoryReader("staff").load_data())
help_tool = QueryEngineTool.from_defaults(
    customer.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2),
    name="help_centre", description="Customer refund, delivery and lamp questions.")
policy_tool = QueryEngineTool.from_defaults(
    staff.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2),
    name="staff_policy", description="Internal refund-approval rules for staff.")

HANDOFF = "is now handling"


def last_user(messages):
    return next((m.content for m in reversed(messages) if m.role == MessageRole.USER), "")


def calls_tool(tool_name):
    def generate(messages, **kwargs):
        answered = [m for m in messages if m.role == MessageRole.TOOL and HANDOFF not in (m.content or "")]
        if answered:
            return ChatMessage(role=MessageRole.ASSISTANT, content=answered[-1].content)
        return ChatMessage(role=MessageRole.ASSISTANT, blocks=[ToolCallBlock(
            tool_call_id="c" + uuid4().hex[:6], tool_name=tool_name,
            tool_kwargs={"input": last_user(messages)})])
    return generate


def routes(messages, **kwargs):
    question = last_user(messages).lower()
    to = "policy_agent" if ("approve" in question or "approval" in question) else "search_agent"
    return ChatMessage(role=MessageRole.ASSISTANT, blocks=[ToolCallBlock(
        tool_call_id="h" + uuid4().hex[:6], tool_name="handoff",
        tool_kwargs={"to_agent": to, "reason": "specialist question"})])


triage = FunctionAgent(name="triage_agent", description="Sends a question to the right specialist.",
    llm=MockFunctionCallingLLM(response_generator=routes), tools=[],
    can_handoff_to=["search_agent", "policy_agent"], system_prompt="Route the question.")
search = FunctionAgent(name="search_agent", description="Answers customer help-centre questions.",
    llm=MockFunctionCallingLLM(response_generator=calls_tool("help_centre")), tools=[help_tool],
    system_prompt="Answer from the help centre.")
policy = FunctionAgent(name="policy_agent", description="Answers staff refund-approval questions.",
    llm=MockFunctionCallingLLM(response_generator=calls_tool("staff_policy")), tools=[policy_tool],
    system_prompt="Answer from the staff policy.")

workflow = AgentWorkflow(agents=[triage, search, policy], root_agent="triage_agent")


async def main():
    for question in ["Can I refund a sale item?", "Who approves a refund over 200?"]:
        result = await workflow.run(user_msg=question)
        print(question)
        print("  ", result)


asyncio.run(main())

Where each answer came from

  • The sale-item question was routed to the search agent, which answered from the customer help files.
  • The approval question was routed to the policy agent, which answered from the staff-only file.
  • The final reply is the specialist's answer, not the handoff note, because the specialist skipped that note and ran its own tool.

One agent vs an AgentWorkflow

ShapeEach agent's jobGood for
Single agent, many toolsEverythingA few tools with clear names
AgentWorkflow, handoffsOne small job eachSeparate knowledge or permissions per agent

When to split into several agents

  • Customer and staff knowledge that must stay in separate indexes.
  • A triage step that routes to the right team before answering.
  • Specialists you want to test and change on their own.
Watch out. After a handoff the specialist receives the transfer note as a tool message. If it treats that note as its answer, the customer sees "now handling the request" instead of a real reply; make the specialist run its own tool first.
Try it yourself
  • Ask a delivery question and confirm it routes to the search agent.
  • Add the word "approval" to a delivery question and watch the route change.
  • Add a third agent for lamp questions and route to it on the word "lamp".

This is what real progress feels like.