FunctionAgent: an agent that queries your index
A FunctionAgent is an agent that decides which tool to call for a question, calls it, and answers from the result.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
So far the query engine ran every time. An agent looks at the question first, then chooses to use a tool. Wrap your index as a tool and the agent can search the shop help centre only when a question needs it.
Wrapping the index as a tool
A QueryEngineTool turns a query engine into something an agent can call. The name and description are what the agent reads when it decides whether to use it.
from llama_index.core.tools import QueryEngineTool
help_tool = QueryEngineTool.from_defaults(
query_engine, # the query engine from earlier lessons
name="help_centre",
description="Answers refund, delivery and lamp questions.",
)The keyless stand-in model
A real agent needs a model that can emit tool calls. Without a key, use MockFunctionCallingLLM with a small generator that says what to do: on the first turn, call the tool with the question; once the tool has run, return its result as the answer.
from llama_index.core.base.llms.types import ChatMessage, MessageRole, ToolCallBlock
def generate(messages, **kwargs):
done = [m for m in messages if m.role == MessageRole.TOOL]
if done: # the tool has run: its result is the answer
return ChatMessage(role=MessageRole.ASSISTANT, content=done[-1].content)
question = messages[-1].content # first turn: call the tool with the question
return ChatMessage(role=MessageRole.ASSISTANT, blocks=[ToolCallBlock(
tool_call_id="call-1", tool_name="help_centre",
tool_kwargs={"input": question})])Passing the question as input is what makes the tool run a real search. The default generator would call the tool with empty arguments, so the retrieval would find nothing.
Building and running the agent
Give the agent the tool and the stand-in model, then run it. agent.run is asynchronous, so it is called inside asyncio.run.
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.core.llms import MockFunctionCallingLLM
agent = FunctionAgent(tools=[help_tool],
llm=MockFunctionCallingLLM(response_generator=generate),
system_prompt="Answer shop questions using the help centre tool.")
answer = await agent.run("How long until my refund money reaches my card?")An agent answering two questions
The whole program. The tool runs a real retrieval over the help files, and the answer is the sentence the tool returned.
import asyncio
from uuid import uuid4
from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.core.base.llms.types import ChatMessage, MessageRole, ToolCallBlock
from llama_index.core.llms import MockFunctionCallingLLM
from llama_index.core.tools import QueryEngineTool
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from extractive_llm import ExtractiveLLM # the stand-in answering model from the answers-with-sources lesson
Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
index = VectorStoreIndex.from_documents(SimpleDirectoryReader("help").load_data())
query_engine = index.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2)
help_tool = QueryEngineTool.from_defaults(
query_engine, name="help_centre",
description="Answers questions about refunds, delivery and lamps from the shop help centre.")
def generate(messages, **kwargs):
done = [m for m in messages if m.role == MessageRole.TOOL]
if done: # tool has run: its result is the answer
return ChatMessage(role=MessageRole.ASSISTANT, content=done[-1].content)
question = messages[-1].content # first turn: call the tool with the question
return ChatMessage(role=MessageRole.ASSISTANT, blocks=[ToolCallBlock(
tool_call_id="call-" + uuid4().hex[:8], tool_name="help_centre",
tool_kwargs={"input": question})])
agent = FunctionAgent(tools=[help_tool], llm=MockFunctionCallingLLM(response_generator=generate),
system_prompt="Answer shop questions using the help centre tool.")
async def main():
for question in ["How long until my refund money reaches my card?", "What is your phone number?"]:
answer = await agent.run(question)
print(question)
print(" ", answer)
asyncio.run(main())What the two runs show
- The first question is passed into the tool, which retrieves from the refund and delivery files and returns the matching sentence.
- The second question has no answer in the documents, so the tool returns the stand-in model's refusal and the agent repeats it.
- The answer changes with the question, which means the tool ran for real; a fake agent would print the same thing both times.
Query engine alone vs a FunctionAgent
| Approach | Who decides to search | When to use |
|---|---|---|
| Query engine | You do, every call | One fixed step: retrieve then answer |
| FunctionAgent | The model, per question | A question that may or may not need the tool, or one of several tools |
When an agent over your index helps
- A chatbot that sometimes answers from documents and sometimes from a live lookup.
- A question that needs a search plus another action, such as opening a ticket.
- Giving the model a choice, so simple questions skip retrieval entirely.
input is a required field it leaves blank. Pass a response_generator that puts the question into input, or the search runs on nothing and finds nothing.Related
- Previous: Measuring retrieval: hit rate on labelled questions
- Next: Agent tools: giving an agent more than search
- See also: Answers with sources: a stand-in model and citations
- Reference: Understanding: agents
- Ask a delivery question and confirm the tool returns a delivery sentence.
- Remove
tool_kwargs={"input": question}and pass empty kwargs. What does the agent answer now? - Change the tool description to mention only refunds and see whether the answer changes.
Every expert started right here.