Testing an agent
Testing an agent is checking its tools, routing and refusals with a scripted model in place of a real one, so every test runs the same way each time.
Last updated: 27 Sep, 2026 · LangChain 1.4
Tests for code, evaluations for answers
There are two different things to check in an LLM app. Tests check your code: that a tool returns the right thing, that the router sends a question to the right place, that a failure is handled. They use fixed inputs and exact expected outputs, and they run without a model. Evaluations check the model's answers, which are never exactly the same twice; The desk on a hosted model runs a small one. This lesson is about the first kind: tests that pin down your own code with a scripted stand-in model, so they are exact and free to run on every change.
LangChain ships GenericFakeChatModel for this: it returns the replies you give it, one per call, including tool calls. The tests run with pytest.
Scripting replies with GenericFakeChatModel
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel
model = GenericFakeChatModel(messages=iter([reply_1, reply_2])) # one reply per callWhen the fake model gets a tool
This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. It lives in orders.py, the file Subagents as tools started, and the tests below import it from there.
from langchain.tools import tool
ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
status = ORDERS.get(order_id)
return f"{order_id} {status}." if status else f"{order_id} is not an order we have."Give the fake a tool and it fails on the first call. Here is the error before the fix.
from orders import lookup_order
from langchain.agents import create_agent
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel
model = GenericFakeChatModel(messages=iter(["done"]))
agent = create_agent(model, tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "B22?"}]})Traceback (most recent call last):
File "main.py", line 7, in <module>
agent.invoke({"messages": [{"role": "user", "content": "B22?"}]})
NotImplementedError
During task with name 'model' and id 'e53ce057-a5ce-73e4-4bf2-32828b7e7c77'It cannot be given tools: its bind_tools raises NotImplementedError, so an agent with any tool fails on its first call. It only works in an agent built with tools=[]. Three lines fix it.
A fake that accepts tools
Subclass it and let bind_tools return the model unchanged. Save it as scripted.py.
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel
class ScriptedModel(GenericFakeChatModel):
def bind_tools(self, tools, **kwargs):
return self # accept tools, keep the scripted repliesScripting a tool call, then an answer
Script two replies: first an AIMessage that asks for the tool, then a final text answer.
from orders import lookup_order
from langchain.messages import AIMessage, ToolCall
from scripted import ScriptedModel
call = ToolCall(name="lookup_order", args={"order_id": "B22"}, id="call_1")
# first the model asks for the tool, then it gives a final answer
model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))Running the scripted model through the agent
Run the scripted model through an agent with the tools lesson's lookup_order and print every message.
from langchain.agents import create_agent
from langchain.messages import AIMessage, ToolCall
from scripted import ScriptedModel
call = ToolCall(name="lookup_order", args={"order_id": "B22"}, id="call_1")
model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))
result = create_agent(model, tools=[lookup_order]).invoke({"messages": [{"role": "user", "content": "B22?"}]})
for message in result["messages"]:
print(f"{message.type:<5} {message.text or message.tool_calls[0]['args']}")human B22?
ai {'order_id': 'B22'}
tool B22 is not an order we have.
ai doneWhat the scripted run proves
- The model's two replies were fixed in advance, so this run is the same every time.
- The tool itself ran: the
toolline islookup_order's real answer, which is what the test checks. - The last line is "done", the second scripted reply, which ends the loop.
Tests with pytest
Install pytest, the test runner these tests use.
pip install "pytest==9.1.1"A scripted specialist
The router lesson's specialists are real Groq agents, and a test cannot assert on a real model's wording. So the tests build their own stand-in specialist: an agent on a ScriptedModel that asks for one tool call and then says "done". Add it to the end of scripted.py.
from langchain.agents import create_agent
from langchain.messages import AIMessage, ToolCall
def scripted_agent(tool, args):
"""An agent whose model asks for one tool call, then says "done"."""
call = ToolCall(name=tool.name, args=args, id="call_1")
model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))
return create_agent(model, tools=[tool])- written in Documents and splitting
- written in Embeddings and a vector store
- written in Retrieval as a tool
- written in Router: sending each question to one agent
View the code here
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter
POLICIES = {
"refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
"\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
"shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
"\n\nExpress shipping arrives the next working day and costs 9 euros.",
"accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}
docs = [Document(page_content=text, metadata={"source": name}) for name, text in POLICIES.items()]
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs)
import re
import zlib
from langchain_core.embeddings import Embeddings
COMMON = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i",
"if", "is", "it", "my", "of", "on", "the", "to", "what", "with", "you", "your"}
class WordEmbeddings(Embeddings):
def embed_query(self, text):
vector = [0.0] * 256
for word in re.findall(r"[a-z]+", text.lower()):
if word not in COMMON:
vector[zlib.crc32(word.rstrip("s").encode()) % 256] += 1.0
return vector
def embed_documents(self, texts):
return [self.embed_query(text) for text in texts]
from langchain.tools import tool
from langchain_core.vectorstores import InMemoryVectorStore
from policies import chunks
from word_embeddings import WordEmbeddings
store = InMemoryVectorStore(WordEmbeddings())
store.add_documents(chunks)
@tool
def search_policies(query: str) -> str:
"""Search the shop's policies on refunds, shipping and accounts.
Pass the customer's question, word for word, as the query."""
found = [doc for doc, score in store.similarity_search_with_score(query, k=2) if score >= 0.3]
if not found:
return "No policy covers this."
return "\n".join(f"[{doc.metadata['source']}] {doc.page_content}" for doc in found)
import re
def route(question):
if re.search(r"\b[A-Z]\d+\b", question): # an order id, like A17
return "orders"
return "policies" # everything else
def answer(question, agents):
name = route(question) # pick one specialist
result = agents[name].invoke({"messages": [{"role": "user", "content": question}]})
return name, result["messages"][-1].text # the question went straight through
The test file
Wrap the checks in pytest functions in test_router.py. It needs orders.py, router.py and search.py next to it, and search.py needs policies.py and word_embeddings.py: the files from the subagents, router, retrieval-tool, documents and embeddings lessons. Each one builds what it needs, then asserts on the result. ask sends one question and returns the messages.
from orders import lookup_order
from router import route
from scripted import scripted_agent
from search import search_policies
def ask(agent, question):
return agent.invoke({"messages": [{"role": "user", "content": question}]})["messages"]
def tool_reply(messages):
return next(m.text for m in messages if m.type == "tool")
def test_unknown_order_is_reported():
messages = ask(scripted_agent(lookup_order, {"order_id": "B22"}), "Where is B22?")
assert tool_reply(messages) == "B22 is not an order we have." # the real tool's replydef test_order_questions_go_to_the_orders_agent():
assert route("Where is A17?") == "orders"
assert route("Is shipping free?") == "policies"
def test_uncovered_questions_are_refused():
question = "Can I pay with bitcoin?"
messages = ask(scripted_agent(search_policies, {"query": question}), question)
assert tool_reply(messages) == "No policy covers this." # the score cut heldThe first and third tests run a scripted model instead of a real one: it asks for one tool call and says "done", and the assertion reads what the real tool returned, the unknown-order message and the retrieval lesson's "No policy covers this.". The second checks the router lesson's route, which is plain Python and needs no model at all. None of them needs a key or a network, so they can run on every change.
Run the file. -q keeps the report short, and -p no:warnings turns off pytest's warnings summary, which would otherwise list deprecation notices from the installed libraries.
pytest -q -p no:warnings test_router.py... [100%] 3 passed in 0.32s
Real model vs scripted fake
| Real model | Scripted fake | |
|---|---|---|
| Same result each run | No | Yes |
| Needs a key or network | Yes | No |
| What it tests | The model too | Your tools, routing and refusals |
| Tool calls | The model decides | You script them |
When to test with a scripted fake
- Checking a tool's output for a known input on every commit.
- Proving routing and refusals behave without paying for a model.
GenericFakeChatModel raises NotImplementedError from bind_tools, so an agent with any tool fails on the first call. Give it a subclass whose bind_tools returns itself before you add tools.Related
- Previous: Router: sending each question to one agent
- Next: Three tools and a model that picks
- Reference: Testing
- Change the expected text in the first test and read how pytest reports the failure.
- Add a test that
route("How do I reset my password?")returns"policies". - Script a model that asks for
lookup_ordertwice and assert there are two tool messages.
Little by little, you're building something great.