Testing an agent
Testing an agent is checking its tools, routing and refusals with no model at all: a scripted fake plays the model's part, so every test runs the same way each time.
Last updated: 27 Sep, 2026 · LangChain 1.4
A test should check your code, not a model's mood. LangChain's documentation recommends GenericFakeChatModel, which returns the replies you give it, one per call, including tool calls.
Scripting replies with GenericFakeChatModel
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel
model = GenericFakeChatModel(messages=iter([reply_1, reply_2])) # one reply per callWhen the fake model gets a tool
Give the fake a tool and it fails on the first call. Here is the error before the fix.
from langchain.agents import create_agent
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel
from tools import lookup_order
model = GenericFakeChatModel(messages=iter(["done"]))
agent = create_agent(model, tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "B22?"}]})It cannot be given tools: its bind_tools raises NotImplementedError, so an agent with any tool fails on its first call. The documentation's example passes tools=[], which is why it works there. Three lines fix it.
A fake that accepts tools
Subclass it and let bind_tools return the model unchanged.
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel
class ScriptedModel(GenericFakeChatModel):
def bind_tools(self, tools, **kwargs):
return self # accept tools, keep the scripted repliesScripting a tool call, then an answer
Script two replies: first an AIMessage that asks for the tool, then a final text answer.
from langchain.messages import AIMessage, ToolCall
from scripted import ScriptedModel
call = ToolCall(name="lookup_order", args={"order_id": "B22"}, id="call_1")
# first the model asks for the tool, then it gives a final answer
model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))Running the scripted model through the agent
Run the scripted model through the agent and print every message.
from langchain.messages import AIMessage, ToolCall
from scripted import ScriptedModel
call = ToolCall(name="lookup_order", args={"order_id": "B22"}, id="call_1")
model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))
result = create_agent(model, tools=[lookup_order]).invoke({"messages": [{"role": "user", "content": "B22?"}]})
for message in result["messages"]:
print(f"{message.type:<5} {message.text or message.tool_calls[0]['args']}")What the scripted run proves
- The model's two replies were fixed in advance, so this run is the same every time.
- The tool itself ran: the
toolline islookup_order's real answer, which is what the test checks. - The last line is "done", the second scripted reply, which ends the loop.
Tests with pytest
Wrap the checks in pytest functions. Each one builds its own agent or calls the router, then asserts on the result.
from langchain.agents import create_agent
from langchain.messages import AIMessage, ToolCall
from router import answer
from scripted import ScriptedModel
from tools import lookup_order
def test_unknown_order_is_reported():
call = ToolCall(name="lookup_order", args={"order_id": "B22"}, id="call_1")
model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))
agent = create_agent(model, tools=[lookup_order])
result = agent.invoke({"messages": [{"role": "user", "content": "Where is B22?"}]})
assert result["messages"][2].text == "B22 is not an order we have."def test_order_questions_go_to_the_orders_agent():
assert answer("Where is A17?") == ("orders", "A17 shipped on 3 March.")
def test_uncovered_questions_are_refused():
name, reply = answer("Can I pay with bitcoin?")
assert name == "policies" and reply.startswith("Our policies do not cover that")The first test uses the scripted model to check the tool's answer for an unknown order. The other two check lesson 31's router and lesson 28's refusal with the stand-in models, which decide the same way on every run. None of them needs a key or a network, so they can run on every change.
pytest -q -p no:warnings test_desk.pyReal model vs scripted fake
| Real model | Scripted fake | |
|---|---|---|
| Same result each run | No | Yes |
| Needs a key or network | Yes | No |
| What it tests | The model too | Your tools, routing and refusals |
| Tool calls | The model decides | You script them |
When to test with a scripted fake
- Checking a tool's output for a known input on every commit.
- Proving routing and refusals behave without paying for a model.
GenericFakeChatModel raises NotImplementedError from bind_tools, so an agent with any tool fails on the first call. Give it a subclass whose bind_tools returns itself before you add tools.Related
- Previous: Routing with a router node
- Next: Three tools and a model that picks
- Reference: Testing
- Change the expected text in the first test and read how pytest reports the failure.
- Add a test that
answer("Is shipping free?")goes to the policies agent. - Script a model that asks for
lookup_ordertwice and assert there are two tool messages.
Little by little, you're building something great.