LangChainLangChain 1.4 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Testing an agent

Testing an agent is checking its tools, routing and refusals with no model at all: a scripted fake plays the model's part, so every test runs the same way each time.

Last updated: 27 Sep, 2026 · LangChain 1.4

A test should check your code, not a model's mood. LangChain's documentation recommends GenericFakeChatModel, which returns the replies you give it, one per call, including tool calls.

Scripting replies with GenericFakeChatModel

python
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel

model = GenericFakeChatModel(messages=iter([reply_1, reply_2]))   # one reply per call

When the fake model gets a tool

Give the fake a tool and it fails on the first call. Here is the error before the fix.

Example
from langchain.agents import create_agent
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel
from tools import lookup_order

model = GenericFakeChatModel(messages=iter(["done"]))
agent = create_agent(model, tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "B22?"}]})

It cannot be given tools: its bind_tools raises NotImplementedError, so an agent with any tool fails on its first call. The documentation's example passes tools=[], which is why it works there. Three lines fix it.

A fake that accepts tools

Subclass it and let bind_tools return the model unchanged.

python
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel


class ScriptedModel(GenericFakeChatModel):
    def bind_tools(self, tools, **kwargs):
        return self          # accept tools, keep the scripted replies

Scripting a tool call, then an answer

Script two replies: first an AIMessage that asks for the tool, then a final text answer.

python
from langchain.messages import AIMessage, ToolCall
from scripted import ScriptedModel

call = ToolCall(name="lookup_order", args={"order_id": "B22"}, id="call_1")
# first the model asks for the tool, then it gives a final answer
model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))

Running the scripted model through the agent

Run the scripted model through the agent and print every message.

Example
from langchain.messages import AIMessage, ToolCall
from scripted import ScriptedModel

call = ToolCall(name="lookup_order", args={"order_id": "B22"}, id="call_1")
model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))
result = create_agent(model, tools=[lookup_order]).invoke({"messages": [{"role": "user", "content": "B22?"}]})

for message in result["messages"]:
    print(f"{message.type:<5} {message.text or message.tool_calls[0]['args']}")

What the scripted run proves

  • The model's two replies were fixed in advance, so this run is the same every time.
  • The tool itself ran: the tool line is lookup_order's real answer, which is what the test checks.
  • The last line is "done", the second scripted reply, which ends the loop.

Tests with pytest

Wrap the checks in pytest functions. Each one builds its own agent or calls the router, then asserts on the result.

python
from langchain.agents import create_agent
from langchain.messages import AIMessage, ToolCall
from router import answer
from scripted import ScriptedModel
from tools import lookup_order


def test_unknown_order_is_reported():
    call = ToolCall(name="lookup_order", args={"order_id": "B22"}, id="call_1")
    model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))
    agent = create_agent(model, tools=[lookup_order])
    result = agent.invoke({"messages": [{"role": "user", "content": "Where is B22?"}]})
    assert result["messages"][2].text == "B22 is not an order we have."
python
def test_order_questions_go_to_the_orders_agent():
    assert answer("Where is A17?") == ("orders", "A17 shipped on 3 March.")


def test_uncovered_questions_are_refused():
    name, reply = answer("Can I pay with bitcoin?")
    assert name == "policies" and reply.startswith("Our policies do not cover that")

The first test uses the scripted model to check the tool's answer for an unknown order. The other two check lesson 31's router and lesson 28's refusal with the stand-in models, which decide the same way on every run. None of them needs a key or a network, so they can run on every change.

Example
pytest -q -p no:warnings test_desk.py

Real model vs scripted fake

Real modelScripted fake
Same result each runNoYes
Needs a key or networkYesNo
What it testsThe model tooYour tools, routing and refusals
Tool callsThe model decidesYou script them

When to test with a scripted fake

  • Checking a tool's output for a known input on every commit.
  • Proving routing and refusals behave without paying for a model.
Watch out. GenericFakeChatModel raises NotImplementedError from bind_tools, so an agent with any tool fails on the first call. Give it a subclass whose bind_tools returns itself before you add tools.
Try it yourself
  • Change the expected text in the first test and read how pytest reports the failure.
  • Add a test that answer("Is shipping free?") goes to the policies agent.
  • Script a model that asks for lookup_order twice and assert there are two tool messages.

Little by little, you're building something great.