LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Testing an agent

Testing an agent is checking its tools, routing and refusals with a scripted model in place of a real one, so every test runs the same way each time.

Last updated: 27 Sep, 2026 · LangChain 1.4

Tests for code, evaluations for answers

There are two different things to check in an LLM app. Tests check your code: that a tool returns the right thing, that the router sends a question to the right place, that a failure is handled. They use fixed inputs and exact expected outputs, and they run without a model. Evaluations check the model's answers, which are never exactly the same twice; The desk on a hosted model runs a small one. This lesson is about the first kind: tests that pin down your own code with a scripted stand-in model, so they are exact and free to run on every change.

LangChain ships GenericFakeChatModel for this: it returns the replies you give it, one per call, including tool calls. The tests run with pytest.

Scripting replies with GenericFakeChatModel

python
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel

model = GenericFakeChatModel(messages=iter([reply_1, reply_2]))   # one reply per call

When the fake model gets a tool

This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. It lives in orders.py, the file Subagents as tools started, and the tests below import it from there.

python
from langchain.tools import tool

ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}


@tool
def lookup_order(order_id: str) -> str:
    """Look up an order's shipping status by its id, such as A17."""
    status = ORDERS.get(order_id)
    return f"{order_id} {status}." if status else f"{order_id} is not an order we have."

Give the fake a tool and it fails on the first call. Here is the error before the fix.

Example
from orders import lookup_order
from langchain.agents import create_agent
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel

model = GenericFakeChatModel(messages=iter(["done"]))
agent = create_agent(model, tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "B22?"}]})

It cannot be given tools: its bind_tools raises NotImplementedError, so an agent with any tool fails on its first call. It only works in an agent built with tools=[]. Three lines fix it.

A fake that accepts tools

Subclass it and let bind_tools return the model unchanged. Save it as scripted.py.

python
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel


class ScriptedModel(GenericFakeChatModel):
    def bind_tools(self, tools, **kwargs):
        return self          # accept tools, keep the scripted replies

Scripting a tool call, then an answer

Script two replies: first an AIMessage that asks for the tool, then a final text answer.

python
from orders import lookup_order
from langchain.messages import AIMessage, ToolCall
from scripted import ScriptedModel

call = ToolCall(name="lookup_order", args={"order_id": "B22"}, id="call_1")
# first the model asks for the tool, then it gives a final answer
model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))

Running the scripted model through the agent

Run the scripted model through an agent with the tools lesson's lookup_order and print every message.

Example
from langchain.agents import create_agent
from langchain.messages import AIMessage, ToolCall
from scripted import ScriptedModel

call = ToolCall(name="lookup_order", args={"order_id": "B22"}, id="call_1")
model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))
result = create_agent(model, tools=[lookup_order]).invoke({"messages": [{"role": "user", "content": "B22?"}]})

for message in result["messages"]:
    print(f"{message.type:<5} {message.text or message.tool_calls[0]['args']}")

What the scripted run proves

  • The model's two replies were fixed in advance, so this run is the same every time.
  • The tool itself ran: the tool line is lookup_order's real answer, which is what the test checks.
  • The last line is "done", the second scripted reply, which ends the loop.

Tests with pytest

Install pytest, the test runner these tests use.

pip install "pytest==9.1.1"

A scripted specialist

The router lesson's specialists are real Groq agents, and a test cannot assert on a real model's wording. So the tests build their own stand-in specialist: an agent on a ScriptedModel that asks for one tool call and then says "done". Add it to the end of scripted.py.

python
from langchain.agents import create_agent
from langchain.messages import AIMessage, ToolCall


def scripted_agent(tool, args):
    """An agent whose model asks for one tool call, then says "done"."""
    call = ToolCall(name=tool.name, args=args, id="call_1")
    model = ScriptedModel(messages=iter([AIMessage("", tool_calls=[call]), "done"]))
    return create_agent(model, tools=[tool])
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
policies.py
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter

POLICIES = {
    "refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
                  "\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
    "shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
                   "\n\nExpress shipping arrives the next working day and costs 9 euros.",
    "accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}

docs = [Document(page_content=text, metadata={"source": name}) for name, text in POLICIES.items()]
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs)
word_embeddings.py
import re
import zlib

from langchain_core.embeddings import Embeddings

COMMON = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i",
          "if", "is", "it", "my", "of", "on", "the", "to", "what", "with", "you", "your"}


class WordEmbeddings(Embeddings):
    def embed_query(self, text):
        vector = [0.0] * 256
        for word in re.findall(r"[a-z]+", text.lower()):
            if word not in COMMON:
                vector[zlib.crc32(word.rstrip("s").encode()) % 256] += 1.0
        return vector

    def embed_documents(self, texts):
        return [self.embed_query(text) for text in texts]
search.py
from langchain.tools import tool
from langchain_core.vectorstores import InMemoryVectorStore
from policies import chunks
from word_embeddings import WordEmbeddings

store = InMemoryVectorStore(WordEmbeddings())
store.add_documents(chunks)


@tool
def search_policies(query: str) -> str:
    """Search the shop's policies on refunds, shipping and accounts.
    Pass the customer's question, word for word, as the query."""
    found = [doc for doc, score in store.similarity_search_with_score(query, k=2) if score >= 0.3]
    if not found:
        return "No policy covers this."
    return "\n".join(f"[{doc.metadata['source']}] {doc.page_content}" for doc in found)
router.py
import re


def route(question):
    if re.search(r"\b[A-Z]\d+\b", question):   # an order id, like A17
        return "orders"
    return "policies"                           # everything else


def answer(question, agents):
    name = route(question)                      # pick one specialist
    result = agents[name].invoke({"messages": [{"role": "user", "content": question}]})
    return name, result["messages"][-1].text    # the question went straight through

The test file

Wrap the checks in pytest functions in test_router.py. It needs orders.py, router.py and search.py next to it, and search.py needs policies.py and word_embeddings.py: the files from the subagents, router, retrieval-tool, documents and embeddings lessons. Each one builds what it needs, then asserts on the result. ask sends one question and returns the messages.

python
from orders import lookup_order
from router import route
from scripted import scripted_agent
from search import search_policies

def ask(agent, question):
    return agent.invoke({"messages": [{"role": "user", "content": question}]})["messages"]


def tool_reply(messages):
    return next(m.text for m in messages if m.type == "tool")


def test_unknown_order_is_reported():
    messages = ask(scripted_agent(lookup_order, {"order_id": "B22"}), "Where is B22?")
    assert tool_reply(messages) == "B22 is not an order we have."   # the real tool's reply
python
def test_order_questions_go_to_the_orders_agent():
    assert route("Where is A17?") == "orders"
    assert route("Is shipping free?") == "policies"


def test_uncovered_questions_are_refused():
    question = "Can I pay with bitcoin?"
    messages = ask(scripted_agent(search_policies, {"query": question}), question)
    assert tool_reply(messages) == "No policy covers this."         # the score cut held

The first and third tests run a scripted model instead of a real one: it asks for one tool call and says "done", and the assertion reads what the real tool returned, the unknown-order message and the retrieval lesson's "No policy covers this.". The second checks the router lesson's route, which is plain Python and needs no model at all. None of them needs a key or a network, so they can run on every change.

Run the file. -q keeps the report short, and -p no:warnings turns off pytest's warnings summary, which would otherwise list deprecation notices from the installed libraries.

Example
pytest -q -p no:warnings test_router.py

Real model vs scripted fake

Real modelScripted fake
Same result each runNoYes
Needs a key or networkYesNo
What it testsThe model tooYour tools, routing and refusals
Tool callsThe model decidesYou script them

When to test with a scripted fake

  • Checking a tool's output for a known input on every commit.
  • Proving routing and refusals behave without paying for a model.
Watch out. GenericFakeChatModel raises NotImplementedError from bind_tools, so an agent with any tool fails on the first call. Give it a subclass whose bind_tools returns itself before you add tools.
Try it yourself
  • Change the expected text in the first test and read how pytest reports the failure.
  • Add a test that route("How do I reset my password?") returns "policies".
  • Script a model that asks for lookup_order twice and assert there are two tool messages.

Little by little, you're building something great.