LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Testing the desk agent

A desk test sends one message on a fresh thread and checks the reply, one test for each promise the desk makes. All five run in under a second and need no key, so you can run them after every change.

Last updated: 27 Sep, 2026 · LangChain 1.4

Every test sends one message on a fresh thread and looks at the reply. The desk runs on DeskModel, so each reply is exact and the file needs no key. The tests run with pytest; install it if you skipped the testing lesson.

pip install "pytest==9.1.1"

A pytest test for the desk

python
# a test sends one message and checks the reply
def test_shipped_order():
    assert "shipped" in ask("ravi", "Where is A17?")

# pytest runs every test_ function in the file
# pytest -q test_desk.py
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
shop_model.py
import re

from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult


class ShopModel(BaseChatModel):
    tools: list = []

    @property
    def _llm_type(self):
        return "shop"

    def bind_tools(self, tools, **kwargs):
        return self.model_copy(update={"tools": tools})   # a copy holding the tools

    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        message = self.decide(messages)                   # the reply comes from decide
        return ChatResult(generations=[ChatGeneration(message=message)])

    def decide(self, messages):
        results = []                                # the tool results at the end
        for m in reversed(messages):
            if not isinstance(m, ToolMessage):
                break
            results.insert(0, m.text)
        if results:                                 # results are back: answer with them
            return AIMessage(" ".join(results))
        text = messages[-1].text
        orders = re.findall(r"\b[A-Z]\d+\b", text)
        tool = "refund_order" if "refund" in text.lower() else "lookup_order"
        if orders and tool in [t.name for t in self.tools]:   # one call per order id
            calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
                     for o in orders]
            return AIMessage("", tool_calls=calls)
        if orders:                                  # that tool is not bound
            return AIMessage(f"I have no way to look up {orders[0]} yet.")
        return AIMessage("Hello. Which order is this about?")
desk_tools.py
from dataclasses import dataclass

from langchain.tools import ToolRuntime, tool

ORDERS = {"A17": ("ravi", "shipped on 3 March"), "C40": ("mei", "waiting for stock")}


@dataclass
class Customer:
    name: str


@tool
def lookup_order(order_id: str, runtime: ToolRuntime[Customer]) -> str:
    """Look up one of the customer's orders by its id, such as A17."""
    owner, status = ORDERS.get(order_id, (None, None))
    if owner != runtime.context.name:
        return f"{order_id} is not one of your orders."
    return f"{order_id} {status}."


@tool
def refund_order(order_id: str, runtime: ToolRuntime[Customer]) -> str:
    """Refund one of the customer's orders in full. This cannot be undone."""
    owner, _ = ORDERS.get(order_id, (None, None))
    if owner != runtime.context.name:
        return f"{order_id} is not one of your orders, so it cannot be refunded."
    return f"Refunded {order_id}."
policies.py
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter

POLICIES = {
    "refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
                  "\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
    "shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
                   "\n\nExpress shipping arrives the next working day and costs 9 euros.",
    "accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}

docs = [Document(page_content=text, metadata={"source": name}) for name, text in POLICIES.items()]
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs)
word_embeddings.py
import re
import zlib

from langchain_core.embeddings import Embeddings

COMMON = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i",
          "if", "is", "it", "my", "of", "on", "the", "to", "what", "with", "you", "your"}


class WordEmbeddings(Embeddings):
    def embed_query(self, text):
        vector = [0.0] * 256
        for word in re.findall(r"[a-z]+", text.lower()):
            if word not in COMMON:
                vector[zlib.crc32(word.rstrip("s").encode()) % 256] += 1.0
        return vector

    def embed_documents(self, texts):
        return [self.embed_query(text) for text in texts]
search.py
from langchain.tools import tool
from langchain_core.vectorstores import InMemoryVectorStore
from policies import chunks
from word_embeddings import WordEmbeddings

store = InMemoryVectorStore(WordEmbeddings())
store.add_documents(chunks)


@tool
def search_policies(query: str) -> str:
    """Search the shop's policies on refunds, shipping and accounts.
    Pass the customer's question, word for word, as the query."""
    found = [doc for doc, score in store.similarity_search_with_score(query, k=2) if score >= 0.3]
    if not found:
        return "No policy covers this."
    return "\n".join(f"[{doc.metadata['source']}] {doc.page_content}" for doc in found)
desk_model.py
import re

from langchain.messages import AIMessage
from shop_model import ShopModel


class DeskModel(ShopModel):
    def decide(self, messages):
        last = messages[-1]
        if last.type == "tool" and last.text == "No policy covers this.":
            return AIMessage("Our policies do not cover that. A person will reply.")
        if last.type == "tool" or re.findall(r"\b[A-Z]\d+\b", last.text):
            return super().decide(messages)
        query = {"name": "search_policies", "args": {"query": last.text}, "id": "call_p"}
        return AIMessage("", tool_calls=[query])
password_check.py
from langchain.agents.middleware import before_agent
from langchain.messages import AIMessage


@before_agent(can_jump_to=["end"])
def no_passwords(state, runtime):
    if "password" in state["messages"][-1].text.lower():
        answer = AIMessage("I cannot help with passwords. Please use the reset link.")
        return {"messages": [answer], "jump_to": "end"}
desk.py
from langchain.agents import create_agent
from langchain.agents.middleware import HumanInTheLoopMiddleware, ModelCallLimitMiddleware, PIIMiddleware
from langgraph.checkpoint.memory import InMemorySaver
from desk_model import DeskModel
from password_check import no_passwords
from search import search_policies
from desk_tools import Customer, lookup_order, refund_order

agent = create_agent(
    DeskModel(),
    system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned. If a tool says an order is not the customer's, say exactly that. Add nothing the tools did not say.",
    tools=[lookup_order, refund_order, search_policies],
    context_schema=Customer,
    middleware=[
        no_passwords,
        PIIMiddleware("credit_card", strategy="mask"),
        ModelCallLimitMiddleware(run_limit=6),
        HumanInTheLoopMiddleware(interrupt_on={"refund_order": True}),
    ],
    checkpointer=InMemorySaver(),
)
chat.py
from langgraph.types import Command
from desk import Customer, agent


def say(who, text, thread):
    config = {"configurable": {"thread_id": thread}}
    result = agent.invoke({"messages": [{"role": "user", "content": text}]}, config,
                          context=Customer(who), version="v2")
    if result.interrupts:
        print(f"{who}: {text}\n  paused for approval: {result.interrupts[0].value['action_requests'][0]['args']}")
        result = agent.invoke(Command(resume={"decisions": [{"type": "approve"}]}), config,
                              context=Customer(who), version="v2")
        text = "(approved)"
    print(f"{who}: {text}\n  desk: {result.value['messages'][-1].text}")

The ask helper

The helper does the repetitive part. It opens a fresh thread for every call, with a random id from uuid4(), so no test leaks into the next, even when the checkpointer keeps its threads in a file between runs. It sends the message as a customer and returns the desk's last reply. The helper and the five tests below go in one file, test_desk.py.

python
from uuid import uuid4

from desk import agent
from desk_tools import Customer


def ask(who, text):
    config = {"configurable": {"thread_id": str(uuid4())}}   # a new thread for every call
    result = agent.invoke({"messages": [{"role": "user", "content": text}]}, config,
                          context=Customer(who), version="v2")
    return result.value["messages"][-1].text

Testing the order tools

Each test is one line of intent. The first two go through the order tools: an owner sees their order, and someone else is refused.

python
def test_an_order_is_answered_to_its_owner():
    assert "shipped" in ask("ravi", "Where is A17?")


def test_another_customers_order_is_refused():
    assert "not one of your orders" in ask("mei", "Where is A17?")

Testing policy answers and refusals

The next three cover the policy answer, the refusal for a question no policy covers, and the guardrail that ends the run before the model is reached.

python
def test_a_policy_question_is_answered_from_the_documents():
    assert "5 working days" in ask("ravi", "How long does a refund take?")


def test_an_uncovered_question_is_refused():
    assert "A person will reply" in ask("ravi", "Do you sell gift cards?")


def test_passwords_never_reach_the_model():
    assert "reset link" in ask("ravi", "What is my password?")

Run the tests

One command runs the file. Each dot is a test that sent its message and saw the reply it expected.

Example
pytest -q -p no:warnings test_desk.py

What the five tests cover

  • Each of the five dots is a test that found the words it expected in the desk's reply.
  • Each promise the desk makes has its own test: the owner check, the policy answer, the refusal for an uncovered question, and the password guardrail.
The support desk, one message at a time
Middleware, in orderRavi askswith his card numberno_passwords, ends the runPII masks the cardcall limit caps the loopapproval holds refundscheckpointer saves the threadShopModelthe stand-inlookup_orderowner checkedThe answeror an honest refusalrefund_orderowner checked toosearch_policiesrefunds, shipping, accounts
Hover or tap a piece to see what it is and which lesson built it.
Follow a message

Pick one to watch it run, step by step.

Each test walks one path through the drawing: two through the order tools, two through the policy search, one answered and one refused, and one that the guardrail ends before the model is reached at all.

When to run the whole file

  • Running the whole file before you change a line, so a red dot tells you which promise you broke.
  • Checking your own logic before a hosted model is added, so a later difference points at the model, not the code.
Watch out. These tests exercise the stand-in model, not a hosted one. Keep them fast and key-free so they run on every change; test against a real provider separately, where the wording of an answer varies from one run to the next.
Try it yourself
  • Add a returns policy to POLICIES and a test that asks about returns.
  • Add a test that Mei asking about C40, her own order, gets "waiting for stock".
  • Replace DeskModel() in desk.py with init_chat_model("groq:openai/gpt-oss-120b", temperature=0), run the tests again, and see which assertion fails and why.

Little by little, you're building something great.