LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Middleware: code around the model

Middleware is code that runs at fixed points in the agent loop, such as before or after every model call, and can read the state, change it, or end the run.

Last updated: 27 Sep, 2026 · LangChain 1.4

What middleware is, the airport example and hooks · from the Updated LangChain Version V1 Crash Course · 110:23 to 115:54

Hooks inside the agent's loop

Middleware is a way to control more tightly what happens inside an agent. It is useful for:

  • Tracking what the agent does, with logging, analytics and debugging.
  • Transforming prompts, choosing tools and formatting output.
  • Adding retries, fallbacks and early stops.
  • Applying rate limits, guardrails and PII detection.

Think of an airport. Before you reach your gate you pass a security check, where your luggage is inspected, then immigration, where your passport is checked, then the boarding pass check. Each stop is a middleware, one, two and three, each running its own check before you move on. An agent without middleware is the plain ReAct loop: the request goes to the model, the model decides whether a tool is needed, the tool runs and gives its result back, and the model answers. Middleware adds trigger points to that loop, called hooks: before the agent, before the model, around tool calls, after the model and after the agent. At each one your code can do something, such as logging or summarization.

Where each middleware hook runs: before_agent and after_agent once per request, before_model and after_model around each model call, the wrap hooks around a single model or tool call.
Where each hook runs

before_agent and after_agent run once per request. before_model and after_model run around every model call, so they fire again on each turn of the loop. The wrap hooks surround a single model or tool call and can change it or run it again. LangChain ships ready-made middleware for common jobs, such as summarization, human approval and PII redaction, and you can write your own.

The video explains the hooks and then moves straight on to a ready-made middleware, summarization. This lesson writes two small hooks of its own first, so you can watch exactly when each one runs; the ready-made ones come in later lessons.

The before_model and after_model hooks

python
from langchain.agents.middleware import before_model, after_model

@before_model              # runs before each model call
def hook(state, runtime):  # the current state and the runtime
    ...                    # return nothing to leave the state as it is

agent = create_agent(model, tools=[...], middleware=[hook])

The before and after hooks

Let's write two small hooks that only print, so we can watch when they run. A decorated function becomes middleware; it receives the current state, with its messages, and the runtime.

python
from langchain.agents.middleware import after_model, before_model

@before_model
def count(state, runtime):
    print("before the model:", len(state["messages"]), "messages")

@after_model
def report(state, runtime):
    last = state["messages"][-1]
    print("after the model: ", last.text or last.tool_calls[0]["name"])

The agent with the hooks

This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Start the file with it.

python
from langchain.tools import tool

ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}


@tool
def lookup_order(order_id: str) -> str:
    """Look up an order's shipping status by its id, such as A17."""
    status = ORDERS.get(order_id)
    return f"{order_id} {status}." if status else f"{order_id} is not an order we have."

Both hooks around a lookup

Attach both hooks to the agent with middleware= and ask one question. There are two model calls, so each hook runs twice.

ExampleAPI key
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
agent = create_agent(model, tools=[lookup_order], middleware=[count, report],
                     system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned.")

agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})

The hooks as steps in the loop

Streaming the same agent shows the hooks as steps of their own, named after the function and the hook.

ExampleAPI key
for step in agent.stream({"messages": [{"role": "user", "content": "Where is A17?"}]}, stream_mode="updates"):
    print("step:", list(step))

What the hooks printed

  • Two model calls, so each hook ran twice: the first saw one message and the model asked for lookup_order; the second saw three and answered.
  • A decorated function becomes middleware. It receives the state and the runtime; returning nothing leaves the state unchanged.
  • Streaming shows the hooks as steps of their own, between the model and tools steps.

before_model vs after_model

before_modelafter_model
RunsBefore the model callAfter the model call
SeesThe messages going inThe message that came back
Typical useTrim or check the inputLog or inspect the reply
How oftenOnce per model callOnce per model call

Where middleware hooks fit

  • Logging or measuring every model call in one place.
  • Checking or trimming the messages before they reach the model.
Watch out. A hook that returns a state update changes the conversation for every later step. Return nothing when you mean to observe, and return an update only when you mean to change what the model sees.
Try it yourself
  • Add an @after_agent hook that prints the number of messages at the end.
  • Return {"messages": []} from count and see whether the conversation changes.
  • Ask a question with no order in it and count how many times each hook runs.

This is what real progress feels like.