Changing the call: wrap_model_call
wrap_model_call is a middleware hook that runs around each model call: it receives the request, can change it, then hands it to a handler that sends it to the model.
Last updated: 27 Sep, 2026 · LangChain 1.4
The state hooks from Middleware: code around the model see the state. A wrap hook sees the request about to go to the model, with its messages, tools, system prompt and model, and a handler that sends it on. Calling handler(request) once is the normal path; not calling it, or calling it twice, is how a hook skips or retries a call.
The wrap_model_call hook
from langchain.agents.middleware import wrap_model_call
@wrap_model_call
def hook(request, handler): # request: the call about to go to the model
# inspect or change request here
return handler(request) # call the model once, return its resultA prompt from the context
Start with a hook that sets the system prompt. @dynamic_prompt turns a function into a wrap hook whose return value becomes the system prompt. Customer is the context class from Runtime context: who is asking.
from dataclasses import dataclass
from langchain.agents.middleware import dynamic_prompt, wrap_model_call
@dataclass
class Customer:
name: str # the customer this run is for
@dynamic_prompt
def with_name(request):
# read the name from the run's context and build the prompt
return f"You help {request.runtime.context.name}, a customer of a small online shop."A wrapper that prints the prompt
A second hook prints the system prompt the model is about to get, then passes the request on unchanged.
@wrap_model_call
def show_prompt(request, handler):
print("system prompt:", request.system_prompt) # what the model will receive
return handler(request) # send it on, unchangedThe agent with both middleware
This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Start the file with it.
from langchain.tools import tool
ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
status = ORDERS.get(order_id)
return f"{order_id} {status}." if status else f"{order_id} is not an order we have."Build the agent with both hooks. Their order in the list is the layering: with_name is first, so it wraps show_prompt.
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEY
agent = create_agent(model, tools=[lookup_order], context_schema=Customer,
middleware=[with_name, show_prompt])The prompt the model receives
Run it for Ravi and read what the model is handed.
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]}, context=Customer("ravi"))system prompt: You help ravi, a customer of a small online shop. system prompt: You help ravi, a customer of a small online shop.
Two model calls, and each got a system prompt written for Ravi. with_name comes first in the list, so it is the outer layer: it set the prompt before show_prompt saw the request.
A different model per request
wrap_model_call can also swap the model itself. request.override returns a copy of the request with one thing changed. Here a message with no order id goes to openai/gpt-oss-20b, a smaller and cheaper model on Groq, while order questions stay on the larger one.
import re
from langchain.agents.middleware import wrap_model_call
from langchain.chat_models import init_chat_model
small = init_chat_model("groq:openai/gpt-oss-20b", temperature=0) # a smaller, cheaper model
@wrap_model_call
def pick_model(request, handler):
if not re.search(r"\b[A-Z]\d+\b", request.messages[-1].text): # no order id
request = request.override(model=small) # swap the model
return handler(request)
Build the agent with this one hook, then send it two messages: small talk, and a message with an order id. Each reply carries the name of the model that wrote it in response_metadata, so print that too.
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEY
agent = create_agent(model, tools=[lookup_order], middleware=[pick_model],
system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned.")for text in ["Hello", "Where is A17?"]:
reply = agent.invoke({"messages": [{"role": "user", "content": text}]})["messages"][-1]
print(reply.response_metadata["model_name"], "->", reply.text)openai/gpt-oss-20b -> Hello! How can I help you today? openai/gpt-oss-120b -> A17 was shipped on 3 March.
Trimming the conversation on every call
trim_messages from Trimming and removing messages shortens a list. Inside a wrap hook it shortens what the model is sent on every call, while the thread keeps everything. start_on="human" makes the kept part begin with a user message, which providers expect. This agent has no tools, so each question is one model call.
from langchain.agents import create_agent
from langchain.agents.middleware import wrap_model_call
from langchain.chat_models import init_chat_model
from langchain_core.messages.utils import trim_messages
from langgraph.checkpoint.memory import InMemorySaver
@wrap_model_call
def last_three(request, handler):
kept = trim_messages(request.messages, strategy="last", token_counter=len,
max_tokens=3, start_on="human")
print(f"thread holds {len(request.messages)}, model sees {len(kept)}")
return handler(request.override(messages=kept))
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEY
agent = create_agent(model, middleware=[last_three], checkpointer=InMemorySaver(),
system_prompt="You are the support assistant for a small online shop. Answer in one short sentence, using only what this conversation says.")
thread = {"configurable": {"thread_id": "ravi-trim"}}
for text in ["My name is Ravi.", "I want to ask about an order.", "What is my name?"]:
result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
print("ai:", result["messages"][-1].text)thread holds 1, model sees 1 ai: Nice to meet you, Ravi. thread holds 3, model sees 3 ai: Sure, what would you like to know about your order? thread holds 5, model sees 3 ai: I don’t know your name from this conversation.
On the third question the thread holds five messages, but the model is sent only the last three, and the one with Ravi's name is not among them. It answers that it does not know, which is the cost of trimming: whatever falls outside the window is gone for the model, even though the checkpointer still has it.
What the wrappers did
- Two prints in the first run. The agent called the model twice, once to ask for the tool and once to reply, and
show_promptran around each call. - The prompt was set before it was seen.
with_nameis first in the list, so it is the outer layer; it wrote Ravi's prompt beforeshow_promptprinted it. request.overridechanges one thing. In the second example the greeting had no order id, sopick_modelsent it to the small model; the message with A17 kept the larger model and ran the tool.- This is model routing. A cheap model handles small talk while a larger one does the real work, and the choice is made per request.
- Trimming happens per call.
trim_messagesin a wrap hook shortened what the model saw on the third question, while the checkpointer still held every message.
wrap_model_call vs the state hooks
| State hooks | wrap_model_call | |
|---|---|---|
| What it sees | The agent state | The request about to go to the model |
| Can change | State keys it returns | The request, via request.override |
| Controls the call | No | Yes: call handler once, not at all, or again |
| Use for | Reading or editing state | Editing the prompt, model or messages; retrying; skipping |
Where wrap_model_call fits
- Setting a system prompt from who the request is for, as
with_namedoes. - Sending small talk to a cheaper model and real work to a larger one.
- Logging or checking every request before it reaches the model.
- Keeping a long conversation short for the model with
trim_messages, while the thread keeps every message.
handler(request) and return its result. Forget to call it and the model never runs; return nothing and the agent has no reply to work with.Related
- Previous: The order middleware runs in
- Next: Tool errors, caught and retried
- Reference: LangChain agent middleware
- Change
with_nameto add the number of messages inrequest.messagesto the prompt. - Make
show_promptskip the model by returning without callinghandler, and read the error. - Swap the order of
with_nameandshow_promptand compare what is printed.
Little by little, you're building something great.