Summarization with SummarizationMiddleware
SummarizationMiddleware is a middleware that replaces older messages in a thread with a short summary, so a long conversation does not send every past message to the model on every call.
Last updated: 27 Sep, 2026 · LangChain 1.4
Six questions, then a summary
A conversation grows by two messages a turn, a question and an answer, and left alone every turn sends the whole history to the model. The video's agent has a model, no tools and an InMemorySaver checkpointer, and gets SummarizationMiddleware through its middleware parameter. That parameter is a list, so you can add as many middleware as you need, one after another. Inside SummarizationMiddleware, three parameters decide when the summary happens and what it keeps:
from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware
from langgraph.checkpoint.memory import InMemorySaver
from langchain_core.messages import HumanMessage
agent=create_agent(
model="groq:openai/gpt-oss-120b",
checkpointer=InMemorySaver(),
middleware=[
SummarizationMiddleware(
model="groq:openai/gpt-oss-120b",
trigger=("messages",10),
keep=("messages",4)
)
]
)model writes the summary. It runs again every time the history passes the limit, so a model that costs less is a good choice. trigger=("messages", 10) means: summarize once the history reaches 10 messages, a small number for the demo where a real chatbot would use a larger one. keep=("messages", 4) means: keep the 4 most recent messages as they are. The config's thread_id, test-1, identifies one user, and every call with it continues the same thread. Each question goes in as a HumanMessage, and the loop prints how many messages the thread holds:
config = {"configurable": {"thread_id": "test-1"}}
questions = [
"What is 2+2?",
"What is 10*5?",
"What is 100/4?",
"What is 15-7?",
"What is 3*3?",
"What is 4*4?",
]
for q in questions:
response = agent.invoke({"messages": [HumanMessage(content=q)]}, config)
print(f"{q:<16} messages: {len(response['messages'])}")
state = agent.get_state(config).values["messages"]
print("FIRST:", state[0].text[:300].replace("\n", " "))What is 2+2? messages: 2 What is 10*5? messages: 4 What is 100/4? messages: 6 What is 15-7? messages: 8 What is 3*3? messages: 10 What is 4*4? messages: 6 FIRST: Here is a summary of the conversation to date: ## SESSION INTENT Provide correct answers to simple arithmetic questions posed by the user. ## SUMMARY - User asked: “What is 2+2?” → AI answered: 4. - User asked: “What is 10*5?” → AI answered: 50. - User asked: “What is 100/4?” → AI answered: 25
Two, four, six, eight, ten: the history grows by two each turn. At ten the trigger fires, and the next count is six, not twelve: the older turns were folded into one summary message, now the first message on the thread. It starts "Here is a summary of the conversation to date:" and lists each question with its answer; the recent messages and the new turn follow it. The video uses gpt-4o-mini to answer and to summarize; the run above uses Groq's openai/gpt-oss-120b for both.
Triggers by tokens or by fraction
Counting messages is one trigger. The video's second agent, with a search_hotels tool, counts tokens instead: summarize past 550 tokens and keep the most recent 200. It estimates tokens at four characters each and asks for hotels in one city after another; the count climbs to 149, 302 and 456, then drops to 396 once a summary replaces the older turns. The third trigger is a fraction of the model's context window: 0.005 is half a percent of it. The small numbers make each trigger fire quickly in a demo; real limits run to thousands of tokens:
SummarizationMiddleware(model="groq:openai/gpt-oss-120b", trigger=("tokens", 550), keep=("tokens", 200))
# a fraction of the model's context window
SummarizationMiddleware(model="groq:openai/gpt-oss-120b", trigger=("fraction", 0.005), keep=("fraction", 0.002))Now the same idea on the shop's support thread. In Short-term memory with a checkpointer every question added four messages to Ravi's thread, and all of them went to the model with the next question. Here both the agent and the summary use the Groq model.
The shop agent with summarization
This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Start the file with it.
from langchain.tools import tool
ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
status = ORDERS.get(order_id)
return f"{order_id} {status}." if status else f"{order_id} is not an order we have."Build the agent with the middleware and a checkpointer, so the thread is remembered across questions. The same model answers the customer and writes the summary. Here it summarizes once the thread reaches six messages and keeps the last two.
from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware
from langchain.chat_models import init_chat_model
from langgraph.checkpoint.memory import InMemorySaver
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEY
summarize = SummarizationMiddleware(model, trigger=("messages", 6), keep=("messages", 2))
agent = create_agent(model, tools=[lookup_order], middleware=[summarize],
system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned.",
checkpointer=InMemorySaver())Three turns held to four messages
Ask three questions on one thread and count the messages after each.
thread = {"configurable": {"thread_id": "ravi-1"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
print(text, "->", len(result["messages"]), "messages")Where is A17? -> 4 messages And C40? -> 4 messages And B22? -> 4 messages
The thread stays at four messages instead of growing to twelve. When the cut would separate a tool call from its result, the middleware moves it so the two stay together.
Reading the summary
Ask the same three questions on a new thread, ravi-2, and this time print the first message afterwards to see what the summary holds.
thread = {"configurable": {"thread_id": "ravi-2"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
print(text, "->", len(result["messages"]), "messages")
print(result["messages"][0].text)Where is A17? -> 4 messages And C40? -> 4 messages And B22? -> 4 messages Here is a summary of the conversation to date: ## SESSION INTENT The user is requesting the current status/location of specific orders: A17 (already answered), C40 (answered), and B22 (pending). ## SUMMARY - Order **A17**: Previously looked up; result was “A17 shipped on 3 March.” - Order **C40**: Looked up via `lookup_order` tool; result was “C40 waiting for stock.” This information was communicated to the user. - Order **B22**: User has just asked for its status; no lookup has been performed yet. ## ARTIFACTS - Tool call `lookup_order` (ID: fc_1c99de13-135b-4edf-a3e0-4c8833f5efcf) with `order_id: "A17"` – returned “A17 shipped on 3 March.” - Tool call `lookup_order` (ID: fc_b088fd11-6728-4f1a-8912-948cc35941ad) with `order_id: "C40"` – returned “C40 waiting for stock.” ## NEXT STEPS 1. Perform a `lookup_order` tool call for `order_id: "B22"`. 2. Relay the returned status/location of order B22 to the user.
The first message is now a human message holding the summary. The summary was written right before the last model call, after the lookup for B22 had been asked for, so it covers A17, C40 and the B22 question itself. The two kept messages are the B22 tool request and its result, and the model's final reply follows them: four in all. The long layout, with sections such as SESSION INTENT, SUMMARY and NEXT STEPS, comes from the middleware's default summary prompt; pass summary_prompt= to write your own.
Without a trigger
Leave out trigger and the middleware never fires, so the thread grows unchecked. Build the agent again with the bare middleware and ask the same three questions on a fresh thread.
summarize = SummarizationMiddleware(model) # no trigger
agent = create_agent(model, tools=[lookup_order], middleware=[summarize],
system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned.",
checkpointer=InMemorySaver())
thread = {"configurable": {"thread_id": "ravi-3"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
print(text, "->", len(result["messages"]), "messages")Where is A17? -> 4 messages And C40? -> 8 messages And B22? -> 12 messages
With no trigger, the middleware never summarizes, and the thread grows by four each time.
What the trigger did
- The thread stayed at four messages. With the trigger, older messages were folded into one summary while the newest two were kept.
- The summary is a human message at the front. It covered A17, C40 and the B22 question; the B22 tool request and its result were kept as they were, followed by the reply.
- A tool call and its result stay together. The middleware moves the cut so a tool message is never split from its call.
- No trigger, no summary. Without one, the thread grew by four every question, to twelve.
With a trigger vs without
| trigger=('messages', 6) | No trigger | |
|---|---|---|
| When it summarizes | Once the thread reaches 6 messages | Never |
| Thread after 3 questions | 4 messages | 12 messages |
| Cost per later call | Small and steady | Grows every turn |
Where summarization fits
- A long support chat where every past message would otherwise be resent each turn.
- Keeping token cost flat as a conversation runs on.
- Holding the gist of earlier turns while keeping the latest exchange in full.
Related
- Previous: Retries and a fallback model
- Next: More built-in middleware to reach for
- Reference: LangChain agent middleware
- Set
keep=("messages", 4)and check what the summary covers. - Use
trigger=("messages", 10)and ask five questions. - Set
trigger=("messages", 8)and check after which question the first summary appears.
Every expert started right here.