Summarization with SummarizationMiddleware
SummarizationMiddleware is a middleware that replaces older messages in a thread with a short summary, so a long conversation does not send every past message to the model on every call.
Last updated: 27 Sep, 2026 · LangChain 1.4
In lesson 13 each question added four messages to Ravi's thread, and all of them were sent with the next question. Summarizing needs a model, so this stand-in lists the orders it finds in the text it is given.
SummarizationMiddleware: trigger and keep
from langchain.agents.middleware import SummarizationMiddleware
summarize = SummarizationMiddleware(
model, # the model that writes the summary
trigger=("messages", 6), # summarize once the thread reaches 6 messages
keep=("messages", 2), # leave the last 2 messages untouched
)The summary model
The summary needs a model. This stand-in reads the text it is handed and lists the orders it can find, so the lesson runs without a real model. Start with the class header.
import re
from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage
from langchain_core.outputs import ChatGeneration, ChatResult
class SummaryModel(BaseChatModel):
@property
def _llm_type(self):
return "summary"Building the summary text
Its _generate pulls the order ids and statuses out of the newest text and returns them as one line.
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
# find every "A17 shipped on 3 March" style phrase in the newest text
found = re.findall(r"\b([A-Z]\d+) (shipped on \d+ \w+|waiting for stock)", messages[-1].text)
summary = "; ".join(f"{order} {status}" for order, status in dict(found).items())
return ChatResult(generations=[ChatGeneration(message=AIMessage(f"Orders discussed: {summary}."))])The agent with summarization
Build the agent with the middleware and a checkpointer, so the thread is remembered across questions. trigger is when it summarizes; keep is what it leaves alone.
from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware
from langgraph.checkpoint.memory import InMemorySaver
from shop_model import ShopModel
from summary_model import SummaryModel
from tools import lookup_order
summarize = SummarizationMiddleware(SummaryModel(), trigger=("messages", 6), keep=("messages", 2))
agent = create_agent(ShopModel(), tools=[lookup_order], middleware=[summarize],
checkpointer=InMemorySaver())Before each model call, the middleware counts the messages. Once there are at least six, the trigger, it asks SummaryModel to summarize all but the last two, the keep, and puts the summary in their place.
Three turns held to four messages
Ask three questions on one thread and count the messages after each.
thread = {"configurable": {"thread_id": "ravi-1"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
print(text, "->", len(result["messages"]), "messages")The thread stays at four messages instead of growing to twelve. When the cut would separate a tool call from its result, the middleware moves it so the two stay together.
Reading the summary
Print the first message to see what the summary holds.
thread = {"configurable": {"thread_id": "ravi-1"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
print(text, "->", len(result["messages"]), "messages")
print(result["messages"][0].text)The first message is now a human message holding the summary, which covers A17 and C40. The newest exchange about B22 is kept as it was.
Without a trigger
Leave out trigger and the middleware never fires, so the thread grows unchecked.
summarize = SummarizationMiddleware(SummaryModel())thread = {"configurable": {"thread_id": "ravi-1"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
print(text, "->", len(result["messages"]), "messages")With no trigger, the middleware never summarizes, and the thread grows by four each time. A trigger can also be a token count, ("tokens", 4000); a dictionary of conditions must all hold, and a list fires when any one does.
What the trigger did
- The thread stayed at four messages. With the trigger, older turns were folded into one summary while the newest two were kept.
- The summary is a human message at the front. It covered A17 and C40; B22 was still recent, so it was left as it was.
- A tool call and its result stay together. The middleware moves the cut so a tool message is never split from its call.
- No trigger, no summary. Without one, the thread grew by four every question, to twelve.
With a trigger vs without
| trigger=('messages', 6) | No trigger | |
|---|---|---|
| When it summarizes | Once the thread reaches 6 messages | Never |
| Thread after 3 questions | 4 messages | 12 messages |
| Cost per later call | Small and steady | Grows every turn |
Where summarization fits
- A long support chat where every past message would otherwise be resent each turn.
- Keeping token cost flat as a conversation runs on.
- Holding the gist of earlier turns while keeping the latest exchange in full.
trigger the middleware never summarizes and the thread keeps growing. Set a message count or token count so it fires.Related
- Previous: Retries and a fallback model
- Next: More built-in middleware to reach for
- Reference: LangChain agent middleware
- Set
keep=("messages", 4)and check what the summary covers. - Use
trigger=("messages", 10)and ask five questions. - Ask about B22 first and read the summary: why is B22 missing from it?
Every expert started right here.