LangChainLangChain 1.4 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Summarization with SummarizationMiddleware

SummarizationMiddleware is a middleware that replaces older messages in a thread with a short summary, so a long conversation does not send every past message to the model on every call.

Last updated: 27 Sep, 2026 · LangChain 1.4

In lesson 13 each question added four messages to Ravi's thread, and all of them were sent with the next question. Summarizing needs a model, so this stand-in lists the orders it finds in the text it is given.

SummarizationMiddleware: trigger and keep

python
from langchain.agents.middleware import SummarizationMiddleware

summarize = SummarizationMiddleware(
    model,                          # the model that writes the summary
    trigger=("messages", 6),        # summarize once the thread reaches 6 messages
    keep=("messages", 2),           # leave the last 2 messages untouched
)

The summary model

The summary needs a model. This stand-in reads the text it is handed and lists the orders it can find, so the lesson runs without a real model. Start with the class header.

python
import re
from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage
from langchain_core.outputs import ChatGeneration, ChatResult

class SummaryModel(BaseChatModel):
    @property
    def _llm_type(self):
        return "summary"

Building the summary text

Its _generate pulls the order ids and statuses out of the newest text and returns them as one line.

python
    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        # find every "A17 shipped on 3 March" style phrase in the newest text
        found = re.findall(r"\b([A-Z]\d+) (shipped on \d+ \w+|waiting for stock)", messages[-1].text)
        summary = "; ".join(f"{order} {status}" for order, status in dict(found).items())
        return ChatResult(generations=[ChatGeneration(message=AIMessage(f"Orders discussed: {summary}."))])

The agent with summarization

Build the agent with the middleware and a checkpointer, so the thread is remembered across questions. trigger is when it summarizes; keep is what it leaves alone.

python
from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware
from langgraph.checkpoint.memory import InMemorySaver
from shop_model import ShopModel
from summary_model import SummaryModel
from tools import lookup_order

summarize = SummarizationMiddleware(SummaryModel(), trigger=("messages", 6), keep=("messages", 2))
agent = create_agent(ShopModel(), tools=[lookup_order], middleware=[summarize],
                     checkpointer=InMemorySaver())

Before each model call, the middleware counts the messages. Once there are at least six, the trigger, it asks SummaryModel to summarize all but the last two, the keep, and puts the summary in their place.

Three turns held to four messages

Ask three questions on one thread and count the messages after each.

Example
thread = {"configurable": {"thread_id": "ravi-1"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
    result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
    print(text, "->", len(result["messages"]), "messages")

The thread stays at four messages instead of growing to twelve. When the cut would separate a tool call from its result, the middleware moves it so the two stay together.

Reading the summary

Print the first message to see what the summary holds.

Example
thread = {"configurable": {"thread_id": "ravi-1"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
    result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
    print(text, "->", len(result["messages"]), "messages")

print(result["messages"][0].text)

The first message is now a human message holding the summary, which covers A17 and C40. The newest exchange about B22 is kept as it was.

Without a trigger

Leave out trigger and the middleware never fires, so the thread grows unchecked.

python
summarize = SummarizationMiddleware(SummaryModel())
Example
thread = {"configurable": {"thread_id": "ravi-1"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
    result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
    print(text, "->", len(result["messages"]), "messages")

With no trigger, the middleware never summarizes, and the thread grows by four each time. A trigger can also be a token count, ("tokens", 4000); a dictionary of conditions must all hold, and a list fires when any one does.

What the trigger did

  • The thread stayed at four messages. With the trigger, older turns were folded into one summary while the newest two were kept.
  • The summary is a human message at the front. It covered A17 and C40; B22 was still recent, so it was left as it was.
  • A tool call and its result stay together. The middleware moves the cut so a tool message is never split from its call.
  • No trigger, no summary. Without one, the thread grew by four every question, to twelve.

With a trigger vs without

trigger=('messages', 6)No trigger
When it summarizesOnce the thread reaches 6 messagesNever
Thread after 3 questions4 messages12 messages
Cost per later callSmall and steadyGrows every turn

Where summarization fits

  • A long support chat where every past message would otherwise be resent each turn.
  • Keeping token cost flat as a conversation runs on.
  • Holding the gist of earlier turns while keeping the latest exchange in full.
Watch out. With no trigger the middleware never summarizes and the thread keeps growing. Set a message count or token count so it fires.
Try it yourself
  • Set keep=("messages", 4) and check what the summary covers.
  • Use trigger=("messages", 10) and ask five questions.
  • Ask about B22 first and read the summary: why is B22 missing from it?

Every expert started right here.