LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Summarization with SummarizationMiddleware

SummarizationMiddleware is a middleware that replaces older messages in a thread with a short summary, so a long conversation does not send every past message to the model on every call.

Last updated: 27 Sep, 2026 · LangChain 1.4

Message-based summarization: six questions on one thread · from the Updated LangChain Version V1 Crash Course · 123:16 to 129:13

Six questions, then a summary

A conversation grows by two messages a turn, a question and an answer, and left alone every turn sends the whole history to the model. The video's agent has a model, no tools and an InMemorySaver checkpointer, and gets SummarizationMiddleware through its middleware parameter. That parameter is a list, so you can add as many middleware as you need, one after another. Inside SummarizationMiddleware, three parameters decide when the summary happens and what it keeps:

python
from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware
from langgraph.checkpoint.memory import InMemorySaver
from langchain_core.messages import HumanMessage

agent=create_agent(
    model="groq:openai/gpt-oss-120b",
    checkpointer=InMemorySaver(),
    middleware=[
        SummarizationMiddleware(
            model="groq:openai/gpt-oss-120b",
            trigger=("messages",10),
            keep=("messages",4)
        )
    ]
)

model writes the summary. It runs again every time the history passes the limit, so a model that costs less is a good choice. trigger=("messages", 10) means: summarize once the history reaches 10 messages, a small number for the demo where a real chatbot would use a larger one. keep=("messages", 4) means: keep the 4 most recent messages as they are. The config's thread_id, test-1, identifies one user, and every call with it continues the same thread. Each question goes in as a HumanMessage, and the loop prints how many messages the thread holds:

ExampleAPI keyFrom the video, run on Groq
config = {"configurable": {"thread_id": "test-1"}}

questions = [
    "What is 2+2?",
    "What is 10*5?",
    "What is 100/4?",
    "What is 15-7?",
    "What is 3*3?",
    "What is 4*4?",
]

for q in questions:
    response = agent.invoke({"messages": [HumanMessage(content=q)]}, config)
    print(f"{q:<16} messages: {len(response['messages'])}")

state = agent.get_state(config).values["messages"]
print("FIRST:", state[0].text[:300].replace("\n", " "))

Two, four, six, eight, ten: the history grows by two each turn. At ten the trigger fires, and the next count is six, not twelve: the older turns were folded into one summary message, now the first message on the thread. It starts "Here is a summary of the conversation to date:" and lists each question with its answer; the recent messages and the new turn follow it. The video uses gpt-4o-mini to answer and to summarize; the run above uses Groq's openai/gpt-oss-120b for both.

Triggers by tokens and by fraction · from the Updated LangChain Version V1 Crash Course · 131:14 to 136:30

Triggers by tokens or by fraction

Counting messages is one trigger. The video's second agent, with a search_hotels tool, counts tokens instead: summarize past 550 tokens and keep the most recent 200. It estimates tokens at four characters each and asks for hotels in one city after another; the count climbs to 149, 302 and 456, then drops to 396 once a summary replaces the older turns. The third trigger is a fraction of the model's context window: 0.005 is half a percent of it. The small numbers make each trigger fire quickly in a demo; real limits run to thousands of tokens:

python
SummarizationMiddleware(model="groq:openai/gpt-oss-120b", trigger=("tokens", 550), keep=("tokens", 200))

# a fraction of the model's context window
SummarizationMiddleware(model="groq:openai/gpt-oss-120b", trigger=("fraction", 0.005), keep=("fraction", 0.002))

Now the same idea on the shop's support thread. In Short-term memory with a checkpointer every question added four messages to Ravi's thread, and all of them went to the model with the next question. Here both the agent and the summary use the Groq model.

The shop agent with summarization

This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Start the file with it.

python
from langchain.tools import tool

ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}


@tool
def lookup_order(order_id: str) -> str:
    """Look up an order's shipping status by its id, such as A17."""
    status = ORDERS.get(order_id)
    return f"{order_id} {status}." if status else f"{order_id} is not an order we have."

Build the agent with the middleware and a checkpointer, so the thread is remembered across questions. The same model answers the customer and writes the summary. Here it summarizes once the thread reaches six messages and keeps the last two.

python
from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware
from langchain.chat_models import init_chat_model
from langgraph.checkpoint.memory import InMemorySaver

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)  # uses your GROQ_API_KEY
summarize = SummarizationMiddleware(model, trigger=("messages", 6), keep=("messages", 2))
agent = create_agent(model, tools=[lookup_order], middleware=[summarize],
                     system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned.",
                     checkpointer=InMemorySaver())

Three turns held to four messages

Ask three questions on one thread and count the messages after each.

ExampleAPI key
thread = {"configurable": {"thread_id": "ravi-1"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
    result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
    print(text, "->", len(result["messages"]), "messages")

The thread stays at four messages instead of growing to twelve. When the cut would separate a tool call from its result, the middleware moves it so the two stay together.

Reading the summary

Ask the same three questions on a new thread, ravi-2, and this time print the first message afterwards to see what the summary holds.

ExampleAPI key
thread = {"configurable": {"thread_id": "ravi-2"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
    result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
    print(text, "->", len(result["messages"]), "messages")

print(result["messages"][0].text)

The first message is now a human message holding the summary. The summary was written right before the last model call, after the lookup for B22 had been asked for, so it covers A17, C40 and the B22 question itself. The two kept messages are the B22 tool request and its result, and the model's final reply follows them: four in all. The long layout, with sections such as SESSION INTENT, SUMMARY and NEXT STEPS, comes from the middleware's default summary prompt; pass summary_prompt= to write your own.

Without a trigger

Leave out trigger and the middleware never fires, so the thread grows unchecked. Build the agent again with the bare middleware and ask the same three questions on a fresh thread.

python
summarize = SummarizationMiddleware(model)          # no trigger
agent = create_agent(model, tools=[lookup_order], middleware=[summarize],
                     system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned.",
                     checkpointer=InMemorySaver())
ExampleAPI key
thread = {"configurable": {"thread_id": "ravi-3"}}
for text in ["Where is A17?", "And C40?", "And B22?"]:
    result = agent.invoke({"messages": [{"role": "user", "content": text}]}, thread)
    print(text, "->", len(result["messages"]), "messages")

With no trigger, the middleware never summarizes, and the thread grows by four each time.

What the trigger did

  • The thread stayed at four messages. With the trigger, older messages were folded into one summary while the newest two were kept.
  • The summary is a human message at the front. It covered A17, C40 and the B22 question; the B22 tool request and its result were kept as they were, followed by the reply.
  • A tool call and its result stay together. The middleware moves the cut so a tool message is never split from its call.
  • No trigger, no summary. Without one, the thread grew by four every question, to twelve.

With a trigger vs without

trigger=('messages', 6)No trigger
When it summarizesOnce the thread reaches 6 messagesNever
Thread after 3 questions4 messages12 messages
Cost per later callSmall and steadyGrows every turn

Where summarization fits

  • A long support chat where every past message would otherwise be resent each turn.
  • Keeping token cost flat as a conversation runs on.
  • Holding the gist of earlier turns while keeping the latest exchange in full.
Try it yourself
  • Set keep=("messages", 4) and check what the summary covers.
  • Use trigger=("messages", 10) and ask five questions.
  • Set trigger=("messages", 8) and check after which question the first summary appears.

Every expert started right here.