LangMemLangMem 0.0.30 · LangGraph 1.2 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Running summaries

A running summary is short-term memory for one long conversation: LangMem's summarize_messages replaces the oldest messages with a model-written summary once the chat passes a token budget, and remembers which messages it already covered.

Last updated: 30 Sep, 2026 · LangMem 0.0.30

Everything so far was long-term memory, carried between conversations. A single conversation has its own problem: every turn adds tokens, and a long chat eventually costs too much or no longer fits the model. Summarizing the older part keeps the recent turns word for word and the rest as a short summary.

Summary buffer memory · from the AI Security Course · 236:30 to 241:15
The video's notebook builds this loop from scratch with the OpenAI SDK: a buffer, a threshold check, and a summarizer call that merges evicted messages into the old summary. LangMem's summarize_messages is that loop as one function, and RunningSummary is the stored summary.

Recent turns verbatim, older turns summarized

The clip's diagram: new messages go into a buffer; when the buffer passes a threshold, the oldest messages are evicted and a summarizer model merges them with the old summary; the prompt is the summary plus the recent messages plus the new one. Like remembering the last few sentences of a long phone call word for word, and only the gist of what was said fifteen minutes ago. It suits support agents that need precise recent context and awareness of the history.

Syntax:

python
result = summarize_messages(messages, running_summary=previous, model=model,
                            max_tokens=120, max_summary_tokens=80,
                            token_counter=count_tokens_approximately)
result.messages         # what to send to the model now
result.running_summary  # store it and pass it back next turn

A conversation long enough to summarize

Messages need ids: the running summary records which ids it covered.

python
turns = [
    ("Hi, I'm Asha. Order A-1001 arrived with a cracked screen.", "Sorry, Asha. I have opened a replacement for A-1001."),
    ("Can the replacement come by Friday? I travel on Saturday.", "Yes, the replacement ships Thursday by express courier."),
    ("Please send it to my office address in Pune.", "Done. The replacement goes to your Pune office."),
    ("Will I get back the shipping I paid?", "Yes, the 99 rupee shipping fee is refunded to your card."),
    ("What happens to the broken phone?", "A courier collects it when the replacement arrives."),
]
messages = []
for i, (question, answer) in enumerate(turns):
    messages += [HumanMessage(question, id=f"q{i}"), AIMessage(answer, id=f"a{i}")]

A budget and a counter

max_tokens is the budget for the messages returned; summarizing starts when the conversation is over it. count_tokens_approximately is LangChain's quick estimate, so no tokenizer has to be downloaded.

python
result = summarize_messages(
    messages,
    running_summary=None,
    model=model,
    max_tokens=120,
    max_summary_tokens=80,
    token_counter=count_tokens_approximately,
)

Summarizing Asha's conversation

ExampleAPI key
from langchain.chat_models import init_chat_model
from langchain_core.messages import AIMessage, HumanMessage
from langchain_core.messages.utils import count_tokens_approximately
from langmem.short_term import summarize_messages

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
turns = [
    ("Hi, I'm Asha. Order A-1001 arrived with a cracked screen.", "Sorry, Asha. I have opened a replacement for A-1001."),
    ("Can the replacement come by Friday? I travel on Saturday.", "Yes, the replacement ships Thursday by express courier."),
    ("Please send it to my office address in Pune.", "Done. The replacement goes to your Pune office."),
    ("Will I get back the shipping I paid?", "Yes, the 99 rupee shipping fee is refunded to your card."),
    ("What happens to the broken phone?", "A courier collects it when the replacement arrives."),
]
messages = []
for i, (question, answer) in enumerate(turns):
    messages += [HumanMessage(question, id=f"q{i}"), AIMessage(answer, id=f"a{i}")]

print("tokens before:", count_tokens_approximately(messages))
result = summarize_messages(
    messages,
    running_summary=None,
    model=model,
    max_tokens=120,
    max_summary_tokens=80,
    token_counter=count_tokens_approximately,
)
print("tokens after:", count_tokens_approximately(result.messages))
for message in result.messages:
    print(f"--- {message.type}")
    print(message.content)
print("summarized ids:", sorted(result.running_summary.summarized_message_ids))

What was kept and what was lost

  • The last turn stayed word for word: the broken-phone question and its answer, as the clip describes.
  • The first turn is missing from the summary, though its ids q0 and a0 are listed as summarized. When the messages to summarize are longer than max_tokens, LangMem sends the summarizer only the last max_tokens of them. Asha's name, the order number and the cracked screen were cut before the model saw them.
  • The summary came out longer than the text it replaced: 213 tokens after, 172 before. max_summary_tokens is a budget LangMem plans with, not a limit it sends to the model; the reference says to pass model.bind(max_tokens=...) to enforce it.
  • It also lists "Asked for a summary of the conversation" as a user request. That line is LangMem's own summary prompt, which the model summarized along with the chat.

A tight budget vs a roomy one

The next turn passes the running summary back, so messages it already covered are not summarized again. This program runs the same next turn with two budgets and asks the question the first summary can no longer answer. max_tokens_before_summary keeps the trigger at 120 tokens while max_tokens=400 lets the summarizer see every older message. The system message keeps the answer to what the conversation says.

ExampleAPI key
from langchain.chat_models import init_chat_model
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage
from langchain_core.messages.utils import count_tokens_approximately
from langmem.short_term import summarize_messages

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
turns = [
    ("Hi, I'm Asha. Order A-1001 arrived with a cracked screen.", "Sorry, Asha. I have opened a replacement for A-1001."),
    ("Can the replacement come by Friday? I travel on Saturday.", "Yes, the replacement ships Thursday by express courier."),
    ("Please send it to my office address in Pune.", "Done. The replacement goes to your Pune office."),
    ("Will I get back the shipping I paid?", "Yes, the 99 rupee shipping fee is refunded to your card."),
    ("What happens to the broken phone?", "A courier collects it when the replacement arrives."),
]
messages = []
for i, (question, answer) in enumerate(turns):
    messages += [HumanMessage(question, id=f"q{i}"), AIMessage(answer, id=f"a{i}")]
rule = SystemMessage("Answer in one short sentence using only facts stated in this conversation. If they are not stated, say you do not know.")
budgets = {
    "tight: max_tokens=120": dict(max_tokens=120, max_summary_tokens=80),
    "roomy: max_tokens=400, summarize at 120": dict(max_tokens=400, max_tokens_before_summary=120, max_summary_tokens=100),
}

for label, budget in budgets.items():
    first = summarize_messages(messages, running_summary=None, model=model, token_counter=count_tokens_approximately, **budget)
    turn = messages + [HumanMessage("Which order was this about, and what was wrong with it?", id="q5")]
    result = summarize_messages(turn, running_summary=first.running_summary, model=model, token_counter=count_tokens_approximately, **budget)
    answer = model.invoke([rule, *result.messages])
    print(label)
    print("  summary reused:", result.running_summary is first.running_summary)
    print("  answer:", answer.text)

Reading the two answers

  • tight: "I do not know." The order number never reached the summary, so the model, told to use only the conversation, could not answer. Without that rule it could have guessed.
  • roomy: "order A-1001, which arrived with a cracked screen". With max_tokens=400 the summarizer saw every older message, so the facts survived.
  • summary reused: True in both. One new question still fit the budget, so summarize_messages returned the same running summary object instead of calling the model again.

Running summary vs long-term memory

Running summaryStore memories
ScopeOne conversation threadEvery conversation of a user
Kept inYour graph state or variablesThe LangGraph store
WrittenWhen the thread passes the token budgetBy a manager or the agent's tools
Searched from other threadsNoYes

When to summarize

  • Support chats that run for dozens of turns.
  • Agents with many tool calls, where tool results fill the context fast.
  • Any model with a small context window.
Watch out. Summaries lose details: the tight run above dropped the order number, and a summary of a summary loses rare but important facts first. Anything the next conversation must know, like an allergy or an order number, belongs in the store as a memory, not only in the summary.
Try it yourself
  • In the first example, pass max_tokens=400 and max_tokens_before_summary=120, and check whether A-1001 is in the summary. (With max_tokens=400 alone the chat is under budget, nothing is summarized and running_summary is None.)
  • Pass max_tokens_before_summary=200 with max_tokens=120 and compare when summarizing starts.
  • Remove the ids from the messages and read the error.

Every expert started right here.