Running summaries
A running summary is short-term memory for one long conversation: LangMem's summarize_messages replaces the oldest messages with a model-written summary once the chat passes a token budget, and remembers which messages it already covered.
Last updated: 30 Sep, 2026 · LangMem 0.0.30
Everything so far was long-term memory, carried between conversations. A single conversation has its own problem: every turn adds tokens, and a long chat eventually costs too much or no longer fits the model. Summarizing the older part keeps the recent turns word for word and the rest as a short summary.
summarize_messages is that loop as one function, and RunningSummary is the stored summary.Recent turns verbatim, older turns summarized
The clip's diagram: new messages go into a buffer; when the buffer passes a threshold, the oldest messages are evicted and a summarizer model merges them with the old summary; the prompt is the summary plus the recent messages plus the new one. Like remembering the last few sentences of a long phone call word for word, and only the gist of what was said fifteen minutes ago. It suits support agents that need precise recent context and awareness of the history.
Syntax:
result = summarize_messages(messages, running_summary=previous, model=model,
max_tokens=120, max_summary_tokens=80,
token_counter=count_tokens_approximately)
result.messages # what to send to the model now
result.running_summary # store it and pass it back next turnA conversation long enough to summarize
Messages need ids: the running summary records which ids it covered.
turns = [
("Hi, I'm Asha. Order A-1001 arrived with a cracked screen.", "Sorry, Asha. I have opened a replacement for A-1001."),
("Can the replacement come by Friday? I travel on Saturday.", "Yes, the replacement ships Thursday by express courier."),
("Please send it to my office address in Pune.", "Done. The replacement goes to your Pune office."),
("Will I get back the shipping I paid?", "Yes, the 99 rupee shipping fee is refunded to your card."),
("What happens to the broken phone?", "A courier collects it when the replacement arrives."),
]
messages = []
for i, (question, answer) in enumerate(turns):
messages += [HumanMessage(question, id=f"q{i}"), AIMessage(answer, id=f"a{i}")]A budget and a counter
max_tokens is the budget for the messages returned; summarizing starts when the conversation is over it. count_tokens_approximately is LangChain's quick estimate, so no tokenizer has to be downloaded.
result = summarize_messages(
messages,
running_summary=None,
model=model,
max_tokens=120,
max_summary_tokens=80,
token_counter=count_tokens_approximately,
)Summarizing Asha's conversation
from langchain.chat_models import init_chat_model
from langchain_core.messages import AIMessage, HumanMessage
from langchain_core.messages.utils import count_tokens_approximately
from langmem.short_term import summarize_messages
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
turns = [
("Hi, I'm Asha. Order A-1001 arrived with a cracked screen.", "Sorry, Asha. I have opened a replacement for A-1001."),
("Can the replacement come by Friday? I travel on Saturday.", "Yes, the replacement ships Thursday by express courier."),
("Please send it to my office address in Pune.", "Done. The replacement goes to your Pune office."),
("Will I get back the shipping I paid?", "Yes, the 99 rupee shipping fee is refunded to your card."),
("What happens to the broken phone?", "A courier collects it when the replacement arrives."),
]
messages = []
for i, (question, answer) in enumerate(turns):
messages += [HumanMessage(question, id=f"q{i}"), AIMessage(answer, id=f"a{i}")]
print("tokens before:", count_tokens_approximately(messages))
result = summarize_messages(
messages,
running_summary=None,
model=model,
max_tokens=120,
max_summary_tokens=80,
token_counter=count_tokens_approximately,
)
print("tokens after:", count_tokens_approximately(result.messages))
for message in result.messages:
print(f"--- {message.type}")
print(message.content)
print("summarized ids:", sorted(result.running_summary.summarized_message_ids))tokens before: 172 tokens after: 213 --- system Summary of the conversation so far: **Conversation Summary** - **User Request:** Asked if the replacement could arrive by Friday (they travel on Saturday). **Assistant Reply:** Confirmed the replacement will ship Thursday via express courier. - **User Request:** Asked to send the replacement to their office address in Pune. **Assistant Reply:** Confirmed the replacement will be sent to the Pune office. - **User Question:** Inquired whether they will get back the shipping fee they paid. **Assistant Reply:** Confirmed the ₹99 shipping fee will be refunded to their card. - **User Request:** Asked for a summary of the conversation. **Assistant Action:** Provided this concise summary. --- human What happens to the broken phone? --- ai A courier collects it when the replacement arrives. summarized ids: ['a0', 'a1', 'a2', 'a3', 'q0', 'q1', 'q2', 'q3']
What was kept and what was lost
- The last turn stayed word for word: the broken-phone question and its answer, as the clip describes.
- The first turn is missing from the summary, though its ids
q0anda0are listed as summarized. When the messages to summarize are longer thanmax_tokens, LangMem sends the summarizer only the lastmax_tokensof them. Asha's name, the order number and the cracked screen were cut before the model saw them. - The summary came out longer than the text it replaced: 213 tokens after, 172 before.
max_summary_tokensis a budget LangMem plans with, not a limit it sends to the model; the reference says to passmodel.bind(max_tokens=...)to enforce it. - It also lists "Asked for a summary of the conversation" as a user request. That line is LangMem's own summary prompt, which the model summarized along with the chat.
A tight budget vs a roomy one
The next turn passes the running summary back, so messages it already covered are not summarized again. This program runs the same next turn with two budgets and asks the question the first summary can no longer answer. max_tokens_before_summary keeps the trigger at 120 tokens while max_tokens=400 lets the summarizer see every older message. The system message keeps the answer to what the conversation says.
from langchain.chat_models import init_chat_model
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage
from langchain_core.messages.utils import count_tokens_approximately
from langmem.short_term import summarize_messages
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
turns = [
("Hi, I'm Asha. Order A-1001 arrived with a cracked screen.", "Sorry, Asha. I have opened a replacement for A-1001."),
("Can the replacement come by Friday? I travel on Saturday.", "Yes, the replacement ships Thursday by express courier."),
("Please send it to my office address in Pune.", "Done. The replacement goes to your Pune office."),
("Will I get back the shipping I paid?", "Yes, the 99 rupee shipping fee is refunded to your card."),
("What happens to the broken phone?", "A courier collects it when the replacement arrives."),
]
messages = []
for i, (question, answer) in enumerate(turns):
messages += [HumanMessage(question, id=f"q{i}"), AIMessage(answer, id=f"a{i}")]
rule = SystemMessage("Answer in one short sentence using only facts stated in this conversation. If they are not stated, say you do not know.")
budgets = {
"tight: max_tokens=120": dict(max_tokens=120, max_summary_tokens=80),
"roomy: max_tokens=400, summarize at 120": dict(max_tokens=400, max_tokens_before_summary=120, max_summary_tokens=100),
}
for label, budget in budgets.items():
first = summarize_messages(messages, running_summary=None, model=model, token_counter=count_tokens_approximately, **budget)
turn = messages + [HumanMessage("Which order was this about, and what was wrong with it?", id="q5")]
result = summarize_messages(turn, running_summary=first.running_summary, model=model, token_counter=count_tokens_approximately, **budget)
answer = model.invoke([rule, *result.messages])
print(label)
print(" summary reused:", result.running_summary is first.running_summary)
print(" answer:", answer.text)tight: max_tokens=120 summary reused: True answer: I do not know. roomy: max_tokens=400, summarize at 120 summary reused: True answer: It was order A‑1001, which arrived with a cracked screen.
Reading the two answers
- tight: "I do not know." The order number never reached the summary, so the model, told to use only the conversation, could not answer. Without that rule it could have guessed.
- roomy: "order A-1001, which arrived with a cracked screen". With
max_tokens=400the summarizer saw every older message, so the facts survived. - summary reused: True in both. One new question still fit the budget, so
summarize_messagesreturned the same running summary object instead of calling the model again.
Running summary vs long-term memory
| Running summary | Store memories | |
|---|---|---|
| Scope | One conversation thread | Every conversation of a user |
| Kept in | Your graph state or variables | The LangGraph store |
| Written | When the thread passes the token budget | By a manager or the agent's tools |
| Searched from other threads | No | Yes |
When to summarize
- Support chats that run for dozens of turns.
- Agents with many tool calls, where tool results fill the context fast.
- Any model with a small context window.
Related
- Previous: Background memory
- Next: Prompt optimization
- Reference: How to manage long context with summarization
- In the first example, pass
max_tokens=400andmax_tokens_before_summary=120, and check whether A-1001 is in the summary. (Withmax_tokens=400alone the chat is under budget, nothing is summarized andrunning_summaryisNone.) - Pass
max_tokens_before_summary=200withmax_tokens=120and compare when summarizing starts. - Remove the ids from the messages and read the error.
Every expert started right here.