Summary memory
Summary memory is a short-term memory technique that replaces the oldest turns of a conversation with a summary written by an LLM and sends that summary together with the most recent turns.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
Sliding window memory drops old turns, and the salary from turn 1 goes with them. Summary memory keeps the meaning of those turns in one paragraph. The price is an extra model call and whatever the paragraph leaves out.
This part of the video starts at 3:40:22. It reads the newspaper picture of 3_summary_memory.ipynb, the context window at turn 10, and the first point: the summary is written by the LLM, which is called abstractive summarisation.
The raw turns leave the prompt, not the system. Keep them in an archive, so that a summary can be checked against them and rebuilt.
How summary memory works
The notebook's picture is "a newspaper that rewrites yesterday's news into a single paragraph each morning, then uses the paragraph as the starting point for today's edition". Old turns are handed to an LLM with a summarising prompt. The paragraph it returns takes their place in the request, and the recent turns stay word for word. The model is not updated by any of this: the summary is ordinary text that your code stores and sends.
This is abstractive summarisation: the model writes new text that captures the content. The other kind, extractive summarisation, copies sentences out of the source. The video's prompt is "Summarise this financial advisory conversation, preserving all key facts: salary, expenses, assets, risk profile, and any advice given."
The summary is not a user message and not an assistant message. The notebook sends it as a second system message, right after the agent's instructions.
This part of the video starts at 3:45:14. It covers the two triggers, turn-based and token-based, and the main risk: a summary is lossy.
The "third compression cycle" in the clip is an illustration, not a rule. How soon a detail drops out depends on the prompt, the model and the conversation, and the drift run further down measures it on one example.
Triggering the summary by a threshold
Summarising after every message would mean one extra model call per turn. The video names two triggers instead: turn-based (summarise when the buffer holds more than N turns) and token-based (summarise when the tokens pass a budget). This lesson uses the turn trigger of the video's summary memory notebook; the token trigger is the next lesson.
The schedule can be worked out without a model. With the notebook's demo settings, the summary fires when 4 turns are held verbatim and keeps the last 2:
def schedule(total_turns, max_turns_before_summary, turns_to_keep_verbatim):
summarised_up_to, recent, calls = 0, [], 0
for turn in range(1, total_turns + 1):
summary = f"summary of turns 1 to {summarised_up_to}" if summarised_up_to else "no summary"
print(f" turn {turn:>2} request: {summary:<25} + verbatim turns {recent + [turn]}")
recent.append(turn) # the reply arrives, the turn is complete
if len(recent) >= max_turns_before_summary:
cut = len(recent) - turns_to_keep_verbatim
summarised_up_to, recent, calls = recent[cut - 1], recent[cut:], calls + 1
return calls
print("Summarise after 4 turns, keep 2 verbatim:")
calls = schedule(10, max_turns_before_summary=4, turns_to_keep_verbatim=2)
print(f" summariser calls in 10 turns: {calls}\n")
print("Summarise after 4 turns, keep 3 verbatim:")
calls = schedule(10, max_turns_before_summary=4, turns_to_keep_verbatim=3)
print(f" summariser calls in 10 turns: {calls}")Summarise after 4 turns, keep 2 verbatim: turn 1 request: no summary + verbatim turns [1] turn 2 request: no summary + verbatim turns [1, 2] turn 3 request: no summary + verbatim turns [1, 2, 3] turn 4 request: no summary + verbatim turns [1, 2, 3, 4] turn 5 request: summary of turns 1 to 2 + verbatim turns [3, 4, 5] turn 6 request: summary of turns 1 to 2 + verbatim turns [3, 4, 5, 6] turn 7 request: summary of turns 1 to 4 + verbatim turns [5, 6, 7] turn 8 request: summary of turns 1 to 4 + verbatim turns [5, 6, 7, 8] turn 9 request: summary of turns 1 to 6 + verbatim turns [7, 8, 9] turn 10 request: summary of turns 1 to 6 + verbatim turns [7, 8, 9, 10] summariser calls in 10 turns: 4 Summarise after 4 turns, keep 3 verbatim: turn 1 request: no summary + verbatim turns [1] turn 2 request: no summary + verbatim turns [1, 2] turn 3 request: no summary + verbatim turns [1, 2, 3] turn 4 request: no summary + verbatim turns [1, 2, 3, 4] turn 5 request: summary of turns 1 to 1 + verbatim turns [2, 3, 4, 5] turn 6 request: summary of turns 1 to 2 + verbatim turns [3, 4, 5, 6] turn 7 request: summary of turns 1 to 3 + verbatim turns [4, 5, 6, 7] turn 8 request: summary of turns 1 to 4 + verbatim turns [5, 6, 7, 8] turn 9 request: summary of turns 1 to 5 + verbatim turns [6, 7, 8, 9] turn 10 request: summary of turns 1 to 6 + verbatim turns [7, 8, 9, 10] summariser calls in 10 turns: 7
- Keep 2 of 4: each summary call folds 2 turns, so 10 turns need 4 calls, and a request holds 3 or 4 verbatim turns.
- Keep 3 of 4: each call folds 1 turn, so from turn 4 on there is a summariser call after every turn: 7 calls in 10 turns.
- The gap between the two numbers sets the cost. A trigger close to the number of kept turns means a summary call on almost every turn.
Lossy compression and drift
Every summary is lossy: the model decides what matters, and it can decide wrongly with nobody watching. The video's example is a user who says they are allergic to equity investments. A detail that is mentioned once may be kept in the first summary and missing a few summaries later. Each new summary is written from the previous one, so a change made once is carried into every later version. That slow movement away from what was said is called summary drift.
This part of the video starts at 3:48:21. It explains summaries of summaries, the four levels, and the FinCoach turn 6 with and without summary memory.
Abstractive and progressive are not two kinds of summary. Abstractive says how a summary is written. Progressive says that an old summary is fed back in and summarised again. The rolling summary of this lesson is both.
Progressive summarisation
In a long session the summary itself grows, so it gets summarised again into a shorter form. The video calls this progressive or hierarchical summarisation and lists four levels, each with less detail and fewer tokens than the one below it. The code in this lesson builds the step from level 0 to level 1.
Building the summary memory class
The notebook 3_summary_memory.ipynb summarises with gpt-4o-mini, a prompt that allows 150 words and max_tokens=300. Here the summariser is the same Groq model as the chat, the prompt keeps the notebook's list of facts in a shorter form with an 80-word limit so that the output stays readable, and the call sets reasoning_effort="low" with a larger max_tokens: gpt-oss checks a word limit in its hidden reasoning, which counts inside max_tokens, and a reply whose reasoning uses up the limit comes back empty.
The summariser prompt
The prompt names the kinds of fact that must be kept. A generic "summarise this" leaves that choice to the model.
FINCOACH_SUMMARY_PROMPT = """You are summarising a financial advisory conversation for FinCoach.
Your summary will be the only memory of these turns in the next API call.
Preserve these facts if mentioned: salary and expenses, existing investments with their
values and maturity dates, risk profile, goals, decisions made, advice FinCoach gave.
Write in third person, as one paragraph of at most 80 words.
Do not add facts that are not in the conversation."""Folding new turns into the old summary
When a summary already exists, it is sent to the summariser together with the new turns. This is what makes the summary progressive.
if existing_summary:
user_content = (f"EXISTING SUMMARY (from previous turns):\n{existing_summary}\n\n"
f"NEW TURNS TO ADD TO THE SUMMARY:\n{turns_text}\n\n"
"Produce an updated summary that incorporates both.")The trigger
After each assistant reply the class counts the verbatim turns. At the threshold it cuts the list in two: the old part goes to the summariser, the last turns stay.
if role == "assistant" and len(self.recent_turns) // 2 >= self.max_turns_before_summary:
keep = self.turns_to_keep_verbatim * 2
old, self.recent_turns = self.recent_turns[:-keep], self.recent_turns[-keep:]
self.current_summary = summarise_turns(old, self.current_summary)The request
The request is the system prompt, the summary as a second system message when there is one, and the verbatim turns.
def get_messages_for_api(self):
messages = [{"role": "system", "content": self.system_prompt}]
if self.current_summary:
messages.append({"role": "system", "content":
"CONVERSATION SUMMARY (earlier turns):\n" + self.current_summary})
return messages + self.recent_turnsRunning the six turns with a summary
These are the six turns that made the window forget in Sliding window memory, with the notebook's settings: summarise after 4 turns, keep 2. Each line says whether a summary was sent and whether the text 1,20,000 is in any verbatim turn. Groq's free tier limits the tokens per minute; if a run stops with a 429 rate limit error, wait a minute and run it again.
import os
import tiktoken
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"
TOKENISER = tiktoken.get_encoding("o200k_base")
FINCOACH_SYSTEM_PROMPT = """You are FinCoach, a personal financial advisor assistant.
You serve users in India who want guidance on savings, investments, budgeting, and financial planning.
Your principles:
- Always personalise advice using information the user has shared in this conversation.
- Be specific with numbers when the user has provided their financial details.
- Flag when you are making assumptions due to missing information.
- Keep responses concise: 3 to 5 sentences unless the user asks for detail.
- Never provide specific buy/sell recommendations on individual stocks.
- Always recommend consulting a SEBI-registered advisor for major financial decisions.
Use all context, including the conversation summary if provided, to give consistent, personalised advice."""
FINCOACH_SUMMARY_PROMPT = """You are summarising a financial advisory conversation for FinCoach.
Your summary will be the only memory of these turns in the next API call.
Preserve these facts if mentioned: salary and expenses, existing investments with their
values and maturity dates, risk profile, goals, decisions made, advice FinCoach gave.
Write in third person, as one paragraph of at most 80 words.
Do not add facts that are not in the conversation."""
def summarise_turns(turns_to_summarise, existing_summary=None):
turns_text = "\n".join(f"{m['role'].upper()}: {m['content']}" for m in turns_to_summarise)
if existing_summary:
user_content = (f"EXISTING SUMMARY (from previous turns):\n{existing_summary}\n\n"
f"NEW TURNS TO ADD TO THE SUMMARY:\n{turns_text}\n\n"
"Produce an updated summary that incorporates both.")
else:
user_content = f"CONVERSATION TURNS TO SUMMARISE:\n{turns_text}\n\nProduce a summary of these turns."
response = client.chat.completions.create(
model=MODEL, max_tokens=1000, temperature=0, reasoning_effort="low",
messages=[{"role": "system", "content": FINCOACH_SUMMARY_PROMPT},
{"role": "user", "content": user_content}])
summary = response.choices[0].message.content.strip()
original = sum(len(TOKENISER.encode(m["content"])) for m in turns_to_summarise)
print(f" [SUMMARISE] {len(turns_to_summarise)} messages ({original} tokens) "
f"-> summary ({len(TOKENISER.encode(summary))} tokens)")
return summary
class SummaryMemory:
def __init__(self, system_prompt, max_turns_before_summary=4, turns_to_keep_verbatim=2):
self.system_prompt = system_prompt
self.max_turns_before_summary = max_turns_before_summary
self.turns_to_keep_verbatim = turns_to_keep_verbatim
self.current_summary = None
self.recent_turns = [] # verbatim messages
self.archive = [] # every message, never sent
def add_message(self, role, content):
message = {"role": role, "content": content}
self.recent_turns.append(message)
self.archive.append(message)
if role == "assistant" and len(self.recent_turns) // 2 >= self.max_turns_before_summary:
keep = self.turns_to_keep_verbatim * 2
old, self.recent_turns = self.recent_turns[:-keep], self.recent_turns[-keep:]
self.current_summary = summarise_turns(old, self.current_summary)
def get_messages_for_api(self):
messages = [{"role": "system", "content": self.system_prompt}]
if self.current_summary:
messages.append({"role": "system", "content":
"CONVERSATION SUMMARY (earlier turns):\n" + self.current_summary})
return messages + self.recent_turns
def chat(memory, user_message):
memory.add_message("user", user_message)
request = memory.get_messages_for_api()
response = client.chat.completions.create(
model=MODEL, max_tokens=1024, temperature=0, messages=request)
reply = response.choices[0].message.content
verbatim = any("1,20,000" in m["content"] for m in memory.recent_turns)
print(f"[Turn {len(memory.archive) // 2 + 1}] messages sent: {len(request)} | "
f"prompt_tokens: {response.usage.prompt_tokens} | "
f"summary sent: {memory.current_summary is not None} | salary in a verbatim turn: {verbatim}")
memory.add_message("assistant", reply)
return reply
summary_memory = SummaryMemory(FINCOACH_SYSTEM_PROMPT, max_turns_before_summary=4, turns_to_keep_verbatim=2)
demo_turns = [
"Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000.",
"My monthly expenses are about ₹60,000 for rent, food, and transport.",
"I have an FD of ₹50,000 that matures in 3 months. I am risk-averse.",
"What is a Systematic Investment Plan and how does it work?",
"What are the different types of mutual funds available in India?",
"What is my exact monthly take-home salary that I told you at the start?",
]
for user_message in demo_turns:
reply = chat(summary_memory, user_message)
if user_message == demo_turns[-1]:
print("FinCoach, turn 6:", reply)
print()
print("Summary after turn 6:", summary_memory.current_summary)[Turn 1] messages sent: 2 | prompt_tokens: 229 | summary sent: False | salary in a verbatim turn: True [Turn 2] messages sent: 4 | prompt_tokens: 470 | summary sent: False | salary in a verbatim turn: True [Turn 3] messages sent: 6 | prompt_tokens: 742 | summary sent: False | salary in a verbatim turn: True [Turn 4] messages sent: 8 | prompt_tokens: 997 | summary sent: False | salary in a verbatim turn: True [SUMMARISE] 4 messages (489 tokens) -> summary (131 tokens) [Turn 5] messages sent: 7 | prompt_tokens: 857 | summary sent: True | salary in a verbatim turn: False [Turn 6] messages sent: 9 | prompt_tokens: 1125 | summary sent: True | salary in a verbatim turn: False [SUMMARISE] 4 messages (468 tokens) -> summary (151 tokens) FinCoach, turn 6: Your take‑home salary is **₹1,20,000 per month**. Summary after turn 6: Chiru earns ₹1,20,000 net monthly, spends ₹60,000 and has ₹60,000 discretionary income. He holds a ₹50,000 fixed deposit maturing in three months and is risk‑averse. FinCoach advised keeping about ₹30,000 in a short‑term FD (3‑6 months) for liquidity and ₹20,000 in a government‑backed debt fund or RD, while continuing to save ₹15‑20 k each month for a 3‑6‑month emergency fund (₹1.8‑3.6 L). After the cushion is built, allocate ₹20‑25 k to a conservative hybrid SIP (≈30 % equity, 70 % debt).
What the summary carried to turn 6
- Turns 1 to 4 run like a buffer: 2, 4, 6 and 8 messages,
prompt_tokensfrom 229 to 997, no summary. - After turn 4 the first cycle fires. Turns 1 and 2, 4 messages of 489 tokens by tiktoken, become a 131-token summary. Turn 5 then sends 7 messages and the API counts 857 prompt tokens, fewer than the 997 of turn 4.
- At turn 6 the salary is in no verbatim turn (the last column is False), a summary is sent, and the reply states ₹1,20,000 per month. The sliding window failed on this same question.
- After turn 6 the second cycle folds turns 3 and 4 into the old summary: 468 tokens of messages by tiktoken, a 151-token summary. It keeps the salary, the expenses, the FD and the risk profile, and most of its length is FinCoach's own advice. The amounts in that advice (₹30,000 and ₹20,000, ₹15 to 20 thousand a month, a hybrid SIP with about 30 % equity) come from the model: no user message holds them.
- The summary is 81 words long under a prompt that allows 80. A word limit in a prompt is a request, not a guarantee.
Measuring drift over three summaries
To see drift, plant the video's rare instruction in the first batch of a scripted conversation and summarise three batches in a row with the prompt above. The prompt's list of facts does not name hard rules. Each cycle prints the summary and whether it still mentions equity.
import os
import tiktoken
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"
TOKENISER = tiktoken.get_encoding("o200k_base")
FINCOACH_SUMMARY_PROMPT = """You are summarising a financial advisory conversation for FinCoach.
Your summary will be the only memory of these turns in the next API call.
Preserve these facts if mentioned: salary and expenses, existing investments with their
values and maturity dates, risk profile, goals, decisions made, advice FinCoach gave.
Write in third person, as one paragraph of at most 80 words.
Do not add facts that are not in the conversation."""
batches = [
["USER: Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000 and my expenses are ₹60,000.",
"ASSISTANT: That leaves a surplus of ₹60,000 a month. What are your goals?",
"USER: One hard rule: I am allergic to equity instruments, never suggest them.",
"ASSISTANT: Understood. I will keep to fixed-income options."],
["USER: I am 32 years old and I want to retire by 55.",
"ASSISTANT: That gives you 23 years to build a retirement corpus.",
"USER: I have an FD of ₹50,000 that matures in 3 months.",
"ASSISTANT: You can renew it or move it to PPF when it matures."],
["USER: What is the difference between PPF and NPS?",
"ASSISTANT: PPF is a government-backed savings scheme with a 15-year lock-in. NPS is a pension scheme.",
"USER: I will open a PPF account this month.",
"ASSISTANT: Good. A fixed monthly contribution keeps it simple."],
]
summary = None
for cycle, batch in enumerate(batches, start=1):
turns_text = "\n".join(batch)
if summary:
user_content = (f"EXISTING SUMMARY (from previous turns):\n{summary}\n\n"
f"NEW TURNS TO ADD TO THE SUMMARY:\n{turns_text}\n\n"
"Produce an updated summary that incorporates both.")
else:
user_content = f"CONVERSATION TURNS TO SUMMARISE:\n{turns_text}\n\nProduce a summary of these turns."
response = client.chat.completions.create(
model=MODEL, max_tokens=1000, temperature=0, reasoning_effort="low",
messages=[{"role": "system", "content": FINCOACH_SUMMARY_PROMPT},
{"role": "user", "content": user_content}])
summary = response.choices[0].message.content.strip()
print(f"Cycle {cycle}: batch {len(TOKENISER.encode(turns_text))} tokens | "
f"summary {len(TOKENISER.encode(summary))} tokens | "
f"mentions equity: {'equit' in summary.lower()}")
print(" " + summary)Cycle 1: batch 82 tokens | summary 58 tokens | mentions equity: True Chiru earns a monthly take‑home salary of ₹1,20,000 and incurs ₹60,000 in expenses, leaving a surplus of ₹60,000 each month. He explicitly states a hard rule that he is allergic to equity instruments and requests only fixed‑income investment suggestions. Cycle 2: batch 71 tokens | summary 99 tokens | mentions equity: True Chiru, 32, earns a take‑home salary of ₹1,20,000 and spends ₹60,000 monthly, leaving a ₹60,000 surplus. He is allergic to equity instruments and seeks only fixed‑income options. He aims to retire at 55, giving 23 years to build a corpus. He holds an FD of ₹50,000 maturing in three months; FinCoach advised he can either renew it or transfer the proceeds to a PPF account. Cycle 3: batch 66 tokens | summary 137 tokens | mentions equity: True Chiru, 32, earns ₹1,20,000 take‑home and spends ₹60,000 monthly, leaving a ₹60,000 surplus. He avoids equities, preferring fixed‑income. He aims to retire at 55 (23 years). He holds an FD of ₹50,000 maturing in three months; FinCoach suggested renewing it or moving the proceeds to a PPF account. Chiru asked about PPF vs NPS; FinCoach explained PPF is a 15‑year government‑backed savings scheme, while NPS is a pension scheme. Chiru decided to open a PPF account this month with regular monthly contributions.
What the three cycles kept
- All three summaries mention equity, so the keyword check prints True every time.
- The wording weakens. Cycle 1 writes "a hard rule that he is allergic to equity instruments", cycle 2 "is allergic to equity instruments", cycle 3 "He avoids equities". The user's "never suggest them" is a rule; by the third summary it reads as a preference. A keyword check does not see that, a reader does.
- The summary grows: 58, 99 and 137 tokens by tiktoken. In cycle 3 the summary is about twice the size of the 66-token batch it absorbed, and at 82 words it is over the 80-word limit.
- This is one run of one model. Another model or prompt can drop the rule earlier or keep it word for word, so measure it on your own conversations and name in the prompt the facts that must survive.
Summary memory vs sliding window memory
| Sliding window | Summary memory | |
|---|---|---|
| Old turns | Dropped from the request | Compressed into a paragraph |
| Early facts | Lost | Kept if the summary names them |
| Extra model calls | None | One per summary cycle |
| Exact wording of old turns | Gone | Gone: the summary is new text |
| Failure you can see | The agent asks again | The agent answers from a summary that changed a fact |
| Auditing | Compare with the archive | Compare the summary with the archive |
Where you use summary memory
- Long advisory or support sessions. Facts from the start (salary, risk profile, goals) have to reach turn 40.
- Coding and research agents. Long runs are compacted into a summary of what was decided and done, so the work can go on in a fresh context.
- Handing a conversation over. A session summary is what the next session or a human agent reads first.
Related
- Previous: Sliding window memory
- Next: Summary buffer memory
- Reference: Groq reasoning models
- In the schedule example, change the second call to
max_turns_before_summary=6, turns_to_keep_verbatim=2: the summariser is called 2 times in 10 turns, and the request of turn 10 holds six verbatim turns. - In the drift run, add "hard rules the user stated" to the list of facts in
FINCOACH_SUMMARY_PROMPT, run it again and compare how each cycle words the equity rule. - In the six-turn run, add
print(len(summary_memory.archive))at the end: it prints 12, since every message is still in the archive after two summary cycles.
This is what real progress feels like.