Summary buffer memory
Summary buffer memory is a short-term memory technique that keeps the most recent messages word for word within a token budget and folds older messages into a running summary.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
Summary memory decides when to summarise by counting turns. A turn can be three tokens or five hundred, so the size of the request is still hard to predict. Summary buffer memory uses the unit you are billed in: tokens.
This part of the video starts at 3:58:23. It reads the opening of 4_summary_buffer_memory.ipynb: recent messages word for word, older history as a summary, the phone-call picture, and how the technique combines a buffer with a summary.
Both parts, the summary and the recent messages, sit in the prompt of one conversation, so both are short-term memory. Long-term memory, which survives the conversation, starts with Vector store memory.
How summary buffer memory works
The notebook's picture is a long phone call with a friend: "You recall the last few sentences almost word-for-word." Of the earlier parts you keep the gist, not the exact words. The memory has two regions in the same way: a buffer of recent messages kept as they were written, and a summary of everything older. Both regions are text in the request; the model itself is not changed.
- A new message is appended to the buffer.
- When the buffer would pass its token budget, the oldest messages are removed from it.
- The removed messages are merged into the running summary by one summariser call.
- The next prompt is assembled as system prompt, summary, then the buffer.
The video's architecture diagram labels the buffer "last K messages". In the notebook's code the buffer is capped in tokens, not in messages, and that is what the class below does.
Why the budget is in tokens
The notebook's reason is short: "A turn where the user pastes a long document costs 500 tokens. A turn where they say "ok" costs 3 tokens." A budget in turns treats the two alike. A budget in tokens follows what the request costs.
Choosing max_buffer_tokens
A small budget overflows often, and every overflow is a summariser call that can lose a detail. A large budget rarely overflows and sends more tokens on every call. The notebook estimates this before any model call. The version here runs its ten user messages through the overflow rule of the class below, assuming every reply is 100 tokens long. It counts overflows only; the size of the summary is not simulated.
import tiktoken
import matplotlib.pyplot as plt
TOKENISER = tiktoken.get_encoding("o200k_base")
analysis_turns = [
"Hi, my monthly salary is ₹1,20,000 and expenses are ₹60,000.",
"I'm 32, risk-averse, and want to retire at 55.",
"I have an FD of ₹50,000 maturing next quarter.",
"What is the difference between PPF, NPS, and ELSS?",
"Which of these suits a conservative investor like me?",
"How much should I invest monthly to build a ₹2 crore corpus by 55?",
"Should I invest my FD maturity amount in a lump sum SIP?",
"What are the tax implications of NPS under the new tax regime?",
"Should I open a PPF account given the 15-year lock-in?",
"Give me a complete monthly investment breakdown.",
]
avg_response_tokens = 100 # assumed length of every reply
def analyse(budget, messages_to_compress=4):
buffer, overflows, peak = [], 0, 0 # the buffer as a list of token counts
for user_msg in analysis_turns:
for size in (len(TOKENISER.encode(user_msg)), avg_response_tokens):
if sum(buffer) + size > budget:
n = min(messages_to_compress, len(buffer)) // 2 * 2
buffer, overflows = buffer[n:], overflows + 1
buffer.append(size)
peak = max(peak, sum(buffer))
return overflows, peak
budgets = [300, 400, 600, 800, 1000, 1500]
results = [analyse(budget) for budget in budgets]
print(f"{'Budget':>6} | {'Overflows':>9} | {'Largest buffer':>14}")
for budget, (overflows, peak) in zip(budgets, results):
print(f"{budget:>6} | {overflows:>9} | {peak:>14}")
plt.figure(figsize=(7.5, 4))
plt.bar([str(b) for b in budgets], [r[0] for r in results], color="#9370DB")
plt.xlabel("max_buffer_tokens")
plt.ylabel("Summariser calls in 10 turns")
plt.yticks(range(0, 5))
plt.title("Summary buffer memory: overflows per token budget")
plt.show()Budget | Overflows | Largest buffer 300 | 4 | 250 400 | 4 | 365 600 | 3 | 592 800 | 2 | 798 1000 | 1 | 934 1500 | 0 | 1142
- 300 and 400 tokens overflow 4 times in 10 turns, 600 tokens 3 times, 800 twice and 1,000 once.
- At 1,500 tokens nothing overflows: the ten turns peak at 1,142 tokens, so the technique behaves like a plain buffer.
- The answer depends on the replies. Change the assumed 100 tokens to the length your model writes before you read the table.
Building the summary buffer class
The notebook 4_summary_buffer_memory.ipynb runs ten turns with max_buffer_tokens=800. Here six of its turns run with a budget of 600, so that the buffer overflows early, on the Groq model of the earlier lessons. The summariser prompt also names hard rules the user stated, which the notebook's prompt for this technique lists as constraints.
The overflow check
The check runs before a message is added: if the buffer plus the new message would pass the budget, the oldest messages are compressed first. Token counts are tiktoken o200k_base counts of the contents.
def add_message(self, role, content):
if self.buffer_tokens() + count(content) > self.max_buffer_tokens:
self.handle_overflow()
message = {"role": role, "content": content}
self.buffer.append(message)
self.archive.append(message)Compressing the oldest messages
An overflow takes the oldest 4 messages (2 turns) out of the buffer, or as many whole turns as the buffer holds, and hands them to the summariser together with the old summary. The printed line compares the tokens that went in, old summary included, with the new summary.
def handle_overflow(self):
n = min(self.messages_to_compress_on_overflow, len(self.buffer)) // 2 * 2
if n == 0:
return
old, self.buffer = self.buffer[:n], self.buffer[n:]
before = sum(count(m["content"]) for m in old) + count(self.summary or "")
self.summary = summarise_messages(old, self.summary)
self.overflow_count += 1Assembling the prompt
The request is built from the same two fields the bookkeeping prints, so the printed numbers describe the request that was sent.
def get_messages_for_api(self):
messages = [{"role": "system", "content": self.system_prompt}]
if self.summary:
messages.append({"role": "system", "content":
"CONVERSATION MEMORY (earlier turns, compressed):\n" + self.summary})
return messages + self.bufferRunning six turns of the video's conversation
The first four turns give facts: salary, expenses, age, risk profile, a fixed deposit and a retirement goal. Turn 5 asks about PPF and NPS, two Indian retirement schemes, and turn 6 asks which one fits "my salary and risk profile". The run makes eight or more model calls in a few seconds, so a 429 rate limit error on Groq's free tier means: wait a minute and run it again.
import os
import tiktoken
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"
TOKENISER = tiktoken.get_encoding("o200k_base")
FINCOACH_SYSTEM_PROMPT = """You are FinCoach, a personal financial advisor assistant.
You serve users in India who want guidance on savings, investments, budgeting, and financial planning.
Your principles:
- Always personalise advice using information the user has shared in this conversation.
- Be specific with numbers when the user has provided their financial details.
- Flag when you are making assumptions due to missing information.
- Keep responses concise: 3 to 5 sentences unless the user asks for detail.
- Never provide specific buy/sell recommendations on individual stocks.
- Always recommend consulting a SEBI-registered advisor for major financial decisions.
Use all context, including the conversation summary if provided, to give consistent, personalised advice."""
FINCOACH_SUMMARY_PROMPT = """You are summarising a financial advisory conversation for FinCoach.
Your summary will be the only memory of earlier turns in the next API call.
Preserve these facts if mentioned: salary and expenses, existing investments with their
values and maturity dates, risk profile, goals, hard rules the user stated, decisions made,
advice FinCoach gave, age and other personal context.
Write in third person, as one paragraph of at most 80 words.
Do not add facts that are not in the conversation."""
def count(text):
return len(TOKENISER.encode(text))
def summarise_messages(messages_to_compress, existing_summary=None):
turns_text = "\n".join(f"{m['role'].upper()}: {m['content']}" for m in messages_to_compress)
if existing_summary:
user_content = (f"EXISTING SUMMARY (from previous turns):\n{existing_summary}\n\n"
f"NEW MESSAGES TO INCORPORATE:\n{turns_text}\n\n"
"Produce an updated summary that merges both.")
else:
user_content = f"MESSAGES TO SUMMARISE:\n{turns_text}\n\nProduce a summary of these messages."
response = client.chat.completions.create(
model=MODEL, max_tokens=1000, temperature=0, reasoning_effort="low",
messages=[{"role": "system", "content": FINCOACH_SUMMARY_PROMPT},
{"role": "user", "content": user_content}])
return response.choices[0].message.content.strip()
class SummaryBufferMemory:
def __init__(self, system_prompt, max_buffer_tokens=600, messages_to_compress_on_overflow=4):
self.system_prompt = system_prompt
self.max_buffer_tokens = max_buffer_tokens
self.messages_to_compress_on_overflow = messages_to_compress_on_overflow
self.buffer = [] # recent messages, verbatim
self.summary = None # everything older, compressed
self.archive = [] # every message, never sent
self.overflow_count = 0
def buffer_tokens(self):
return sum(count(m["content"]) for m in self.buffer)
def add_message(self, role, content):
if self.buffer_tokens() + count(content) > self.max_buffer_tokens:
self.handle_overflow()
message = {"role": role, "content": content}
self.buffer.append(message)
self.archive.append(message)
def handle_overflow(self):
n = min(self.messages_to_compress_on_overflow, len(self.buffer)) // 2 * 2
if n == 0:
return
old, self.buffer = self.buffer[:n], self.buffer[n:]
before = sum(count(m["content"]) for m in old) + count(self.summary or "")
self.summary = summarise_messages(old, self.summary)
self.overflow_count += 1
print(f" [COMPRESS] {n} messages + old summary ({before} tokens) "
f"-> new summary ({count(self.summary)} tokens)")
def get_messages_for_api(self):
messages = [{"role": "system", "content": self.system_prompt}]
if self.summary:
messages.append({"role": "system", "content":
"CONVERSATION MEMORY (earlier turns, compressed):\n" + self.summary})
return messages + self.buffer
def chat(memory, user_message):
memory.add_message("user", user_message)
request = memory.get_messages_for_api()
response = client.chat.completions.create(
model=MODEL, max_tokens=1024, temperature=0, messages=request)
reply = response.choices[0].message.content
print(f"[Turn {len(memory.archive) // 2 + 1}] prompt_tokens: {response.usage.prompt_tokens} | "
f"summary: {count(memory.summary or '')} tokens | "
f"buffer: {memory.buffer_tokens()}/{memory.max_buffer_tokens} tokens in "
f"{len(memory.buffer)} messages | overflows: {memory.overflow_count}")
memory.add_message("assistant", reply)
return reply, request
sb_memory = SummaryBufferMemory(FINCOACH_SYSTEM_PROMPT, max_buffer_tokens=600,
messages_to_compress_on_overflow=4)
demo_turns = [
"Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000.",
"My monthly expenses are ₹60,000. I'm 32 years old and risk-averse.",
"I have an FD of ₹50,000 maturing in 3 months and no other investments.",
"My financial goal is to retire by 55 with a comfortable corpus.",
"What is the difference between PPF and NPS for retirement planning?",
"Based on my salary and risk profile, which one would you recommend?",
]
for user_message in demo_turns:
reply, request = chat(sb_memory, user_message)
print()
print("Second system message at turn 6:", request[1]["content"])
print()
print("FinCoach, turn 6:", reply)[Turn 1] prompt_tokens: 229 | summary: 0 tokens | buffer: 19/600 tokens in 1 messages | overflows: 0 [Turn 2] prompt_tokens: 471 | summary: 0 tokens | buffer: 251/600 tokens in 3 messages | overflows: 0 [Turn 3] prompt_tokens: 800 | summary: 0 tokens | buffer: 570/600 tokens in 5 messages | overflows: 0 [COMPRESS] 4 messages + old summary (550 tokens) -> new summary (141 tokens) [Turn 4] prompt_tokens: 622 | summary: 141 tokens | buffer: 243/600 tokens in 3 messages | overflows: 1 [COMPRESS] 2 messages + old summary (370 tokens) -> new summary (148 tokens) [Turn 5] prompt_tokens: 789 | summary: 148 tokens | buffer: 403/600 tokens in 3 messages | overflows: 2 [COMPRESS] 2 messages + old summary (537 tokens) -> new summary (180 tokens) [Turn 6] prompt_tokens: 830 | summary: 180 tokens | buffer: 412/600 tokens in 3 messages | overflows: 3 Second system message at turn 6: CONVERSATION MEMORY (earlier turns, compressed): Chiru, 32, earns ₹1.20 L net monthly and spends ₹60 k, leaving a ₹60 k surplus. He has a ₹50 k FD maturing in 3 months and is risk‑averse with no debts. Goal: retire at 55 with a comfortable corpus of ₹4‑5 cr (≈₹5 cr target). Needed SIP: about ₹55‑60 k/month. Plan: build a 3‑month emergency fund (~₹1.8 L) using the FD plus ₹15‑20 k/month; then allocate the remaining surplus (~₹40‑45 k) to ₹20 k low‑risk debt (short‑term debt funds/PPF), ₹15 k conservative hybrid fund, and ₹10‑15 k NPS Tier I, increasing contributions as salary grows. FinCoach, turn 6: Given your ₹1.20 L net salary, ₹60 k monthly surplus and a risk‑averse stance, I’d suggest using **PPF as the core retirement vehicle** (≈₹20 k / month) because it guarantees a tax‑free 7‑7.5 % return and has the safest credit risk. Complement it with a **modest NPS contribution** (≈₹10‑12 k / month) to capture the higher, market‑linked returns of the equity‑bond mix while still enjoying extra tax deductions. The remaining surplus can stay in low‑risk debt or hybrid funds as you’ve planned, helping you reach the ₹4‑5 cr corpus by age 55. *Assumption:* I’m assuming you have no other employer‑pension or retirement scheme; please verify the exact split with a SEBI‑registered financial advisor.
What the request held at turn 6
- Turns 1 to 3 have no summary. The buffer fills to 19, 251 and 570 of its 600 tokens.
- Three overflows happen before turn 6. Storing reply 3 would pass 600 tokens, so turns 1 and 2 (550 tokens) become a 141-token summary. Replies 4 and 5 are long, and each triggers another overflow that folds one more turn: 370 tokens into 148, then 537 into 180.
- The printed numbers describe the request. At turn 6 the system prompt is 136 tokens, the memory block 191 (the 180-token summary plus its heading) and the buffer 412: 739 tokens by tiktoken, 830 billed with the chat template.
- Turn 6 is answered from both regions. The salary and the risk-averse profile come from the summary, PPF and NPS from turn 5, which is still in the buffer word for word.
- The summary is not a neutral record. It says "no debts", which no user message states, and it gives the goal as a corpus of 4 to 5 crore rupees, while the user asked for "a comfortable corpus" without a figure. Advice from the replies now reads as facts about the user.
- The reply quotes a rate that no message supplied. It says PPF "guarantees a tax-free 7-7.5 % return". Nothing in the request holds that figure, and a rate of this kind changes over time, so it is output to check against an official source, not a fact.
Summary buffer memory vs summary memory
| Summary memory | Summary buffer memory | |
|---|---|---|
| Trigger | The turn count passes a threshold | The token budget of the recent buffer is exceeded |
| Recent zone | A fixed number of turns, word for word | Any number of messages, capped in tokens |
| Control lever | max_turns_before_summary | max_buffer_tokens |
| Cost of a request | Depends on how long the turns are | Bounded: system prompt + summary + budget |
| Extra model calls | One per summary cycle | One per overflow |
Where you use summary buffer memory
- Customer support agents and long-running chatbots. The video names these: they need the exact recent exchange and an awareness of what came before.
- Agent frameworks. LangChain 1.4 ships this pattern as
SummarizationMiddleware, with atriggerand akeepsetting in tokens, messages or a fraction of the context window. - Sessions with uneven messages. Pasted documents and one-word replies in the same chat are where a turn count misleads and a token budget does not.
Related
- Previous: Summary memory
- Next: Token buffer memory
- Reference: LangChain short-term memory
- In the budget example, change
avg_response_tokens = 100to150: a 300-token budget now overflows 9 times in 10 turns, and even 1,500 tokens overflows once. - In the budget example, change
messages_to_compress=4to2inanalyse: a 600-token budget overflows 5 times instead of 3, because each overflow frees less room. - In the live run, set
max_buffer_tokens=2000: nothing overflows in six turns, no summary is written, andrequest[1]is then the first user message.
Every expert started right here.