Token buffer memory
Token buffer memory is a short-term memory technique that keeps the newest messages of a conversation within a fixed token budget and drops the oldest messages when a new one would exceed it.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
Summary buffer memory summarises what overflows. A token buffer throws it away. No summary means no extra model call and no wait, and it also means the dropped messages are gone from the prompt for good.
This part of the video starts at 4:21:53. It reads from 5_token_buffer_memory.ipynb: how a token buffer differs from a summary buffer (partial loss against total loss), the first point (no compression calls, no latency spikes), and the cost formula.
A sliding window fixes the number of turns; a token buffer fixes the number of tokens. Trimming changes only the list that is sent: the model is the same before and after.
How a token buffer trims
The notebook's sentence is: maintain a message buffer that never exceeds a fixed token budget, and when a new message would breach the limit, drop the oldest messages until there is room. Its picture is a ticket tape of fixed length: new messages are printed on the right, and when the tape is full the left end is torn off until the new message fits.
The formula bounds the content tokens that your tokenizer counts. The billed prompt is a little larger, because the provider adds its chat template, and the output is billed on top. So the budget is a ceiling you can plan with, not the exact bill.
Evicting one message or a pair
The notebook's tape has a 500-token budget and five messages of 60, 80, 55, 90 and 70 tokens. A 120-token reply arrives, then a 65-token user message. The function below replays it twice: dropping single messages, and dropping a user message together with its reply.
def evict_to_fit(buffer, incoming, budget, evict_in_pairs):
name, size = incoming
total = sum(tokens for _, tokens in buffer)
print(f" {name} ({size}t) arrives: {total} + {size} = {total + size}")
while buffer and total + size > budget:
dropped = [buffer.pop(0)]
if evict_in_pairs and buffer and buffer[0][0].endswith("-ai"):
dropped.append(buffer.pop(0)) # the reply leaves with its question
total = sum(tokens for _, tokens in buffer)
print(f" over {budget}: drop " + " + ".join(f"{n} ({t}t)" for n, t in dropped)
+ f" -> {total + size}")
buffer.append(incoming)
return buffer
tape = [("T1-user", 60), ("T1-ai", 80), ("T2-user", 55), ("T2-ai", 90), ("T3-user", 70)]
for evict_in_pairs in (False, True):
print(f"evict_in_pairs={evict_in_pairs}")
buffer = list(tape)
for incoming in [("T3-ai", 120), ("T4-user", 65)]:
buffer = evict_to_fit(buffer, incoming, budget=500, evict_in_pairs=evict_in_pairs)
print(f" buffer starts with {buffer[0][0]}, total {sum(t for _, t in buffer)} tokens\n")evict_in_pairs=False T3-ai (120t) arrives: 355 + 120 = 475 T4-user (65t) arrives: 475 + 65 = 540 over 500: drop T1-user (60t) -> 480 buffer starts with T1-ai, total 480 tokens evict_in_pairs=True T3-ai (120t) arrives: 355 + 120 = 475 T4-user (65t) arrives: 475 + 65 = 540 over 500: drop T1-user (60t) + T1-ai (80t) -> 400 buffer starts with T2-user, total 400 tokens
- The 120-token reply fits: 355 + 120 = 475, under 500.
- The 65-token message does not: 475 + 65 = 540.
- Dropping one message is enough for the budget: without T1-user the total is 480. The buffer now starts with T1-ai, a reply whose question is gone.
- Dropping the pair leaves 400 tokens and a buffer that starts with a user message. The notebook's class has an
evict_in_pairsflag for this, on by default.
Building the token buffer class
Evict, then append
add_message drops from the front while the buffer plus the new message is over the budget. With evict_in_pairs, a reply that would be left at the front goes out with its question. There is no model call anywhere in the class.
def add_message(self, role, content):
while self.buffer and self.buffer_tokens() + count(content) > self.max_buffer_tokens:
self.evicted.append(self.buffer.pop(0))
if self.evict_in_pairs and self.buffer and self.buffer[0]["role"] == "assistant":
self.evicted.append(self.buffer.pop(0))
self.buffer.append({"role": role, "content": content})Buffer, window and token buffer on the same turns
The eight scripted turns of Sliding window memory run through all three techniques that need no model: a full buffer, a window of 3 turns and a token buffer of 120 tokens. Each row is the request of one turn, in tiktoken o200k_base tokens of the contents.
from collections import deque
import tiktoken
import matplotlib.pyplot as plt
TOKENISER = tiktoken.get_encoding("o200k_base")
conversation = [ # (user message, scripted reply): fixed text, no model call
("Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000.",
"Hi Chiru! A take-home of ₹1,20,000 is a solid base to plan from."),
("My monthly expenses are about ₹60,000 for rent, food, and transport.",
"That leaves a surplus of ₹60,000 a month."),
("I have an FD of ₹50,000 that matures in 3 months. I am risk-averse.",
"Noted. A risk-averse plan keeps most of the money in fixed-income options."),
("What is a Systematic Investment Plan and how does it work?",
"A Systematic Investment Plan, or SIP, invests a fixed amount in a mutual fund at a "
"regular interval, usually every month. The amount is debited from your bank account "
"automatically. Each instalment buys units at that day's price, so you buy more units "
"when prices are low and fewer when they are high. Over time this averages your cost."),
("What are the different types of mutual funds available in India?",
"The main types are equity funds, debt funds, hybrid funds, index funds and liquid funds."),
("Which of these suits a conservative investor?",
"Debt funds and liquid funds carry the least risk of these."),
("How long should I stay invested?",
"For debt funds, plan for at least three years."),
("What is my exact monthly take-home salary that I told you at the start?", ""),
]
def count(text):
return len(TOKENISER.encode(text))
class TokenBufferMemory:
def __init__(self, max_buffer_tokens, evict_in_pairs=True):
self.max_buffer_tokens = max_buffer_tokens
self.evict_in_pairs = evict_in_pairs
self.buffer = [] # messages sent on the next call
self.evicted = [] # dropped from the buffer, kept for the record
def buffer_tokens(self):
return sum(count(m["content"]) for m in self.buffer)
def add_message(self, role, content):
while self.buffer and self.buffer_tokens() + count(content) > self.max_buffer_tokens:
self.evicted.append(self.buffer.pop(0))
if self.evict_in_pairs and self.buffer and self.buffer[0]["role"] == "assistant":
self.evicted.append(self.buffer.pop(0))
self.buffer.append({"role": role, "content": content})
memory = TokenBufferMemory(max_buffer_tokens=120)
window = deque(maxlen=3)
sizes = {"Buffer: every turn": [], "Sliding window: last 3 turns": [], "Token buffer: 120 tokens": []}
print("Turn | buffer | window | token buffer | token buffer starts with")
for n, (user_message, reply) in enumerate(conversation, start=1):
memory.add_message("user", user_message)
window.append((user_message,))
everything = conversation[:n - 1] + [(user_message,)]
sizes["Buffer: every turn"].append(sum(count(t) for turn in everything for t in turn))
sizes["Sliding window: last 3 turns"].append(sum(count(t) for turn in window for t in turn))
sizes["Token buffer: 120 tokens"].append(memory.buffer_tokens())
first = memory.buffer[0]
print(f"{n:>4} | {sizes['Buffer: every turn'][-1]:>6} | {sizes['Sliding window: last 3 turns'][-1]:>6} | "
f"{memory.buffer_tokens():>12} | {first['role']}: {first['content'][:34]}")
window[-1] = (user_message, reply)
if reply:
memory.add_message("assistant", reply)
print("Messages evicted from the token buffer:", len(memory.evicted))
turns = range(1, len(conversation) + 1)
plt.figure(figsize=(8, 4.2))
for (label, values), color, marker in zip(sizes.items(), ["#d64541", "#3a6fd8", "#2e9e5b"], "os^"):
plt.plot(turns, values, marker=marker, color=color, label=label)
plt.axhline(120, color="#8a8a8a", linestyle="--", linewidth=1)
plt.xlabel("Turn")
plt.ylabel("Conversation tokens in the request")
plt.title("Three ways to bound the same eight turns")
plt.legend()
plt.grid(alpha=0.3)
plt.show()Turn | buffer | window | token buffer | token buffer starts with 1 | 19 | 19 | 19 | user: Hi! I'm Chiru. My monthly take-hom 2 | 58 | 58 | 58 | user: Hi! I'm Chiru. My monthly take-hom 3 | 93 | 93 | 93 | user: Hi! I'm Chiru. My monthly take-hom 4 | 124 | 83 | 83 | user: My monthly expenses are about ₹60, 5 | 206 | 136 | 95 | user: What is a Systematic Investment Pl 6 | 233 | 122 | 39 | user: What are the different types of mu 7 | 252 | 58 | 58 | user: What are the different types of mu 8 | 279 | 54 | 85 | user: What are the different types of mu Messages evicted from the token buffer: 8
What the three lines show
- The buffer passes 120 tokens at turn 4 and reaches 279 at turn 8.
- The window of 3 turns goes over 120 twice: 136 tokens at turn 5 and 122 at turn 6, because turn 4 has a long reply. A turn count does not bound tokens.
- The token buffer never passes its budget: its largest request is 95 tokens. It often sits well below 120, since whole messages are dropped: at turn 6 it holds 39 tokens.
- The salary is gone from the token buffer at turn 4, when the first pair is dropped, and 8 messages have been evicted by the end. What is dropped is not summarised anywhere.
Token buffer memory vs the other four techniques
| Technique | What is sent | What is lost | Extra model calls | Control unit |
|---|---|---|---|---|
| Conversation buffer | Every message | Nothing | None | None: it grows until the context window |
| Sliding window | The last k turns | Everything older, completely | None | Turns |
| Summary | A summary and the last few turns | Details the summary leaves out | One per summary cycle | Turns |
| Summary buffer | A summary and recent messages under a token budget | Details the summary leaves out | One per overflow | Tokens |
| Token buffer | The newest messages under a token budget | Everything older, completely | None | Tokens |
The notebook compares the token buffer with its two neighbours in the same terms: against the sliding window the control unit changes from turns to tokens, and against the summary buffer the loss changes from partial (lossy compression) to total (hard eviction), with no summariser call and no latency spike in exchange.
LangChain's classic memory classes today
The video names LangChain's ConversationTokenBufferMemory as the same technique. The five classic classes (ConversationBufferMemory, ConversationBufferWindowMemory, ConversationSummaryMemory, ConversationSummaryBufferMemory, ConversationTokenBufferMemory) used to be imported from langchain.memory. In langchain 1.4.3 that module does not exist:
from langchain.memory import ConversationTokenBufferMemoryTraceback (most recent call last):
File "main.py", line 1, in <module>
from langchain.memory import ConversationTokenBufferMemory
ModuleNotFoundError: No module named 'langchain.memory'The classes now live in a separate package, langchain-classic, where each one is marked as deprecated since LangChain 0.3.1 and due for removal in 2.0.0. Their definitions differ a little from the notebooks of the video, so check the unit before you compare numbers:
| Classic class | What it sends | Current way in LangChain 1.x |
|---|---|---|
ConversationBufferMemory | All messages | An agent with a checkpointer, which stores the message list per thread |
ConversationBufferWindowMemory | The last k exchanges, so 2k messages | trim_messages before the model call |
ConversationSummaryMemory | A summary only, rewritten after every turn, with no verbatim turns | SummarizationMiddleware |
ConversationSummaryBufferMemory | A summary plus verbatim messages under max_token_limit | SummarizationMiddleware |
ConversationTokenBufferMemory | Messages under max_token_limit, dropped one at a time | trim_messages with a token counter |
trim_messages from langchain_core is the token buffer as a function. It takes a token counter, so the tiktoken count used in this part plugs in directly. strategy="last" keeps the newest messages, include_system=True keeps the system message, and start_on="human" does what evict_in_pairs did above.
import tiktoken
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage, trim_messages
TOKENISER = tiktoken.get_encoding("o200k_base")
def tiktoken_counter(messages):
return sum(len(TOKENISER.encode(m.content)) for m in messages)
history = [
SystemMessage("You are FinCoach, a personal financial advisor assistant."),
HumanMessage("Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000."),
AIMessage("Hi Chiru! A take-home of ₹1,20,000 is a solid base to plan from."),
HumanMessage("My monthly expenses are about ₹60,000 for rent, food, and transport."),
AIMessage("That leaves a surplus of ₹60,000 a month."),
HumanMessage("I have an FD of ₹50,000 that matures in 3 months. I am risk-averse."),
AIMessage("Noted. A risk-averse plan keeps most of the money in fixed-income options."),
HumanMessage("Which fund categories suit a conservative investor?"),
]
print("Whole history:", len(history), "messages,", tiktoken_counter(history), "tokens")
for start_on in (None, "human"):
trimmed = trim_messages(history, max_tokens=80, token_counter=tiktoken_counter,
strategy="last", include_system=True, start_on=start_on)
print(f"\nstart_on={start_on!r}: {len(trimmed)} messages, {tiktoken_counter(trimmed)} tokens")
for m in trimmed:
print(f" {m.type:>6}: {m.content[:50]}")Whole history: 8 messages, 130 tokens
start_on=None: 5 messages, 72 tokens
system: You are FinCoach, a personal financial advisor ass
ai: That leaves a surplus of ₹60,000 a month.
human: I have an FD of ₹50,000 that matures in 3 months.
ai: Noted. A risk-averse plan keeps most of the money
human: Which fund categories suit a conservative investor
start_on='human': 4 messages, 60 tokens
system: You are FinCoach, a personal financial advisor ass
human: I have an FD of ₹50,000 that matures in 3 months.
ai: Noted. A risk-averse plan keeps most of the money
human: Which fund categories suit a conservative investor- The whole history is 130 tokens in 8 messages, over the 80-token limit.
- Without
start_onthe trimmed list is 72 tokens and starts, after the system message, with an AI reply whose question was cut. - With
start_on="human"that reply is dropped as well: 60 tokens, starting with a human message.
Where you use token buffer memory
- Voice and real-time assistants. The video's point is that there is no compression call and no latency spike: eviction is a list operation.
- The active window of a larger system. The video pairs it with external retrieval: the token buffer for recency, a vector store for history.
- Hard cost ceilings. When each request must stay under a known number of tokens whatever the user pastes.
Related
- Previous: Summary buffer memory
- Next: Vector store memory
- Reference: LangChain short-term memory
- See also: the LangChain tutorial
- In the tape example, change
budget=500tobudget=450: now the 120-token reply already forces an eviction, and both modes end at 400 tokens. - In the comparison, use
TokenBufferMemory(max_buffer_tokens=120, evict_in_pairs=False): at turn 4 the buffer starts with the assistant reply "Hi Chiru! A take-home of ₹1,20,000", a reply without its question that still repeats the salary. - In the
trim_messagesexample, changemax_tokens=80to40: withstart_on="human"only the system message and the last question remain, 19 tokens.
Little by little, you're building something great.