AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Token buffer memory

Token buffer memory is a short-term memory technique that keeps the newest messages of a conversation within a fixed token budget and drops the oldest messages when a new one would exceed it.

Last updated: 09 Oct, 2026 · OpenAI SDK 3.3

Summary buffer memory summarises what overflows. A token buffer throws it away. No summary means no extra model call and no wait, and it also means the dropped messages are gone from the prompt for good.

Hard eviction and the cost formula · from the Complete AI Security Course in 8 Hours video · 4:21:53 to 4:23:20

This part of the video starts at 4:21:53. It reads from 5_token_buffer_memory.ipynb: how a token buffer differs from a summary buffer (partial loss against total loss), the first point (no compression calls, no latency spikes), and the cost formula.

A sliding window fixes the number of turns; a token buffer fixes the number of tokens. Trimming changes only the list that is sent: the model is the same before and after.

How a token buffer trims

The notebook's sentence is: maintain a message buffer that never exceeds a fixed token budget, and when a new message would breach the limit, drop the oldest messages until there is room. Its picture is a ticket tape of fixed length: new messages are printed on the right, and when the tape is full the left end is torn off until the new message fits.

The notebook's cost model for a token buffer

The formula bounds the content tokens that your tokenizer counts. The billed prompt is a little larger, because the provider adds its chat template, and the output is billed on top. So the budget is a ceiling you can plan with, not the exact bill.

Three tapes of message blocks against a 500-token line. Before eviction the seven messages total 540 tokens. Dropping one message, T1-user, leaves 480 tokens starting with the reply T1-ai. Dropping the pair T1-user and T1-ai leaves 400 tokens starting with T2-user.

Evicting one message or a pair

The notebook's tape has a 500-token budget and five messages of 60, 80, 55, 90 and 70 tokens. A 120-token reply arrives, then a 65-token user message. The function below replays it twice: dropping single messages, and dropping a user message together with its reply.

ExampleFrom the video, run on Python 3.12
def evict_to_fit(buffer, incoming, budget, evict_in_pairs):
    name, size = incoming
    total = sum(tokens for _, tokens in buffer)
    print(f"  {name} ({size}t) arrives: {total} + {size} = {total + size}")
    while buffer and total + size > budget:
        dropped = [buffer.pop(0)]
        if evict_in_pairs and buffer and buffer[0][0].endswith("-ai"):
            dropped.append(buffer.pop(0))     # the reply leaves with its question
        total = sum(tokens for _, tokens in buffer)
        print(f"  over {budget}: drop " + " + ".join(f"{n} ({t}t)" for n, t in dropped)
              + f" -> {total + size}")
    buffer.append(incoming)
    return buffer

tape = [("T1-user", 60), ("T1-ai", 80), ("T2-user", 55), ("T2-ai", 90), ("T3-user", 70)]
for evict_in_pairs in (False, True):
    print(f"evict_in_pairs={evict_in_pairs}")
    buffer = list(tape)
    for incoming in [("T3-ai", 120), ("T4-user", 65)]:
        buffer = evict_to_fit(buffer, incoming, budget=500, evict_in_pairs=evict_in_pairs)
    print(f"  buffer starts with {buffer[0][0]}, total {sum(t for _, t in buffer)} tokens\n")
  • The 120-token reply fits: 355 + 120 = 475, under 500.
  • The 65-token message does not: 475 + 65 = 540.
  • Dropping one message is enough for the budget: without T1-user the total is 480. The buffer now starts with T1-ai, a reply whose question is gone.
  • Dropping the pair leaves 400 tokens and a buffer that starts with a user message. The notebook's class has an evict_in_pairs flag for this, on by default.

Building the token buffer class

Evict, then append

add_message drops from the front while the buffer plus the new message is over the budget. With evict_in_pairs, a reply that would be left at the front goes out with its question. There is no model call anywhere in the class.

python
    def add_message(self, role, content):
        while self.buffer and self.buffer_tokens() + count(content) > self.max_buffer_tokens:
            self.evicted.append(self.buffer.pop(0))
            if self.evict_in_pairs and self.buffer and self.buffer[0]["role"] == "assistant":
                self.evicted.append(self.buffer.pop(0))
        self.buffer.append({"role": role, "content": content})

Buffer, window and token buffer on the same turns

The eight scripted turns of Sliding window memory run through all three techniques that need no model: a full buffer, a window of 3 turns and a token buffer of 120 tokens. Each row is the request of one turn, in tiktoken o200k_base tokens of the contents.

ExampleRun with tiktoken 0.14
from collections import deque
import tiktoken
import matplotlib.pyplot as plt

TOKENISER = tiktoken.get_encoding("o200k_base")
conversation = [        # (user message, scripted reply): fixed text, no model call
    ("Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000.",
     "Hi Chiru! A take-home of ₹1,20,000 is a solid base to plan from."),
    ("My monthly expenses are about ₹60,000 for rent, food, and transport.",
     "That leaves a surplus of ₹60,000 a month."),
    ("I have an FD of ₹50,000 that matures in 3 months. I am risk-averse.",
     "Noted. A risk-averse plan keeps most of the money in fixed-income options."),
    ("What is a Systematic Investment Plan and how does it work?",
     "A Systematic Investment Plan, or SIP, invests a fixed amount in a mutual fund at a "
     "regular interval, usually every month. The amount is debited from your bank account "
     "automatically. Each instalment buys units at that day's price, so you buy more units "
     "when prices are low and fewer when they are high. Over time this averages your cost."),
    ("What are the different types of mutual funds available in India?",
     "The main types are equity funds, debt funds, hybrid funds, index funds and liquid funds."),
    ("Which of these suits a conservative investor?",
     "Debt funds and liquid funds carry the least risk of these."),
    ("How long should I stay invested?",
     "For debt funds, plan for at least three years."),
    ("What is my exact monthly take-home salary that I told you at the start?", ""),
]

def count(text):
    return len(TOKENISER.encode(text))

class TokenBufferMemory:
    def __init__(self, max_buffer_tokens, evict_in_pairs=True):
        self.max_buffer_tokens = max_buffer_tokens
        self.evict_in_pairs = evict_in_pairs
        self.buffer = []                      # messages sent on the next call
        self.evicted = []                     # dropped from the buffer, kept for the record

    def buffer_tokens(self):
        return sum(count(m["content"]) for m in self.buffer)

    def add_message(self, role, content):
        while self.buffer and self.buffer_tokens() + count(content) > self.max_buffer_tokens:
            self.evicted.append(self.buffer.pop(0))
            if self.evict_in_pairs and self.buffer and self.buffer[0]["role"] == "assistant":
                self.evicted.append(self.buffer.pop(0))
        self.buffer.append({"role": role, "content": content})

memory = TokenBufferMemory(max_buffer_tokens=120)
window = deque(maxlen=3)
sizes = {"Buffer: every turn": [], "Sliding window: last 3 turns": [], "Token buffer: 120 tokens": []}
print("Turn | buffer | window | token buffer | token buffer starts with")
for n, (user_message, reply) in enumerate(conversation, start=1):
    memory.add_message("user", user_message)
    window.append((user_message,))
    everything = conversation[:n - 1] + [(user_message,)]
    sizes["Buffer: every turn"].append(sum(count(t) for turn in everything for t in turn))
    sizes["Sliding window: last 3 turns"].append(sum(count(t) for turn in window for t in turn))
    sizes["Token buffer: 120 tokens"].append(memory.buffer_tokens())
    first = memory.buffer[0]
    print(f"{n:>4} | {sizes['Buffer: every turn'][-1]:>6} | {sizes['Sliding window: last 3 turns'][-1]:>6} | "
          f"{memory.buffer_tokens():>12} | {first['role']}: {first['content'][:34]}")
    window[-1] = (user_message, reply)
    if reply:
        memory.add_message("assistant", reply)
print("Messages evicted from the token buffer:", len(memory.evicted))

turns = range(1, len(conversation) + 1)
plt.figure(figsize=(8, 4.2))
for (label, values), color, marker in zip(sizes.items(), ["#d64541", "#3a6fd8", "#2e9e5b"], "os^"):
    plt.plot(turns, values, marker=marker, color=color, label=label)
plt.axhline(120, color="#8a8a8a", linestyle="--", linewidth=1)
plt.xlabel("Turn")
plt.ylabel("Conversation tokens in the request")
plt.title("Three ways to bound the same eight turns")
plt.legend()
plt.grid(alpha=0.3)
plt.show()
Three lines over eight turns with a dashed line at 120 tokens. The buffer climbs to 279 tokens. The sliding window rises above the dashed line at turns 5 and 6, to 136 and 122 tokens. The token buffer stays under it on every turn, between 19 and 95 tokens.

What the three lines show

  • The buffer passes 120 tokens at turn 4 and reaches 279 at turn 8.
  • The window of 3 turns goes over 120 twice: 136 tokens at turn 5 and 122 at turn 6, because turn 4 has a long reply. A turn count does not bound tokens.
  • The token buffer never passes its budget: its largest request is 95 tokens. It often sits well below 120, since whole messages are dropped: at turn 6 it holds 39 tokens.
  • The salary is gone from the token buffer at turn 4, when the first pair is dropped, and 8 messages have been evicted by the end. What is dropped is not summarised anywhere.

Token buffer memory vs the other four techniques

TechniqueWhat is sentWhat is lostExtra model callsControl unit
Conversation bufferEvery messageNothingNoneNone: it grows until the context window
Sliding windowThe last k turnsEverything older, completelyNoneTurns
SummaryA summary and the last few turnsDetails the summary leaves outOne per summary cycleTurns
Summary bufferA summary and recent messages under a token budgetDetails the summary leaves outOne per overflowTokens
Token bufferThe newest messages under a token budgetEverything older, completelyNoneTokens

The notebook compares the token buffer with its two neighbours in the same terms: against the sliding window the control unit changes from turns to tokens, and against the summary buffer the loss changes from partial (lossy compression) to total (hard eviction), with no summariser call and no latency spike in exchange.

LangChain's classic memory classes today

The video names LangChain's ConversationTokenBufferMemory as the same technique. The five classic classes (ConversationBufferMemory, ConversationBufferWindowMemory, ConversationSummaryMemory, ConversationSummaryBufferMemory, ConversationTokenBufferMemory) used to be imported from langchain.memory. In langchain 1.4.3 that module does not exist:

ExampleRun on langchain 1.4.3
from langchain.memory import ConversationTokenBufferMemory

The classes now live in a separate package, langchain-classic, where each one is marked as deprecated since LangChain 0.3.1 and due for removal in 2.0.0. Their definitions differ a little from the notebooks of the video, so check the unit before you compare numbers:

Classic classWhat it sendsCurrent way in LangChain 1.x
ConversationBufferMemoryAll messagesAn agent with a checkpointer, which stores the message list per thread
ConversationBufferWindowMemoryThe last k exchanges, so 2k messagestrim_messages before the model call
ConversationSummaryMemoryA summary only, rewritten after every turn, with no verbatim turnsSummarizationMiddleware
ConversationSummaryBufferMemoryA summary plus verbatim messages under max_token_limitSummarizationMiddleware
ConversationTokenBufferMemoryMessages under max_token_limit, dropped one at a timetrim_messages with a token counter

trim_messages from langchain_core is the token buffer as a function. It takes a token counter, so the tiktoken count used in this part plugs in directly. strategy="last" keeps the newest messages, include_system=True keeps the system message, and start_on="human" does what evict_in_pairs did above.

ExampleRun on langchain-core 1.6
import tiktoken
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage, trim_messages

TOKENISER = tiktoken.get_encoding("o200k_base")

def tiktoken_counter(messages):
    return sum(len(TOKENISER.encode(m.content)) for m in messages)

history = [
    SystemMessage("You are FinCoach, a personal financial advisor assistant."),
    HumanMessage("Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000."),
    AIMessage("Hi Chiru! A take-home of ₹1,20,000 is a solid base to plan from."),
    HumanMessage("My monthly expenses are about ₹60,000 for rent, food, and transport."),
    AIMessage("That leaves a surplus of ₹60,000 a month."),
    HumanMessage("I have an FD of ₹50,000 that matures in 3 months. I am risk-averse."),
    AIMessage("Noted. A risk-averse plan keeps most of the money in fixed-income options."),
    HumanMessage("Which fund categories suit a conservative investor?"),
]
print("Whole history:", len(history), "messages,", tiktoken_counter(history), "tokens")

for start_on in (None, "human"):
    trimmed = trim_messages(history, max_tokens=80, token_counter=tiktoken_counter,
                            strategy="last", include_system=True, start_on=start_on)
    print(f"\nstart_on={start_on!r}: {len(trimmed)} messages, {tiktoken_counter(trimmed)} tokens")
    for m in trimmed:
        print(f"  {m.type:>6}: {m.content[:50]}")
  • The whole history is 130 tokens in 8 messages, over the 80-token limit.
  • Without start_on the trimmed list is 72 tokens and starts, after the system message, with an AI reply whose question was cut.
  • With start_on="human" that reply is dropped as well: 60 tokens, starting with a human message.

Where you use token buffer memory

  • Voice and real-time assistants. The video's point is that there is no compression call and no latency spike: eviction is a list operation.
  • The active window of a larger system. The video pairs it with external retrieval: the token buffer for recency, a vector store for history.
  • Hard cost ceilings. When each request must stay under a known number of tokens whatever the user pastes.
Watch out. A token buffer drops facts without a trace: the user's salary can leave the prompt while the chat goes on. Keep facts that must survive (salary, constraints, goals) outside the buffer, in a profile or a store that is added to every prompt. Entity memory builds one.
Try it yourself
  • In the tape example, change budget=500 to budget=450: now the 120-token reply already forces an eviction, and both modes end at 400 tokens.
  • In the comparison, use TokenBufferMemory(max_buffer_tokens=120, evict_in_pairs=False): at turn 4 the buffer starts with the assistant reply "Hi Chiru! A take-home of ₹1,20,000", a reply without its question that still repeats the salary.
  • In the trim_messages example, change max_tokens=80 to 40: with start_on="human" only the system message and the last question remain, 19 tokens.

Little by little, you're building something great.