Sliding window memory
Sliding window memory is a short-term memory technique that sends only the last few turns of a conversation to the model and leaves older turns out of the request.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
Conversation buffer memory resends everything, so its requests grow without limit. A sliding window caps the number of turns in a request. What it gives up is the early part of the conversation.
This part of the video starts at 3:16:54. It reads the first lines of 2_sliding_window_memory.ipynb: the window in one sentence, and what a turn is.
A window bounds the number of turns in a request. The token count of a request still moves with the length of the messages inside the window, as the plot below shows. Nothing in the model changes either way: the only thing that differs from turn to turn is which messages your code sends.
How the window slides
The notebook's sentence is: "Instead of sending the entire conversation history to the LLM, send only the last N turns." A turn is one user message plus one assistant reply, so a window of 3 turns holds at most 6 messages. When a new turn arrives and the window is full, the oldest turn falls off, like a window sliding forward through the conversation.
The video's FinCoach example uses this window of 3 turns. The user gives the salary in turn 1, the expenses in turn 2 and a fixed deposit in turn 3. Turn 4 asks about SIPs, and turn 1 leaves the window: the salary is no longer in the request.
Turns, the archive and the trade-off
This part of the video starts at 3:18:16. The window is counted in turns, not tokens, and old messages are evicted from the active window, not deleted.
The unit of the window has to be stated every time. In the video's notebook it counts turns. Other code counts messages, so a "window of 4" can mean four turns or two.
In the notebook an evicted message goes to an archive list, so it can be saved or searched later. That is a choice of the code, not a property of the technique: a bare deque(maxlen=k) drops its oldest item silently. The archive is a Python list in memory, and it is lost on a restart unless you write it to a database.
This part of the video starts at 3:19:36. It asks whether evicted messages could be embedded and kept in a vector store, which is possible at the cost of an embedding model and a store, and then states the trade-off of the window.
The video names the trade-off recency against completeness: a buffer keeps everything, a window keeps only the recent past. Its example is a salary given in turn 1 and asked about in turn 15 with a window of 10 turns: the salary is gone from the request, and the agent has to ask again. Embedding the evicted messages into a store is the hybrid that Vector store memory comes back to.
Counting the window on a fixed conversation
To see the window without a model, take eight turns with scripted replies and count, for each turn, what a buffer would send and what a window of 3 turns sends. The counts are tiktoken o200k_base tokens of the message contents, without the system prompt. A request goes out before its reply exists, so the newest turn in the window has a user message only.
from collections import deque
import tiktoken
import matplotlib.pyplot as plt
TOKENISER = tiktoken.get_encoding("o200k_base")
conversation = [ # (user message, scripted reply): fixed text, no model call
("Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000.",
"Hi Chiru! A take-home of ₹1,20,000 is a solid base to plan from."),
("My monthly expenses are about ₹60,000 for rent, food, and transport.",
"That leaves a surplus of ₹60,000 a month."),
("I have an FD of ₹50,000 that matures in 3 months. I am risk-averse.",
"Noted. A risk-averse plan keeps most of the money in fixed-income options."),
("What is a Systematic Investment Plan and how does it work?",
"A Systematic Investment Plan, or SIP, invests a fixed amount in a mutual fund at a "
"regular interval, usually every month. The amount is debited from your bank account "
"automatically. Each instalment buys units at that day's price, so you buy more units "
"when prices are low and fewer when they are high. Over time this averages your cost."),
("What are the different types of mutual funds available in India?",
"The main types are equity funds, debt funds, hybrid funds, index funds and liquid funds."),
("Which of these suits a conservative investor?",
"Debt funds and liquid funds carry the least risk of these."),
("How long should I stay invested?",
"For debt funds, plan for at least three years."),
("What is my exact monthly take-home salary that I told you at the start?", ""),
]
def tokens(turns):
return sum(len(TOKENISER.encode(text)) for turn in turns for text in turn)
window = deque(maxlen=3) # the last 3 turns, the open one included
buffer_sizes, window_sizes = [], []
print("Turn | buffer: turns, tokens | window: turns, tokens | salary in window")
for n, (user_message, reply) in enumerate(conversation, start=1):
window.append((user_message,)) # the request goes out before the reply exists
everything = conversation[:n - 1] + [(user_message,)]
buffer_sizes.append(tokens(everything))
window_sizes.append(tokens(window))
first = n - len(window) + 1
salary = any("1,20,000" in text for turn in window for text in turn)
print(f"{n:>4} | 1 to {n}, {buffer_sizes[-1]:>4} | {first} to {n}, {window_sizes[-1]:>4}"
f" | {salary}")
window[-1] = (user_message, reply) # the reply completes the turn
turns = range(1, len(conversation) + 1)
plt.figure(figsize=(8, 4.2))
plt.plot(turns, buffer_sizes, marker="o", color="#d64541", label="Buffer: every turn")
plt.plot(turns, window_sizes, marker="s", color="#3a6fd8", label="Sliding window: last 3 turns")
plt.xlabel("Turn")
plt.ylabel("Conversation tokens in the request")
plt.title("Buffer vs sliding window on the same eight turns")
plt.legend()
plt.grid(alpha=0.3)
plt.show()Turn | buffer: turns, tokens | window: turns, tokens | salary in window 1 | 1 to 1, 19 | 1 to 1, 19 | True 2 | 1 to 2, 58 | 1 to 2, 58 | True 3 | 1 to 3, 93 | 1 to 3, 93 | True 4 | 1 to 4, 124 | 2 to 4, 83 | False 5 | 1 to 5, 206 | 3 to 5, 136 | False 6 | 1 to 6, 233 | 4 to 6, 122 | False 7 | 1 to 7, 252 | 5 to 7, 58 | False 8 | 1 to 8, 279 | 6 to 8, 54 | False
What the window does to the request
- Up to turn 3 the two are the same: 19, 58 and 93 tokens. The window is not full yet.
- From turn 4 the window stops growing with the conversation. At turn 8 the buffer sends 279 tokens and the window 54.
- The window is not constant in tokens. It holds 83 tokens at turn 4 and 136 at turn 5, because turn 4 has a long reply. A window fixes the number of turns, not the number of tokens.
- The salary leaves at turn 4. The last column turns False as soon as turn 1 falls off.
Which turns can still see a fact
The notebook gives the rule in one line: a fact stated in turn f is visible in turn q when q − f is smaller than the window size. The example checks the rule against a real deque on the ten cases the notebook prints.
from collections import deque
def fact_still_visible(fact_turn, query_turn, window_size):
return (query_turn - fact_turn) < window_size
def visible_in_a_real_window(fact_turn, query_turn, window_size):
window = deque(maxlen=window_size)
for turn in range(1, query_turn + 1): # the query turn is in the window too
window.append(turn)
return fact_turn in window
cases = [(5, 1, 6), (5, 1, 3), (5, 3, 8), (5, 3, 5), (10, 1, 6),
(10, 1, 11), (10, 3, 8), (10, 3, 15), (20, 1, 15), (20, 1, 22)]
print("Window | fact turn -> query turn | in window")
for window_size, fact_turn, query_turn in cases:
rule = fact_still_visible(fact_turn, query_turn, window_size)
assert rule == visible_in_a_real_window(fact_turn, query_turn, window_size)
print(f"{window_size:>6} | {fact_turn:>9} -> {query_turn:<10} | {'yes' if rule else 'EVICTED'}")
for window_size in (5, 10, 20):
last = max(q for q in range(1, 100) if fact_still_visible(1, q, window_size))
print(f"Window {window_size}: a fact from turn 1 is last visible at turn {last}")Window | fact turn -> query turn | in window
5 | 1 -> 6 | EVICTED
5 | 1 -> 3 | yes
5 | 3 -> 8 | EVICTED
5 | 3 -> 5 | yes
10 | 1 -> 6 | yes
10 | 1 -> 11 | EVICTED
10 | 3 -> 8 | yes
10 | 3 -> 15 | EVICTED
20 | 1 -> 15 | yes
20 | 1 -> 22 | EVICTED
Window 5: a fact from turn 1 is last visible at turn 5
Window 10: a fact from turn 1 is last visible at turn 10
Window 20: a fact from turn 1 is last visible at turn 20A window of 10 turns shows a fact from turn 1 up to turn 10, that is, for 9 more turns, and loses it at turn 11. A window of 5 loses it at turn 6, a window of 20 at turn 21.
This part of the video starts at 3:25:04. It runs six turns through the notebook's SlidingWindowMemory with a window of 3 turns and reads the [EVICT] lines and the reply to turn 6.
In the notebook's saved run the reply to turn 4 is a general explanation of SIPs, and the forgetting shows at turn 6. The notebook's window is a deque of six messages, so a request loses one message before the call and one after it, and a full window starts with an assistant reply. The class below keeps whole turns instead, so a request always starts with a user message.
Building the window class
A deque of turns
deque(maxlen=window_size) is the whole mechanism: appending to a full deque removes the item at the other end. A user message opens a new turn, and the assistant reply is added to that same turn. Every message also goes to archive, which is never sent.
class SlidingWindowMemory:
def __init__(self, system_prompt, window_size=3):
self.system_prompt = system_prompt
self.window = deque(maxlen=window_size) # whole turns, oldest first
self.archive = [] # every message, never sent
def add_message(self, role, content):
message = {"role": role, "content": content}
self.archive.append(message)
if role == "user":
self.window.append([message]) # a new turn; a full deque drops its oldest
else:
self.window[-1].append(message) # the reply completes the newest turnThe request from the window
The request is the system prompt followed by the messages of the turns still in the window.
def get_messages_for_api(self):
system = {"role": "system", "content": self.system_prompt}
return [system] + [message for turn in self.window for message in turn]A check on what was sent
chat() is the same four steps as before. One extra line looks for the text 1,20,000 in the messages of the request, so the output says for each turn whether the salary figure was sent.
salary_sent = any("1,20,000" in m["content"] for m in request[1:])Running the six turns from the video
The six user messages are the notebook's. Turns 4 and 5 are general questions, chosen so that the replies have no reason to repeat the salary. Turn 6 asks for it.
import os
from collections import deque
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"
FINCOACH_SYSTEM_PROMPT = """You are FinCoach, a personal financial advisor assistant.
You serve users in India who want guidance on savings, investments, budgeting, and financial planning.
Your principles:
- Always personalise advice using information the user has shared in this conversation.
- Be specific with numbers when the user has provided their financial details.
- Flag when you are making assumptions due to missing information.
- Keep responses concise: 3 to 5 sentences unless the user asks for detail.
- Never provide specific buy/sell recommendations on individual stocks.
- Always recommend consulting a SEBI-registered advisor for major financial decisions.
Use all context in the conversation history to provide personalised, consistent advice."""
class SlidingWindowMemory:
def __init__(self, system_prompt, window_size=3):
self.system_prompt = system_prompt
self.window = deque(maxlen=window_size) # whole turns, oldest first
self.archive = [] # every message, never sent
def add_message(self, role, content):
message = {"role": role, "content": content}
self.archive.append(message)
if role == "user":
self.window.append([message]) # a new turn; a full deque drops its oldest
else:
self.window[-1].append(message) # the reply completes the newest turn
def get_messages_for_api(self):
system = {"role": "system", "content": self.system_prompt}
return [system] + [message for turn in self.window for message in turn]
def chat(memory, user_message):
memory.add_message("user", user_message)
request = memory.get_messages_for_api()
response = client.chat.completions.create(
model=MODEL, max_tokens=1024, temperature=0, messages=request)
reply = response.choices[0].message.content
memory.add_message("assistant", reply)
turn = len(memory.archive) // 2
first = turn - len(memory.window) + 1
salary_sent = any("1,20,000" in m["content"] for m in request[1:])
print(f"[Turn {turn}] window: turns {first} to {turn} | messages sent: {len(request)} | "
f"prompt_tokens: {response.usage.prompt_tokens} | salary in the request: {salary_sent}")
return reply
sliding_memory = SlidingWindowMemory(FINCOACH_SYSTEM_PROMPT, window_size=3)
demo_turns = [
"Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000.",
"My monthly expenses are about ₹60,000 for rent, food, and transport.",
"I have an FD of ₹50,000 that matures in 3 months. I am risk-averse.",
"What is a Systematic Investment Plan and how does it work?",
"What are the different types of mutual funds available in India?",
"What is my exact monthly take-home salary that I told you at the start?",
]
for user_message in demo_turns:
reply = chat(sliding_memory, user_message)
print()
print("User, turn 6:", demo_turns[-1])
print("FinCoach, turn 6:", reply)
print()
print("Messages in the archive:", len(sliding_memory.archive))[Turn 1] window: turns 1 to 1 | messages sent: 2 | prompt_tokens: 225 | salary in the request: True [Turn 2] window: turns 1 to 2 | messages sent: 4 | prompt_tokens: 438 | salary in the request: True [Turn 3] window: turns 1 to 3 | messages sent: 6 | prompt_tokens: 769 | salary in the request: True [Turn 4] window: turns 2 to 4 | messages sent: 6 | prompt_tokens: 862 | salary in the request: False [Turn 5] window: turns 3 to 5 | messages sent: 6 | prompt_tokens: 793 | salary in the request: False [Turn 6] window: turns 4 to 6 | messages sent: 6 | prompt_tokens: 749 | salary in the request: False User, turn 6: What is my exact monthly take-home salary that I told you at the start? FinCoach, turn 6: I don’t have a record of your monthly take‑home salary from our earlier messages— you haven’t shared that figure yet. If you let me know the exact amount, I can tailor the budgeting and investment suggestions to your income. Feel free to provide the number, and we’ll continue from there. Messages in the archive: 12
What turn 6 shows
- The request stops at 6 messages. Turns 3 to 6 each send the system prompt and the turns in the window.
- The tokens still move. With a full window
prompt_tokens, the API's usage figure, is 862, 793 and 749 at turns 4, 5 and 6: the number of turns is fixed, the length of the replies is not. - The salary figure leaves the request at turn 4, when turn 1 falls off, and the column stays False: no later message repeated it.
- At turn 6 the model says it has no record of the salary and adds that the user has not shared that figure yet. The second half is wrong from the user's side: the figure was given in turn 1 and evicted. A model cannot tell a message that was evicted from one that was never sent.
- The archive holds all 12 messages. Nothing was deleted; the six oldest were not sent.
Sliding window memory vs conversation buffer memory
| Conversation buffer | Sliding window | |
|---|---|---|
| What is sent | Every turn | The last k turns |
| Messages per request | Grow every turn | Fixed once the window is full |
| Tokens per request | Grow every turn | Bounded by k turns, but vary with message length |
| Early facts | Kept | Lost when their turn leaves the window |
| Extra model calls | None | None |
| Tuning | None | The window size k |
Where you use sliding window memory
- Chats where only the recent exchange matters. A coding helper answering follow-up questions about the last snippet rarely needs turn 1.
- The short-term layer under a long-term store. The video's verdict is that a window is almost always present in production and almost never alone: it is paired with a store that keeps the facts that must not be lost.
- Latency-sensitive paths. Eviction is a list operation, with no model call.
Related
- Previous: Conversation buffer memory
- Next: Summary memory
- Reference: collections.deque
- In the fixed conversation, change
deque(maxlen=3)todeque(maxlen=5): the salary stays in the window through turn 5 and leaves at turn 6, and the window of turn 8 holds 168 tokens instead of 54. - Add
(3, 1, 4)and(3, 1, 3)tocasesin the rule example: with a window of 3, a fact from turn 1 is still in the window at turn 3 and EVICTED at turn 4, as in the diagram. - In the live run, set
window_size=6: every turn stays in the window, the last column reads True on all six lines, and the model can state the salary.
Slow is fine. Stopping is the only problem.