AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Sliding window memory

Sliding window memory is a short-term memory technique that sends only the last few turns of a conversation to the model and leaves older turns out of the request.

Last updated: 09 Oct, 2026 · OpenAI SDK 3.3

Conversation buffer memory resends everything, so its requests grow without limit. A sliding window caps the number of turns in a request. What it gives up is the early part of the conversation.

Sliding window memory in one sentence · from the Complete AI Security Course in 8 Hours video · 3:16:54 to 3:17:33

This part of the video starts at 3:16:54. It reads the first lines of 2_sliding_window_memory.ipynb: the window in one sentence, and what a turn is.

A window bounds the number of turns in a request. The token count of a request still moves with the length of the messages inside the window, as the plot below shows. Nothing in the model changes either way: the only thing that differs from turn to turn is which messages your code sends.

How the window slides

The notebook's sentence is: "Instead of sending the entire conversation history to the LLM, send only the last N turns." A turn is one user message plus one assistant reply, so a window of 3 turns holds at most 6 messages. When a new turn arrives and the window is full, the oldest turn falls off, like a window sliding forward through the conversation.

Four rows of turn chips for the requests of turns 3 to 6 with a window of 3 turns. The request of turn 3 holds T1, T2, T3. The request of turn 4 holds T2, T3, T4 and T1 is greyed out, turn 5 holds T3 to T5 and turn 6 holds T4 to T6.

The video's FinCoach example uses this window of 3 turns. The user gives the salary in turn 1, the expenses in turn 2 and a fixed deposit in turn 3. Turn 4 asks about SIPs, and turn 1 leaves the window: the salary is no longer in the request.

Turns, the archive and the trade-off

A window counted in turns, and evicted messages · from the Complete AI Security Course in 8 Hours video · 3:18:16 to 3:18:58

This part of the video starts at 3:18:16. The window is counted in turns, not tokens, and old messages are evicted from the active window, not deleted.

The unit of the window has to be stated every time. In the video's notebook it counts turns. Other code counts messages, so a "window of 4" can mean four turns or two.

In the notebook an evicted message goes to an archive list, so it can be saved or searched later. That is a choice of the code, not a property of the technique: a bare deque(maxlen=k) drops its oldest item silently. The archive is a Python list in memory, and it is lost on a restart unless you write it to a database.

Archiving evicted messages and the trade-off · from the Complete AI Security Course in 8 Hours video · 3:19:36 to 3:21:14

This part of the video starts at 3:19:36. It asks whether evicted messages could be embedded and kept in a vector store, which is possible at the cost of an embedding model and a store, and then states the trade-off of the window.

The video names the trade-off recency against completeness: a buffer keeps everything, a window keeps only the recent past. Its example is a salary given in turn 1 and asked about in turn 15 with a window of 10 turns: the salary is gone from the request, and the agent has to ask again. Embedding the evicted messages into a store is the hybrid that Vector store memory comes back to.

Counting the window on a fixed conversation

To see the window without a model, take eight turns with scripted replies and count, for each turn, what a buffer would send and what a window of 3 turns sends. The counts are tiktoken o200k_base tokens of the message contents, without the system prompt. A request goes out before its reply exists, so the newest turn in the window has a user message only.

ExampleRun with tiktoken 0.14
from collections import deque
import tiktoken
import matplotlib.pyplot as plt

TOKENISER = tiktoken.get_encoding("o200k_base")
conversation = [        # (user message, scripted reply): fixed text, no model call
    ("Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000.",
     "Hi Chiru! A take-home of ₹1,20,000 is a solid base to plan from."),
    ("My monthly expenses are about ₹60,000 for rent, food, and transport.",
     "That leaves a surplus of ₹60,000 a month."),
    ("I have an FD of ₹50,000 that matures in 3 months. I am risk-averse.",
     "Noted. A risk-averse plan keeps most of the money in fixed-income options."),
    ("What is a Systematic Investment Plan and how does it work?",
     "A Systematic Investment Plan, or SIP, invests a fixed amount in a mutual fund at a "
     "regular interval, usually every month. The amount is debited from your bank account "
     "automatically. Each instalment buys units at that day's price, so you buy more units "
     "when prices are low and fewer when they are high. Over time this averages your cost."),
    ("What are the different types of mutual funds available in India?",
     "The main types are equity funds, debt funds, hybrid funds, index funds and liquid funds."),
    ("Which of these suits a conservative investor?",
     "Debt funds and liquid funds carry the least risk of these."),
    ("How long should I stay invested?",
     "For debt funds, plan for at least three years."),
    ("What is my exact monthly take-home salary that I told you at the start?", ""),
]

def tokens(turns):
    return sum(len(TOKENISER.encode(text)) for turn in turns for text in turn)

window = deque(maxlen=3)                      # the last 3 turns, the open one included
buffer_sizes, window_sizes = [], []
print("Turn | buffer: turns, tokens | window: turns, tokens | salary in window")
for n, (user_message, reply) in enumerate(conversation, start=1):
    window.append((user_message,))            # the request goes out before the reply exists
    everything = conversation[:n - 1] + [(user_message,)]
    buffer_sizes.append(tokens(everything))
    window_sizes.append(tokens(window))
    first = n - len(window) + 1
    salary = any("1,20,000" in text for turn in window for text in turn)
    print(f"{n:>4} | 1 to {n}, {buffer_sizes[-1]:>4}         | {first} to {n}, {window_sizes[-1]:>4}"
          f"         | {salary}")
    window[-1] = (user_message, reply)        # the reply completes the turn

turns = range(1, len(conversation) + 1)
plt.figure(figsize=(8, 4.2))
plt.plot(turns, buffer_sizes, marker="o", color="#d64541", label="Buffer: every turn")
plt.plot(turns, window_sizes, marker="s", color="#3a6fd8", label="Sliding window: last 3 turns")
plt.xlabel("Turn")
plt.ylabel("Conversation tokens in the request")
plt.title("Buffer vs sliding window on the same eight turns")
plt.legend()
plt.grid(alpha=0.3)
plt.show()
Two lines over eight turns. The buffer line climbs from 19 to 279 tokens. The sliding window line follows it to 93 tokens at turn 3, then moves between 54 and 136 tokens.

What the window does to the request

  • Up to turn 3 the two are the same: 19, 58 and 93 tokens. The window is not full yet.
  • From turn 4 the window stops growing with the conversation. At turn 8 the buffer sends 279 tokens and the window 54.
  • The window is not constant in tokens. It holds 83 tokens at turn 4 and 136 at turn 5, because turn 4 has a long reply. A window fixes the number of turns, not the number of tokens.
  • The salary leaves at turn 4. The last column turns False as soon as turn 1 falls off.

Which turns can still see a fact

The notebook gives the rule in one line: a fact stated in turn f is visible in turn q when q − f is smaller than the window size. The example checks the rule against a real deque on the ten cases the notebook prints.

ExampleFrom the video, run on Python 3.12
from collections import deque

def fact_still_visible(fact_turn, query_turn, window_size):
    return (query_turn - fact_turn) < window_size

def visible_in_a_real_window(fact_turn, query_turn, window_size):
    window = deque(maxlen=window_size)
    for turn in range(1, query_turn + 1):     # the query turn is in the window too
        window.append(turn)
    return fact_turn in window

cases = [(5, 1, 6), (5, 1, 3), (5, 3, 8), (5, 3, 5), (10, 1, 6),
         (10, 1, 11), (10, 3, 8), (10, 3, 15), (20, 1, 15), (20, 1, 22)]
print("Window | fact turn -> query turn | in window")
for window_size, fact_turn, query_turn in cases:
    rule = fact_still_visible(fact_turn, query_turn, window_size)
    assert rule == visible_in_a_real_window(fact_turn, query_turn, window_size)
    print(f"{window_size:>6} | {fact_turn:>9} -> {query_turn:<10} | {'yes' if rule else 'EVICTED'}")

for window_size in (5, 10, 20):
    last = max(q for q in range(1, 100) if fact_still_visible(1, q, window_size))
    print(f"Window {window_size}: a fact from turn 1 is last visible at turn {last}")

A window of 10 turns shows a fact from turn 1 up to turn 10, that is, for 9 more turns, and loses it at turn 11. A window of 5 loses it at turn 6, a window of 20 at turn 21.

The forgetting demo · from the Complete AI Security Course in 8 Hours video · 3:25:04 to 3:28:04

This part of the video starts at 3:25:04. It runs six turns through the notebook's SlidingWindowMemory with a window of 3 turns and reads the [EVICT] lines and the reply to turn 6.

In the notebook's saved run the reply to turn 4 is a general explanation of SIPs, and the forgetting shows at turn 6. The notebook's window is a deque of six messages, so a request loses one message before the call and one after it, and a full window starts with an assistant reply. The class below keeps whole turns instead, so a request always starts with a user message.

Building the window class

A deque of turns

deque(maxlen=window_size) is the whole mechanism: appending to a full deque removes the item at the other end. A user message opens a new turn, and the assistant reply is added to that same turn. Every message also goes to archive, which is never sent.

python
class SlidingWindowMemory:
    def __init__(self, system_prompt, window_size=3):
        self.system_prompt = system_prompt
        self.window = deque(maxlen=window_size)   # whole turns, oldest first
        self.archive = []                         # every message, never sent

    def add_message(self, role, content):
        message = {"role": role, "content": content}
        self.archive.append(message)
        if role == "user":
            self.window.append([message])         # a new turn; a full deque drops its oldest
        else:
            self.window[-1].append(message)       # the reply completes the newest turn

The request from the window

The request is the system prompt followed by the messages of the turns still in the window.

python
    def get_messages_for_api(self):
        system = {"role": "system", "content": self.system_prompt}
        return [system] + [message for turn in self.window for message in turn]

A check on what was sent

chat() is the same four steps as before. One extra line looks for the text 1,20,000 in the messages of the request, so the output says for each turn whether the salary figure was sent.

python
salary_sent = any("1,20,000" in m["content"] for m in request[1:])

Running the six turns from the video

The six user messages are the notebook's. Turns 4 and 5 are general questions, chosen so that the replies have no reason to repeat the salary. Turn 6 asks for it.

ExampleAPI keyFrom the video, run on Groq
import os
from collections import deque
from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"

FINCOACH_SYSTEM_PROMPT = """You are FinCoach, a personal financial advisor assistant.
You serve users in India who want guidance on savings, investments, budgeting, and financial planning.

Your principles:
- Always personalise advice using information the user has shared in this conversation.
- Be specific with numbers when the user has provided their financial details.
- Flag when you are making assumptions due to missing information.
- Keep responses concise: 3 to 5 sentences unless the user asks for detail.
- Never provide specific buy/sell recommendations on individual stocks.
- Always recommend consulting a SEBI-registered advisor for major financial decisions.

Use all context in the conversation history to provide personalised, consistent advice."""

class SlidingWindowMemory:
    def __init__(self, system_prompt, window_size=3):
        self.system_prompt = system_prompt
        self.window = deque(maxlen=window_size)   # whole turns, oldest first
        self.archive = []                         # every message, never sent

    def add_message(self, role, content):
        message = {"role": role, "content": content}
        self.archive.append(message)
        if role == "user":
            self.window.append([message])         # a new turn; a full deque drops its oldest
        else:
            self.window[-1].append(message)       # the reply completes the newest turn

    def get_messages_for_api(self):
        system = {"role": "system", "content": self.system_prompt}
        return [system] + [message for turn in self.window for message in turn]

def chat(memory, user_message):
    memory.add_message("user", user_message)
    request = memory.get_messages_for_api()
    response = client.chat.completions.create(
        model=MODEL, max_tokens=1024, temperature=0, messages=request)
    reply = response.choices[0].message.content
    memory.add_message("assistant", reply)
    turn = len(memory.archive) // 2
    first = turn - len(memory.window) + 1
    salary_sent = any("1,20,000" in m["content"] for m in request[1:])
    print(f"[Turn {turn}] window: turns {first} to {turn} | messages sent: {len(request)} | "
          f"prompt_tokens: {response.usage.prompt_tokens} | salary in the request: {salary_sent}")
    return reply

sliding_memory = SlidingWindowMemory(FINCOACH_SYSTEM_PROMPT, window_size=3)
demo_turns = [
    "Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000.",
    "My monthly expenses are about ₹60,000 for rent, food, and transport.",
    "I have an FD of ₹50,000 that matures in 3 months. I am risk-averse.",
    "What is a Systematic Investment Plan and how does it work?",
    "What are the different types of mutual funds available in India?",
    "What is my exact monthly take-home salary that I told you at the start?",
]
for user_message in demo_turns:
    reply = chat(sliding_memory, user_message)
print()
print("User, turn 6:", demo_turns[-1])
print("FinCoach, turn 6:", reply)
print()
print("Messages in the archive:", len(sliding_memory.archive))

What turn 6 shows

  • The request stops at 6 messages. Turns 3 to 6 each send the system prompt and the turns in the window.
  • The tokens still move. With a full window prompt_tokens, the API's usage figure, is 862, 793 and 749 at turns 4, 5 and 6: the number of turns is fixed, the length of the replies is not.
  • The salary figure leaves the request at turn 4, when turn 1 falls off, and the column stays False: no later message repeated it.
  • At turn 6 the model says it has no record of the salary and adds that the user has not shared that figure yet. The second half is wrong from the user's side: the figure was given in turn 1 and evicted. A model cannot tell a message that was evicted from one that was never sent.
  • The archive holds all 12 messages. Nothing was deleted; the six oldest were not sent.

Sliding window memory vs conversation buffer memory

Conversation bufferSliding window
What is sentEvery turnThe last k turns
Messages per requestGrow every turnFixed once the window is full
Tokens per requestGrow every turnBounded by k turns, but vary with message length
Early factsKeptLost when their turn leaves the window
Extra model callsNoneNone
TuningNoneThe window size k

Where you use sliding window memory

  • Chats where only the recent exchange matters. A coding helper answering follow-up questions about the last snippet rarely needs turn 1.
  • The short-term layer under a long-term store. The video's verdict is that a window is almost always present in production and almost never alone: it is paired with a store that keeps the facts that must not be lost.
  • Latency-sensitive paths. Eviction is a list operation, with no model call.
Watch out. A fact that left the window is gone only if no later message repeats it. An assistant reply that quotes the salary carries it forward, which is why the video's demo asks two general questions before the recall question. Check what is in the request, not what you think was evicted.
Try it yourself
  • In the fixed conversation, change deque(maxlen=3) to deque(maxlen=5): the salary stays in the window through turn 5 and leaves at turn 6, and the window of turn 8 holds 168 tokens instead of 54.
  • Add (3, 1, 4) and (3, 1, 3) to cases in the rule example: with a window of 3, a fact from turn 1 is still in the window at turn 3 and EVICTED at turn 4, as in the diagram.
  • In the live run, set window_size=6: every turn stays in the window, the last column reads True on all six lines, and the model can state the salary.

Slow is fine. Stopping is the only problem.