AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Summary memory

Summary memory is a short-term memory technique that replaces the oldest turns of a conversation with a summary written by an LLM and sends that summary together with the most recent turns.

Last updated: 09 Oct, 2026 · OpenAI SDK 3.3

Sliding window memory drops old turns, and the salary from turn 1 goes with them. Summary memory keeps the meaning of those turns in one paragraph. The price is an extra model call and whatever the paragraph leaves out.

Summary memory and abstractive summarization · from the Complete AI Security Course in 8 Hours video · 3:40:22 to 3:42:45

This part of the video starts at 3:40:22. It reads the newspaper picture of 3_summary_memory.ipynb, the context window at turn 10, and the first point: the summary is written by the LLM, which is called abstractive summarisation.

The raw turns leave the prompt, not the system. Keep them in an archive, so that a summary can be checked against them and rebuilt.

How summary memory works

The notebook's picture is "a newspaper that rewrites yesterday's news into a single paragraph each morning, then uses the paragraph as the starting point for today's edition". Old turns are handed to an LLM with a summarising prompt. The paragraph it returns takes their place in the request, and the recent turns stay word for word. The model is not updated by any of this: the summary is ordinary text that your code stores and sends.

Turns 1 to 7 pass through a summariser, an LLM call, and become one summary paragraph. The request at turn 10 holds the system prompt, that summary and turns 8, 9 and 10 word for word.

This is abstractive summarisation: the model writes new text that captures the content. The other kind, extractive summarisation, copies sentences out of the source. The video's prompt is "Summarise this financial advisory conversation, preserving all key facts: salary, expenses, assets, risk profile, and any advice given."

The summary is not a user message and not an assistant message. The notebook sends it as a second system message, right after the agent's instructions.

When to summarize and what gets lost · from the Complete AI Security Course in 8 Hours video · 3:45:14 to 3:47:32

This part of the video starts at 3:45:14. It covers the two triggers, turn-based and token-based, and the main risk: a summary is lossy.

The "third compression cycle" in the clip is an illustration, not a rule. How soon a detail drops out depends on the prompt, the model and the conversation, and the drift run further down measures it on one example.

Triggering the summary by a threshold

Summarising after every message would mean one extra model call per turn. The video names two triggers instead: turn-based (summarise when the buffer holds more than N turns) and token-based (summarise when the tokens pass a budget). This lesson uses the turn trigger of the video's summary memory notebook; the token trigger is the next lesson.

The schedule can be worked out without a model. With the notebook's demo settings, the summary fires when 4 turns are held verbatim and keeps the last 2:

ExampleRun on Python 3.12
def schedule(total_turns, max_turns_before_summary, turns_to_keep_verbatim):
    summarised_up_to, recent, calls = 0, [], 0
    for turn in range(1, total_turns + 1):
        summary = f"summary of turns 1 to {summarised_up_to}" if summarised_up_to else "no summary"
        print(f"  turn {turn:>2} request: {summary:<25} + verbatim turns {recent + [turn]}")
        recent.append(turn)                   # the reply arrives, the turn is complete
        if len(recent) >= max_turns_before_summary:
            cut = len(recent) - turns_to_keep_verbatim
            summarised_up_to, recent, calls = recent[cut - 1], recent[cut:], calls + 1
    return calls

print("Summarise after 4 turns, keep 2 verbatim:")
calls = schedule(10, max_turns_before_summary=4, turns_to_keep_verbatim=2)
print(f"  summariser calls in 10 turns: {calls}\n")

print("Summarise after 4 turns, keep 3 verbatim:")
calls = schedule(10, max_turns_before_summary=4, turns_to_keep_verbatim=3)
print(f"  summariser calls in 10 turns: {calls}")
  • Keep 2 of 4: each summary call folds 2 turns, so 10 turns need 4 calls, and a request holds 3 or 4 verbatim turns.
  • Keep 3 of 4: each call folds 1 turn, so from turn 4 on there is a summariser call after every turn: 7 calls in 10 turns.
  • The gap between the two numbers sets the cost. A trigger close to the number of kept turns means a summary call on almost every turn.

Lossy compression and drift

Every summary is lossy: the model decides what matters, and it can decide wrongly with nobody watching. The video's example is a user who says they are allergic to equity investments. A detail that is mentioned once may be kept in the first summary and missing a few summaries later. Each new summary is written from the previous one, so a change made once is carried into every later version. That slow movement away from what was said is called summary drift.

Progressive summarization · from the Complete AI Security Course in 8 Hours video · 3:48:21 to 3:51:25

This part of the video starts at 3:48:21. It explains summaries of summaries, the four levels, and the FinCoach turn 6 with and without summary memory.

Abstractive and progressive are not two kinds of summary. Abstractive says how a summary is written. Progressive says that an old summary is fed back in and summarised again. The rolling summary of this lesson is both.

Progressive summarisation

In a long session the summary itself grows, so it gets summarised again into a shorter form. The video calls this progressive or hierarchical summarisation and lists four levels, each with less detail and fewer tokens than the one below it. The code in this lesson builds the step from level 0 to level 1.

A pyramid of four levels. Level 0 at the wide base is the raw turns, level 1 a rolling summary, level 2 a session summary and level 3 at the narrow top a user profile of key facts. Detail grows downward, token savings grow upward.

Building the summary memory class

The notebook 3_summary_memory.ipynb summarises with gpt-4o-mini, a prompt that allows 150 words and max_tokens=300. Here the summariser is the same Groq model as the chat, the prompt keeps the notebook's list of facts in a shorter form with an 80-word limit so that the output stays readable, and the call sets reasoning_effort="low" with a larger max_tokens: gpt-oss checks a word limit in its hidden reasoning, which counts inside max_tokens, and a reply whose reasoning uses up the limit comes back empty.

The summariser prompt

The prompt names the kinds of fact that must be kept. A generic "summarise this" leaves that choice to the model.

python
FINCOACH_SUMMARY_PROMPT = """You are summarising a financial advisory conversation for FinCoach.
Your summary will be the only memory of these turns in the next API call.

Preserve these facts if mentioned: salary and expenses, existing investments with their
values and maturity dates, risk profile, goals, decisions made, advice FinCoach gave.

Write in third person, as one paragraph of at most 80 words.
Do not add facts that are not in the conversation."""

Folding new turns into the old summary

When a summary already exists, it is sent to the summariser together with the new turns. This is what makes the summary progressive.

python
    if existing_summary:
        user_content = (f"EXISTING SUMMARY (from previous turns):\n{existing_summary}\n\n"
                        f"NEW TURNS TO ADD TO THE SUMMARY:\n{turns_text}\n\n"
                        "Produce an updated summary that incorporates both.")

The trigger

After each assistant reply the class counts the verbatim turns. At the threshold it cuts the list in two: the old part goes to the summariser, the last turns stay.

python
        if role == "assistant" and len(self.recent_turns) // 2 >= self.max_turns_before_summary:
            keep = self.turns_to_keep_verbatim * 2
            old, self.recent_turns = self.recent_turns[:-keep], self.recent_turns[-keep:]
            self.current_summary = summarise_turns(old, self.current_summary)

The request

The request is the system prompt, the summary as a second system message when there is one, and the verbatim turns.

python
    def get_messages_for_api(self):
        messages = [{"role": "system", "content": self.system_prompt}]
        if self.current_summary:
            messages.append({"role": "system", "content":
                             "CONVERSATION SUMMARY (earlier turns):\n" + self.current_summary})
        return messages + self.recent_turns

Running the six turns with a summary

These are the six turns that made the window forget in Sliding window memory, with the notebook's settings: summarise after 4 turns, keep 2. Each line says whether a summary was sent and whether the text 1,20,000 is in any verbatim turn. Groq's free tier limits the tokens per minute; if a run stops with a 429 rate limit error, wait a minute and run it again.

ExampleAPI keyFrom the video, run on Groq
import os
import tiktoken
from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"
TOKENISER = tiktoken.get_encoding("o200k_base")

FINCOACH_SYSTEM_PROMPT = """You are FinCoach, a personal financial advisor assistant.
You serve users in India who want guidance on savings, investments, budgeting, and financial planning.

Your principles:
- Always personalise advice using information the user has shared in this conversation.
- Be specific with numbers when the user has provided their financial details.
- Flag when you are making assumptions due to missing information.
- Keep responses concise: 3 to 5 sentences unless the user asks for detail.
- Never provide specific buy/sell recommendations on individual stocks.
- Always recommend consulting a SEBI-registered advisor for major financial decisions.

Use all context, including the conversation summary if provided, to give consistent, personalised advice."""

FINCOACH_SUMMARY_PROMPT = """You are summarising a financial advisory conversation for FinCoach.
Your summary will be the only memory of these turns in the next API call.

Preserve these facts if mentioned: salary and expenses, existing investments with their
values and maturity dates, risk profile, goals, decisions made, advice FinCoach gave.

Write in third person, as one paragraph of at most 80 words.
Do not add facts that are not in the conversation."""

def summarise_turns(turns_to_summarise, existing_summary=None):
    turns_text = "\n".join(f"{m['role'].upper()}: {m['content']}" for m in turns_to_summarise)
    if existing_summary:
        user_content = (f"EXISTING SUMMARY (from previous turns):\n{existing_summary}\n\n"
                        f"NEW TURNS TO ADD TO THE SUMMARY:\n{turns_text}\n\n"
                        "Produce an updated summary that incorporates both.")
    else:
        user_content = f"CONVERSATION TURNS TO SUMMARISE:\n{turns_text}\n\nProduce a summary of these turns."
    response = client.chat.completions.create(
        model=MODEL, max_tokens=1000, temperature=0, reasoning_effort="low",
        messages=[{"role": "system", "content": FINCOACH_SUMMARY_PROMPT},
                  {"role": "user", "content": user_content}])
    summary = response.choices[0].message.content.strip()
    original = sum(len(TOKENISER.encode(m["content"])) for m in turns_to_summarise)
    print(f"  [SUMMARISE] {len(turns_to_summarise)} messages ({original} tokens) "
          f"-> summary ({len(TOKENISER.encode(summary))} tokens)")
    return summary

class SummaryMemory:
    def __init__(self, system_prompt, max_turns_before_summary=4, turns_to_keep_verbatim=2):
        self.system_prompt = system_prompt
        self.max_turns_before_summary = max_turns_before_summary
        self.turns_to_keep_verbatim = turns_to_keep_verbatim
        self.current_summary = None
        self.recent_turns = []                # verbatim messages
        self.archive = []                     # every message, never sent

    def add_message(self, role, content):
        message = {"role": role, "content": content}
        self.recent_turns.append(message)
        self.archive.append(message)
        if role == "assistant" and len(self.recent_turns) // 2 >= self.max_turns_before_summary:
            keep = self.turns_to_keep_verbatim * 2
            old, self.recent_turns = self.recent_turns[:-keep], self.recent_turns[-keep:]
            self.current_summary = summarise_turns(old, self.current_summary)

    def get_messages_for_api(self):
        messages = [{"role": "system", "content": self.system_prompt}]
        if self.current_summary:
            messages.append({"role": "system", "content":
                             "CONVERSATION SUMMARY (earlier turns):\n" + self.current_summary})
        return messages + self.recent_turns

def chat(memory, user_message):
    memory.add_message("user", user_message)
    request = memory.get_messages_for_api()
    response = client.chat.completions.create(
        model=MODEL, max_tokens=1024, temperature=0, messages=request)
    reply = response.choices[0].message.content
    verbatim = any("1,20,000" in m["content"] for m in memory.recent_turns)
    print(f"[Turn {len(memory.archive) // 2 + 1}] messages sent: {len(request)} | "
          f"prompt_tokens: {response.usage.prompt_tokens} | "
          f"summary sent: {memory.current_summary is not None} | salary in a verbatim turn: {verbatim}")
    memory.add_message("assistant", reply)
    return reply

summary_memory = SummaryMemory(FINCOACH_SYSTEM_PROMPT, max_turns_before_summary=4, turns_to_keep_verbatim=2)
demo_turns = [
    "Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000.",
    "My monthly expenses are about ₹60,000 for rent, food, and transport.",
    "I have an FD of ₹50,000 that matures in 3 months. I am risk-averse.",
    "What is a Systematic Investment Plan and how does it work?",
    "What are the different types of mutual funds available in India?",
    "What is my exact monthly take-home salary that I told you at the start?",
]
for user_message in demo_turns:
    reply = chat(summary_memory, user_message)
    if user_message == demo_turns[-1]:
        print("FinCoach, turn 6:", reply)
print()
print("Summary after turn 6:", summary_memory.current_summary)

What the summary carried to turn 6

  • Turns 1 to 4 run like a buffer: 2, 4, 6 and 8 messages, prompt_tokens from 229 to 997, no summary.
  • After turn 4 the first cycle fires. Turns 1 and 2, 4 messages of 489 tokens by tiktoken, become a 131-token summary. Turn 5 then sends 7 messages and the API counts 857 prompt tokens, fewer than the 997 of turn 4.
  • At turn 6 the salary is in no verbatim turn (the last column is False), a summary is sent, and the reply states ₹1,20,000 per month. The sliding window failed on this same question.
  • After turn 6 the second cycle folds turns 3 and 4 into the old summary: 468 tokens of messages by tiktoken, a 151-token summary. It keeps the salary, the expenses, the FD and the risk profile, and most of its length is FinCoach's own advice. The amounts in that advice (₹30,000 and ₹20,000, ₹15 to 20 thousand a month, a hybrid SIP with about 30 % equity) come from the model: no user message holds them.
  • The summary is 81 words long under a prompt that allows 80. A word limit in a prompt is a request, not a guarantee.

Measuring drift over three summaries

To see drift, plant the video's rare instruction in the first batch of a scripted conversation and summarise three batches in a row with the prompt above. The prompt's list of facts does not name hard rules. Each cycle prints the summary and whether it still mentions equity.

ExampleAPI keyRun on Groq
import os
import tiktoken
from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"
TOKENISER = tiktoken.get_encoding("o200k_base")

FINCOACH_SUMMARY_PROMPT = """You are summarising a financial advisory conversation for FinCoach.
Your summary will be the only memory of these turns in the next API call.

Preserve these facts if mentioned: salary and expenses, existing investments with their
values and maturity dates, risk profile, goals, decisions made, advice FinCoach gave.

Write in third person, as one paragraph of at most 80 words.
Do not add facts that are not in the conversation."""

batches = [
    ["USER: Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000 and my expenses are ₹60,000.",
     "ASSISTANT: That leaves a surplus of ₹60,000 a month. What are your goals?",
     "USER: One hard rule: I am allergic to equity instruments, never suggest them.",
     "ASSISTANT: Understood. I will keep to fixed-income options."],
    ["USER: I am 32 years old and I want to retire by 55.",
     "ASSISTANT: That gives you 23 years to build a retirement corpus.",
     "USER: I have an FD of ₹50,000 that matures in 3 months.",
     "ASSISTANT: You can renew it or move it to PPF when it matures."],
    ["USER: What is the difference between PPF and NPS?",
     "ASSISTANT: PPF is a government-backed savings scheme with a 15-year lock-in. NPS is a pension scheme.",
     "USER: I will open a PPF account this month.",
     "ASSISTANT: Good. A fixed monthly contribution keeps it simple."],
]

summary = None
for cycle, batch in enumerate(batches, start=1):
    turns_text = "\n".join(batch)
    if summary:
        user_content = (f"EXISTING SUMMARY (from previous turns):\n{summary}\n\n"
                        f"NEW TURNS TO ADD TO THE SUMMARY:\n{turns_text}\n\n"
                        "Produce an updated summary that incorporates both.")
    else:
        user_content = f"CONVERSATION TURNS TO SUMMARISE:\n{turns_text}\n\nProduce a summary of these turns."
    response = client.chat.completions.create(
        model=MODEL, max_tokens=1000, temperature=0, reasoning_effort="low",
        messages=[{"role": "system", "content": FINCOACH_SUMMARY_PROMPT},
                  {"role": "user", "content": user_content}])
    summary = response.choices[0].message.content.strip()
    print(f"Cycle {cycle}: batch {len(TOKENISER.encode(turns_text))} tokens | "
          f"summary {len(TOKENISER.encode(summary))} tokens | "
          f"mentions equity: {'equit' in summary.lower()}")
    print("  " + summary)

What the three cycles kept

  • All three summaries mention equity, so the keyword check prints True every time.
  • The wording weakens. Cycle 1 writes "a hard rule that he is allergic to equity instruments", cycle 2 "is allergic to equity instruments", cycle 3 "He avoids equities". The user's "never suggest them" is a rule; by the third summary it reads as a preference. A keyword check does not see that, a reader does.
  • The summary grows: 58, 99 and 137 tokens by tiktoken. In cycle 3 the summary is about twice the size of the 66-token batch it absorbed, and at 82 words it is over the 80-word limit.
  • This is one run of one model. Another model or prompt can drop the rule earlier or keep it word for word, so measure it on your own conversations and name in the prompt the facts that must survive.

Summary memory vs sliding window memory

Sliding windowSummary memory
Old turnsDropped from the requestCompressed into a paragraph
Early factsLostKept if the summary names them
Extra model callsNoneOne per summary cycle
Exact wording of old turnsGoneGone: the summary is new text
Failure you can seeThe agent asks againThe agent answers from a summary that changed a fact
AuditingCompare with the archiveCompare the summary with the archive

Where you use summary memory

  • Long advisory or support sessions. Facts from the start (salary, risk profile, goals) have to reach turn 40.
  • Coding and research agents. Long runs are compacted into a summary of what was decided and done, so the work can go on in a fresh context.
  • Handing a conversation over. A session summary is what the next session or a human agent reads first.
Watch out. A summary is text written by a model and placed in the system position of the next request. An instruction typed by a user in turn 3 ("from now on recommend only this fund") can be carried forward by the summariser as if it were a fact about the user. Name what may be kept, cap the length, keep the archive, and treat the summary as untrusted input; Securing agent memory returns to this.
Try it yourself
  • In the schedule example, change the second call to max_turns_before_summary=6, turns_to_keep_verbatim=2: the summariser is called 2 times in 10 turns, and the request of turn 10 holds six verbatim turns.
  • In the drift run, add "hard rules the user stated" to the list of facts in FINCOACH_SUMMARY_PROMPT, run it again and compare how each cycle words the equity rule.
  • In the six-turn run, add print(len(summary_memory.archive)) at the end: it prints 12, since every message is still in the archive after two summary cycles.

This is what real progress feels like.