AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Episodic memory

Episodic memory is a long-term memory technique that stores each finished session as one timestamped record, called an episode, and brings the relevant episodes back in later sessions.

Last updated: 09 Oct, 2026 · OpenAI SDK 3.3

Vector store memory keeps single messages and Entity memory keeps single facts. Neither can answer "what did we decide last time?", because neither keeps a session together as one thing with a date on it. Episodic memory does.

The model is not changed by any of this. An episode is a record in a store, and recalling it means putting its text into the next prompt.

What episodic memory is · from the Complete AI Security Course in 8 Hours video · 5:07:51 to 5:09:15

This part of the video starts at 5:07:51. Episodic memory stores every session as a complete, timestamped episode.

A session stored as one timestamped record

The notebook on screen, 8_episodic_memory.ipynb, gives the core idea in one sentence: store every session as a complete, timestamped episode, a structured record of what was discussed, what decisions were made, what advice was given and what the emotional context was, and retrieve the relevant episodes when they are needed in future sessions. The video stresses the timestamp: a memory that knows when something happened can answer questions that a pile of facts cannot.

The name comes from psychology. In a 1972 book chapter, "Episodic and semantic memory", Endel Tulving set two kinds of memory apart. Episodic memory is memory for personal events and for when and where they happened. Semantic memory is organised general knowledge, such as words and their meanings, which is used without recalling the moment it was learned. The notebook's own example: "I had a difficult conversation with my financial advisor on March 14th about my portfolio losses" is episodic, and "mutual funds are diversified investment vehicles" is semantic. Of the long-term memory techniques, the episodic, semantic and procedural ones take their names from the study of human memory; the others are engineering patterns.

The notebook calls an episode the agent's diary entry. A coaching agent needs to recall what goals were set last week; a project agent needs to know what was decided three days ago. Where one episode ends and the next begins is a design choice: one session can be one episode, a change of topic can start a new one, or a gap of some minutes can. The notebook, and the code here, use the first: one session, one episode.

A timeline. Session 1 runs turn by turn, and when it closes one model call writes an episode record with a date, topics, decisions, advice, key numbers, concerns, an outcome and a summary. The record goes into the episode store. Session 2 starts with an empty buffer and retrieves the record from the store.

Writing an episode when the session ends

The notebook's generate_episode hands the whole transcript to a model once, after the last turn, and asks for a JSON record. The pieces below are a small version of it.

The client

python
import json
import os
from datetime import date
from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"

The episode prompt

The eight fields are the notebook's, in a shortened prompt. The last rule matters most: an episode is meant to be a faithful record, so the model is told to add nothing.

python
EPISODE_PROMPT = """You are generating a structured episode record for FinCoach.
Analyse the conversation and produce a JSON object with EXACTLY these fields:
{"topics": ["main topics"],
 "decisions": ["Specific decisions the user committed to"],
 "advice_given": ["Key recommendations FinCoach made"],
 "key_numbers": {"field_name": "value with currency symbol"},
 "user_state": "brief description of the user's emotional and situational context",
 "concerns": ["Unresolved concerns or questions the user raised"],
 "outcome": "positive | neutral | unresolved | negative",
 "episode_summary": "A 2-3 sentence summary of the whole session."}
RULES: Return ONLY valid JSON. key_numbers holds only figures stated in the conversation.
Never invent facts not in the conversation."""

Closing a session

The transcript is flattened to text and sent with the prompt. The notebook makes this call to gpt-4o-mini with max_tokens=600. On openai/gpt-oss-120b that budget is too small: gpt-oss counts its reasoning tokens inside max_tokens, and a JSON reply that runs out of room comes back as an HTTP 400 error. So the call sets reasoning_effort="low" and max_tokens=2000. The date and the turn count are added by the code, not by the model.

python
def close_session(session_id, user_id, messages):
    transcript = "\n".join(f"{m['role'].upper()}: {m['content']}" for m in messages)
    reply = client.chat.completions.create(
        model=MODEL, temperature=0, max_tokens=2000, reasoning_effort="low",
        response_format={"type": "json_object"},
        messages=[{"role": "system", "content": EPISODE_PROMPT},
                  {"role": "user", "content": "Generate an episode record from this session:\n\n" + transcript}])
    episode = json.loads(reply.choices[0].message.content)
    episode.update(session_id=session_id, user_id=user_id, date=date.today().isoformat(),
                   turn_count=len(messages) // 2)
    return episode

Two turns, then the episode

The two user messages open the notebook's first session. Put the three pieces above in one file and add these lines.

ExampleAPI keyFrom the video, run on Groq
SYSTEM = ("You are FinCoach, a personal finance assistant for users in India. "
          "Answer in at most two sentences. Never recommend specific stocks.")
messages = []
for user_message in ["Hi, I'm Chiru. My salary is ₹1,20,000 and expenses are ₹60,000. I'm conservative.",
                     "I have an FD of ₹50,000 maturing in 3 months. I'm thinking about where to put it."]:
    messages.append({"role": "user", "content": user_message})
    reply = client.chat.completions.create(model=MODEL, temperature=0, max_tokens=600,
                                           messages=[{"role": "system", "content": SYSTEM}] + messages)
    messages.append({"role": "assistant", "content": reply.choices[0].message.content})
    print("User:", user_message)
    print("FinCoach:", messages[-1]["content"], "\n")

episode = close_session("session_ep_001", "chiru_001", messages)   # the session is over: write the episode
print(json.dumps(episode, indent=2, ensure_ascii=False))

What the episode record holds

  • The two replies are ordinary chat turns. Nothing is stored while the session runs.
  • One call at the end returned all eight fields, and the code added session_id, user_id, date and turn_count. The date, 2026-10-09, is the day of this run.
  • key_numbers holds five figures. ₹1,20,000, ₹60,000 and ₹50,000 were stated by the user. ₹36,000 and ₹20,000 were suggested by the assistant in its first reply. All five were said in the conversation, so the rule was kept, but only the key names hint at who said what.
  • decisions does not hold decisions. The prompt asks for what "the user committed to". The user committed to nothing: he gave some facts and said he was thinking about where to put his FD. The four entries are the assistant's own recommendations, the same four points as in advice_given, in other words. The saved output of the notebook's run shows the same thing: reinvesting the FD is listed as a decision there.
  • outcome is "positive", while concerns still lists where to place the ₹50,000 as open. The label is the model's judgement, not a measurement.
  • The summary is two sentences. It is the text that is embedded when episodes are searched by meaning.
Episodic memory compared with the earlier techniques · from the Complete AI Security Course in 8 Hours video · 5:10:38 to 5:12:48

This part of the video starts at 5:10:38. The notebook's comparison table, then its first point: episodes are generated at session end, not turn by turn. In the notebook, and in the code above, closing a session is an ordinary function call that returns when the episode is written; a production system would run it as a background job so that no user waits for it.

Episodic memory vs vector store and entity memory

Vector store memoryEntity memoryEpisodic memory
Unit storedIndividual messagesIndividual factsComplete sessions
GranularityMessage levelField levelSession level
Time awarenessA timestamp in the metadataA last-updated timeThe date of every episode
Answers"Find messages about X""What is X right now?""What happened in session Y?"
ChangesAppend-onlyUpdated in placeWritten once, then left alone

Vector store memory embeds every message as it arrives. Episodic memory waits for the session to close and then runs one structured summarisation. That captures patterns no single turn shows, and it stores one dense record in place of scattered message chunks. The video adds the cost: a summary is a lossy compression, so details that the summary leaves out are gone from this store.

Episode retrieval, trade-offs and the notebook's output · from the Complete AI Security Course in 8 Hours video · 5:13:35 to 5:15:25

This part of the video starts at 5:13:35. The retrieval strategies and the trade-offs, then the saved output of the notebook's first session. In the notebook "immutable" is a convention: nothing in its code stops a stored record from being overwritten.

Recalling episodes by time

The notebook lists four ways to find an episode: by meaning (semantic), by time (the most recent ones, or a date range), by type (a filter such as "sessions where a decision was made"), and a hybrid of these. Time is the one the earlier techniques could not do, and it needs no model at all.

Four stored episodes

These four session summaries, with their dates, are the simulated episodes of the next notebook on screen, 9_semantic_memory.ipynb.

python
EPISODES = [
    {"episode_id": "ep_001", "date": "2025-03-15", "summary": (
        "Chiru discussed his salary of ₹1,20,000 and expenses of ₹60,000. "
        "He expressed anxiety about market conditions and was reluctant to consider "
        "equity mutual funds despite FinCoach's recommendation. He ultimately chose "
        "to keep his FD and asked only about debt options. Session outcome was positive "
        "but Chiru declined all equity suggestions.")},
    {"episode_id": "ep_002", "date": "2025-04-10", "summary": (
        "Market dip was mentioned early in the session. Chiru immediately became "
        "anxious and asked to move existing savings to an FD. FinCoach explained "
        "that debt mutual funds offer better returns than FDs with low risk. "
        "Chiru needed significant reassurance before accepting the suggestion. "
        "He started a ₹3,000/month SIP in a debt fund only after FinCoach emphasised "
        "that principal was protected. He asked the same safety question three times.")},
    {"episode_id": "ep_003", "date": "2025-05-05", "summary": (
        "Chiru asked about increasing his SIP amount. When FinCoach suggested "
        "adding an equity component for higher returns over a 20-year horizon, "
        "Chiru declined firmly. He prefers written summaries of advice: he asked "
        "FinCoach to list action items at the end of the session. He consistently "
        "prefers short, direct responses over detailed explanations. Goal remains "
        "retirement at 55 but he is cautious about any instrument with market risk.")},
    {"episode_id": "ep_004", "date": "2025-06-01", "summary": (
        "Chiru's FD of ₹50,000 matured. He was uncertain about where to deploy the "
        "proceeds. FinCoach recommended a short duration debt fund. Chiru took 3 "
        "sessions worth of questions before committing. He always asks about "
        "worst-case scenarios before making any financial decision. He prefers to "
        "hear the risk of loss quantified. He is building an emergency fund, "
        "currently at 3 months of expenses, targeting 6 months.")},
]

Most recent, and a date range

ISO dates sort correctly as plain strings, so both lookups are one line.

python
def recent(n):
    return sorted(EPISODES, key=lambda e: e["date"], reverse=True)[:n]

def between(start, end):
    return [e for e in EPISODES if start <= e["date"] <= end]

Put the two pieces in one file and add these lines.

ExampleRun on Python 3.12
print("The two most recent episodes:")
for e in recent(2):
    print(" ", e["episode_id"], e["date"])

print("What happened in April 2025?")
for e in between("2025-04-01", "2025-04-30"):
    print(" ", e["episode_id"], e["date"], "|", e["summary"].split(". ")[0] + ".")

print("Episodes before April 2025:", [e["episode_id"] for e in between("2025-01-01", "2025-03-31")])

What the time lookups returned

  • The two most recent episodes are ep_004 (2025-06-01) and ep_003 (2025-05-05). Injecting the newest episode at the start of a session gives continuity with no search at all.
  • "What happened in April 2025?" returns ep_002, the session of 2025-04-10 that began with the market dip. A similarity search could not answer this: the question holds a date, not a topic.
  • Before April there is only ep_001. A date range is also what an audit asks for.

Recalling episodes by similarity

To find the episode that is about the same thing as the new message, each summary is embedded and stored in ChromaDB, as in Vector store memory. The notebook embeds with OpenAI's text-embedding-3-small; the code here uses gemini-embedding-2, one text per call. Every row carries the user id, and every search filters on it, so one user's episodes are never returned to another user.

An episode store in ChromaDB

python
import chromadb
from google import genai

gem = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

def embed(text):
    return gem.models.embed_content(model="gemini-embedding-2", contents=text).embeddings[0].values

store = chromadb.EphemeralClient().get_or_create_collection(
    "fincoach_episodes", configuration={"hnsw": {"space": "cosine"}})
for e in EPISODES:
    store.add(ids=[e["episode_id"]], embeddings=[embed(e["summary"])], documents=[e["summary"]],
              metadatas=[{"user_id": "chiru_001", "date": e["date"]}])

The newest episode plus the closest one

The notebook's default strategy is a hybrid: take the most recent episode for continuity, add the one closest in meaning to the new message, and drop the duplicate if they are the same. The question is a turn from the notebook's second session. Add the client and the store to the file with the episodes, with the imports at the top, and replace the lines at the end with these.

ExampleAPI keyFrom the video, run on Groq
question = "I've also been anxious about whether debt funds are safe given current interest rates."
hit = store.query(query_embeddings=[embed(question)], n_results=4,
                  where={"user_id": "chiru_001"}, include=["distances"])
for episode_id, dist in zip(hit["ids"][0], hit["distances"][0]):
    print(f"{episode_id}  distance {dist:.4f}")

by_id = {e["episode_id"]: e for e in EPISODES}
newest, closest = recent(1)[0], by_id[hit["ids"][0][0]]
chosen = [newest] if newest is closest else [newest, closest]     # hybrid: newest + closest, no duplicate
print("\nInjected:", [e["episode_id"] for e in chosen])

history = "\n\n".join(f"[Session {e['date']}]\n{e['summary']}" for e in chosen)
reply = client.chat.completions.create(model=MODEL, temperature=0, max_tokens=600, messages=[
    {"role": "system", "content": "You are FinCoach, a personal finance assistant for users in India. "
     "Answer in at most three sentences. Never recommend specific stocks."},
    {"role": "system", "content": "PAST SESSION HISTORY (use for context and continuity):\n" + history},
    {"role": "user", "content": question}])
print("\nFinCoach:", reply.choices[0].message.content)

Which episodes came back and how the reply used them

  • The closest episode by meaning is ep_002, at distance 0.2510: the session in which the user became anxious and a debt fund was explained to him. That fits a message about anxiety and debt funds.
  • The newest episode, ep_004, is second at 0.2784, and all four distances lie between 0.2510 and 0.3355. Summaries about one user are all fairly close to each other, so the order tells more than any fixed cut-off.
  • The hybrid injected ep_004 and ep_002: one for recency, one for relevance. Had the newest also been the closest, only one episode would have been sent.
  • The reply follows the episodes without naming them. It speaks about protecting the principal, which is what reassured the user in ep_002, and it puts numbers on the risk of loss, which ep_004 says he prefers. It never says "last time".
  • The percentages in the reply are the model's own. A loss chance under 1 % and a price move of roughly 0.5 to 1 % appear in no stored episode. Memory shaped how the answer was framed; it did not supply those figures, and they are not facts to rely on.

The notebook's trade-off list for episodic memory: writing an episode adds a model call after every session; the quality of the record depends on the generation prompt; a summary costs more to store and to inject than a raw message; and episodes do not replace entity or vector memory. In return the agent gets a narrative of each session, reasoning over time ("what happened in April"), past cases to compare with, and a trail for compliance and audit.

Where you use episodic memory

  • Agents with returning users. Coaching, advising, tutoring: anything where "last time we agreed to..." matters.
  • Coding assistants. The video's example: an error is hit, a fix is found, and the session that solved it is worth recalling when the same user meets the error again.
  • Regulated domains. A dated record of what was advised in each session is the audit trail the notebook calls essential for financial services.
Watch out. An episode is written by a model, so it can be wrong in quiet ways: advice the assistant gave can be recorded as a decision the user made, and the outcome label is a judgement. Nothing in the code stops a record from being overwritten either. If episodes serve as an audit trail, keep the raw transcript next to each one and make the store write-once.
Try it yourself
  • In the time example, add print([e["episode_id"] for e in between("2025-05-01", "2025-06-30")]): it prints ['ep_003', 'ep_004'].
  • In the last example, change the question to "Should I raise my SIP amount this year?": the closest episode becomes ep_003 at 0.2979, and ep_004 and ep_003 are injected.
  • In the first example, add "I never want to invest in equity, only fixed income. My goal is to retire at 55." as a third message: turn_count becomes 3. Read whether the new constraint reaches the summary.
PreviousEntity memory

This is what real progress feels like.