Episodic memory
Episodic memory is a long-term memory technique that stores each finished session as one timestamped record, called an episode, and brings the relevant episodes back in later sessions.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
Vector store memory keeps single messages and Entity memory keeps single facts. Neither can answer "what did we decide last time?", because neither keeps a session together as one thing with a date on it. Episodic memory does.
The model is not changed by any of this. An episode is a record in a store, and recalling it means putting its text into the next prompt.
This part of the video starts at 5:07:51. Episodic memory stores every session as a complete, timestamped episode.
A session stored as one timestamped record
The notebook on screen, 8_episodic_memory.ipynb, gives the core idea in one sentence: store every session as a complete, timestamped episode, a structured record of what was discussed, what decisions were made, what advice was given and what the emotional context was, and retrieve the relevant episodes when they are needed in future sessions. The video stresses the timestamp: a memory that knows when something happened can answer questions that a pile of facts cannot.
The name comes from psychology. In a 1972 book chapter, "Episodic and semantic memory", Endel Tulving set two kinds of memory apart. Episodic memory is memory for personal events and for when and where they happened. Semantic memory is organised general knowledge, such as words and their meanings, which is used without recalling the moment it was learned. The notebook's own example: "I had a difficult conversation with my financial advisor on March 14th about my portfolio losses" is episodic, and "mutual funds are diversified investment vehicles" is semantic. Of the long-term memory techniques, the episodic, semantic and procedural ones take their names from the study of human memory; the others are engineering patterns.
The notebook calls an episode the agent's diary entry. A coaching agent needs to recall what goals were set last week; a project agent needs to know what was decided three days ago. Where one episode ends and the next begins is a design choice: one session can be one episode, a change of topic can start a new one, or a gap of some minutes can. The notebook, and the code here, use the first: one session, one episode.
Writing an episode when the session ends
The notebook's generate_episode hands the whole transcript to a model once, after the last turn, and asks for a JSON record. The pieces below are a small version of it.
The client
import json
import os
from datetime import date
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"The episode prompt
The eight fields are the notebook's, in a shortened prompt. The last rule matters most: an episode is meant to be a faithful record, so the model is told to add nothing.
EPISODE_PROMPT = """You are generating a structured episode record for FinCoach.
Analyse the conversation and produce a JSON object with EXACTLY these fields:
{"topics": ["main topics"],
"decisions": ["Specific decisions the user committed to"],
"advice_given": ["Key recommendations FinCoach made"],
"key_numbers": {"field_name": "value with currency symbol"},
"user_state": "brief description of the user's emotional and situational context",
"concerns": ["Unresolved concerns or questions the user raised"],
"outcome": "positive | neutral | unresolved | negative",
"episode_summary": "A 2-3 sentence summary of the whole session."}
RULES: Return ONLY valid JSON. key_numbers holds only figures stated in the conversation.
Never invent facts not in the conversation."""Closing a session
The transcript is flattened to text and sent with the prompt. The notebook makes this call to gpt-4o-mini with max_tokens=600. On openai/gpt-oss-120b that budget is too small: gpt-oss counts its reasoning tokens inside max_tokens, and a JSON reply that runs out of room comes back as an HTTP 400 error. So the call sets reasoning_effort="low" and max_tokens=2000. The date and the turn count are added by the code, not by the model.
def close_session(session_id, user_id, messages):
transcript = "\n".join(f"{m['role'].upper()}: {m['content']}" for m in messages)
reply = client.chat.completions.create(
model=MODEL, temperature=0, max_tokens=2000, reasoning_effort="low",
response_format={"type": "json_object"},
messages=[{"role": "system", "content": EPISODE_PROMPT},
{"role": "user", "content": "Generate an episode record from this session:\n\n" + transcript}])
episode = json.loads(reply.choices[0].message.content)
episode.update(session_id=session_id, user_id=user_id, date=date.today().isoformat(),
turn_count=len(messages) // 2)
return episodeTwo turns, then the episode
The two user messages open the notebook's first session. Put the three pieces above in one file and add these lines.
SYSTEM = ("You are FinCoach, a personal finance assistant for users in India. "
"Answer in at most two sentences. Never recommend specific stocks.")
messages = []
for user_message in ["Hi, I'm Chiru. My salary is ₹1,20,000 and expenses are ₹60,000. I'm conservative.",
"I have an FD of ₹50,000 maturing in 3 months. I'm thinking about where to put it."]:
messages.append({"role": "user", "content": user_message})
reply = client.chat.completions.create(model=MODEL, temperature=0, max_tokens=600,
messages=[{"role": "system", "content": SYSTEM}] + messages)
messages.append({"role": "assistant", "content": reply.choices[0].message.content})
print("User:", user_message)
print("FinCoach:", messages[-1]["content"], "\n")
episode = close_session("session_ep_001", "chiru_001", messages) # the session is over: write the episode
print(json.dumps(episode, indent=2, ensure_ascii=False))User: Hi, I'm Chiru. My salary is ₹1,20,000 and expenses are ₹60,000. I'm conservative.
FinCoach: With a ₹1.2 L monthly income and ₹60 k expenses, aim to save at least 30 % (≈₹36 k) each month, directing ₹20 k to an emergency fund (3‑6 months of expenses) and the remainder into low‑risk options like a debt‑oriented SIP or fixed‑deposit ladder. Also, track spending regularly and review your budget quarterly to stay on track with your conservative risk profile.
User: I have an FD of ₹50,000 maturing in 3 months. I'm thinking about where to put it.
FinCoach: Since you’re conservative and the horizon is only three months, consider rolling the ₹50 k into a high‑interest savings account, a short‑term liquid‑fund, or a new short‑term FD ladder to keep the money safe while earning a modest return.
{
"topics": [
"salary and expenses",
"budgeting",
"emergency fund",
"fixed deposit maturity",
"low‑risk investment options"
],
"decisions": [
"save at least 30% of income (~₹36 k) each month",
"allocate ₹20 k to an emergency fund",
"invest remaining surplus in low‑risk options like a debt‑oriented SIP or FD ladder",
"roll the maturing ₹50 k FD into a high‑interest savings account, short‑term liquid fund, or new short‑term FD ladder"
],
"advice_given": [
"save ~30% of monthly income",
"build a 3‑6 month emergency fund",
"use low‑risk instruments for surplus",
"consider short‑term liquid or FD options for the ₹50 k maturing in 3 months"
],
"key_numbers": {
"salary": "₹1,20,000",
"expenses": "₹60,000",
"suggested_savings": "₹36,000",
"emergency_fund_allocation": "₹20,000",
"fd_maturity_amount": "₹50,000"
},
"user_state": "Conservative investor focused on safety, seeking guidance on allocating surplus income and a soon‑to‑mature fixed deposit",
"concerns": [
"best placement for the ₹50,000 FD maturing in 3 months"
],
"outcome": "positive",
"episode_summary": "Chiru, a conservative saver with a ₹1.2 L monthly salary and ₹60 k expenses, was advised to save 30% of income, build an emergency fund, and use low‑risk investments for surplus. For the ₹50 k FD maturing soon, short‑term liquid or FD options were recommended.",
"session_id": "session_ep_001",
"user_id": "chiru_001",
"date": "2026-10-09",
"turn_count": 2
}What the episode record holds
- The two replies are ordinary chat turns. Nothing is stored while the session runs.
- One call at the end returned all eight fields, and the code added
session_id,user_id,dateandturn_count. The date, 2026-10-09, is the day of this run. key_numbersholds five figures. ₹1,20,000, ₹60,000 and ₹50,000 were stated by the user. ₹36,000 and ₹20,000 were suggested by the assistant in its first reply. All five were said in the conversation, so the rule was kept, but only the key names hint at who said what.decisionsdoes not hold decisions. The prompt asks for what "the user committed to". The user committed to nothing: he gave some facts and said he was thinking about where to put his FD. The four entries are the assistant's own recommendations, the same four points as inadvice_given, in other words. The saved output of the notebook's run shows the same thing: reinvesting the FD is listed as a decision there.outcomeis "positive", whileconcernsstill lists where to place the ₹50,000 as open. The label is the model's judgement, not a measurement.- The summary is two sentences. It is the text that is embedded when episodes are searched by meaning.
This part of the video starts at 5:10:38. The notebook's comparison table, then its first point: episodes are generated at session end, not turn by turn. In the notebook, and in the code above, closing a session is an ordinary function call that returns when the episode is written; a production system would run it as a background job so that no user waits for it.
Episodic memory vs vector store and entity memory
| Vector store memory | Entity memory | Episodic memory | |
|---|---|---|---|
| Unit stored | Individual messages | Individual facts | Complete sessions |
| Granularity | Message level | Field level | Session level |
| Time awareness | A timestamp in the metadata | A last-updated time | The date of every episode |
| Answers | "Find messages about X" | "What is X right now?" | "What happened in session Y?" |
| Changes | Append-only | Updated in place | Written once, then left alone |
Vector store memory embeds every message as it arrives. Episodic memory waits for the session to close and then runs one structured summarisation. That captures patterns no single turn shows, and it stores one dense record in place of scattered message chunks. The video adds the cost: a summary is a lossy compression, so details that the summary leaves out are gone from this store.
This part of the video starts at 5:13:35. The retrieval strategies and the trade-offs, then the saved output of the notebook's first session. In the notebook "immutable" is a convention: nothing in its code stops a stored record from being overwritten.
Recalling episodes by time
The notebook lists four ways to find an episode: by meaning (semantic), by time (the most recent ones, or a date range), by type (a filter such as "sessions where a decision was made"), and a hybrid of these. Time is the one the earlier techniques could not do, and it needs no model at all.
Four stored episodes
These four session summaries, with their dates, are the simulated episodes of the next notebook on screen, 9_semantic_memory.ipynb.
EPISODES = [
{"episode_id": "ep_001", "date": "2025-03-15", "summary": (
"Chiru discussed his salary of ₹1,20,000 and expenses of ₹60,000. "
"He expressed anxiety about market conditions and was reluctant to consider "
"equity mutual funds despite FinCoach's recommendation. He ultimately chose "
"to keep his FD and asked only about debt options. Session outcome was positive "
"but Chiru declined all equity suggestions.")},
{"episode_id": "ep_002", "date": "2025-04-10", "summary": (
"Market dip was mentioned early in the session. Chiru immediately became "
"anxious and asked to move existing savings to an FD. FinCoach explained "
"that debt mutual funds offer better returns than FDs with low risk. "
"Chiru needed significant reassurance before accepting the suggestion. "
"He started a ₹3,000/month SIP in a debt fund only after FinCoach emphasised "
"that principal was protected. He asked the same safety question three times.")},
{"episode_id": "ep_003", "date": "2025-05-05", "summary": (
"Chiru asked about increasing his SIP amount. When FinCoach suggested "
"adding an equity component for higher returns over a 20-year horizon, "
"Chiru declined firmly. He prefers written summaries of advice: he asked "
"FinCoach to list action items at the end of the session. He consistently "
"prefers short, direct responses over detailed explanations. Goal remains "
"retirement at 55 but he is cautious about any instrument with market risk.")},
{"episode_id": "ep_004", "date": "2025-06-01", "summary": (
"Chiru's FD of ₹50,000 matured. He was uncertain about where to deploy the "
"proceeds. FinCoach recommended a short duration debt fund. Chiru took 3 "
"sessions worth of questions before committing. He always asks about "
"worst-case scenarios before making any financial decision. He prefers to "
"hear the risk of loss quantified. He is building an emergency fund, "
"currently at 3 months of expenses, targeting 6 months.")},
]Most recent, and a date range
ISO dates sort correctly as plain strings, so both lookups are one line.
def recent(n):
return sorted(EPISODES, key=lambda e: e["date"], reverse=True)[:n]
def between(start, end):
return [e for e in EPISODES if start <= e["date"] <= end]Put the two pieces in one file and add these lines.
print("The two most recent episodes:")
for e in recent(2):
print(" ", e["episode_id"], e["date"])
print("What happened in April 2025?")
for e in between("2025-04-01", "2025-04-30"):
print(" ", e["episode_id"], e["date"], "|", e["summary"].split(". ")[0] + ".")
print("Episodes before April 2025:", [e["episode_id"] for e in between("2025-01-01", "2025-03-31")])The two most recent episodes: ep_004 2025-06-01 ep_003 2025-05-05 What happened in April 2025? ep_002 2025-04-10 | Market dip was mentioned early in the session. Episodes before April 2025: ['ep_001']
What the time lookups returned
- The two most recent episodes are ep_004 (2025-06-01) and ep_003 (2025-05-05). Injecting the newest episode at the start of a session gives continuity with no search at all.
- "What happened in April 2025?" returns ep_002, the session of 2025-04-10 that began with the market dip. A similarity search could not answer this: the question holds a date, not a topic.
- Before April there is only ep_001. A date range is also what an audit asks for.
Recalling episodes by similarity
To find the episode that is about the same thing as the new message, each summary is embedded and stored in ChromaDB, as in Vector store memory. The notebook embeds with OpenAI's text-embedding-3-small; the code here uses gemini-embedding-2, one text per call. Every row carries the user id, and every search filters on it, so one user's episodes are never returned to another user.
An episode store in ChromaDB
import chromadb
from google import genai
gem = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
def embed(text):
return gem.models.embed_content(model="gemini-embedding-2", contents=text).embeddings[0].values
store = chromadb.EphemeralClient().get_or_create_collection(
"fincoach_episodes", configuration={"hnsw": {"space": "cosine"}})
for e in EPISODES:
store.add(ids=[e["episode_id"]], embeddings=[embed(e["summary"])], documents=[e["summary"]],
metadatas=[{"user_id": "chiru_001", "date": e["date"]}])The newest episode plus the closest one
The notebook's default strategy is a hybrid: take the most recent episode for continuity, add the one closest in meaning to the new message, and drop the duplicate if they are the same. The question is a turn from the notebook's second session. Add the client and the store to the file with the episodes, with the imports at the top, and replace the lines at the end with these.
question = "I've also been anxious about whether debt funds are safe given current interest rates."
hit = store.query(query_embeddings=[embed(question)], n_results=4,
where={"user_id": "chiru_001"}, include=["distances"])
for episode_id, dist in zip(hit["ids"][0], hit["distances"][0]):
print(f"{episode_id} distance {dist:.4f}")
by_id = {e["episode_id"]: e for e in EPISODES}
newest, closest = recent(1)[0], by_id[hit["ids"][0][0]]
chosen = [newest] if newest is closest else [newest, closest] # hybrid: newest + closest, no duplicate
print("\nInjected:", [e["episode_id"] for e in chosen])
history = "\n\n".join(f"[Session {e['date']}]\n{e['summary']}" for e in chosen)
reply = client.chat.completions.create(model=MODEL, temperature=0, max_tokens=600, messages=[
{"role": "system", "content": "You are FinCoach, a personal finance assistant for users in India. "
"Answer in at most three sentences. Never recommend specific stocks."},
{"role": "system", "content": "PAST SESSION HISTORY (use for context and continuity):\n" + history},
{"role": "user", "content": question}])
print("\nFinCoach:", reply.choices[0].message.content)ep_002 distance 0.2510 ep_004 distance 0.2784 ep_001 distance 0.2879 ep_003 distance 0.3355 Injected: ['ep_004', 'ep_002'] FinCoach: Debt funds, especially short‑duration or ultra‑short‑duration ones, typically hold high‑quality bonds whose credit risk is low, so the chance of losing principal is under 1 % in normal market conditions. Even if rates rise, the impact on these funds is modest because they have a low‑duration profile, limiting price volatility to roughly 0.5‑1 % for a 100‑basis‑point rate hike. If you keep the fund for at least a year, the expected return usually stays above the prevailing FD rates while preserving your capital.
Which episodes came back and how the reply used them
- The closest episode by meaning is ep_002, at distance 0.2510: the session in which the user became anxious and a debt fund was explained to him. That fits a message about anxiety and debt funds.
- The newest episode, ep_004, is second at 0.2784, and all four distances lie between 0.2510 and 0.3355. Summaries about one user are all fairly close to each other, so the order tells more than any fixed cut-off.
- The hybrid injected ep_004 and ep_002: one for recency, one for relevance. Had the newest also been the closest, only one episode would have been sent.
- The reply follows the episodes without naming them. It speaks about protecting the principal, which is what reassured the user in ep_002, and it puts numbers on the risk of loss, which ep_004 says he prefers. It never says "last time".
- The percentages in the reply are the model's own. A loss chance under 1 % and a price move of roughly 0.5 to 1 % appear in no stored episode. Memory shaped how the answer was framed; it did not supply those figures, and they are not facts to rely on.
The notebook's trade-off list for episodic memory: writing an episode adds a model call after every session; the quality of the record depends on the generation prompt; a summary costs more to store and to inject than a raw message; and episodes do not replace entity or vector memory. In return the agent gets a narrative of each session, reasoning over time ("what happened in April"), past cases to compare with, and a trail for compliance and audit.
Where you use episodic memory
- Agents with returning users. Coaching, advising, tutoring: anything where "last time we agreed to..." matters.
- Coding assistants. The video's example: an error is hit, a fix is found, and the session that solved it is worth recalling when the same user meets the error again.
- Regulated domains. A dated record of what was advised in each session is the audit trail the notebook calls essential for financial services.
Related
- Previous: Entity memory
- Next: Semantic memory
- Reference: LangMem: types of memory
- In the time example, add
print([e["episode_id"] for e in between("2025-05-01", "2025-06-30")]): it prints['ep_003', 'ep_004']. - In the last example, change the question to
"Should I raise my SIP amount this year?": the closest episode becomes ep_003 at 0.2979, and ep_004 and ep_003 are injected. - In the first example, add
"I never want to invest in equity, only fixed income. My goal is to retire at 55."as a third message:turn_countbecomes 3. Read whether the new constraint reaches the summary.
This is what real progress feels like.