Semantic memory
Semantic memory is a long-term memory technique that distils general, reusable facts about the user from many past sessions and stores them without the date or the session they came from.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
Episodic memory records what happened in each session. After a dozen sessions an agent should not have to re-read a dozen episodes to know that this user gets nervous when markets fall. Semantic memory keeps that conclusion as one short sentence.
The facts are text in a store, and using them means adding that text to the prompt. No model weight is retrained or updated.
This part of the video starts at 5:16:05. Semantic memory extracts and stores facts from conversation, which builds persistent knowledge.
Facts without the moment they were learned
The video's example from everyday life: you know that Paris is the capital of France, but you probably cannot remember the exact moment you learned it. The fact stayed and the episode faded. That is semantic memory.
For an agent the video gives this case. A user mentions their favourite language, Python, in session one, their team size in session two and their deployment target in session three. Without semantic memory every session starts from zero. With it the agent builds a growing profile of distilled facts: over time it stops asking the same questions and anticipates needs from what it already knows.
Semantic memory here means what the agent has learned about the user from conversations. Looking facts up in documents, as a RAG application does, is a different thing: that is retrieval from a knowledge base, not memory of the user.
This part of the video starts at 5:18:42. The notebook's table of the two memory types, with one example each.
Episodic memory vs semantic memory
The notebook on screen, 9_semantic_memory.ipynb, puts one line of each kind next to the other. The episodic line is a diary entry: "On June 12, Chiru was anxious about markets and chose debt funds." The semantic line is an encyclopedia entry about the same person: "Chiru consistently panics during market volatility and needs reassurance before deciding." The first was true on a date. The second holds whenever you read it.
| Episodic memory | Semantic memory | |
|---|---|---|
| Unit stored | A complete session record | One distilled fact |
| Time-bound | Yes: "on June 12" | No: holds across sessions |
| Source | One session | A pattern across many sessions |
| Changes | Written once | Meant to change as patterns change; the small store below only adds and strengthens |
| Answers | "What happened in April?" | "What is generally true about this user?" |
| Token cost | Medium: episode summaries | Low: short fact statements |
The notebook's one-line summary: episodic memory records what happened, semantic memory extracts what it means. It also sets semantic memory apart from Entity memory. An entity profile holds what the user stated, as exact values in fixed fields, such as a salary. Semantic memory holds what the agent observed, as free sentences about behaviour, such as "avoids equity discussions".
A fact store with confidence
Each fact carries a confidence between 0 and 1. The notebook's store raises it by 0.1 when a later distillation produces the same fact again, and only facts at or above a threshold of 0.65 are put into the prompt. The pieces below are a small version of its SemanticMemoryStore.
Merging new facts into the store
As in the notebook, a fact's id is a hash of the first 80 characters of its statement. Two facts are the same fact only if their text is the same.
import hashlib
facts = {} # fact id -> fact
def merge(new_facts):
added = 0
for f in new_facts:
fact_id = hashlib.md5(f["statement"][:80].encode()).hexdigest()[:12]
if fact_id in facts: # the same sentence again: confidence goes up
facts[fact_id]["confidence"] = round(min(1.0, facts[fact_id]["confidence"] + 0.1), 2)
else:
facts[fact_id] = dict(f)
added += 1
return added, len(new_facts) - added # (new facts, strengthened facts)The profile as prompt text
def profile_text(threshold=0.65):
reliable = [f for f in facts.values() if f["confidence"] >= threshold]
reliable.sort(key=lambda f: f["confidence"], reverse=True)
return "SEMANTIC PROFILE of the user:\n" + "\n".join(
f"- {f['statement']} [{f['category']}, confidence {f['confidence']}]" for f in reliable)Two batches of facts
The statements below are taken from the saved output of the notebook's run; the confidence of the last one is set to 0.6 by hand to show the threshold. No model is called. Put the two pieces above in one file and add these lines.
first = [{"statement": "The user prefers to avoid market risk in his investments.",
"category": "risk", "confidence": 0.8},
{"statement": "The user prefers short, direct responses over detailed explanations.",
"category": "communication", "confidence": 0.8}]
later = [{"statement": "The user prefers to avoid market risk in his investments.",
"category": "risk", "confidence": 0.8},
{"statement": "The user prefers to invest in fixed deposits and debt options over equity investments.",
"category": "financial", "confidence": 0.8},
{"statement": "The user is showing a growing interest in equity investments, indicating a shift from previous avoidance.",
"category": "financial", "confidence": 0.8},
{"statement": "The user is building an emergency fund and has a specific target for its size.",
"category": "financial", "confidence": 0.6}]
print("first batch (added, strengthened):", merge(first))
print("second batch (added, strengthened):", merge(later))
print(len(facts), "facts stored\n")
print(profile_text())first batch (added, strengthened): (2, 0) second batch (added, strengthened): (3, 1) 5 facts stored SEMANTIC PROFILE of the user: - The user prefers to avoid market risk in his investments. [risk, confidence 0.9] - The user prefers short, direct responses over detailed explanations. [communication, confidence 0.8] - The user prefers to invest in fixed deposits and debt options over equity investments. [financial, confidence 0.8] - The user is showing a growing interest in equity investments, indicating a shift from previous avoidance. [financial, confidence 0.8]
What the merge kept, strengthened and missed
- The first batch adds 2 facts.
- The second batch strengthens 1 and adds 3. The market-risk sentence arrives again, word for word, and its confidence goes from 0.8 to 0.9.
- The emergency-fund fact, at 0.6, is stored but not shown. It is under the 0.65 threshold, so 5 facts are stored and 4 reach the prompt.
- A paraphrase counts as a new fact. "Prefers to invest in fixed deposits and debt options over equity" says nearly the same as "prefers to avoid market risk", but the text differs, so nothing is strengthened.
- Two facts that contradict each other are both in the profile: the user avoids market risk, and the user shows a growing interest in equity. A comparison of text cannot see the conflict, and nothing is weakened. The saved output of the notebook's run ends the same way, with both facts injected at 0.80.
Catching a paraphrase takes a comparison by meaning: embed each statement, as in Vector store memory, and treat a new fact that is very close to a stored one as the same fact. Catching a conflict takes a model call that is shown both sentences and asked which one holds now.
Distilling facts from episodes
The facts come from a model call that reads several episode summaries at once and writes down what holds across them. The notebook runs it after every second episode, with gpt-4o-mini and max_tokens=800. The code here runs it once over four episodes with openai/gpt-oss-120b. Since gpt-oss counts its reasoning tokens inside max_tokens and a JSON reply that is cut off is an HTTP 400 error, the call uses reasoning_effort="low" and max_tokens=2000.
The client
import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"Four episode summaries
These are the notebook's four simulated sessions with Chiru.
EPISODES = [
{"episode_id": "ep_001", "date": "2025-03-15", "summary": (
"Chiru discussed his salary of ₹1,20,000 and expenses of ₹60,000. "
"He expressed anxiety about market conditions and was reluctant to consider "
"equity mutual funds despite FinCoach's recommendation. He ultimately chose "
"to keep his FD and asked only about debt options. Session outcome was positive "
"but Chiru declined all equity suggestions.")},
{"episode_id": "ep_002", "date": "2025-04-10", "summary": (
"Market dip was mentioned early in the session. Chiru immediately became "
"anxious and asked to move existing savings to an FD. FinCoach explained "
"that debt mutual funds offer better returns than FDs with low risk. "
"Chiru needed significant reassurance before accepting the suggestion. "
"He started a ₹3,000/month SIP in a debt fund only after FinCoach emphasised "
"that principal was protected. He asked the same safety question three times.")},
{"episode_id": "ep_003", "date": "2025-05-05", "summary": (
"Chiru asked about increasing his SIP amount. When FinCoach suggested "
"adding an equity component for higher returns over a 20-year horizon, "
"Chiru declined firmly. He prefers written summaries of advice: he asked "
"FinCoach to list action items at the end of the session. He consistently "
"prefers short, direct responses over detailed explanations. Goal remains "
"retirement at 55 but he is cautious about any instrument with market risk.")},
{"episode_id": "ep_004", "date": "2025-06-01", "summary": (
"Chiru's FD of ₹50,000 matured. He was uncertain about where to deploy the "
"proceeds. FinCoach recommended a short duration debt fund. Chiru took 3 "
"sessions worth of questions before committing. He always asks about "
"worst-case scenarios before making any financial decision. He prefers to "
"hear the risk of loss quantified. He is building an emergency fund, "
"currently at 3 months of expenses, targeting 6 months.")},
]The distillation prompt
The prompt is a shortened form of the notebook's. It asks for general truths, forbids dates and amounts, and ties the confidence to how many sessions show the pattern.
DISTILLATION_PROMPT = """You are a semantic memory distiller for FinCoach.
Read a set of past session summaries and extract GENERAL TRUTHS about the user.
These are not facts about specific sessions. They are patterns that hold across sessions.
Return a JSON object with this structure:
{"facts": [{"statement": "A single general truth about the user, with no dates or sessions",
"category": "behavioural | financial | risk | preference | communication",
"confidence": 0.6 to 0.9}]}
CONFIDENCE: 0.6 pattern seen in 1 session, 0.7 in 2 sessions, 0.8 in 3 or more, 0.9 in all sessions.
RULES: Return ONLY valid JSON. No specific dates, session numbers or amounts.
Extract only patterns that appear in the summaries. Maximum 6 facts.
Write in third person: 'The user...'"""The distil function
def distil(episodes):
text = "\n\n".join(f"[Session {i} on {e['date']}]:\n{e['summary']}" for i, e in enumerate(episodes, 1))
reply = client.chat.completions.create(
model=MODEL, temperature=0, max_tokens=2000, reasoning_effort="low",
response_format={"type": "json_object"},
messages=[{"role": "system", "content": DISTILLATION_PROMPT},
{"role": "user", "content": f"Extract general semantic facts from these "
f"{len(episodes)} sessions:\n\n{text}"}])
return json.loads(reply.choices[0].message.content)["facts"]From four episodes to a profile, and a reply that uses it
The question is the first turn of the notebook's live conversation. It is asked twice, once with only the base instructions and once with the profile added as a second system message. Add the client, the episodes, the prompt and distil to the file with the fact store, with the imports at the top, and replace the lines at the end with these.
added, strengthened = merge(distil(EPISODES)) # one model call for all four episodes
print(f"{len(EPISODES)} episodes -> {added} new facts, {strengthened} strengthened\n")
print(profile_text())
SYSTEM = ("You are FinCoach, a personal finance assistant for users in India. "
"Answer in at most three sentences. Never recommend specific stocks.")
question = "Hi, I have ₹50,000 to invest. What should I do?"
for label, memory in [("WITHOUT THE PROFILE", []),
("WITH THE PROFILE", [{"role": "system", "content": profile_text()}])]:
reply = client.chat.completions.create(
model=MODEL, temperature=0, max_tokens=600,
messages=[{"role": "system", "content": SYSTEM}] + memory + [{"role": "user", "content": question}])
print(f"\n{label}: {reply.choices[0].message.content}")4 episodes -> 6 new facts, 0 strengthened SEMANTIC PROFILE of the user: - The user is highly risk‑averse and consistently avoids equity investments, preferring debt‑only or fixed‑deposit options [risk, confidence 0.9] - The user requires extensive reassurance and repeatedly asks about worst‑case loss scenarios before committing to any product [behavioural, confidence 0.9] - The user prefers concise, direct responses and written summaries of advice rather than detailed explanations [communication, confidence 0.9] - The user is focused on building a safety‑first emergency fund, targeting six months of expenses [financial, confidence 0.9] - The user is comfortable with systematic investment plans in debt funds when the principal is explicitly protected [preference, confidence 0.9] - The user repeatedly asks the same safety‑related question multiple times, indicating a need for repeated confirmation [behavioural, confidence 0.9] WITHOUT THE PROFILE: With ₹50,000, first ensure you have an emergency fund (≈3‑6 months of expenses) in a liquid savings account, then allocate the remaining amount across low‑cost diversified options such as an equity‑linked SIP in a broad market index fund and a short‑term debt fund for stability. Keep your investment horizon and risk tolerance in mind—longer horizons can handle more equity exposure, while a shorter horizon may favor debt or hybrid funds. Review and rebalance your portfolio periodically (e.g., annually) to stay aligned with your goals. WITH THE PROFILE: Put the entire ₹50,000 into a liquid/ultra‑short‑term debt fund or a bank fixed deposit of 6‑12 months, which preserves your principal and offers modest interest. If you prefer a systematic approach, you can split ₹30,000 into a short‑term FD and invest the remaining ₹20,000 as a monthly SIP in a capital‑protected debt fund, which historically shows negligible risk of loss. In the worst‑case scenario, the FD guarantees return of principal plus interest, and the debt‑fund’s NAV is designed to stay at or above the initial amount, keeping your capital essentially safe.
What was distilled and how it changed the reply
- Four episodes became 6 facts in one call, each a sentence about the user with no date. Avoiding equity, asking about the worst case and wanting concise answers are patterns that run through the four summaries.
- Every fact carries confidence 0.9. The prompt reserves 0.9 for a pattern seen in all sessions, and the emergency fund is mentioned in one of the four. The number is the model's estimate, not a count.
- One fact kept a number, "six months of expenses", although the prompt asked for no amounts. Rules in a prompt are requests, not guarantees.
- Without the profile the reply is generic: an emergency fund, then an equity-linked SIP in an index fund next to a debt fund. Equity is what this user has turned down before.
- With the profile the reply changes: a debt fund or a fixed deposit, no equity, and the worst case named without being asked, which is the second fact in the profile.
- The second reply also overreaches. It says a debt fund's NAV "is designed to stay at or above the initial amount". That is not true of debt funds, whose value can fall. A profile that says the user needs reassurance pushed the model to reassure more than the facts allow. Memory changes the tone of a reply as well as its content, so the reply still needs checking, as in Input and output rails.
Where you use semantic memory
- Long relationships. An advisor, tutor or support agent that has seen the same user many times and should act on patterns, not on single remarks.
- Keeping the prompt small. A handful of fact sentences costs far fewer tokens than the episodes they were distilled from.
- As one layer of several. The video's advice for production: never a single memory layer, always a hybrid. Facts that the user states go to an entity profile, sessions to episodic memory, patterns to semantic memory.
Related
- Previous: Episodic memory
- Next: Procedural memory
- Reference: LangMem: types of memory
- In the two-batch example, print
profile_text(threshold=0.85): only the market-risk fact, at 0.9, is left. - Add
print(merge(later))at the end and print the profile again: the result is(0, 4), the market-risk fact reaches 1.0, and the emergency-fund fact, now at 0.7, enters the profile. - In
later, change "in his investments" to "in investments": the second batch returns(4, 0)and the profile holds both versions of the sentence.
Every expert started right here.