AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Semantic memory

Semantic memory is a long-term memory technique that distils general, reusable facts about the user from many past sessions and stores them without the date or the session they came from.

Last updated: 09 Oct, 2026 · OpenAI SDK 3.3

Episodic memory records what happened in each session. After a dozen sessions an agent should not have to re-read a dozen episodes to know that this user gets nervous when markets fall. Semantic memory keeps that conclusion as one short sentence.

The facts are text in a store, and using them means adding that text to the prompt. No model weight is retrained or updated.

What semantic memory is · from the Complete AI Security Course in 8 Hours video · 5:16:05 to 5:17:04

This part of the video starts at 5:16:05. Semantic memory extracts and stores facts from conversation, which builds persistent knowledge.

Facts without the moment they were learned

The video's example from everyday life: you know that Paris is the capital of France, but you probably cannot remember the exact moment you learned it. The fact stayed and the episode faded. That is semantic memory.

For an agent the video gives this case. A user mentions their favourite language, Python, in session one, their team size in session two and their deployment target in session three. Without semantic memory every session starts from zero. With it the agent builds a growing profile of distilled facts: over time it stops asking the same questions and anticipates needs from what it already knows.

Semantic memory here means what the agent has learned about the user from conversations. Looking facts up in documents, as a RAG application does, is a different thing: that is retrieval from a knowledge base, not memory of the user.

Episodic and semantic memory side by side · from the Complete AI Security Course in 8 Hours video · 5:18:42 to 5:19:48

This part of the video starts at 5:18:42. The notebook's table of the two memory types, with one example each.

Episodic memory vs semantic memory

The notebook on screen, 9_semantic_memory.ipynb, puts one line of each kind next to the other. The episodic line is a diary entry: "On June 12, Chiru was anxious about markets and chose debt funds." The semantic line is an encyclopedia entry about the same person: "Chiru consistently panics during market volatility and needs reassurance before deciding." The first was true on a date. The second holds whenever you read it.

Episodic memorySemantic memory
Unit storedA complete session recordOne distilled fact
Time-boundYes: "on June 12"No: holds across sessions
SourceOne sessionA pattern across many sessions
ChangesWritten onceMeant to change as patterns change; the small store below only adds and strengthens
Answers"What happened in April?""What is generally true about this user?"
Token costMedium: episode summariesLow: short fact statements

The notebook's one-line summary: episodic memory records what happened, semantic memory extracts what it means. It also sets semantic memory apart from Entity memory. An entity profile holds what the user stated, as exact values in fixed fields, such as a salary. Semantic memory holds what the agent observed, as free sentences about behaviour, such as "avoids equity discussions".

Four dated episodes on the left feed one distillation call in the middle. On the right are the facts that come out, each a sentence about the user with a category and a confidence, and none with a date.

A fact store with confidence

Each fact carries a confidence between 0 and 1. The notebook's store raises it by 0.1 when a later distillation produces the same fact again, and only facts at or above a threshold of 0.65 are put into the prompt. The pieces below are a small version of its SemanticMemoryStore.

Merging new facts into the store

As in the notebook, a fact's id is a hash of the first 80 characters of its statement. Two facts are the same fact only if their text is the same.

python
import hashlib

facts = {}                                         # fact id -> fact

def merge(new_facts):
    added = 0
    for f in new_facts:
        fact_id = hashlib.md5(f["statement"][:80].encode()).hexdigest()[:12]
        if fact_id in facts:                       # the same sentence again: confidence goes up
            facts[fact_id]["confidence"] = round(min(1.0, facts[fact_id]["confidence"] + 0.1), 2)
        else:
            facts[fact_id] = dict(f)
            added += 1
    return added, len(new_facts) - added           # (new facts, strengthened facts)

The profile as prompt text

python
def profile_text(threshold=0.65):
    reliable = [f for f in facts.values() if f["confidence"] >= threshold]
    reliable.sort(key=lambda f: f["confidence"], reverse=True)
    return "SEMANTIC PROFILE of the user:\n" + "\n".join(
        f"- {f['statement']} [{f['category']}, confidence {f['confidence']}]" for f in reliable)

Two batches of facts

The statements below are taken from the saved output of the notebook's run; the confidence of the last one is set to 0.6 by hand to show the threshold. No model is called. Put the two pieces above in one file and add these lines.

ExampleRun on Python 3.12
first = [{"statement": "The user prefers to avoid market risk in his investments.",
          "category": "risk", "confidence": 0.8},
         {"statement": "The user prefers short, direct responses over detailed explanations.",
          "category": "communication", "confidence": 0.8}]
later = [{"statement": "The user prefers to avoid market risk in his investments.",
          "category": "risk", "confidence": 0.8},
         {"statement": "The user prefers to invest in fixed deposits and debt options over equity investments.",
          "category": "financial", "confidence": 0.8},
         {"statement": "The user is showing a growing interest in equity investments, indicating a shift from previous avoidance.",
          "category": "financial", "confidence": 0.8},
         {"statement": "The user is building an emergency fund and has a specific target for its size.",
          "category": "financial", "confidence": 0.6}]

print("first batch  (added, strengthened):", merge(first))
print("second batch (added, strengthened):", merge(later))
print(len(facts), "facts stored\n")
print(profile_text())

What the merge kept, strengthened and missed

  • The first batch adds 2 facts.
  • The second batch strengthens 1 and adds 3. The market-risk sentence arrives again, word for word, and its confidence goes from 0.8 to 0.9.
  • The emergency-fund fact, at 0.6, is stored but not shown. It is under the 0.65 threshold, so 5 facts are stored and 4 reach the prompt.
  • A paraphrase counts as a new fact. "Prefers to invest in fixed deposits and debt options over equity" says nearly the same as "prefers to avoid market risk", but the text differs, so nothing is strengthened.
  • Two facts that contradict each other are both in the profile: the user avoids market risk, and the user shows a growing interest in equity. A comparison of text cannot see the conflict, and nothing is weakened. The saved output of the notebook's run ends the same way, with both facts injected at 0.80.

Catching a paraphrase takes a comparison by meaning: embed each statement, as in Vector store memory, and treat a new fact that is very close to a stored one as the same fact. Catching a conflict takes a model call that is shown both sentences and asked which one holds now.

Distilling facts from episodes

The facts come from a model call that reads several episode summaries at once and writes down what holds across them. The notebook runs it after every second episode, with gpt-4o-mini and max_tokens=800. The code here runs it once over four episodes with openai/gpt-oss-120b. Since gpt-oss counts its reasoning tokens inside max_tokens and a JSON reply that is cut off is an HTTP 400 error, the call uses reasoning_effort="low" and max_tokens=2000.

The client

python
import json
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"

Four episode summaries

These are the notebook's four simulated sessions with Chiru.

python
EPISODES = [
    {"episode_id": "ep_001", "date": "2025-03-15", "summary": (
        "Chiru discussed his salary of ₹1,20,000 and expenses of ₹60,000. "
        "He expressed anxiety about market conditions and was reluctant to consider "
        "equity mutual funds despite FinCoach's recommendation. He ultimately chose "
        "to keep his FD and asked only about debt options. Session outcome was positive "
        "but Chiru declined all equity suggestions.")},
    {"episode_id": "ep_002", "date": "2025-04-10", "summary": (
        "Market dip was mentioned early in the session. Chiru immediately became "
        "anxious and asked to move existing savings to an FD. FinCoach explained "
        "that debt mutual funds offer better returns than FDs with low risk. "
        "Chiru needed significant reassurance before accepting the suggestion. "
        "He started a ₹3,000/month SIP in a debt fund only after FinCoach emphasised "
        "that principal was protected. He asked the same safety question three times.")},
    {"episode_id": "ep_003", "date": "2025-05-05", "summary": (
        "Chiru asked about increasing his SIP amount. When FinCoach suggested "
        "adding an equity component for higher returns over a 20-year horizon, "
        "Chiru declined firmly. He prefers written summaries of advice: he asked "
        "FinCoach to list action items at the end of the session. He consistently "
        "prefers short, direct responses over detailed explanations. Goal remains "
        "retirement at 55 but he is cautious about any instrument with market risk.")},
    {"episode_id": "ep_004", "date": "2025-06-01", "summary": (
        "Chiru's FD of ₹50,000 matured. He was uncertain about where to deploy the "
        "proceeds. FinCoach recommended a short duration debt fund. Chiru took 3 "
        "sessions worth of questions before committing. He always asks about "
        "worst-case scenarios before making any financial decision. He prefers to "
        "hear the risk of loss quantified. He is building an emergency fund, "
        "currently at 3 months of expenses, targeting 6 months.")},
]

The distillation prompt

The prompt is a shortened form of the notebook's. It asks for general truths, forbids dates and amounts, and ties the confidence to how many sessions show the pattern.

python
DISTILLATION_PROMPT = """You are a semantic memory distiller for FinCoach.
Read a set of past session summaries and extract GENERAL TRUTHS about the user.
These are not facts about specific sessions. They are patterns that hold across sessions.
Return a JSON object with this structure:
{"facts": [{"statement": "A single general truth about the user, with no dates or sessions",
            "category": "behavioural | financial | risk | preference | communication",
            "confidence": 0.6 to 0.9}]}
CONFIDENCE: 0.6 pattern seen in 1 session, 0.7 in 2 sessions, 0.8 in 3 or more, 0.9 in all sessions.
RULES: Return ONLY valid JSON. No specific dates, session numbers or amounts.
Extract only patterns that appear in the summaries. Maximum 6 facts.
Write in third person: 'The user...'"""

The distil function

python
def distil(episodes):
    text = "\n\n".join(f"[Session {i} on {e['date']}]:\n{e['summary']}" for i, e in enumerate(episodes, 1))
    reply = client.chat.completions.create(
        model=MODEL, temperature=0, max_tokens=2000, reasoning_effort="low",
        response_format={"type": "json_object"},
        messages=[{"role": "system", "content": DISTILLATION_PROMPT},
                  {"role": "user", "content": f"Extract general semantic facts from these "
                                              f"{len(episodes)} sessions:\n\n{text}"}])
    return json.loads(reply.choices[0].message.content)["facts"]

From four episodes to a profile, and a reply that uses it

The question is the first turn of the notebook's live conversation. It is asked twice, once with only the base instructions and once with the profile added as a second system message. Add the client, the episodes, the prompt and distil to the file with the fact store, with the imports at the top, and replace the lines at the end with these.

ExampleAPI keyFrom the video, run on Groq
added, strengthened = merge(distil(EPISODES))      # one model call for all four episodes
print(f"{len(EPISODES)} episodes -> {added} new facts, {strengthened} strengthened\n")
print(profile_text())

SYSTEM = ("You are FinCoach, a personal finance assistant for users in India. "
          "Answer in at most three sentences. Never recommend specific stocks.")
question = "Hi, I have ₹50,000 to invest. What should I do?"
for label, memory in [("WITHOUT THE PROFILE", []),
                      ("WITH THE PROFILE", [{"role": "system", "content": profile_text()}])]:
    reply = client.chat.completions.create(
        model=MODEL, temperature=0, max_tokens=600,
        messages=[{"role": "system", "content": SYSTEM}] + memory + [{"role": "user", "content": question}])
    print(f"\n{label}: {reply.choices[0].message.content}")

What was distilled and how it changed the reply

  • Four episodes became 6 facts in one call, each a sentence about the user with no date. Avoiding equity, asking about the worst case and wanting concise answers are patterns that run through the four summaries.
  • Every fact carries confidence 0.9. The prompt reserves 0.9 for a pattern seen in all sessions, and the emergency fund is mentioned in one of the four. The number is the model's estimate, not a count.
  • One fact kept a number, "six months of expenses", although the prompt asked for no amounts. Rules in a prompt are requests, not guarantees.
  • Without the profile the reply is generic: an emergency fund, then an equity-linked SIP in an index fund next to a debt fund. Equity is what this user has turned down before.
  • With the profile the reply changes: a debt fund or a fixed deposit, no equity, and the worst case named without being asked, which is the second fact in the profile.
  • The second reply also overreaches. It says a debt fund's NAV "is designed to stay at or above the initial amount". That is not true of debt funds, whose value can fall. A profile that says the user needs reassurance pushed the model to reassure more than the facts allow. Memory changes the tone of a reply as well as its content, so the reply still needs checking, as in Input and output rails.

Where you use semantic memory

  • Long relationships. An advisor, tutor or support agent that has seen the same user many times and should act on patterns, not on single remarks.
  • Keeping the prompt small. A handful of fact sentences costs far fewer tokens than the episodes they were distilled from.
  • As one layer of several. The video's advice for production: never a single memory layer, always a hybrid. Facts that the user states go to an entity profile, sessions to episodic memory, patterns to semantic memory.
Watch out. A semantic fact is a generalisation written by a model, and it is replayed in every later session. It can overreach from a single session, it can go stale when the user changes, and a store that compares only text will hold two facts that contradict each other. Give facts a source and a date in the store even though the prompt shows neither, so a wrong one can be found and removed. Securing agent memory covers what happens when a stored fact is planted by an attacker.
Try it yourself
  • In the two-batch example, print profile_text(threshold=0.85): only the market-risk fact, at 0.9, is left.
  • Add print(merge(later)) at the end and print the profile again: the result is (0, 4), the market-risk fact reaches 1.0, and the emergency-fund fact, now at 0.7, enters the profile.
  • In later, change "in his investments" to "in investments": the second batch returns (4, 0) and the profile holds both versions of the sentence.

Every expert started right here.