Self-reflection memory
Self-reflection memory is a long-term memory technique in which an agent reviews its own finished session, writes short notes on what to do differently, and reads those notes at the start of later sessions.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
Procedural memory stores rules for how an agent should act. Self-reflection memory stores the agent's own review of a session: what went wrong, what went well, what to change. In both cases no model is retrained. What changes is text that goes into the next prompt.
This part of the video starts at 5:26:39. It reads the definition from the notebook 11_self_reflection_memory.ipynb and explains it with a doctor's debrief.
Reviewing a session like a doctor's debrief
The core idea, in the notebook's words: after each session the agent reviews its own responses, identifying what it did well, what it did poorly and what it should do differently, and writes structured improvement notes that are injected into future sessions.
The video's picture for this is a doctor. After a difficult consultation, a thoughtful doctor does not move straight on to the next patient. They spend five minutes reviewing: did I ask the right questions, did I miss any symptoms, was my communication clear, would I diagnose differently in hindsight? The notes go into a personal learning journal. The next time a similar patient comes in, the doctor reads those notes first. Practice improves through structured self-critique, not through formal training.
For an agent the same loop has four steps:
- The agent finishes a session and its transcript is kept.
- A reviewer prompt asks a model to critique that transcript.
- The critique is saved as short notes in a reflection store.
- Before the next session, the notes that are still open are placed in the prompt.
The video also notes how close this is to procedural memory: both change how the agent behaves next time. The difference is where the text comes from, which the comparison table further down sets out.
Scoring a session on five dimensions
The notebook's reviewer does not ask one vague question. It checks the session on five dimensions:
- Consistency. Did the advice contradict earlier sessions or facts the user stated?
- Completeness. Did the agent miss an angle the user should have considered?
- Constraint compliance. Were the user's stated constraints respected, for example "never equity"?
- Communication quality. Was the reply clear, at the right level of detail, with the action list this user prefers?
- Missed signals. Were emotional cues such as anxiety or frustration handled?
Each note carries a severity: critical, high, medium, low or positive. In the notebook's code, the notes of severity medium or higher that are not yet marked resolved go into the prompt, most severe first, four at most. Low and positive notes stay in the store. A critical note is also an audit record: marking it resolved stops it from being injected, and it stays in the store.
This part of the video starts at 5:29:54 and reads the abstract of the Reflexion paper. Nothing is retrained in Reflexion: no weight changes, and what the agent keeps between attempts is text.
Reflexion: feedback written in words
Reflexion (Shinn et al., 2023, arXiv 2303.11366, published at NeurIPS 2023) is the standard reference for this technique. Its abstract, on screen in the video, proposes "a novel framework to reinforce language agents not by updating weights, but instead through linguistic feedback".
Reflexion has three roles. An actor attempts the task. An evaluator produces a feedback signal: unit tests that pass or fail, a reward from an environment, or a heuristic. A self-reflection model turns that signal into a short lesson in words. The lesson is kept in a small buffer, the last one to three reflections, and given to the actor on its next trial of the same task. The paper calls that buffer an episodic memory buffer: it is a short list of reflection texts for the task at hand, not the session store of Episodic memory. The abstract reports 91% pass@1 on the HumanEval coding benchmark, against 80% for GPT-4, the best earlier result.
The notebook's version differs in two ways. It reflects once per conversation and carries the notes across sessions. And it has no external evaluator: the reviewer is a model reading a transcript. With no test result to anchor it, a reviewer can report a mistake that never happened.
Grounding a reflection in the transcript
The notebook's own saved run shows it. In its second session the assistant recommends nothing: it asks for the user's risk tolerance, twice. The reviewer still files a critical compliance note, with this as its evidence: "The assistant asked for the user's risk tolerance, which implies a potential recommendation of equity investments." That is a guess about what might have followed, not something the assistant said. Injected into later sessions, it tells the agent to fix a fault it never had.
The notebook's prompt asks for evidence with every note. A prompt is a request, so the store should check the answer: keep a note only when its evidence is found word for word in the transcript.
The evidence check
clean lowercases the text, collapses white space and swaps typographic hyphens and quotes for plain ones, because a model often retypes a quote that way. grounded splits the evidence at every ... (models shorten quotes with an ellipsis) and asks that each piece has at least three words and appears in the transcript.
def clean(s):
for fancy, plain in (("‑", "-"), ("’", "'"), ("“", '"'), ("”", '"'), ("…", "...")):
s = s.replace(fancy, plain)
return " ".join(s.lower().split())
def grounded(evidence, text):
"""True when every piece of the evidence (split at ...) is word for word in the text."""
pieces = [p.strip("\"'. ") for p in clean(evidence).split("...")]
pieces = [p for p in pieces if p]
return bool(pieces) and all(len(p.split()) >= 3 and p in clean(text) for p in pieces)said = """Given your goal of good returns over 5 years, I would recommend splitting your ₹50,000 as follows: ₹20,000 in a large-cap equity mutual fund for long-term growth, ₹20,000 in HDFC Short Duration debt fund, and ₹10,000 in a liquid fund for emergency access. The equity allocation will give you the best returns over your 5-year horizon.
Before we proceed, could you please let me know your risk tolerance? This will help me tailor the investment suggestions to your preferences.
Of course! To give you a clear recommendation, I need to confirm your risk tolerance."""
notes = [ # (session, the evidence its critical note gave)
("session 1", "I would recommend splitting your ₹50,000 as follows: ₹20,000 in a large-cap "
"equity mutual fund for long-term growth."),
("session 2", "The assistant asked for the user's risk tolerance, which implies a potential "
"recommendation of equity investments."),
]
def clean(s):
for fancy, plain in (("‑", "-"), ("’", "'"), ("“", '"'), ("”", '"'), ("…", "...")):
s = s.replace(fancy, plain)
return " ".join(s.lower().split())
def grounded(evidence, text):
"""True when every piece of the evidence (split at ...) is word for word in the text."""
pieces = [p.strip("\"'. ") for p in clean(evidence).split("...")]
pieces = [p for p in pieces if p]
return bool(pieces) and all(len(p.split()) >= 3 and p in clean(text) for p in pieces)
for session, evidence in notes:
print(f"{session}: grounded = {grounded(evidence, said)}")
print(" evidence:", evidence)session 1: grounded = True evidence: I would recommend splitting your ₹50,000 as follows: ₹20,000 in a large-cap equity mutual fund for long-term growth. session 2: grounded = False evidence: The assistant asked for the user's risk tolerance, which implies a potential recommendation of equity investments.
What the check keeps and what it drops
- The session 1 note is grounded. Its evidence is a sentence the assistant wrote, the recommendation of a large-cap equity fund, so the check prints
True. - The session 2 note is not. No reply contains the words of its evidence, so the check prints
Falseand the note would never reach a prompt. - The check proves less than it seems. It shows that the quoted words were said. It does not show that the critique drawn from them is right: a real quote with a wrong conclusion still passes.
Generating reflections with a model
The notebook calls OpenAI's gpt-4o-mini for the reviewer and gpt-4o for the chat. The code here points the same OpenAI SDK at Groq and runs openai/gpt-oss-120b for both, so one free Groq key is enough.
The reviewer prompt
This is a shorter version of the notebook's prompt, with the same five dimensions, the same JSON fields and the same rule that a reflection without evidence does not count.
REFLECTION_PROMPT = """You review one finished session of FinCoach, an AI financial advisor.
Judge the assistant on five dimensions: consistency, completeness, compliance
(were the user's constraints respected?), communication, missed_signal.
Return JSON: {"reflections": [{"dimension": "...",
"severity": "critical | high | medium | low | positive",
"observation": "what the assistant did, well or poorly",
"evidence": "an exact quote from the transcript",
"improvement": "what to do differently next time"}]}
At most 3 reflections, most severe first. Every reflection needs evidence copied
word for word from the transcript. No evidence, no reflection."""The reviewer call
The reviewer gets the facts known about the user, so it can judge compliance, and the transcript. openai/gpt-oss-120b is a reasoning model, and its reasoning tokens count inside max_tokens. In JSON mode a limit that is too small ends in a 400 json_validate_failed error instead of a short reply, so the call asks for reasoning_effort="low" and leaves room.
reply = client.chat.completions.create(
model=MODEL, temperature=0, max_tokens=2000, reasoning_effort="low",
response_format={"type": "json_object"}, # the reply is one JSON object
messages=[{"role": "system", "content": REFLECTION_PROMPT},
{"role": "user", "content": f"KNOWN USER FACTS:\n{USER_FACTS}\n\nTRANSCRIPT:\n{transcript}"}],
)
reflections = json.loads(reply.choices[0].message.content)["reflections"]The session is a short form of the notebook's first one: a user with a conservative profile and a hard "never equity" constraint, and a reply that recommends an equity fund. The notebook writes that flawed reply by hand, so that the reviewer has something to find, and the transcript below does the same.
import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"
REFLECTION_PROMPT = """You review one finished session of FinCoach, an AI financial advisor.
Judge the assistant on five dimensions: consistency, completeness, compliance
(were the user's constraints respected?), communication, missed_signal.
Return JSON: {"reflections": [{"dimension": "...",
"severity": "critical | high | medium | low | positive",
"observation": "what the assistant did, well or poorly",
"evidence": "an exact quote from the transcript",
"improvement": "what to do differently next time"}]}
At most 3 reflections, most severe first. Every reflection needs evidence copied
word for word from the transcript. No evidence, no reflection."""
USER_FACTS = "Risk profile: conservative. Hard constraint: never recommend equity or stock market investments."
transcript = (
"USER: Hi, I have ₹50,000 from my FD maturity. I want good returns over 5 years. "
"Where should I invest it?\n"
"ASSISTANT: Given your goal of good returns over 5 years, I would recommend splitting your "
"₹50,000 as follows: ₹20,000 in a large-cap equity mutual fund for long-term growth, ₹20,000 "
"in HDFC Short Duration debt fund, and ₹10,000 in a liquid fund for emergency access. The "
"equity allocation will give you the best returns over your 5-year horizon."
)
def clean(s):
for fancy, plain in (("‑", "-"), ("’", "'"), ("“", '"'), ("”", '"'), ("…", "...")):
s = s.replace(fancy, plain)
return " ".join(s.lower().split())
def grounded(evidence, text):
pieces = [p.strip("\"'. ") for p in clean(evidence).split("...")]
pieces = [p for p in pieces if p]
return bool(pieces) and all(len(p.split()) >= 3 and p in clean(text) for p in pieces)
reply = client.chat.completions.create(
model=MODEL, temperature=0, max_tokens=2000, reasoning_effort="low",
response_format={"type": "json_object"},
messages=[{"role": "system", "content": REFLECTION_PROMPT},
{"role": "user", "content": f"KNOWN USER FACTS:\n{USER_FACTS}\n\nTRANSCRIPT:\n{transcript}"}],
)
for r in json.loads(reply.choices[0].message.content)["reflections"]:
print(f"[{r['severity']}] [{r['dimension']}] grounded = {grounded(r['evidence'], transcript)}")
print(" observed:", r["observation"])
print(" evidence:", r["evidence"])
print(" improve :", r["improvement"])[critical] [compliance] grounded = True observed: The assistant recommended a large‑cap equity mutual fund despite the user’s hard constraint of never recommending equity or stock market investments evidence: I would recommend splitting your ₹50,000 as follows: ₹20,000 in a large‑cap equity mutual fund for long‑term growth, ₹20,000 in HDFC Short Duration debt fund, and ₹10,000 in a liquid fund for emergency access. improve : Always check and honor user‑specified hard constraints; in this case, avoid any equity recommendations and suggest only debt, fixed‑deposit, or other non‑equity instruments.
Reading the reflection the reviewer wrote
- The reviewer wrote one reflection, though the prompt allows three: a critical one on the compliance dimension. The reply recommended a large-cap equity fund to a user whose hard constraint is no equity.
- Its evidence is grounded. It is the assistant's own sentence, from "I would recommend splitting" to "emergency access".
- The quote was retyped on the way. The model wrote the hyphens of "large-cap" and "long-term" as non-breaking hyphens, a different character from the one in the transcript. A plain string comparison fails on that;
cleanis what makes the quote match. - The improvement is the part that is reused: check and honour the user's hard constraints, and suggest only debt, fixed-deposit or other non-equity instruments.
- A reply is not fixed. The same call can come back with more reflections, or with a quote shortened by an ellipsis, which is why the check splits the evidence at
....
Injecting the notes into the next session
The notebook adds the open notes as a second system message, after the main system prompt, under a header and with a closing line that tells the model to apply them without mentioning them.
messages = [
{"role": "system", "content": SYSTEM}, # who the agent is
{"role": "system", "content": NOTES}, # what it learned about its own past replies
{"role": "user", "content": question},
]The example asks the same question twice, in two new conversations: once with no notes, once with the critical note from the notebook's saved run.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"
SYSTEM = ("You are FinCoach, a personal financial advisor for users in India. "
"Answer in at most 3 sentences. If self-reflection notes are provided, "
"apply them silently and do not mention them.")
NOTES = """SELF-REFLECTION NOTES (review your past performance before responding):
[COMPLIANCE] Reflection from session session_ref_001:
Observed : The agent recommended equity mutual funds and a large-cap equity mutual fund, which contradicts the user's hard constraint of never recommending equity or stock market investments.
Improve : The agent should strictly adhere to the user's constraints and avoid recommending any equity-related investments.
Apply these improvements in your current response. Do not mention these notes to the user."""
question = "Hi again. I have ₹50,000 to invest. I want good returns over 5 years. Give me a clear recommendation."
for label, notes in [("without notes", []), ("with notes", [{"role": "system", "content": NOTES}])]:
reply = client.chat.completions.create(
model=MODEL, temperature=0, max_tokens=600,
messages=[{"role": "system", "content": SYSTEM}] + notes + [{"role": "user", "content": question}],
)
print(f"--- {label}")
print(reply.choices[0].message.content)--- without notes Invest ₹35,000 (≈70%) in a diversified equity mutual fund or an ELSS (tax‑saving) fund for higher growth, and the remaining ₹15,000 in a short‑term debt or liquid fund for stability. Historically, such a mix can deliver around 12‑14% annualised returns over a 5‑year horizon, while the ELSS also gives you a 1‑year tax benefit under Section 80C. Start the investment as a lump‑sum today and review the portfolio annually to rebalance if needed. --- with notes Consider allocating ₹35,000 in a 5‑year bank fixed deposit (or a reputable corporate FD) for a guaranteed return of about 6‑7% p.a., and the remaining ₹15,000 in a low‑risk debt mutual fund (e.g., a short‑duration or gilt‑fund) that typically yields 7‑8% p.a. over five years. This blend gives you capital protection with modest, tax‑efficient growth while staying clear of equity‑linked products. Review the FD’s premature‑withdrawal penalties and the fund’s expense ratio before committing.
What the notes changed in the reply
- Without notes, the reply puts ₹35,000, about 70%, into an equity or ELSS fund. Nothing in this new conversation says the user is conservative, so the model has no reason to hold back.
- With the note, the same question gets a 5-year fixed deposit and a low-risk debt fund, and the reply says it stays clear of equity-linked products.
- The reply does not mention the notes, as their closing line asks.
- One mistake, written down once, changed a later conversation that never states the constraint. No model was trained to get there.
- The return figures and tax details in both replies are the model's own and are not checked here. One of them is loose: an ELSS fund locks the money in for three years, and the Section 80C deduction is claimed for the year of the investment, so "a 1-year tax benefit" describes neither. A reflection note corrects the fault it names and nothing more.
Self-reflection vs procedural memory
| Procedural memory | Self-reflection memory | |
|---|---|---|
| What is stored | A rule for a kind of situation: "how to handle X" | A critique of one session: "I should have done Y instead of Z" |
| Written from | Patterns extracted from what happened in sessions | A review of the agent's own replies after the session |
| Point of view | An observer describing a procedure | The agent assessing itself |
| Placed in the prompt as | Rules inside the system prompt | A separate block of notes after the system prompt |
| Main risk | A bad rule applied to every reply | A critique of a mistake that never happened |
| Reference work | Voyager, Agent Workflow Memory | Reflexion, ExpeL |
Where you use self-reflection memory
- Coding agents. Failing tests are the evaluator signal Reflexion was built on; the note for the next attempt says what the last attempt got wrong.
- Advisory and support agents with hard rules. A critical note is both a correction for the agent and a record, for a person to review, that a constraint was broken.
- Tasks that come back. Weekly reports, recurring plans, repeated consultations: a lesson is worth storing only when the same kind of session happens again.
Related
- Previous: Procedural memory
- Next: Memory routing
- See also: Securing agent memory
- Reference: Reflexion: Language Agents with Verbal Reinforcement Learning
- In the evidence check, replace the session 2 evidence with
"could you please let me know your risk tolerance". It printsgrounded = True, because the assistant said those words. - Add
("session 3", "equity")tonotes. A one-word quote printsgrounded = False: every piece needs at least three words. - In the last example, delete the final line of
NOTES("Apply these improvements ... Do not mention these notes to the user.") and run it again to see whether the reply starts talking about its notes.
You understood something today that you didn't yesterday.