Procedural memory
Procedural memory is a long-term memory technique that stores rules and workflows for how the agent itself should act, and adds them to the system prompt as instructions.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
Semantic memory is knowledge about the user. Procedural memory is knowledge for the agent: not "this user is risk-averse" but "always quantify the worst case before the user asks". It is the one memory type that changes the agent's instructions instead of its context.
The words "learns" and "self-improving" are often used for this technique. They describe text. A list of procedures grows, and the ones that have worked are written into the system prompt of the next session. The model is never retrained and none of its weights change; the notebook says so itself: "outcome-based learning without retraining".
This part of the video starts at 5:20:11. Procedural memory keeps updating the agent's system instructions. Where the clip says the agent learns and updates its behaviour, what changes is the text of the system prompt; the model is not retrained.
Instructions the agent writes for itself
System instructions are how you tell an agent to work, think and act. Normally a developer writes them once. With procedural memory the agent adds to them from experience. The notebook on screen, 10_procedural_memory.ipynb, defines it this way: procedural memory stores reusable step-by-step workflows in a skill library, so agents can retrieve proven procedures instead of reasoning from scratch.
The human parallel in the video is riding a bicycle. You do it without thinking, and you do not recall the session in which you learned to balance. You know how. In psychology this is the third kind of long-term memory next to episodic and semantic memory: Tulving's 1972 chapter set those two apart, and his 1985 paper "How many memory systems are there?" describes three systems: procedural memory (knowing how), semantic memory and episodic memory.
Semantic memory vs procedural memory
| Semantic memory | Procedural memory | |
|---|---|---|
| Stores | Facts about the user | Rules for how the agent should act |
| Example | "Chiru is risk-averse" | "Always quantify worst-case before user asks" |
| Subject | The user | The agent's own behaviour |
| Updates | When the user's behaviour changes | When a workflow succeeds or fails |
| Enters the prompt as | Context about the user | Instructions in the system prompt |
The notebook names three forms a procedure can take:
- A workflow: steps for a recurring task. "When handling a FD maturity scenario: Step 1, check risk profile. Step 2, present 3 options. Step 3, quantify risk for each."
- A rule: a condition and an action. "IF user shows anxiety about markets THEN lead with reassurance before giving the recommendation."
- An interaction pattern: a learned way of communicating. "This user prefers numbered lists."
All three end up in the same place. The notebook's first point: procedural memory lives in the system prompt. Episodes are injected as context blocks and semantic facts as facts about the user, which the model reads as background. Procedures are injected as instructions, which the model reads as directives.
Confidence that follows outcomes
The notebook's second point is that procedures are learned from outcomes, not only stated. Each procedure has a confidence. After a session in which it was applied, the confidence goes up by 0.08 if the session went well and down by 0.15 if it did not. Only procedures at or above 0.55 are written into the prompt, at most six of them. These four numbers are the notebook's.
A procedure record
from dataclasses import dataclass
@dataclass
class Procedure:
title: str
trigger: str # when the procedure applies
description: str # the instruction, written to the agent
confidence: float = 0.6
successes: int = 0
failures: int = 0Recording an outcome
This method belongs to the Procedure class.
def record_outcome(self, positive):
if positive:
self.successes += 1
self.confidence = min(1.0, self.confidence + 0.08)
else:
self.failures += 1
self.confidence = max(0.0, self.confidence - 0.15)Building the system prompt
The base instructions come first. The procedures that pass the threshold are sorted by confidence and appended under a heading.
BASE_PROMPT = ("You are FinCoach, a personal financial advisor assistant for users in India. "
"Never recommend specific stocks. Answer in at most 120 words.")
def build_system_prompt(procedures, threshold=0.55, limit=6):
active = sorted((p for p in procedures if p.confidence >= threshold),
key=lambda p: p.confidence, reverse=True)[:limit]
if not active:
return BASE_PROMPT
lines = [f"{i}. {p.title}\n Trigger: {p.trigger}\n {p.description}" for i, p in enumerate(active, 1)]
return (BASE_PROMPT + "\n\nLEARNED OPERATIONAL PROCEDURES (apply these in relevant situations):\n"
+ "\n".join(lines))One procedure that works and one that does not
The first procedure and the "push equity" rule are the notebook's own examples, and the sequence of two good sessions, one bad one and two more good ones is the one its confidence demo runs. No model is called. Put the three pieces above in one file and add these lines.
import matplotlib.pyplot as plt
worst_case = Procedure("Quantify the worst case", "When the user asks about an investment",
"Always quantify the worst-case scenario before the user asks.")
push_equity = Procedure("Push equity", "When the user has a long horizon",
"Recommend adding an equity component for higher returns.")
paths = {}
for proc, outcomes in [(worst_case, [True, True, False, True, True]), (push_equity, [False, False])]:
paths[proc.title] = [proc.confidence]
for positive in outcomes:
proc.record_outcome(positive)
paths[proc.title].append(proc.confidence)
print(f"{proc.title}: " + " -> ".join(f"{c:.2f}" for c in paths[proc.title])
+ f" (successes {proc.successes}, failures {proc.failures})")
print()
print(build_system_prompt([worst_case, push_equity]))
plt.figure(figsize=(7, 3.2))
for title, path in paths.items():
plt.plot(range(len(path)), path, marker="o", label=title)
plt.axhline(0.55, color="gray", linestyle="--", label="threshold 0.55")
plt.xlabel("sessions in which the procedure was applied")
plt.ylabel("confidence")
plt.ylim(0, 1)
plt.title("Confidence follows outcomes: +0.08 on success, -0.15 on failure")
plt.legend(loc="lower right")
plt.show()Quantify the worst case: 0.60 -> 0.68 -> 0.76 -> 0.61 -> 0.69 -> 0.77 (successes 4, failures 1) Push equity: 0.60 -> 0.45 -> 0.30 (successes 0, failures 2) You are FinCoach, a personal financial advisor assistant for users in India. Never recommend specific stocks. Answer in at most 120 words. LEARNED OPERATIONAL PROCEDURES (apply these in relevant situations): 1. Quantify the worst case Trigger: When the user asks about an investment Always quantify the worst-case scenario before the user asks.
What the two confidence paths did
- "Quantify the worst case" moves 0.60, 0.68, 0.76, 0.61, 0.69, 0.77. Each good session adds 0.08. The one bad session takes away 0.15, nearly two successes' worth.
- It never falls below 0.55, so it is written into the system prompt of every session.
- "Push equity" moves 0.60, 0.45, 0.30. After its first failure it is already under the threshold.
- The printed system prompt holds the base instructions and one procedure. The failed rule is still in the list, but the model never sees it.
- This is all the "learning" there is: a number on a record, and a filter on which records are pasted into the prompt.
Learning procedures from a session
Where do procedures come from? In the notebook, a model reads the transcript of a finished session and writes down rules worth keeping. The notebook calls gpt-4o-mini with max_tokens=600; the code here calls openai/gpt-oss-120b on Groq. gpt-oss counts its reasoning tokens inside max_tokens, and a JSON reply that is cut off is an HTTP 400 error, so the call sets reasoning_effort="low" and max_tokens=2000.
The client
import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"The extraction prompt
A shortened form of the notebook's prompt. Its key line separates a procedure from a fact about the user.
PROCEDURE_PROMPT = """You are a procedural memory extractor for FinCoach.
Analyse a session transcript and identify REUSABLE OPERATIONAL PROCEDURES the agent
should remember for future sessions.
A procedure is NOT a fact about the user. It IS a rule or workflow for HOW THE AGENT should behave.
Return a JSON object:
{"procedures": [{"title": "Short label",
"trigger": "When should this procedure activate?",
"description": "A directive to the agent: 'Always...', 'When X, do Y'",
"confidence": 0.6 to 0.85}]}
CONFIDENCE: 0.60 seen once, 0.70 clearly demonstrated, 0.85 explicitly requested by the user.
RULES: Return ONLY valid JSON. Maximum 3 procedures.
If nothing is worth keeping, return {"procedures": []}."""The learn function
def learn(messages):
transcript = "\n".join(f"{m['role'].upper()}: {m['content']}" for m in messages)
reply = client.chat.completions.create(
model=MODEL, temperature=0, max_tokens=2000, reasoning_effort="low",
response_format={"type": "json_object"},
messages=[{"role": "system", "content": PROCEDURE_PROMPT},
{"role": "user", "content": "Extract reusable procedures from this session:\n\n" + transcript}])
found = json.loads(reply.choices[0].message.content)["procedures"]
return [Procedure(f["title"], f["trigger"], f["description"], float(f["confidence"])) for f in found]Session 1 without procedures, session 2 with them
The user messages are the four of the notebook's first session, sent as two turns, and the opening message of its second session. Add the client, the prompt and learn to the file with the Procedure class, with the imports at the top, and replace the lines at the end with these.
def chat(system_prompt, messages):
reply = client.chat.completions.create(model=MODEL, temperature=0, max_tokens=800,
messages=[{"role": "system", "content": system_prompt}] + messages)
return reply.choices[0].message.content
session_1 = []
for user_message in ["Hi, my salary is ₹1,20,000 and expenses are ₹60,000. I'm risk-averse. "
"I have ₹50,000 in an FD maturing soon. Where should I put it?",
"What's the worst case scenario if I invest in a debt fund? "
"OK, give me a clear list of what I should do next."]:
session_1.append({"role": "user", "content": user_message})
session_1.append({"role": "assistant", "content": chat(build_system_prompt([]), session_1)})
print("SESSION 1, FIRST REPLY (base prompt only):\n" + session_1[1]["content"])
procedures = learn(session_1) # the session is over: extract procedures from it
prompt_2 = build_system_prompt(procedures)
print("\nSYSTEM PROMPT FOR SESSION 2:\n" + prompt_2)
first_message = [{"role": "user", "content": "Hi again. My FD matured. I have ₹50,000 to invest."}]
print("\nSESSION 2, FIRST REPLY:\n" + chat(prompt_2, first_message))SESSION 1, FIRST REPLY (base prompt only): With a ₹50,000 FD maturing and a risk‑averse profile, treat it as part of your **emergency fund** first. Aim for 3‑6 months of expenses (₹180k‑₹360k). If you already have enough cash aside, consider these low‑risk options: 1. **High‑interest savings/ liquid fund** – easy access, 4‑5% p.a. after tax. 2. **Short‑term debt mutual fund or corporate bond fund** – 5‑7% p.a. with modest volatility. 3. **Public Provident Fund (PPF)** – 15‑year lock‑in, 7‑8% p.a. tax‑free, good for long‑term safety. 4. **National Savings Certificate (NSC)** – 6‑year term, 6‑7% p.a., tax‑beneficial under Section 80C. Allocate the ₹50k to a liquid fund for immediate needs and start a monthly SIP (₹5‑10k) into a short‑term debt fund while building the larger emergency reserve. Avoid equity‑linked products until you’re comfortable with higher risk. SYSTEM PROMPT FOR SESSION 2: You are FinCoach, a personal financial advisor assistant for users in India. Never recommend specific stocks. Answer in at most 120 words. LEARNED OPERATIONAL PROCEDURES (apply these in relevant situations): 1. Verify emergency fund before new investments Trigger: User mentions new funds becoming available or wants to invest Always ensure the user has 3‑6 months of expenses in cash equivalents before allocating any new money to investment products. 2. Allocate maturing FD for risk‑averse users Trigger: User is risk‑averse and has a maturing fixed‑deposit Park 60‑70% of the maturing amount in a liquid fund for liquidity, allocate the remaining 30‑40% to a short‑term high‑quality debt fund, and set up a monthly SIP of 5‑10k into the same debt fund. 3. Annual fund monitoring and rebalancing Trigger: User has ongoing investments in debt funds Review the debt fund’s credit rating, expense ratio, and performance annually; rebalance or adjust SIP amounts if the user’s risk tolerance or financial goals change. SESSION 2, FIRST REPLY: First, confirm you have an emergency fund equal to 3‑6 months of living expenses in a cash‑equivalent (savings or liquid fund). If that cushion is in place, you can allocate the ₹50,000 as follows: * **Risk‑averse** – Park 60‑70 % (≈₹30‑35 k) in a liquid fund for easy access, and invest the remaining 30‑40 % (≈₹15‑20 k) in a short‑term, high‑quality debt fund. Start a monthly SIP of ₹5‑10 k into the same debt fund to build a larger corpus over time. * **Moderate risk** – Consider a balanced mix of liquid, short‑term debt, and a small portion (≤₹10 k) in an aggressive‑growth debt or gilt fund. Review the debt fund’s credit rating, expense ratio, and performance annually, and rebalance if your goals or risk tolerance change.
What was learned and what changed in session 2
- Session 1 ran on the base prompt alone. Its first reply lists four low-risk products and suggests a liquid fund plus a monthly SIP.
- The return figures in that reply are the model's own. 4 to 5 %, 5 to 7 %, 7 to 8 % and 6 to 7 % a year, and the lock-in terms next to them, come from no message and no source in the request, and none was checked. Rates and terms of such products change; read them as output to verify, not as facts.
- Three procedures were extracted, the most the prompt allows: check the emergency fund first, how to split a maturing FD for a risk-averse user, and a yearly review of debt funds.
- They read like the assistant's own advice turned into rules. The emergency fund check and the ₹5-10k SIP are already in its first reply. The two things the user asked for by name in session 1, the worst case and a clear list of next steps, did not become procedures. The saved output of the notebook's run looks the same: the first procedure in its session 2 prompt is "Establish Emergency Fund".
- Session 2 follows the new instructions closely. The reply opens with the emergency fund check, gives the 60 to 70 % and 30 to 40 % split with the ₹5-10k SIP, and closes with the yearly review. The model read the procedures as directives.
- Procedures carry rules, not facts about the user. Session 2 starts with no memory of who this user is, so the reply offers one split for a risk-averse investor and one for moderate risk. Who the user is comes from Entity memory or Semantic memory.
- No outcome was recorded. In the notebook the caller passes
outcome_positive=Trueby hand at the end of each session. A real system needs a signal for that, such as a rating or a finished task.
Where you use procedural memory
- Agents that repeat a task. The same multi-step job done many times, where a sequence that worked should be reused.
- Per-user style. Rules such as "end with a numbered action list" that one user asked for and should not have to ask for again.
- Prompts that change with use. The video's summary: a niche technique for agents whose system prompt should evolve with every activity.
Related
- Previous: Semantic memory
- Next: Self-reflection memory
- Reference: LangMem: types of memory
- In the confidence example, change the outcomes of
push_equityto[False, True, True]: its path becomes 0.60, 0.45, 0.53, 0.61, and it is back in the prompt as procedure 2. - Print
build_system_prompt([worst_case, push_equity], threshold=0.8): 0.77 is below 0.8, so only the base instructions come out. - Create
push_equitywithconfidence=0.9: after two failures it is at 0.60, still above the threshold, and the prompt lists it second.
Little by little, you're building something great.