AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Procedural memory

Procedural memory is a long-term memory technique that stores rules and workflows for how the agent itself should act, and adds them to the system prompt as instructions.

Last updated: 09 Oct, 2026 · OpenAI SDK 3.3

Semantic memory is knowledge about the user. Procedural memory is knowledge for the agent: not "this user is risk-averse" but "always quantify the worst case before the user asks". It is the one memory type that changes the agent's instructions instead of its context.

The words "learns" and "self-improving" are often used for this technique. They describe text. A list of procedures grows, and the ones that have worked are written into the system prompt of the next session. The model is never retrained and none of its weights change; the notebook says so itself: "outcome-based learning without retraining".

What procedural memory is · from the Complete AI Security Course in 8 Hours video · 5:20:11 to 5:23:03

This part of the video starts at 5:20:11. Procedural memory keeps updating the agent's system instructions. Where the clip says the agent learns and updates its behaviour, what changes is the text of the system prompt; the model is not retrained.

Instructions the agent writes for itself

System instructions are how you tell an agent to work, think and act. Normally a developer writes them once. With procedural memory the agent adds to them from experience. The notebook on screen, 10_procedural_memory.ipynb, defines it this way: procedural memory stores reusable step-by-step workflows in a skill library, so agents can retrieve proven procedures instead of reasoning from scratch.

The human parallel in the video is riding a bicycle. You do it without thinking, and you do not recall the session in which you learned to balance. You know how. In psychology this is the third kind of long-term memory next to episodic and semantic memory: Tulving's 1972 chapter set those two apart, and his 1985 paper "How many memory systems are there?" describes three systems: procedural memory (knowing how), semantic memory and episodic memory.

Semantic memory vs procedural memory

Semantic memoryProcedural memory
StoresFacts about the userRules for how the agent should act
Example"Chiru is risk-averse""Always quantify worst-case before user asks"
SubjectThe userThe agent's own behaviour
UpdatesWhen the user's behaviour changesWhen a workflow succeeds or fails
Enters the prompt asContext about the userInstructions in the system prompt

The notebook names three forms a procedure can take:

  • A workflow: steps for a recurring task. "When handling a FD maturity scenario: Step 1, check risk profile. Step 2, present 3 options. Step 3, quantify risk for each."
  • A rule: a condition and an action. "IF user shows anxiety about markets THEN lead with reassurance before giving the recommendation."
  • An interaction pattern: a learned way of communicating. "This user prefers numbered lists."

All three end up in the same place. The notebook's first point: procedural memory lives in the system prompt. Episodes are injected as context blocks and semantic facts as facts about the user, which the model reads as background. Procedures are injected as instructions, which the model reads as directives.

One model call with its three inputs. Episodic memory and semantic memory enter as context blocks that the model reads as background. Procedural memory is written into the system prompt, next to the base instructions, and is read as a directive. An arrow from the session outcome back to the procedure list shows confidence going up after a success and down after a failure.

Confidence that follows outcomes

The notebook's second point is that procedures are learned from outcomes, not only stated. Each procedure has a confidence. After a session in which it was applied, the confidence goes up by 0.08 if the session went well and down by 0.15 if it did not. Only procedures at or above 0.55 are written into the prompt, at most six of them. These four numbers are the notebook's.

A procedure record

python
from dataclasses import dataclass

@dataclass
class Procedure:
    title: str
    trigger: str                # when the procedure applies
    description: str            # the instruction, written to the agent
    confidence: float = 0.6
    successes: int = 0
    failures: int = 0

Recording an outcome

This method belongs to the Procedure class.

python
    def record_outcome(self, positive):
        if positive:
            self.successes += 1
            self.confidence = min(1.0, self.confidence + 0.08)
        else:
            self.failures += 1
            self.confidence = max(0.0, self.confidence - 0.15)

Building the system prompt

The base instructions come first. The procedures that pass the threshold are sorted by confidence and appended under a heading.

python
BASE_PROMPT = ("You are FinCoach, a personal financial advisor assistant for users in India. "
               "Never recommend specific stocks. Answer in at most 120 words.")

def build_system_prompt(procedures, threshold=0.55, limit=6):
    active = sorted((p for p in procedures if p.confidence >= threshold),
                    key=lambda p: p.confidence, reverse=True)[:limit]
    if not active:
        return BASE_PROMPT
    lines = [f"{i}. {p.title}\n   Trigger: {p.trigger}\n   {p.description}" for i, p in enumerate(active, 1)]
    return (BASE_PROMPT + "\n\nLEARNED OPERATIONAL PROCEDURES (apply these in relevant situations):\n"
            + "\n".join(lines))

One procedure that works and one that does not

The first procedure and the "push equity" rule are the notebook's own examples, and the sequence of two good sessions, one bad one and two more good ones is the one its confidence demo runs. No model is called. Put the three pieces above in one file and add these lines.

ExampleRun on Python 3.12
import matplotlib.pyplot as plt

worst_case = Procedure("Quantify the worst case", "When the user asks about an investment",
                       "Always quantify the worst-case scenario before the user asks.")
push_equity = Procedure("Push equity", "When the user has a long horizon",
                        "Recommend adding an equity component for higher returns.")

paths = {}
for proc, outcomes in [(worst_case, [True, True, False, True, True]), (push_equity, [False, False])]:
    paths[proc.title] = [proc.confidence]
    for positive in outcomes:
        proc.record_outcome(positive)
        paths[proc.title].append(proc.confidence)
    print(f"{proc.title}: " + " -> ".join(f"{c:.2f}" for c in paths[proc.title])
          + f"   (successes {proc.successes}, failures {proc.failures})")

print()
print(build_system_prompt([worst_case, push_equity]))

plt.figure(figsize=(7, 3.2))
for title, path in paths.items():
    plt.plot(range(len(path)), path, marker="o", label=title)
plt.axhline(0.55, color="gray", linestyle="--", label="threshold 0.55")
plt.xlabel("sessions in which the procedure was applied")
plt.ylabel("confidence")
plt.ylim(0, 1)
plt.title("Confidence follows outcomes: +0.08 on success, -0.15 on failure")
plt.legend(loc="lower right")
plt.show()
A line chart of confidence over sessions. Quantify the worst case starts at 0.60 and moves to 0.68, 0.76, 0.61, 0.69 and 0.77, always above the dashed threshold line at 0.55. Push equity starts at 0.60 and falls to 0.45 and 0.30, below the threshold.

What the two confidence paths did

  • "Quantify the worst case" moves 0.60, 0.68, 0.76, 0.61, 0.69, 0.77. Each good session adds 0.08. The one bad session takes away 0.15, nearly two successes' worth.
  • It never falls below 0.55, so it is written into the system prompt of every session.
  • "Push equity" moves 0.60, 0.45, 0.30. After its first failure it is already under the threshold.
  • The printed system prompt holds the base instructions and one procedure. The failed rule is still in the list, but the model never sees it.
  • This is all the "learning" there is: a number on a record, and a filter on which records are pasted into the prompt.

Learning procedures from a session

Where do procedures come from? In the notebook, a model reads the transcript of a finished session and writes down rules worth keeping. The notebook calls gpt-4o-mini with max_tokens=600; the code here calls openai/gpt-oss-120b on Groq. gpt-oss counts its reasoning tokens inside max_tokens, and a JSON reply that is cut off is an HTTP 400 error, so the call sets reasoning_effort="low" and max_tokens=2000.

The client

python
import json
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"

The extraction prompt

A shortened form of the notebook's prompt. Its key line separates a procedure from a fact about the user.

python
PROCEDURE_PROMPT = """You are a procedural memory extractor for FinCoach.
Analyse a session transcript and identify REUSABLE OPERATIONAL PROCEDURES the agent
should remember for future sessions.
A procedure is NOT a fact about the user. It IS a rule or workflow for HOW THE AGENT should behave.
Return a JSON object:
{"procedures": [{"title": "Short label",
                 "trigger": "When should this procedure activate?",
                 "description": "A directive to the agent: 'Always...', 'When X, do Y'",
                 "confidence": 0.6 to 0.85}]}
CONFIDENCE: 0.60 seen once, 0.70 clearly demonstrated, 0.85 explicitly requested by the user.
RULES: Return ONLY valid JSON. Maximum 3 procedures.
If nothing is worth keeping, return {"procedures": []}."""

The learn function

python
def learn(messages):
    transcript = "\n".join(f"{m['role'].upper()}: {m['content']}" for m in messages)
    reply = client.chat.completions.create(
        model=MODEL, temperature=0, max_tokens=2000, reasoning_effort="low",
        response_format={"type": "json_object"},
        messages=[{"role": "system", "content": PROCEDURE_PROMPT},
                  {"role": "user", "content": "Extract reusable procedures from this session:\n\n" + transcript}])
    found = json.loads(reply.choices[0].message.content)["procedures"]
    return [Procedure(f["title"], f["trigger"], f["description"], float(f["confidence"])) for f in found]

Session 1 without procedures, session 2 with them

The user messages are the four of the notebook's first session, sent as two turns, and the opening message of its second session. Add the client, the prompt and learn to the file with the Procedure class, with the imports at the top, and replace the lines at the end with these.

ExampleAPI keyFrom the video, run on Groq
def chat(system_prompt, messages):
    reply = client.chat.completions.create(model=MODEL, temperature=0, max_tokens=800,
                                           messages=[{"role": "system", "content": system_prompt}] + messages)
    return reply.choices[0].message.content

session_1 = []
for user_message in ["Hi, my salary is ₹1,20,000 and expenses are ₹60,000. I'm risk-averse. "
                     "I have ₹50,000 in an FD maturing soon. Where should I put it?",
                     "What's the worst case scenario if I invest in a debt fund? "
                     "OK, give me a clear list of what I should do next."]:
    session_1.append({"role": "user", "content": user_message})
    session_1.append({"role": "assistant", "content": chat(build_system_prompt([]), session_1)})
print("SESSION 1, FIRST REPLY (base prompt only):\n" + session_1[1]["content"])

procedures = learn(session_1)                      # the session is over: extract procedures from it
prompt_2 = build_system_prompt(procedures)
print("\nSYSTEM PROMPT FOR SESSION 2:\n" + prompt_2)

first_message = [{"role": "user", "content": "Hi again. My FD matured. I have ₹50,000 to invest."}]
print("\nSESSION 2, FIRST REPLY:\n" + chat(prompt_2, first_message))

What was learned and what changed in session 2

  • Session 1 ran on the base prompt alone. Its first reply lists four low-risk products and suggests a liquid fund plus a monthly SIP.
  • The return figures in that reply are the model's own. 4 to 5 %, 5 to 7 %, 7 to 8 % and 6 to 7 % a year, and the lock-in terms next to them, come from no message and no source in the request, and none was checked. Rates and terms of such products change; read them as output to verify, not as facts.
  • Three procedures were extracted, the most the prompt allows: check the emergency fund first, how to split a maturing FD for a risk-averse user, and a yearly review of debt funds.
  • They read like the assistant's own advice turned into rules. The emergency fund check and the ₹5-10k SIP are already in its first reply. The two things the user asked for by name in session 1, the worst case and a clear list of next steps, did not become procedures. The saved output of the notebook's run looks the same: the first procedure in its session 2 prompt is "Establish Emergency Fund".
  • Session 2 follows the new instructions closely. The reply opens with the emergency fund check, gives the 60 to 70 % and 30 to 40 % split with the ₹5-10k SIP, and closes with the yearly review. The model read the procedures as directives.
  • Procedures carry rules, not facts about the user. Session 2 starts with no memory of who this user is, so the reply offers one split for a risk-averse investor and one for moderate risk. Who the user is comes from Entity memory or Semantic memory.
  • No outcome was recorded. In the notebook the caller passes outcome_positive=True by hand at the end of each session. A real system needs a signal for that, such as a rating or a finished task.

Where you use procedural memory

  • Agents that repeat a task. The same multi-step job done many times, where a sequence that worked should be reused.
  • Per-user style. Rules such as "end with a numbered action list" that one user asked for and should not have to ask for again.
  • Prompts that change with use. The video's summary: a niche technique for agents whose system prompt should evolve with every activity.
Watch out. A bad procedure that keeps being reinforced causes the same error in every session, and two procedures that contradict each other leave the model to pick one. The model treats the system prompt as its instructions, so a procedure is also a way in: if a sentence from a user can become a rule, that user has edited the agent's instructions. Review what is learned before it is injected. Securing agent memory shows the attack.
Try it yourself
  • In the confidence example, change the outcomes of push_equity to [False, True, True]: its path becomes 0.60, 0.45, 0.53, 0.61, and it is back in the prompt as procedure 2.
  • Print build_system_prompt([worst_case, push_equity], threshold=0.8): 0.77 is below 0.8, so only the base instructions come out.
  • Create push_equity with confidence=0.9: after two failures it is at 0.60, still above the threshold, and the prompt lists it second.

Little by little, you're building something great.