AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Memory routing

Memory routing is the step in which an agent classifies each incoming message by intent and then reads from, and writes to, only the memory stores that intent needs.

Last updated: 09 Oct, 2026 · OpenAI SDK 3.3

An agent can hold a recent buffer and six long-term stores at once: Vector store memory, Entity memory, Episodic memory, Semantic memory, Procedural memory and Self-reflection memory. Pasting all of them into every prompt costs tokens and buries the one memory the message needs. A router picks per message.

What memory routing is · from the Complete AI Security Course in 8 Hours video · 5:33:14 to 5:34:50

This part of the video starts at 5:33:14. It introduces routing with the examples of the notebook 12_memory_routing.ipynb.

Sending each message to the right store

The notebook's picture is an air traffic controller. When a plane arrives, the controller does not send it to all runways at once. They look at the plane's type, destination and size, and route it to one runway and one gate. In the same way each incoming message has a type and an intent, and the router dispatches it to the stores that fit, not to all of them. Working out the intent is called intent classification.

The video walks through the notebook's examples:

  • "What is my current salary?" goes to the entity store: a structured, current fact.
  • "What did we decide last April?" goes to the episodic store: a question about a past session.
  • "How does a SIP work?" goes to the vector store: general knowledge discussed before.
  • A message about a job change goes to the entity store as an update.
  • "Never recommend equity to me again" goes to the procedural store: a hard constraint to write.
  • "I'm worried about market volatility" goes to the semantic store, which holds the user's behaviour patterns.
A user message goes through regex rules, where a rule with confidence 0.75 or more gives the intent and otherwise an LLM classifier returns one or more intents; a routing table maps the intents to the stores to read and write, three of the seven stores are read on every turn, and a fact update also reads the entity profile and writes the entity profile and the vector store.

Intents, stores and the routing table

The notebook defines eight intents and seven stores. Three stores are read on every turn, whatever the intent: the recent buffer (the last few turns), the procedural store (rules) and the reflection store (notes). The routing table then adds stores per intent:

IntentExampleAlso readsWrites to
fact_query"What is my salary?"entity profilenothing
history_query"What did we decide last April?"episodic store, vector storenothing
knowledge_query"How does a SIP work?"vector storevector store
advice_request"What should I do with my FD?"entity profile, semantic store, episodic store, vector storevector store
fact_update"My new salary is ..."entity profileentity profile, vector store
constraint_update"Never recommend equity to me"nothing moreprocedural store, entity profile
emotional_signal"I'm worried about markets"semantic store, entity profilesemantic store
general_chat"Thanks!"nothing morenothing

Reads and writes are routed separately. A fact question only reads. A fact update reads the current value and then writes the new one.

Estimating what routing saves

The notebook gives each store a token estimate set by hand (300 for the recent buffer, 200 for three retrieved messages, and so on) and compares the routed stores with reading all seven.

ExampleThe notebook's store estimates, added up in Python 3.12
ALWAYS_READ = {"recent_buffer", "procedural_store", "reflection_store"}
READS = {   # the extra stores each intent reads, from the notebook's ROUTING_TABLE
    "fact_query":        {"entity_profile"},
    "history_query":     {"episodic_store", "vector_store"},
    "knowledge_query":   {"vector_store"},
    "advice_request":    {"entity_profile", "semantic_store", "episodic_store", "vector_store"},
    "fact_update":       {"entity_profile"},
    "constraint_update": {"procedural_store"},
    "emotional_signal":  {"semantic_store", "entity_profile"},
    "general_chat":      set(),
}
TOKENS = {"recent_buffer": 300, "vector_store": 200, "entity_profile": 150, "episodic_store": 250,
          "semantic_store": 150, "procedural_store": 200, "reflection_store": 150}

everything = sum(TOKENS.values())
print(f"all 7 stores: {everything} tokens, the 3 always-read stores: {sum(TOKENS[s] for s in ALWAYS_READ)} tokens")
print()
for intent, extra in READS.items():
    routed = sum(TOKENS[s] for s in ALWAYS_READ | extra)
    print(f"{intent:<18} reads {len(ALWAYS_READ | extra)} stores, {routed:>4} tokens, "
          f"saves {everything - routed:>3} ({(everything - routed) / everything:.0%})")

What the estimates say

  • Reading everything costs 1400 estimated tokens a turn, and the three always-read stores alone cost 650.
  • A fact question reads 4 stores, 800 tokens, a saving of 600, or 43%.
  • An advice request reads all 7 stores and saves 0. Routing does nothing for the messages that need the most context.
  • The best case is 54%, for a constraint update or small talk, which read only the three always-read stores.
  • These are estimates, not measurements. On the video's screen the context assembled for the first demo question is about 220 tokens, counted with tiktoken's o200k_base encoding, where the table above says 800. Measure with the provider's usage.prompt_tokens before quoting a saving.
The routing demo · from the Complete AI Security Course in 8 Hours video · 5:37:07 to 5:38:18

This part of the video starts at 5:37:07 and reads two entries of the routing demo's output. The reply on screen to the first question is Rs 1,50,000.

Reading the routing demo

The notebook's demo sends 8 queries and prints a routing decision before each reply. For "What is my current monthly salary?" the decision is rule_based with confidence 90%, intent fact_query, read from entity_profile, procedural_store, recent_buffer and reflection_store, write to none. The reply is "Your current monthly salary is Rs 1,50,000."

For "Never suggest cryptocurrency to me under any circumstances." the decision is rule_based at 80%, intent constraint_update, write to entity_profile and procedural_store, and the agent answers "Understood. I will not suggest cryptocurrency investments to you." Look at that route again: a sentence with "never" or "always" in it is written into the store that feeds every later system prompt. Securing agent memory shows what an attacker does with that.

Routing by rules

The router tries rules first: a list of regular expressions, each tied to an intent and a confidence. Rules cost nothing and return the same route every time.

Patterns and the best match

Each rule is a pattern, an intent and a confidence. route_by_rules collects every rule that matches and returns the one with the highest confidence; when two tie, the one listed first wins.

python
RULES = [   # (pattern, intent, confidence)
    (r"\b(what is|what's|tell me|show me)\b.*(my|current).*(salary|income|risk|goal)", "fact_query", 0.90),
    (r"\b(last time|last session|last month|did we|we decided|april|march)", "history_query", 0.85),
    (r"\b(never|always|don't|do not|i prefer|make sure|remember that)", "constraint_update", 0.80),
]

def route_by_rules(message):
    """The best-scoring rule that matches, as (intent, confidence), or None."""
    hits = [(intent, conf) for pattern, intent, conf in RULES if re.search(pattern, message.lower().strip())]
    return max(hits, key=lambda hit: hit[1]) if hits else None

The example runs all eight rules on the six messages from the video, on two more from the notebook and on the first message of its fan-out cell, a long one. The patterns are the notebook's, with shorter word lists.

ExampleFrom the video, run on Python 3.12
import re

RULES = [   # (pattern, intent, confidence): the notebook's rules, with shorter word lists
    (r"\b(what is|what's|tell me|show me)\b.*(my|current).*(salary|income|risk|goal)", "fact_query", 0.90),
    (r"\b(last time|last session|last month|did we|we decided|april|march)", "history_query", 0.85),
    (r"\b(never|always|don't|do not|i prefer|make sure|remember that)", "constraint_update", 0.80),
    (r"\b(my (new |current )?(salary|income|job|employer) is|i (now|recently|just) (work|earn|got|moved|changed)|i changed)",
     "fact_update", 0.85),
    (r"\b(what is|how does|explain|difference between)\b.*(sip|mutual fund|fd|nps|ppf|equity|debt|tax)",
     "knowledge_query", 0.85),
    (r"\b(worried|anxious|scared|nervous|afraid|not sure|risky|safe)", "emotional_signal", 0.80),
    (r"\b(what should|should i|recommend|advice|suggest|where to invest)", "advice_request", 0.80),
    (r"^(thanks|thank you|great|ok|okay|got it|bye|hello)\.?$", "general_chat", 0.95),
]

def route_by_rules(message):
    """The best-scoring rule that matches, as (intent, confidence), or None."""
    hits = [(intent, conf) for pattern, intent, conf in RULES if re.search(pattern, message.lower().strip())]
    return max(hits, key=lambda hit: hit[1]) if hits else None

for message in ["What is my current salary?", "What did we decide last April?", "How does a SIP work?",
                "I just changed jobs to TCS", "Never recommend equity to me again",
                "I'm worried about market volatility",
                "How does a Systematic Investment Plan work?", "Thanks!",
                "I changed jobs to TCS and my new salary is Rs 1,50,000. "
                "Given my conservative profile, what should I invest in now?"]:
    print(f"{str(route_by_rules(message)):<28} {message}")

What the rules matched and missed

  • The six messages from the video get the intents the notebook names: fact_query, history_query, knowledge_query, fact_update, constraint_update and emotional_signal.
  • "How does a Systematic Investment Plan work?" matches nothing. The knowledge rule lists sip, not the long form.
  • "Thanks!" matches nothing. The last pattern ends in \.?$, which allows a full stop after the word and no exclamation mark.
  • The long message gets one intent, fact_update. It also asks for advice, and the advice rule does match, at 0.80. The fact rule's 0.85 wins and the second intent is dropped: the rule path returns one intent per message.

Routing with an LLM

When no rule matches, or the best rule is under the confidence threshold of 0.75, the notebook asks a model to classify the message and reply in JSON. The prompt in the example is the notebook's own. The notebook sends it to gpt-4o-mini with max_tokens=150; here it goes to openai/gpt-oss-120b on Groq. openai/gpt-oss-120b is a reasoning model, and its reasoning tokens count inside max_tokens. In JSON mode a limit that is too small ends in a 400 json_validate_failed error instead of a short reply, so the call asks for reasoning_effort="low" and leaves room.

python
def route_by_llm(message):
    """Ask the model for the intents of one message; returns the parsed JSON."""
    reply = client.chat.completions.create(
        model=MODEL, temperature=0, max_tokens=1000, reasoning_effort="low",
        response_format={"type": "json_object"},
        messages=[{"role": "system", "content": LLM_ROUTING_PROMPT},
                  {"role": "user", "content": f"Classify: {message}"}],
    )
    return json.loads(reply.choices[0].message.content)
ExampleAPI keyFrom the video, run on Groq
import json
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"

LLM_ROUTING_PROMPT = """You are a memory routing classifier for FinCoach, a financial advisor agent.
Classify the user message into one or more of these intents:

- fact_query       : asking for a specific personal fact (salary, risk profile, etc.)
- history_query    : asking about a past session, decision, or event
- knowledge_query  : asking about a general financial concept or product
- advice_request   : asking for a recommendation or investment plan
- fact_update      : providing new personal information
- constraint_update: stating a hard rule, preference, or constraint
- emotional_signal : expressing anxiety, worry, or uncertainty
- general_chat     : small talk, thanks, or non-financial content

Return JSON: {"intents": ["intent1", "intent2"], "confidence": 0.0-1.0, "reasoning": "brief"}
Multiple intents allowed for complex messages. Return ONLY valid JSON."""

def route_by_llm(message):
    """Ask the model for the intents of one message; returns the parsed JSON."""
    reply = client.chat.completions.create(
        model=MODEL, temperature=0, max_tokens=1000, reasoning_effort="low",
        response_format={"type": "json_object"},
        messages=[{"role": "system", "content": LLM_ROUTING_PROMPT},
                  {"role": "user", "content": f"Classify: {message}"}],
    )
    return json.loads(reply.choices[0].message.content)

for message in ["Thanks!",
                "I changed jobs to TCS and my new salary is Rs 1,50,000. "
                "Given my conservative profile, what should I invest in now?"]:
    print(message)
    print("  ", route_by_llm(message))

What the LLM router returned

  • "Thanks!" is general_chat: the message no rule could match gets its intent.
  • The long message gets two intents, fact_update and advice_request. The router then reads the stores of both rows of the routing table together, which here is all seven, and writes to the entity profile and the vector store. This is fan-out, and only the LLM path does it.
  • The confidence of 0.99 is the model's own statement, printed for both messages. It is not a measured probability.

Measuring routing accuracy

A router fails silently. A wrong route still produces a fluent reply, built on the wrong memories. The way to see it is a labelled set: messages paired with the intent a person would give them, run through the router and counted.

Rules first, then the model

python
hit = route_by_rules(message)
if hit and hit[1] >= 0.75:                          # a confident rule wins, no model call
    intents = [hit[0]]
else:                                               # no rule, or a weak one: ask the model
    intents = route_by_llm(message)["intents"]

The example below continues the two above. Put these above its lines, in this order: from the rules example, import re, RULES and route_by_rules (everything before its for loop); from the LLM example, the three imports, client, MODEL, LLM_ROUTING_PROMPT and route_by_llm (everything before its for loop). A route counts as right when the labelled intent is among the intents returned.

ExampleAPI keyRun on Groq
labelled = [   # (message, the intent a person would give it)
    ("What is my current monthly salary?", "fact_query"),
    ("What did we decide last April about my investments?", "history_query"),
    ("How does a Systematic Investment Plan work?", "knowledge_query"),
    ("I changed jobs to TCS and my new salary is Rs 1,50,000.", "fact_update"),
    ("Never suggest cryptocurrency to me under any circumstances.", "constraint_update"),
    ("I'm worried about the current interest rate environment.", "emotional_signal"),
    ("Given my profile, what should I do with Rs 1 lakh I have available?", "advice_request"),
    ("Thanks!", "general_chat"),
    ("My bonus comes in March, should I invest it?", "advice_request"),
    ("I always get anxious when markets fall.", "emotional_signal"),
]

right = {"rules": 0, "llm": 0}
used = {"rules": 0, "llm": 0}
for message, label in labelled:
    hit = route_by_rules(message)
    if hit and hit[1] >= 0.75:                      # a confident rule wins, no model call
        method, intents = "rules", [hit[0]]
    else:
        method, intents = "llm", route_by_llm(message)["intents"]
    ok = label in intents
    used[method] += 1
    right[method] += ok
    print(f"{'ok  ' if ok else 'MISS'} {method:<5} {str(intents):<24} {message}")

total = sum(right.values())
print(f"\nrules: {right['rules']} of {used['rules']} right, llm: {right['llm']} of {used['llm']} right")
print(f"routing accuracy: {total} of {len(labelled)} = {total / len(labelled):.0%}")

What the ten routes show

  • 8 of the 10 routes are right: 80%.
  • Rules decided 8 messages and got 6 right. The model was asked for 2 and got both right. Ten messages cost two model calls.
  • "My bonus comes in March, should I invest it?" went to history_query. The word "March" is in the history pattern, at 0.85, ahead of the advice rule at 0.80. The message is about the future.
  • "I always get anxious when markets fall." went to constraint_update. The constraint rule ("always") and the emotional rule ("anxious") both match at 0.80, and the constraint rule is listed first. With the routing table above, a sentence about a feeling would be written into the procedural store as a rule.
  • The model never saw the two misses. A rule that matches at 0.75 or more ends the routing, so rules-first saves calls and hides its own mistakes. Only a labelled set shows them.
  • Ten messages is a small sample. One message moves the result by 10 points. A real set needs every intent several times, in the words your users write.

Rule-based vs LLM routing

RulesLLM classifier
Cost per messageNoneOne model call
Same message, same routeAlwaysUsually, at temperature 0
Several intents in one messageNo: the single best matchYes: a list of intents
Typical failureA keyword out of context; wording that no pattern listsA wrong label stated with high confidence; a call that fails or times out
How you improve itEdit a pattern and re-run the labelled setEdit the prompt and re-run the labelled set

Where you use memory routing

  • Agents with three or more stores. Past that point, reading everything on every turn crowds the prompt with memories the message does not need.
  • Deciding what may be written. The write side of the table is the gate to long-term memory: a greeting writes nothing, a fact update writes to the entity profile and the vector store.
  • Token and latency budgets. Small talk reads only the always-read stores; only advice requests pay for the full context.
Watch out. A misroute does not raise an error. The agent answers anyway, from the wrong memories or with a wrong write. Keep a labelled set of real messages and re-run it after every change to a pattern or a prompt. When a message is ambiguous, read from more stores, not fewer, and be stricter about writes than about reads.
Try it yourself
  • In the rules example, change the end of the last pattern from \.?$ to [.!]?$. "Thanks!" now prints ('general_chat', 0.95).
  • Add systematic investment plan| in front of mutual fund in the knowledge pattern. The long-form question now prints ('knowledge_query', 0.85).
  • Remove |march from the history pattern and call route_by_rules("My bonus comes in March, should I invest it?"). It returns ('advice_request', 0.8).

Slow is fine. Stopping is the only problem.