Memory routing
Memory routing is the step in which an agent classifies each incoming message by intent and then reads from, and writes to, only the memory stores that intent needs.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
An agent can hold a recent buffer and six long-term stores at once: Vector store memory, Entity memory, Episodic memory, Semantic memory, Procedural memory and Self-reflection memory. Pasting all of them into every prompt costs tokens and buries the one memory the message needs. A router picks per message.
This part of the video starts at 5:33:14. It introduces routing with the examples of the notebook 12_memory_routing.ipynb.
Sending each message to the right store
The notebook's picture is an air traffic controller. When a plane arrives, the controller does not send it to all runways at once. They look at the plane's type, destination and size, and route it to one runway and one gate. In the same way each incoming message has a type and an intent, and the router dispatches it to the stores that fit, not to all of them. Working out the intent is called intent classification.
The video walks through the notebook's examples:
- "What is my current salary?" goes to the entity store: a structured, current fact.
- "What did we decide last April?" goes to the episodic store: a question about a past session.
- "How does a SIP work?" goes to the vector store: general knowledge discussed before.
- A message about a job change goes to the entity store as an update.
- "Never recommend equity to me again" goes to the procedural store: a hard constraint to write.
- "I'm worried about market volatility" goes to the semantic store, which holds the user's behaviour patterns.
Intents, stores and the routing table
The notebook defines eight intents and seven stores. Three stores are read on every turn, whatever the intent: the recent buffer (the last few turns), the procedural store (rules) and the reflection store (notes). The routing table then adds stores per intent:
| Intent | Example | Also reads | Writes to |
|---|---|---|---|
fact_query | "What is my salary?" | entity profile | nothing |
history_query | "What did we decide last April?" | episodic store, vector store | nothing |
knowledge_query | "How does a SIP work?" | vector store | vector store |
advice_request | "What should I do with my FD?" | entity profile, semantic store, episodic store, vector store | vector store |
fact_update | "My new salary is ..." | entity profile | entity profile, vector store |
constraint_update | "Never recommend equity to me" | nothing more | procedural store, entity profile |
emotional_signal | "I'm worried about markets" | semantic store, entity profile | semantic store |
general_chat | "Thanks!" | nothing more | nothing |
Reads and writes are routed separately. A fact question only reads. A fact update reads the current value and then writes the new one.
Estimating what routing saves
The notebook gives each store a token estimate set by hand (300 for the recent buffer, 200 for three retrieved messages, and so on) and compares the routed stores with reading all seven.
ALWAYS_READ = {"recent_buffer", "procedural_store", "reflection_store"}
READS = { # the extra stores each intent reads, from the notebook's ROUTING_TABLE
"fact_query": {"entity_profile"},
"history_query": {"episodic_store", "vector_store"},
"knowledge_query": {"vector_store"},
"advice_request": {"entity_profile", "semantic_store", "episodic_store", "vector_store"},
"fact_update": {"entity_profile"},
"constraint_update": {"procedural_store"},
"emotional_signal": {"semantic_store", "entity_profile"},
"general_chat": set(),
}
TOKENS = {"recent_buffer": 300, "vector_store": 200, "entity_profile": 150, "episodic_store": 250,
"semantic_store": 150, "procedural_store": 200, "reflection_store": 150}
everything = sum(TOKENS.values())
print(f"all 7 stores: {everything} tokens, the 3 always-read stores: {sum(TOKENS[s] for s in ALWAYS_READ)} tokens")
print()
for intent, extra in READS.items():
routed = sum(TOKENS[s] for s in ALWAYS_READ | extra)
print(f"{intent:<18} reads {len(ALWAYS_READ | extra)} stores, {routed:>4} tokens, "
f"saves {everything - routed:>3} ({(everything - routed) / everything:.0%})")all 7 stores: 1400 tokens, the 3 always-read stores: 650 tokens fact_query reads 4 stores, 800 tokens, saves 600 (43%) history_query reads 5 stores, 1100 tokens, saves 300 (21%) knowledge_query reads 4 stores, 850 tokens, saves 550 (39%) advice_request reads 7 stores, 1400 tokens, saves 0 (0%) fact_update reads 4 stores, 800 tokens, saves 600 (43%) constraint_update reads 3 stores, 650 tokens, saves 750 (54%) emotional_signal reads 5 stores, 950 tokens, saves 450 (32%) general_chat reads 3 stores, 650 tokens, saves 750 (54%)
What the estimates say
- Reading everything costs 1400 estimated tokens a turn, and the three always-read stores alone cost 650.
- A fact question reads 4 stores, 800 tokens, a saving of 600, or 43%.
- An advice request reads all 7 stores and saves 0. Routing does nothing for the messages that need the most context.
- The best case is 54%, for a constraint update or small talk, which read only the three always-read stores.
- These are estimates, not measurements. On the video's screen the context assembled for the first demo question is about 220 tokens, counted with
tiktoken'so200k_baseencoding, where the table above says 800. Measure with the provider'susage.prompt_tokensbefore quoting a saving.
This part of the video starts at 5:37:07 and reads two entries of the routing demo's output. The reply on screen to the first question is Rs 1,50,000.
Reading the routing demo
The notebook's demo sends 8 queries and prints a routing decision before each reply. For "What is my current monthly salary?" the decision is rule_based with confidence 90%, intent fact_query, read from entity_profile, procedural_store, recent_buffer and reflection_store, write to none. The reply is "Your current monthly salary is Rs 1,50,000."
For "Never suggest cryptocurrency to me under any circumstances." the decision is rule_based at 80%, intent constraint_update, write to entity_profile and procedural_store, and the agent answers "Understood. I will not suggest cryptocurrency investments to you." Look at that route again: a sentence with "never" or "always" in it is written into the store that feeds every later system prompt. Securing agent memory shows what an attacker does with that.
Routing by rules
The router tries rules first: a list of regular expressions, each tied to an intent and a confidence. Rules cost nothing and return the same route every time.
Patterns and the best match
Each rule is a pattern, an intent and a confidence. route_by_rules collects every rule that matches and returns the one with the highest confidence; when two tie, the one listed first wins.
RULES = [ # (pattern, intent, confidence)
(r"\b(what is|what's|tell me|show me)\b.*(my|current).*(salary|income|risk|goal)", "fact_query", 0.90),
(r"\b(last time|last session|last month|did we|we decided|april|march)", "history_query", 0.85),
(r"\b(never|always|don't|do not|i prefer|make sure|remember that)", "constraint_update", 0.80),
]
def route_by_rules(message):
"""The best-scoring rule that matches, as (intent, confidence), or None."""
hits = [(intent, conf) for pattern, intent, conf in RULES if re.search(pattern, message.lower().strip())]
return max(hits, key=lambda hit: hit[1]) if hits else NoneThe example runs all eight rules on the six messages from the video, on two more from the notebook and on the first message of its fan-out cell, a long one. The patterns are the notebook's, with shorter word lists.
import re
RULES = [ # (pattern, intent, confidence): the notebook's rules, with shorter word lists
(r"\b(what is|what's|tell me|show me)\b.*(my|current).*(salary|income|risk|goal)", "fact_query", 0.90),
(r"\b(last time|last session|last month|did we|we decided|april|march)", "history_query", 0.85),
(r"\b(never|always|don't|do not|i prefer|make sure|remember that)", "constraint_update", 0.80),
(r"\b(my (new |current )?(salary|income|job|employer) is|i (now|recently|just) (work|earn|got|moved|changed)|i changed)",
"fact_update", 0.85),
(r"\b(what is|how does|explain|difference between)\b.*(sip|mutual fund|fd|nps|ppf|equity|debt|tax)",
"knowledge_query", 0.85),
(r"\b(worried|anxious|scared|nervous|afraid|not sure|risky|safe)", "emotional_signal", 0.80),
(r"\b(what should|should i|recommend|advice|suggest|where to invest)", "advice_request", 0.80),
(r"^(thanks|thank you|great|ok|okay|got it|bye|hello)\.?$", "general_chat", 0.95),
]
def route_by_rules(message):
"""The best-scoring rule that matches, as (intent, confidence), or None."""
hits = [(intent, conf) for pattern, intent, conf in RULES if re.search(pattern, message.lower().strip())]
return max(hits, key=lambda hit: hit[1]) if hits else None
for message in ["What is my current salary?", "What did we decide last April?", "How does a SIP work?",
"I just changed jobs to TCS", "Never recommend equity to me again",
"I'm worried about market volatility",
"How does a Systematic Investment Plan work?", "Thanks!",
"I changed jobs to TCS and my new salary is Rs 1,50,000. "
"Given my conservative profile, what should I invest in now?"]:
print(f"{str(route_by_rules(message)):<28} {message}")('fact_query', 0.9) What is my current salary?
('history_query', 0.85) What did we decide last April?
('knowledge_query', 0.85) How does a SIP work?
('fact_update', 0.85) I just changed jobs to TCS
('constraint_update', 0.8) Never recommend equity to me again
('emotional_signal', 0.8) I'm worried about market volatility
None How does a Systematic Investment Plan work?
None Thanks!
('fact_update', 0.85) I changed jobs to TCS and my new salary is Rs 1,50,000. Given my conservative profile, what should I invest in now?What the rules matched and missed
- The six messages from the video get the intents the notebook names:
fact_query,history_query,knowledge_query,fact_update,constraint_updateandemotional_signal. - "How does a Systematic Investment Plan work?" matches nothing. The knowledge rule lists
sip, not the long form. - "Thanks!" matches nothing. The last pattern ends in
\.?$, which allows a full stop after the word and no exclamation mark. - The long message gets one intent,
fact_update. It also asks for advice, and the advice rule does match, at 0.80. The fact rule's 0.85 wins and the second intent is dropped: the rule path returns one intent per message.
Routing with an LLM
When no rule matches, or the best rule is under the confidence threshold of 0.75, the notebook asks a model to classify the message and reply in JSON. The prompt in the example is the notebook's own. The notebook sends it to gpt-4o-mini with max_tokens=150; here it goes to openai/gpt-oss-120b on Groq. openai/gpt-oss-120b is a reasoning model, and its reasoning tokens count inside max_tokens. In JSON mode a limit that is too small ends in a 400 json_validate_failed error instead of a short reply, so the call asks for reasoning_effort="low" and leaves room.
def route_by_llm(message):
"""Ask the model for the intents of one message; returns the parsed JSON."""
reply = client.chat.completions.create(
model=MODEL, temperature=0, max_tokens=1000, reasoning_effort="low",
response_format={"type": "json_object"},
messages=[{"role": "system", "content": LLM_ROUTING_PROMPT},
{"role": "user", "content": f"Classify: {message}"}],
)
return json.loads(reply.choices[0].message.content)import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"
LLM_ROUTING_PROMPT = """You are a memory routing classifier for FinCoach, a financial advisor agent.
Classify the user message into one or more of these intents:
- fact_query : asking for a specific personal fact (salary, risk profile, etc.)
- history_query : asking about a past session, decision, or event
- knowledge_query : asking about a general financial concept or product
- advice_request : asking for a recommendation or investment plan
- fact_update : providing new personal information
- constraint_update: stating a hard rule, preference, or constraint
- emotional_signal : expressing anxiety, worry, or uncertainty
- general_chat : small talk, thanks, or non-financial content
Return JSON: {"intents": ["intent1", "intent2"], "confidence": 0.0-1.0, "reasoning": "brief"}
Multiple intents allowed for complex messages. Return ONLY valid JSON."""
def route_by_llm(message):
"""Ask the model for the intents of one message; returns the parsed JSON."""
reply = client.chat.completions.create(
model=MODEL, temperature=0, max_tokens=1000, reasoning_effort="low",
response_format={"type": "json_object"},
messages=[{"role": "system", "content": LLM_ROUTING_PROMPT},
{"role": "user", "content": f"Classify: {message}"}],
)
return json.loads(reply.choices[0].message.content)
for message in ["Thanks!",
"I changed jobs to TCS and my new salary is Rs 1,50,000. "
"Given my conservative profile, what should I invest in now?"]:
print(message)
print(" ", route_by_llm(message))Thanks!
{'intents': ['general_chat'], 'confidence': 0.99, 'reasoning': 'User expresses gratitude, which is small talk/non‑financial content.'}
I changed jobs to TCS and my new salary is Rs 1,50,000. Given my conservative profile, what should I invest in now?
{'intents': ['fact_update', 'advice_request'], 'confidence': 0.99, 'reasoning': 'User provides new personal salary info (fact_update) and asks for investment recommendation (advice_request).'}What the LLM router returned
- "Thanks!" is
general_chat: the message no rule could match gets its intent. - The long message gets two intents,
fact_updateandadvice_request. The router then reads the stores of both rows of the routing table together, which here is all seven, and writes to the entity profile and the vector store. This is fan-out, and only the LLM path does it. - The confidence of 0.99 is the model's own statement, printed for both messages. It is not a measured probability.
Measuring routing accuracy
A router fails silently. A wrong route still produces a fluent reply, built on the wrong memories. The way to see it is a labelled set: messages paired with the intent a person would give them, run through the router and counted.
Rules first, then the model
hit = route_by_rules(message)
if hit and hit[1] >= 0.75: # a confident rule wins, no model call
intents = [hit[0]]
else: # no rule, or a weak one: ask the model
intents = route_by_llm(message)["intents"]The example below continues the two above. Put these above its lines, in this order: from the rules example, import re, RULES and route_by_rules (everything before its for loop); from the LLM example, the three imports, client, MODEL, LLM_ROUTING_PROMPT and route_by_llm (everything before its for loop). A route counts as right when the labelled intent is among the intents returned.
labelled = [ # (message, the intent a person would give it)
("What is my current monthly salary?", "fact_query"),
("What did we decide last April about my investments?", "history_query"),
("How does a Systematic Investment Plan work?", "knowledge_query"),
("I changed jobs to TCS and my new salary is Rs 1,50,000.", "fact_update"),
("Never suggest cryptocurrency to me under any circumstances.", "constraint_update"),
("I'm worried about the current interest rate environment.", "emotional_signal"),
("Given my profile, what should I do with Rs 1 lakh I have available?", "advice_request"),
("Thanks!", "general_chat"),
("My bonus comes in March, should I invest it?", "advice_request"),
("I always get anxious when markets fall.", "emotional_signal"),
]
right = {"rules": 0, "llm": 0}
used = {"rules": 0, "llm": 0}
for message, label in labelled:
hit = route_by_rules(message)
if hit and hit[1] >= 0.75: # a confident rule wins, no model call
method, intents = "rules", [hit[0]]
else:
method, intents = "llm", route_by_llm(message)["intents"]
ok = label in intents
used[method] += 1
right[method] += ok
print(f"{'ok ' if ok else 'MISS'} {method:<5} {str(intents):<24} {message}")
total = sum(right.values())
print(f"\nrules: {right['rules']} of {used['rules']} right, llm: {right['llm']} of {used['llm']} right")
print(f"routing accuracy: {total} of {len(labelled)} = {total / len(labelled):.0%}")ok rules ['fact_query'] What is my current monthly salary? ok rules ['history_query'] What did we decide last April about my investments? ok llm ['knowledge_query'] How does a Systematic Investment Plan work? ok rules ['fact_update'] I changed jobs to TCS and my new salary is Rs 1,50,000. ok rules ['constraint_update'] Never suggest cryptocurrency to me under any circumstances. ok rules ['emotional_signal'] I'm worried about the current interest rate environment. ok rules ['advice_request'] Given my profile, what should I do with Rs 1 lakh I have available? ok llm ['general_chat'] Thanks! MISS rules ['history_query'] My bonus comes in March, should I invest it? MISS rules ['constraint_update'] I always get anxious when markets fall. rules: 6 of 8 right, llm: 2 of 2 right routing accuracy: 8 of 10 = 80%
What the ten routes show
- 8 of the 10 routes are right: 80%.
- Rules decided 8 messages and got 6 right. The model was asked for 2 and got both right. Ten messages cost two model calls.
- "My bonus comes in March, should I invest it?" went to
history_query. The word "March" is in the history pattern, at 0.85, ahead of the advice rule at 0.80. The message is about the future. - "I always get anxious when markets fall." went to
constraint_update. The constraint rule ("always") and the emotional rule ("anxious") both match at 0.80, and the constraint rule is listed first. With the routing table above, a sentence about a feeling would be written into the procedural store as a rule. - The model never saw the two misses. A rule that matches at 0.75 or more ends the routing, so rules-first saves calls and hides its own mistakes. Only a labelled set shows them.
- Ten messages is a small sample. One message moves the result by 10 points. A real set needs every intent several times, in the words your users write.
Rule-based vs LLM routing
| Rules | LLM classifier | |
|---|---|---|
| Cost per message | None | One model call |
| Same message, same route | Always | Usually, at temperature 0 |
| Several intents in one message | No: the single best match | Yes: a list of intents |
| Typical failure | A keyword out of context; wording that no pattern lists | A wrong label stated with high confidence; a call that fails or times out |
| How you improve it | Edit a pattern and re-run the labelled set | Edit the prompt and re-run the labelled set |
Where you use memory routing
- Agents with three or more stores. Past that point, reading everything on every turn crowds the prompt with memories the message does not need.
- Deciding what may be written. The write side of the table is the gate to long-term memory: a greeting writes nothing, a fact update writes to the entity profile and the vector store.
- Token and latency budgets. Small talk reads only the always-read stores; only advice requests pay for the full context.
Related
- Previous: Self-reflection memory
- Next: Forgetting and decay in agent memory
- See also: Securing agent memory
- Reference: re, regular expression operations
- In the rules example, change the end of the last pattern from
\.?$to[.!]?$. "Thanks!" now prints('general_chat', 0.95). - Add
systematic investment plan|in front ofmutual fundin the knowledge pattern. The long-form question now prints('knowledge_query', 0.85). - Remove
|marchfrom the history pattern and callroute_by_rules("My bonus comes in March, should I invest it?"). It returns('advice_request', 0.8).
Slow is fine. Stopping is the only problem.