The extraction request
The extraction request is the message list and tool list that a LangMem memory manager sends to the chat model: a short system message, your instructions around the conversation, and one tool per memory schema.
Last updated: 30 Sep, 2026 · LangMem 0.0.30
Reading the request explains everything the manager does, including why the default call failed on gpt-oss in Extracting memories. A LangChain callback handler can print it, because LangChain calls the handler with the exact messages and tools right before each model call.
Syntax:
class ShowRequest(BaseCallbackHandler):
def on_chat_model_start(self, serialized, messages, **kwargs):
... # messages[0] is the list about to be sent
manager.invoke({"messages": conversation}, config={"callbacks": [ShowRequest()]})A handler that prints the request
on_chat_model_start receives the messages as a list of lists, one per prompt, and the call's parameters, which include the tools in the OpenAI format.
from langchain_core.callbacks import BaseCallbackHandler
class ShowRequest(BaseCallbackHandler):
def on_chat_model_start(self, serialized, messages, **kwargs):
tools = kwargs["invocation_params"].get("tools", [])
print("tools:", [tool["function"]["name"] for tool in tools])
for message in messages[0]:
print(f"--- {message.type}")
print(message.content)Passing the handler to a call
Any LangChain runnable takes callbacks in its config. LangMem passes the config down to the model call.
manager.invoke(
{"messages": [{"role": "user", "content": "Please email me."}]},
config={"callbacks": [ShowRequest()]},
)The request for a new conversation
from langchain.chat_models import init_chat_model
from langchain_core.callbacks import BaseCallbackHandler
from langmem import create_memory_manager
class ShowRequest(BaseCallbackHandler):
def on_chat_model_start(self, serialized, messages, **kwargs):
tools = kwargs["invocation_params"].get("tools", [])
print("tools:", [tool["function"]["name"] for tool in tools])
for message in messages[0]:
print(f"--- {message.type}")
print(message.content)
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
manager = create_memory_manager(model, instructions="Extract what helps support this customer. Record everything in a single Memory call.")
manager.invoke(
{"messages": [{"role": "user", "content": "Please email me."}]},
config={"callbacks": [ShowRequest()]},
)tools: ['Memory'] --- system You are a memory subroutine for an AI. --- human Extract what helps support this customer. Record everything in a single Memory call. Enrich, prune, and organize memories based on any new information. If an existing memory is incorrect or outdated, update it based on the new information. All operations must be done in single parallel multi-tool call. Avoid duplicate extractions. <session_d5efa83f-82f5-47ce-a339-021ba24492b9> ================================ Human Message ================================= Please email me. </session_d5efa83f-82f5-47ce-a339-021ba24492b9>
Reading the request
- One tool, Memory. Its arguments are the fields of the
Memoryschema, one text field calledcontent. - A one-line system message that tells the model it is a memory subroutine.
- A human message with your instructions first, then LangMem's rules, then the conversation inside a
<session_...>tag with a random id. - The sentence behind the 400. The rules end with "All operations must be done in single parallel multi-tool call." gpt-oss sends one call per reply, which is why the instructions ask for a single
Memorycall.
The request with existing memories
Pass existing, a list of (id, memory) pairs, and the request changes. Memory is the default schema; here one memory is already known:
from langchain.chat_models import init_chat_model
from langchain_core.callbacks import BaseCallbackHandler
from langmem import create_memory_manager
from langmem.knowledge.extraction import Memory
class ShowRequest(BaseCallbackHandler):
def on_chat_model_start(self, serialized, messages, **kwargs):
tools = kwargs["invocation_params"].get("tools", [])
print("tools:", [tool["function"]["name"] for tool in tools])
for message in messages[0]:
print(f"--- {message.type}")
print(message.content)
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
manager = create_memory_manager(model, instructions="Extract what helps support this customer. Record everything in a single Memory call.")
existing = [("m-1", Memory(content="Asha wants us to email her."))]
manager.invoke(
{"messages": [{"role": "user", "content": "Text me instead."}], "existing": existing},
config={"callbacks": [ShowRequest()]},
)tools: ['PatchDoc', 'Memory']
--- system
You are a memory subroutine for an AI.
Generate JSONPatches to update the existing schema instances. If you need to extract or insert *new* instances of the schemas, call the relevant function(s).
<existing>
<instance id=m-1 schema_type="Memory">
{'content': 'Asha wants us to email her.'}
</instance>
</existing>
--- human
Extract what helps support this customer. Record everything in a single Memory call.
Enrich, prune, and organize memories based on any new information. If an existing memory is incorrect or outdated, update it based on the new information. All operations must be done in single parallel multi-tool call. Avoid duplicate extractions.
<session_56cd84f6-7701-4f5d-a2ed-374274691b74>
================================ Human Message =================================
Text me instead.
</session_56cd84f6-7701-4f5d-a2ed-374274691b74>What existing memories add
- A PatchDoc tool appears next to
Memory. It edits an existing memory by id with JSON Patch operations; Updates and deletes uses it. - The system message lists each memory in an
<instance id=m-1>block, so the model can refer to it by id. - Your instructions stay first in the human message; the conversation stays in its session tag.
Request with and without existing memories
| New conversation only | With existing memories | |
|---|---|---|
| Tools | Memory | PatchDoc and Memory |
| System message | One line | Adds the JSON Patch rules and every existing memory |
| What the model can do | Create memories | Create memories or patch old ones |
When to print the request
- When a manager keeps or skips something you did not expect: read what the model was told.
- Before you write instructions, to see where they land in the prompt.
- When changing models, to check the new one gets the same tools.
<instance> blocks in the prompt; the store manager in Store managers searches first and passes only the related few.Related
- Previous: Extracting memories
- Next: Chat models
- Reference: LangChain callbacks
- Pass
enable_deletes=Trueto the manager with existing memories and look for a new tool. - Remove
instructionsand read LangMem's default prompt. - Print the whole tool dictionary,
print(tools), to see the JSON schema ofMemory.
Slow is fine. Stopping is the only problem.