LangMemLangMem 0.0.30 · LangGraph 1.2 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

The extraction request

The extraction request is the message list and tool list that a LangMem memory manager sends to the chat model: a short system message, your instructions around the conversation, and one tool per memory schema.

Last updated: 30 Sep, 2026 · LangMem 0.0.30

Reading the request explains everything the manager does, including why the default call failed on gpt-oss in Extracting memories. A LangChain callback handler can print it, because LangChain calls the handler with the exact messages and tools right before each model call.

Syntax:

python
class ShowRequest(BaseCallbackHandler):
    def on_chat_model_start(self, serialized, messages, **kwargs):
        ...  # messages[0] is the list about to be sent

manager.invoke({"messages": conversation}, config={"callbacks": [ShowRequest()]})

A handler that prints the request

on_chat_model_start receives the messages as a list of lists, one per prompt, and the call's parameters, which include the tools in the OpenAI format.

python
from langchain_core.callbacks import BaseCallbackHandler


class ShowRequest(BaseCallbackHandler):
    def on_chat_model_start(self, serialized, messages, **kwargs):
        tools = kwargs["invocation_params"].get("tools", [])
        print("tools:", [tool["function"]["name"] for tool in tools])
        for message in messages[0]:
            print(f"--- {message.type}")
            print(message.content)

Passing the handler to a call

Any LangChain runnable takes callbacks in its config. LangMem passes the config down to the model call.

python
manager.invoke(
    {"messages": [{"role": "user", "content": "Please email me."}]},
    config={"callbacks": [ShowRequest()]},
)

The request for a new conversation

ExampleAPI key
from langchain.chat_models import init_chat_model
from langchain_core.callbacks import BaseCallbackHandler
from langmem import create_memory_manager

class ShowRequest(BaseCallbackHandler):
    def on_chat_model_start(self, serialized, messages, **kwargs):
        tools = kwargs["invocation_params"].get("tools", [])
        print("tools:", [tool["function"]["name"] for tool in tools])
        for message in messages[0]:
            print(f"--- {message.type}")
            print(message.content)

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
manager = create_memory_manager(model, instructions="Extract what helps support this customer. Record everything in a single Memory call.")
manager.invoke(
    {"messages": [{"role": "user", "content": "Please email me."}]},
    config={"callbacks": [ShowRequest()]},
)

Reading the request

  • One tool, Memory. Its arguments are the fields of the Memory schema, one text field called content.
  • A one-line system message that tells the model it is a memory subroutine.
  • A human message with your instructions first, then LangMem's rules, then the conversation inside a <session_...> tag with a random id.
  • The sentence behind the 400. The rules end with "All operations must be done in single parallel multi-tool call." gpt-oss sends one call per reply, which is why the instructions ask for a single Memory call.

The request with existing memories

Pass existing, a list of (id, memory) pairs, and the request changes. Memory is the default schema; here one memory is already known:

ExampleAPI key
from langchain.chat_models import init_chat_model
from langchain_core.callbacks import BaseCallbackHandler
from langmem import create_memory_manager
from langmem.knowledge.extraction import Memory

class ShowRequest(BaseCallbackHandler):
    def on_chat_model_start(self, serialized, messages, **kwargs):
        tools = kwargs["invocation_params"].get("tools", [])
        print("tools:", [tool["function"]["name"] for tool in tools])
        for message in messages[0]:
            print(f"--- {message.type}")
            print(message.content)

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
manager = create_memory_manager(model, instructions="Extract what helps support this customer. Record everything in a single Memory call.")
existing = [("m-1", Memory(content="Asha wants us to email her."))]
manager.invoke(
    {"messages": [{"role": "user", "content": "Text me instead."}], "existing": existing},
    config={"callbacks": [ShowRequest()]},
)

What existing memories add

  • A PatchDoc tool appears next to Memory. It edits an existing memory by id with JSON Patch operations; Updates and deletes uses it.
  • The system message lists each memory in an <instance id=m-1> block, so the model can refer to it by id.
  • Your instructions stay first in the human message; the conversation stays in its session tag.

Request with and without existing memories

New conversation onlyWith existing memories
ToolsMemoryPatchDoc and Memory
System messageOne lineAdds the JSON Patch rules and every existing memory
What the model can doCreate memoriesCreate memories or patch old ones

When to print the request

  • When a manager keeps or skips something you did not expect: read what the model was told.
  • Before you write instructions, to see where they land in the prompt.
  • When changing models, to check the new one gets the same tools.
Watch out. The request grows with every existing memory you pass. Passing a thousand memories puts a thousand <instance> blocks in the prompt; the store manager in Store managers searches first and passes only the related few.
Try it yourself
  • Pass enable_deletes=True to the manager with existing memories and look for a new tool.
  • Remove instructions and read LangMem's default prompt.
  • Print the whole tool dictionary, print(tools), to see the JSON schema of Memory.

Slow is fine. Stopping is the only problem.