Extracting memories
create_memory_manager is a LangMem function that turns a chat model into a memory extractor: give it a conversation and it returns the memories the model decided to keep, each with an id.
Last updated: 30 Sep, 2026 · LangMem 0.0.30
This is the core of LangMem. Everything later in the course, the store, the tools, the background executor, wraps this one step: a model reads a conversation and writes down what is worth remembering. What it writes down is semantic memory, the facts a later conversation needs.
create_memory_manager does the extraction and the updating, with a Groq model.Facts that last across sessions
The clip's example: a user says their favourite language is Python in session one, their team size in session two and their deployment target in session three. Without semantic memory, each session starts from zero. With it, the agent keeps a growing set of distilled facts and stops asking the same questions. A fact is kept independent of when it was learned, and it is updated when it changes.
Syntax:
from langmem import create_memory_manager
manager = create_memory_manager(model, instructions="What to keep")
memories = manager.invoke({"messages": conversation}) # a list of ExtractedMemoryThe conversation
A conversation is a list of messages in the role-and-content format chat APIs use. Asha's first message to the shop:
conversation = [
{"role": "user", "content": "Hi, my name is Asha. Order A-1001 arrived broken."},
{"role": "assistant", "content": "Sorry to hear that. How should we contact you?"},
{"role": "user", "content": "Please email me, I work nights."},
]The model and the manager
create_memory_manager takes the chat model from Installation and setup. The docs call it with only a model, so try that first:
from langchain.chat_models import init_chat_model
from langmem import create_memory_manager
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
manager = create_memory_manager(model)The default manager on gpt-oss
from langchain.chat_models import init_chat_model
from langmem import create_memory_manager
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
manager = create_memory_manager(model)
conversation = [
{"role": "user", "content": "Hi, my name is Asha. Order A-1001 arrived broken."},
{"role": "assistant", "content": "Sorry to hear that. How should we contact you?"},
{"role": "user", "content": "Please email me, I work nights."},
]
memories = manager.invoke({"messages": conversation})Traceback (most recent call last):
File "main.py", line 11, in <module>
memories = manager.invoke({"messages": conversation})
groq.BadRequestError: Error code: 400 - {'error': {'message': 'Failed to parse tool call arguments as JSON', 'type': 'invalid_request_error', 'code': 'tool_use_failed', 'failed_generation': '{"name": "Memory", "arguments": [\n {\n "content": "User\'s name is Asha."\n },\n {\n "content": "Order A-1001 arrived broken."\n },\n {\n "content": "User prefers to be contacted via email."\n },\n {\n "content": "User works nights, indicating they may be more responsive to email outside typical business hours."\n }\n]"}'}}
During task with name 'extract' and id 'd88a0ceb-eb89-ee8f-bdc9-de1530f12601'Groq rejected the model's reply with 400. LangMem's prompt asks the model to record every memory "in single parallel multi-tool call", one Memory tool call per fact, all in one reply. gpt-oss on Groq sends one tool call per reply, so it tried to pack all the facts into that one call as a list, and Groq could not parse the list as the tool's arguments. It does not fail every time: sometimes the model sends a single fact and the call succeeds, with one memory. The extraction request prints the prompt with that sentence in it.
Asking for a single Memory call
Instructions are how you tell LangMem what to keep. They also fix the failure: ask for everything in one Memory call, and gpt-oss writes one complete memory instead of trying to send a list.
manager = create_memory_manager(
model,
instructions="Extract what helps support this customer. Record everything in a single Memory call.",
)Extracting Asha's memory
from langchain.chat_models import init_chat_model
from langmem import create_memory_manager
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
manager = create_memory_manager(
model,
instructions="Extract what helps support this customer. Record everything in a single Memory call.",
)
conversation = [
{"role": "user", "content": "Hi, my name is Asha. Order A-1001 arrived broken."},
{"role": "assistant", "content": "Sorry to hear that. How should we contact you?"},
{"role": "user", "content": "Please email me, I work nights."},
]
memories = manager.invoke({"messages": conversation})
for memory in memories:
print(type(memory).__name__, "|", type(memory.content).__name__, "|", len(memory.id), "characters in the id")
print(memory.content.content)ExtractedMemory | Memory | 36 characters in the id Customer Asha reported that order A-1001 arrived broken. Preferred contact method: email. Note: works nights, so email contact preferred for night hours.
What the manager returned
- A list of ExtractedMemory. Each has an
idand acontent. The id is a UUID, 36 characters, so a later update can say which memory it changes. - content is a Memory.
Memoryis LangMem's default schema, a Pydantic model with one text field, also calledcontent. That is why the text is atmemory.content.content. - The model chose the wording. Nothing in the code says "name" or "email"; the model kept the name, the broken order and email in one memory, as the instructions asked.
- It also kept an inference. "email contact preferred for night hours" is the model's reading of "I work nights"; Asha did not say it.
- Nothing was saved anywhere. The manager is a function from a conversation to memories. Keeping them is your job, or a store's, in Store managers.
Default prompt vs your instructions
| No instructions | With instructions | |
|---|---|---|
| Prompt | A long default about semantic, episodic and procedural memory | Your text, then LangMem's short rules |
| What gets kept | Whatever the model finds noteworthy | What your app says it needs |
| On gpt-oss | Often fails with a 400 | One complete memory per call |
When to extract memories
- After a support chat, to keep the customer's contact preference and open problems.
- In a tutoring app, to keep what a student already knows and where they struggled.
- In any assistant that should stop asking returning users the same questions.
Related
- Previous: Installation and setup
- Next: The extraction request
- Reference: How to extract semantic memories
- Add
{"role": "user", "content": "I live in Pune."}to the conversation and extract again. - Change the instructions to keep only contact preferences, and compare the memory.
- Print
memories[0].content.model_dump()to see the memory as a dictionary.
You understood something today that you didn't yesterday.