LangMemLangMem 0.0.30 · LangGraph 1.2 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Extracting memories

create_memory_manager is a LangMem function that turns a chat model into a memory extractor: give it a conversation and it returns the memories the model decided to keep, each with an id.

Last updated: 30 Sep, 2026 · LangMem 0.0.30

This is the core of LangMem. Everything later in the course, the store, the tools, the background executor, wraps this one step: a model reads a conversation and writes down what is worth remembering. What it writes down is semantic memory, the facts a later conversation needs.

Semantic memory · from the AI Security Course · 315:51 to 319:55
The video's notebook builds semantic memory from scratch with the OpenAI SDK: an extraction prompt, embeddings, duplicate and contradiction checks. On this page LangMem's create_memory_manager does the extraction and the updating, with a Groq model.

Facts that last across sessions

The clip's example: a user says their favourite language is Python in session one, their team size in session two and their deployment target in session three. Without semantic memory, each session starts from zero. With it, the agent keeps a growing set of distilled facts and stops asking the same questions. A fact is kept independent of when it was learned, and it is updated when it changes.

Syntax:

python
from langmem import create_memory_manager

manager = create_memory_manager(model, instructions="What to keep")
memories = manager.invoke({"messages": conversation})  # a list of ExtractedMemory

The conversation

A conversation is a list of messages in the role-and-content format chat APIs use. Asha's first message to the shop:

python
conversation = [
    {"role": "user", "content": "Hi, my name is Asha. Order A-1001 arrived broken."},
    {"role": "assistant", "content": "Sorry to hear that. How should we contact you?"},
    {"role": "user", "content": "Please email me, I work nights."},
]

The model and the manager

create_memory_manager takes the chat model from Installation and setup. The docs call it with only a model, so try that first:

python
from langchain.chat_models import init_chat_model
from langmem import create_memory_manager

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
manager = create_memory_manager(model)

The default manager on gpt-oss

ExampleAPI key
from langchain.chat_models import init_chat_model
from langmem import create_memory_manager

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
manager = create_memory_manager(model)
conversation = [
    {"role": "user", "content": "Hi, my name is Asha. Order A-1001 arrived broken."},
    {"role": "assistant", "content": "Sorry to hear that. How should we contact you?"},
    {"role": "user", "content": "Please email me, I work nights."},
]
memories = manager.invoke({"messages": conversation})

Groq rejected the model's reply with 400. LangMem's prompt asks the model to record every memory "in single parallel multi-tool call", one Memory tool call per fact, all in one reply. gpt-oss on Groq sends one tool call per reply, so it tried to pack all the facts into that one call as a list, and Groq could not parse the list as the tool's arguments. It does not fail every time: sometimes the model sends a single fact and the call succeeds, with one memory. The extraction request prints the prompt with that sentence in it.

Asking for a single Memory call

Instructions are how you tell LangMem what to keep. They also fix the failure: ask for everything in one Memory call, and gpt-oss writes one complete memory instead of trying to send a list.

python
manager = create_memory_manager(
    model,
    instructions="Extract what helps support this customer. Record everything in a single Memory call.",
)

Extracting Asha's memory

ExampleAPI key
from langchain.chat_models import init_chat_model
from langmem import create_memory_manager

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
manager = create_memory_manager(
    model,
    instructions="Extract what helps support this customer. Record everything in a single Memory call.",
)
conversation = [
    {"role": "user", "content": "Hi, my name is Asha. Order A-1001 arrived broken."},
    {"role": "assistant", "content": "Sorry to hear that. How should we contact you?"},
    {"role": "user", "content": "Please email me, I work nights."},
]
memories = manager.invoke({"messages": conversation})
for memory in memories:
    print(type(memory).__name__, "|", type(memory.content).__name__, "|", len(memory.id), "characters in the id")
    print(memory.content.content)

What the manager returned

  • A list of ExtractedMemory. Each has an id and a content. The id is a UUID, 36 characters, so a later update can say which memory it changes.
  • content is a Memory. Memory is LangMem's default schema, a Pydantic model with one text field, also called content. That is why the text is at memory.content.content.
  • The model chose the wording. Nothing in the code says "name" or "email"; the model kept the name, the broken order and email in one memory, as the instructions asked.
  • It also kept an inference. "email contact preferred for night hours" is the model's reading of "I work nights"; Asha did not say it.
  • Nothing was saved anywhere. The manager is a function from a conversation to memories. Keeping them is your job, or a store's, in Store managers.

Default prompt vs your instructions

No instructionsWith instructions
PromptA long default about semantic, episodic and procedural memoryYour text, then LangMem's short rules
What gets keptWhatever the model finds noteworthyWhat your app says it needs
On gpt-ossOften fails with a 400One complete memory per call

When to extract memories

  • After a support chat, to keep the customer's contact preference and open problems.
  • In a tutoring app, to keep what a student already knows and where they struggled.
  • In any assistant that should stop asking returning users the same questions.
Watch out. Model output varies. Run the example twice and the memory is worded differently each time, and a run can leave out a fact. Code that reads memories should not depend on exact wording.
Try it yourself
  • Add {"role": "user", "content": "I live in Pune."} to the conversation and extract again.
  • Change the instructions to keep only contact preferences, and compare the memory.
  • Print memories[0].content.model_dump() to see the memory as a dictionary.

You understood something today that you didn't yesterday.