Chat models
A chat model is the part of LangMem that does the reading: any LangChain chat model, passed as an object or as a "provider:model" string, that answers LangMem's request with tool calls.
Last updated: 30 Sep, 2026 · LangMem 0.0.30
LangMem has no model of its own. Every manager, optimizer and summarizer takes a model argument, and what it can do depends on that model: how well it follows instructions and how it makes tool calls.
Syntax:
create_memory_manager("groq:openai/gpt-oss-120b") # a string: LangMem calls init_chat_model for you
create_memory_manager(model) # an object you configured yourselfA model string
A string is passed to LangChain's init_chat_model with no other settings. It is the shortest way to write it.
manager = create_memory_manager("groq:openai/gpt-oss-120b", instructions="Extract what helps support this customer. Record everything in a single Memory call.")A model object
An object lets you set options first, such as temperature or a smaller model. The course passes an object so every run uses temperature=0.
model = init_chat_model("groq:openai/gpt-oss-20b", temperature=0)
manager = create_memory_manager(model, instructions="Extract what helps support this customer. Record everything in a single Memory call.")Running without a key
Without GROQ_API_KEY the model cannot be created, so the manager fails before anything is sent:
from langmem import create_memory_manager
manager = create_memory_manager("groq:openai/gpt-oss-120b")
manager.invoke({"messages": [{"role": "user", "content": "Please email me."}]})Traceback (most recent call last):
File "main.py", line 3, in <module>
manager = create_memory_manager("groq:openai/gpt-oss-120b")
groq.GroqError: The api_key client option must be set either by passing api_key to the client or by setting the GROQ_API_KEY environment variableThe error comes from the line that creates the manager: LangMem builds the Groq client there, and the client needs the key. The message names the variable to set; set it as in Installation and setup and the same code runs.
The smaller model on the same conversation
When the 120b model's daily limit is used up, gpt-oss-20b has a separate budget. The same manager, with only the model string changed:
from langchain.chat_models import init_chat_model
from langmem import create_memory_manager
model = init_chat_model("groq:openai/gpt-oss-20b", temperature=0)
manager = create_memory_manager(model, instructions="Extract what helps support this customer. Record everything in a single Memory call.")
conversation = [
{"role": "user", "content": "Hi, my name is Asha. Order A-1001 arrived broken."},
{"role": "assistant", "content": "Sorry to hear that. How should we contact you?"},
{"role": "user", "content": "Please email me, I work nights."},
]
for memory in manager.invoke({"messages": conversation}):
print(memory.content.content)Customer Asha (order A-1001) reported that the item arrived broken. She prefers to be contacted via email and works nights, so any follow‑up communication should be sent to her email and scheduled for nighttime hours.
What changed with the smaller model
- The code did not change, only the model string. Every LangMem function works the same way with any chat model.
- The same three facts came out: the broken order, email, and the night shifts, as with the 120b model in Extracting memories.
- It added an instruction Asha never gave: follow-ups "scheduled for nighttime hours". Both models inferred something from "I work nights"; read memories as the model's interpretation, not a transcript.
One tool call per reply vs parallel tool calls
| gpt-oss on Groq | Models with parallel tool calls | |
|---|---|---|
| Tool calls in one reply | One | Several |
| LangMem's default prompt | Often fails with a 400 | Works as the docs show |
| What to do | Ask for a single Memory call, or use a profile schema | Nothing extra |
| Examples | openai/gpt-oss-120b, openai/gpt-oss-20b | The docs' examples use Anthropic's Claude |
Choosing a model for memory
- For extraction, pick a model that is good at tool calling; the prose quality of its replies matters less.
- Use the same model in development and production. A manager tuned on one model can keep different things on another.
- Keep a cheaper model for background extraction if the main agent needs the larger one.
model_not_found or decommissioned, check the provider's models page and change the string.Related
- Previous: The extraction request
- Next: Collections
- Reference: init_chat_model
- Pass the model as the string
"groq:openai/gpt-oss-120b"with the instructions and extract Asha's memory. - Set
temperature=1and run the extraction three times. - Print
model.model_namefor the object you created.
This is what real progress feels like.