Background memory
Background memory is memory formation that happens after the reply: LangMem's ReflectionExecutor runs a store manager on a worker thread after a delay, so extracting memories never slows the conversation down.
Last updated: 30 Sep, 2026 · LangMem 0.0.30
A store manager call takes a model round trip, often several seconds. Running it before every reply makes the customer wait; running it after every message also wastes calls, because a customer typing five quick messages would trigger five extractions over half-finished context. ReflectionExecutor solves both: it waits, and a newer submit for the same thread replaces one still waiting.
Syntax:
executor = ReflectionExecutor(manager, store=store)
future = executor.submit({"messages": messages}, config=config, after_seconds=60)Wrapping a store manager
The executor takes the store manager from Store managers and the store it writes to. Used in a with block, it shuts its worker thread down at the end.
manager = create_memory_store_manager(model, namespace=("memories", "{user_id}"), instructions=INSTRUCTIONS, store=store)
with ReflectionExecutor(manager, store=store) as executor:
...Submitting a conversation
submit returns at once with a Future. The extraction runs after_seconds later. Outside a LangGraph node, the config must carry the ids, including thread_id, which the executor uses to find earlier pending work for the same conversation.
config = {"configurable": {"user_id": "meera", "thread_id": "chat-9"}}
future = executor.submit({"messages": messages}, config=config, after_seconds=1)Extracting after the reply
import time
from langchain.chat_models import init_chat_model
from langgraph.store.memory import InMemoryStore
from langmem import ReflectionExecutor, create_memory_store_manager
INSTRUCTIONS = "Extract what helps support this customer. Record everything in a single Memory call."
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
store = InMemoryStore(index={"dims": 3072, "embed": "google_genai:gemini-embedding-2"})
manager = create_memory_store_manager(model, namespace=("memories", "{user_id}"), instructions=INSTRUCTIONS, store=store)
config = {"configurable": {"user_id": "meera", "thread_id": "chat-9"}}
with ReflectionExecutor(manager, store=store) as executor:
started = time.time()
first = executor.submit({"messages": [{"role": "user", "content": "I moved to Pune last month."}]}, config=config, after_seconds=2)
second = executor.submit({"messages": [
{"role": "user", "content": "I moved to Pune last month."},
{"role": "user", "content": "Also, deliveries should go to my office."},
]}, config=config, after_seconds=2)
print(f"submit returned after {time.time() - started:.1f}s; memories now: {len(store.search(('memories', 'meera')))}")
second.result()
print("first cancelled:", first.cancelled())
for item in store.search(("memories", "meera")):
print("saved:", item.value["content"]["content"])submit returned after 0.0s; memories now: 0 first cancelled: True saved: Customer relocated to Pune last month; future deliveries should be sent to the customer's office address.
What the executor did
- submit returned after 0.0s with no memories saved yet. The extraction had not started; a reply to the customer could go out at this point.
- first cancelled: True. The second submit came for the same
thread_idwhile the first was still waiting, so the executor cancelled the first. One extraction ran, over both messages. - One memory, with both facts: the move to Pune and deliveries to the office.
Hot path vs background
| In the hot path | In the background | |
|---|---|---|
| When memories are written | During the reply | After a delay, on another thread |
| Customer waits for it | Yes | No |
| Context the model sees | The conversation so far | The whole conversation once it goes quiet |
| LangMem pieces | Memory tools | ReflectionExecutor with a store manager |
When to extract in the background
- Chat apps where reply speed matters more than a memory existing within the same second.
- Long conversations, where extracting once when the user goes quiet costs one call instead of dozens.
- Episodes, which only make sense once a conversation is over.
Related
- Previous: Agent memory
- Next: Running summaries
- Reference: Delayed background memory processing
- Submit for two different
thread_idvalues and check that neither is cancelled. - Remove
second.result()and print the memories inside thewithblock. - Set
after_seconds=0and see whether the first submit still gets cancelled.
This is what real progress feels like.