LangMemLangMem 0.0.30 · LangGraph 1.2 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Background memory

Background memory is memory formation that happens after the reply: LangMem's ReflectionExecutor runs a store manager on a worker thread after a delay, so extracting memories never slows the conversation down.

Last updated: 30 Sep, 2026 · LangMem 0.0.30

A store manager call takes a model round trip, often several seconds. Running it before every reply makes the customer wait; running it after every message also wastes calls, because a customer typing five quick messages would trigger five extractions over half-finished context. ReflectionExecutor solves both: it waits, and a newer submit for the same thread replaces one still waiting.

Syntax:

python
executor = ReflectionExecutor(manager, store=store)
future = executor.submit({"messages": messages}, config=config, after_seconds=60)

Wrapping a store manager

The executor takes the store manager from Store managers and the store it writes to. Used in a with block, it shuts its worker thread down at the end.

python
manager = create_memory_store_manager(model, namespace=("memories", "{user_id}"), instructions=INSTRUCTIONS, store=store)
with ReflectionExecutor(manager, store=store) as executor:
    ...

Submitting a conversation

submit returns at once with a Future. The extraction runs after_seconds later. Outside a LangGraph node, the config must carry the ids, including thread_id, which the executor uses to find earlier pending work for the same conversation.

python
config = {"configurable": {"user_id": "meera", "thread_id": "chat-9"}}
future = executor.submit({"messages": messages}, config=config, after_seconds=1)

Extracting after the reply

ExampleAPI key
import time

from langchain.chat_models import init_chat_model
from langgraph.store.memory import InMemoryStore
from langmem import ReflectionExecutor, create_memory_store_manager

INSTRUCTIONS = "Extract what helps support this customer. Record everything in a single Memory call."

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
store = InMemoryStore(index={"dims": 3072, "embed": "google_genai:gemini-embedding-2"})
manager = create_memory_store_manager(model, namespace=("memories", "{user_id}"), instructions=INSTRUCTIONS, store=store)

config = {"configurable": {"user_id": "meera", "thread_id": "chat-9"}}
with ReflectionExecutor(manager, store=store) as executor:
    started = time.time()
    first = executor.submit({"messages": [{"role": "user", "content": "I moved to Pune last month."}]}, config=config, after_seconds=2)
    second = executor.submit({"messages": [
        {"role": "user", "content": "I moved to Pune last month."},
        {"role": "user", "content": "Also, deliveries should go to my office."},
    ]}, config=config, after_seconds=2)
    print(f"submit returned after {time.time() - started:.1f}s; memories now: {len(store.search(('memories', 'meera')))}")
    second.result()
    print("first cancelled:", first.cancelled())

for item in store.search(("memories", "meera")):
    print("saved:", item.value["content"]["content"])

What the executor did

  • submit returned after 0.0s with no memories saved yet. The extraction had not started; a reply to the customer could go out at this point.
  • first cancelled: True. The second submit came for the same thread_id while the first was still waiting, so the executor cancelled the first. One extraction ran, over both messages.
  • One memory, with both facts: the move to Pune and deliveries to the office.

Hot path vs background

In the hot pathIn the background
When memories are writtenDuring the replyAfter a delay, on another thread
Customer waits for itYesNo
Context the model seesThe conversation so farThe whole conversation once it goes quiet
LangMem piecesMemory toolsReflectionExecutor with a store manager

When to extract in the background

  • Chat apps where reply speed matters more than a memory existing within the same second.
  • Long conversations, where extracting once when the user goes quiet costs one call instead of dozens.
  • Episodes, which only make sense once a conversation is over.
Watch out. The executor's worker is a thread in your process. In a serverless function the process can stop before a delayed extraction runs, and the memory is lost; LangMem's docs point to LangGraph Platform's remote executor for that case.
Try it yourself
  • Submit for two different thread_id values and check that neither is cancelled.
  • Remove second.result() and print the memories inside the with block.
  • Set after_seconds=0 and see whether the first submit still gets cancelled.
PreviousAgent memory

This is what real progress feels like.