Long-term memory: recalling past sessions
A memory block is a store of past messages that a later turn can search, so an agent recalls a fact a customer gave earlier.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
Short-term memory keeps the last few turns. Once they scroll out, they are gone. A VectorMemoryBlock embeds old turns and retrieves the ones relevant to the new question, reusing the same embedding model as the rest of the course.
Building a vector memory block
The block needs a vector store that keeps the text, so an ephemeral Chroma collection stands in for a real one. The embedding model is the same local model used for retrieval.
import chromadb
from llama_index.core.memory import VectorMemoryBlock
from llama_index.vector_stores.chroma import ChromaVectorStore
store = ChromaVectorStore(chroma_collection=chromadb.EphemeralClient().create_collection("recall"))
recall = VectorMemoryBlock(name="recall", vector_store=store,
embed_model=Settings.embed_model, similarity_top_k=1)The memory that flushes to long-term storage
Memory.from_defaults holds recent turns in a short buffer and moves older ones into the block. A small buffer here pushes the older facts into long-term memory quickly.
from llama_index.core.memory import Memory
memory = Memory.from_defaults(
session_id="cust-42",
token_limit=300, # small buffer, so older turns move to the block
memory_blocks=[recall],
insert_method="user",
chat_history_token_ratio=0.01, # keep almost nothing short-term
token_flush_size=1,
)Storing turns, then asking later
Store a few turns from an earlier session, then ask a new question. memory.aget(input=...) returns the recalled turns wrapped in a memory block.
await memory.aput_messages([...past turns...])
messages = await memory.aget(input="Where should my order be delivered?")
# the returned messages include a <recall> section with the matching turnRecalling the right fact for each question
The whole program stores three facts, then asks two questions. Each question pulls back the fact that matches it, not whatever was stored last.
import asyncio
import re
import chromadb
from llama_index.core import Settings
from llama_index.core.base.llms.types import ChatMessage
from llama_index.core.memory import Memory, VectorMemoryBlock
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.vector_stores.chroma import ChromaVectorStore
Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
store = ChromaVectorStore(chroma_collection=chromadb.EphemeralClient().create_collection("recall"))
recall = VectorMemoryBlock(name="recall", vector_store=store, embed_model=Settings.embed_model, similarity_top_k=1)
memory = Memory.from_defaults(
session_id="cust-42",
token_limit=300, # a small buffer, so older turns move to long-term memory
memory_blocks=[recall],
insert_method="user",
chat_history_token_ratio=0.01, # keep almost nothing short-term; flush the rest
token_flush_size=1,
)
def first_recalled(messages):
for message in messages:
found = re.search(r"<recall>(.*?)</recall>", message.content or "", re.S)
if found:
said = re.findall(r"role='user'>(.*?)</message>", found.group(1))
return said[0] if said else ""
return "nothing recalled"
async def main():
await memory.aput_messages([
ChatMessage(role="user", content="Please deliver order A1001 to my office, not my home."),
ChatMessage(role="assistant", content="Done, A1001 will go to your office."),
ChatMessage(role="user", content="I prefer email over phone for updates."),
ChatMessage(role="assistant", content="Noted, email it is."),
ChatMessage(role="user", content="My favourite is the LMP-310 floor lamp."),
ChatMessage(role="assistant", content="Good to know."),
])
for question in ["Where should my order be delivered?", "How do you contact me?"]:
messages = await memory.aget(input=question)
print(question)
print(" recalled:", first_recalled(messages))
asyncio.run(main())Why each question recalled a different turn
- The delivery question is closest in meaning to the office-address turn, so that turn comes back.
- The contact question matches the email-preference turn instead, which proves the recall is by meaning, not by order.
- The lamp turn stayed in the short buffer and was not needed, so neither answer mentions it.
Short-term buffer vs a vector memory block
| Memory | Holds | Recall |
|---|---|---|
| Short-term buffer | The last few turns | Always included, until they scroll out |
| VectorMemoryBlock | All flushed turns, embedded | Only the turns close in meaning to the question |
Where long-term memory earns its place
- Remembering a customer's address or preferences between sessions.
- A long chat where early details would otherwise scroll away.
- Recalling a past decision without replaying the whole history to the model.
SimpleVectorStore is refused: it holds vectors but not the messages. Use a store like Chroma that keeps both.Related
- Previous: Agent tools: giving an agent more than search
- Next: Workflows: a custom RAG pipeline with @step
- See also: Embeddings: text as numbers that carry meaning
- Reference: Agent memory
- Add a fourth fact and a question that should recall it.
- Raise
similarity_top_kto 2 and see both facts a question can bring back. - Ask a question unrelated to every stored fact and read what comes back.
You understood something today that you didn't yesterday.