LlamaIndexllama-index-core 0.14 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Long-term memory: recalling past sessions

A memory block is a store of past messages that a later turn can search, so an agent recalls a fact a customer gave earlier.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

Short-term memory keeps the last few turns. Once they scroll out, they are gone. A VectorMemoryBlock embeds old turns and retrieves the ones relevant to the new question, reusing the same embedding model as the rest of the course.

Building a vector memory block

The block needs a vector store that keeps the text, so an ephemeral Chroma collection stands in for a real one. The embedding model is the same local model used for retrieval.

python
import chromadb
from llama_index.core.memory import VectorMemoryBlock
from llama_index.vector_stores.chroma import ChromaVectorStore

store = ChromaVectorStore(chroma_collection=chromadb.EphemeralClient().create_collection("recall"))
recall = VectorMemoryBlock(name="recall", vector_store=store,
                           embed_model=Settings.embed_model, similarity_top_k=1)

The memory that flushes to long-term storage

Memory.from_defaults holds recent turns in a short buffer and moves older ones into the block. A small buffer here pushes the older facts into long-term memory quickly.

python
from llama_index.core.memory import Memory

memory = Memory.from_defaults(
    session_id="cust-42",
    token_limit=300,                # small buffer, so older turns move to the block
    memory_blocks=[recall],
    insert_method="user",
    chat_history_token_ratio=0.01,  # keep almost nothing short-term
    token_flush_size=1,
)

Storing turns, then asking later

Store a few turns from an earlier session, then ask a new question. memory.aget(input=...) returns the recalled turns wrapped in a memory block.

python
await memory.aput_messages([...past turns...])
messages = await memory.aget(input="Where should my order be delivered?")
# the returned messages include a <recall> section with the matching turn

Recalling the right fact for each question

The whole program stores three facts, then asks two questions. Each question pulls back the fact that matches it, not whatever was stored last.

Example
import asyncio
import re

import chromadb
from llama_index.core import Settings
from llama_index.core.base.llms.types import ChatMessage
from llama_index.core.memory import Memory, VectorMemoryBlock
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.vector_stores.chroma import ChromaVectorStore

Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")

store = ChromaVectorStore(chroma_collection=chromadb.EphemeralClient().create_collection("recall"))
recall = VectorMemoryBlock(name="recall", vector_store=store, embed_model=Settings.embed_model, similarity_top_k=1)
memory = Memory.from_defaults(
    session_id="cust-42",
    token_limit=300,                # a small buffer, so older turns move to long-term memory
    memory_blocks=[recall],
    insert_method="user",
    chat_history_token_ratio=0.01,  # keep almost nothing short-term; flush the rest
    token_flush_size=1,
)


def first_recalled(messages):
    for message in messages:
        found = re.search(r"<recall>(.*?)</recall>", message.content or "", re.S)
        if found:
            said = re.findall(r"role='user'>(.*?)</message>", found.group(1))
            return said[0] if said else ""
    return "nothing recalled"


async def main():
    await memory.aput_messages([
        ChatMessage(role="user", content="Please deliver order A1001 to my office, not my home."),
        ChatMessage(role="assistant", content="Done, A1001 will go to your office."),
        ChatMessage(role="user", content="I prefer email over phone for updates."),
        ChatMessage(role="assistant", content="Noted, email it is."),
        ChatMessage(role="user", content="My favourite is the LMP-310 floor lamp."),
        ChatMessage(role="assistant", content="Good to know."),
    ])
    for question in ["Where should my order be delivered?", "How do you contact me?"]:
        messages = await memory.aget(input=question)
        print(question)
        print("  recalled:", first_recalled(messages))


asyncio.run(main())

Why each question recalled a different turn

  • The delivery question is closest in meaning to the office-address turn, so that turn comes back.
  • The contact question matches the email-preference turn instead, which proves the recall is by meaning, not by order.
  • The lamp turn stayed in the short buffer and was not needed, so neither answer mentions it.

Short-term buffer vs a vector memory block

MemoryHoldsRecall
Short-term bufferThe last few turnsAlways included, until they scroll out
VectorMemoryBlockAll flushed turns, embeddedOnly the turns close in meaning to the question

Where long-term memory earns its place

  • Remembering a customer's address or preferences between sessions.
  • A long chat where early details would otherwise scroll away.
  • Recalling a past decision without replaying the whole history to the model.
Watch out. A vector memory block needs a store that keeps the text, so the in-memory SimpleVectorStore is refused: it holds vectors but not the messages. Use a store like Chroma that keeps both.
Try it yourself
  • Add a fourth fact and a question that should recall it.
  • Raise similarity_top_k to 2 and see both facts a question can bring back.
  • Ask a question unrelated to every stored fact and read what comes back.

You understood something today that you didn't yesterday.