Vector store memory
Vector store memory is a long-term memory technique that saves every conversation turn as an embedding in a vector database and, on each new turn, retrieves the saved turns whose meaning is closest to the new message.
Last updated: 09 Oct, 2026 · ChromaDB 1.5
Every technique up to Token buffer memory keeps the conversation in a Python list. The list lives in RAM, it is gone when the session ends, and it is cut by position: the oldest turns fall out first. Vector store memory keeps the turns in a database instead, so they are still there next week, and it looks them up by meaning, not by age.
Nothing in the model is retrained or updated by this. The model's weights stay as they were. What changes is the rows in the store and the text that is placed in the next prompt.
This part of the video starts at 4:28:19. It opens the long-term half of the memory module: a vector store converts every conversation turn into a vector.
Turning a conversation into vectors
The video starts from the plain idea of a vector. On a two-dimensional plane, mark a point with an x value and a y value and draw a line from the origin to it: that line is a two-dimensional vector, and two numbers describe it. An embedding is the same thing with many more numbers. An embedding model reads a piece of text and returns a long list of numbers that stands for its meaning. Texts that mean similar things get lists that point in similar directions.
How many numbers there are is a property of the embedding model. More numbers cost more storage and more search time, and they do not make retrieval better by themselves.
This part of the video starts at 4:30:48. The vectors are kept in a persistent database, and at the start of every new turn the most semantically relevant past messages are retrieved.
The video walks one question through the flow. The user asks "Should I rebalance my portfolio?". That query is embedded into a vector. The store is searched for the stored vectors closest to it, with an approximate nearest neighbour (ANN) lookup, which finds the closest matches fast without comparing against every row. With k = 3 the top three messages come back. They are placed in the prompt next to the current conversation, the LLM is called, and the answer is personal because it has read the three memories.
The three memories in the picture are the ones on the notebook's flow chart. The write path runs at the end of a turn, the read path at the start of the next one.
Measuring meaning with cosine similarity
The notebook on screen, 6_vector_store_memory.ipynb, explains embeddings with three sentences: "My salary is ₹1,20,000", "I earn ₹1.2 lakh per month" and "I enjoy playing cricket". The first two say the same thing in different words, the third is about something else. A search for "what is my income" should land near the two salary sentences.
Closeness is measured with cosine similarity: the cosine of the angle between two vectors. It is 1 when they point the same way and 0 when they are at a right angle. A vector database reports the same thing the other way round, as a distance, where smaller is closer.
The video's notebook embeds with OpenAI's text-embedding-3-small, which returns 1536 numbers per text. The code here uses Gemini's gemini-embedding-2, which runs on the Gemini key and returns 3072 numbers. It takes one text per call.
Embedding one text
import os
from google import genai
gem = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
reply = gem.models.embed_content(model="gemini-embedding-2", contents="what is my income")
vector = reply.embeddings[0].values # a list of floatsThe notebook's three sentences, measured
import os
import numpy as np
import matplotlib.pyplot as plt
from google import genai
gem = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
def embed(text):
# one text per call; the reply holds one vector
reply = gem.models.embed_content(model="gemini-embedding-2", contents=text)
return np.array(reply.embeddings[0].values)
sentences = ["My salary is ₹1,20,000", "I earn ₹1.2 lakh per month", "I enjoy playing cricket"]
query = "what is my income"
q = embed(query)
vectors = [embed(s) for s in sentences]
print("numbers per vector:", len(q), "| length of the vector:", round(float(np.linalg.norm(q)), 4))
print("first 5 numbers of the first sentence:", np.round(vectors[0][:5], 4))
sims = [float(q @ v) for v in vectors] # the vectors have length 1, so the dot product is the cosine
for s, sim in zip(sentences, sims):
print(f"similarity {sim:.4f} distance {1 - sim:.4f} {s}")
print("the two salary sentences to each other:", round(float(vectors[0] @ vectors[1]), 4))
plt.figure(figsize=(7, 2.6))
plt.barh(sentences[::-1], sims[::-1], color=["#8a8a8a", "#3a6fd8", "#3a6fd8"])
plt.xlim(0, 1)
plt.xlabel('cosine similarity to "what is my income"')
plt.title("Similarity measured with gemini-embedding-2")
plt.show()numbers per vector: 3072 | length of the vector: 1.0 first 5 numbers of the first sentence: [-0.0122 -0.0116 0.0165 0.0066 0.0045] similarity 0.7437 distance 0.2563 My salary is ₹1,20,000 similarity 0.6825 distance 0.3175 I earn ₹1.2 lakh per month similarity 0.5394 distance 0.4606 I enjoy playing cricket the two salary sentences to each other: 0.8672
What the measured similarities show
- Each text became 3072 numbers, and the vector has length 1.0. For vectors of length 1 the dot product
q @ vis already the cosine similarity. - The numbers mean nothing one by one. The first five of the salary sentence are -0.0122, -0.0116, 0.0165, 0.0066 and 0.0045. Meaning shows only when two vectors are compared.
- Both salary sentences are close to the query: 0.7437 and 0.6825, although neither contains the word "income". To each other they score 0.8672.
- The cricket sentence still scores 0.5394. An unrelated sentence does not land near 0 with this model. What separates related from unrelated here is a gap of about 0.14, so a cut-off has to be chosen from measured values.
- Distance is 1 minus similarity: 0.2563, 0.3175 and 0.4606. Smaller means closer.
This part of the video starts at 4:35:31. ChromaDB stores each embedding along with the original text and its metadata.
What ChromaDB keeps for each memory
ChromaDB is an open-source vector database, and the notebook is built on it. "Vector store" and "vector database" name the same kind of system here. The video lists the four things a row holds:
- An id. A unique identifier for the row.
- The chunk. The original text. The clip explains chunks with a 12-page PDF cut up by a text splitter; in conversation memory the chunk is one user message.
- The embedding. The vector of that text.
- The metadata. Extra fields such as the user id, the session and a timestamp, used to filter a search.
Reading distances from ChromaDB
Two facts about ChromaDB decide every number a memory lookup returns. It always reports a distance, never a similarity. And the distance it uses is set when the collection is created. Three tiny vectors make both visible without any API key.
import chromadb
import numpy as np
vectors = {"a": [1.0, 0.0, 0.0], "b": [0.6, 0.8, 0.0], "c": [0.0, 0.0, 2.0]}
query = [2.0, 0.0, 0.0]
db = chromadb.EphemeralClient() # in memory, nothing is written to disk
plain = db.get_or_create_collection("metric_default")
cosine = db.get_or_create_collection("metric_cosine", configuration={"hnsw": {"space": "cosine"}})
for col in (plain, cosine):
col.add(ids=list(vectors), embeddings=list(vectors.values()))
hit = col.query(query_embeddings=[query], n_results=3)
print(f"space={col.configuration['hnsw']['space']:6}", hit["ids"][0],
[round(d, 4) for d in hit["distances"][0]])
q = np.array(query)
for name, v in vectors.items():
v = np.array(v)
sim = float(q @ v / (np.linalg.norm(q) * np.linalg.norm(v)))
print(f"{name}: cosine similarity {sim:.1f} | 1 - similarity {1 - sim:.1f} | squared L2 {((q - v) ** 2).sum():.1f}")space=l2 ['a', 'b', 'c'] [1.0, 2.6, 8.0] space=cosine ['a', 'b', 'c'] [0.0, 0.4, 1.0] a: cosine similarity 1.0 | 1 - similarity 0.0 | squared L2 1.0 b: cosine similarity 0.6 | 1 - similarity 0.4 | squared L2 2.6 c: cosine similarity 0.0 | 1 - similarity 1.0 | squared L2 8.0
What the two collections returned
- The collection created with no setting reports
space=l2. Its distances 1.0, 2.6 and 8.0 are squared straight-line (L2) distances, which the NumPy lines confirm. Cosine has to be asked for. - The collection created with
"space": "cosine"returns 0.0, 0.4 and 1.0: 1 minus the cosine similarities 1.0, 0.6 and 0.0. - Vector a is at cosine distance 0.0 although it is half as long as the query. Cosine looks only at direction. L2 also counts length, and puts the same pair 1.0 apart.
- Both numbers are distances, on different scales. A cut-off of 0.7 keeps a and b in the cosine collection and nothing in the L2 one.
Building vector store memory on ChromaDB
The notebook's VectorStoreMemory class does two jobs: it writes every user message into the store, and before each model call it reads back the closest ones. The pieces below are a small version of it. The notebook keeps its store in a folder on disk with chromadb.PersistentClient(path=...), which is what makes the memory outlive a restart. The code here uses chromadb.EphemeralClient(), an in-memory store, so that running it leaves no files behind; swapping the one line back gives the persistent store.
The clients and the embed helper
The chat model is called with the OpenAI SDK pointed at Groq, the same SDK the notebook uses. The notebook's models are gpt-4o and OpenAI embeddings; here they are openai/gpt-oss-120b and gemini-embedding-2.
import os
import chromadb
from google import genai
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
MODEL = "openai/gpt-oss-120b"
gem = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
def embed(text):
# one text per call; the reply holds one vector of 3072 numbers
return gem.models.embed_content(model="gemini-embedding-2", contents=text).embeddings[0].valuesA collection that measures cosine distance
store = chromadb.EphemeralClient().get_or_create_collection(
"fincoach_memory", configuration={"hnsw": {"space": "cosine"}})The write path: one message in
Each message is stored with its vector, its text and the id of the user it belongs to. The vector is passed in embeddings=; the store never embeds anything itself.
def remember(user_id, n, text):
store.add(ids=[f"{user_id}-{n}"], embeddings=[embed(text)], documents=[text],
metadatas=[{"user_id": user_id, "session": "session_1"}])The read path: the closest messages of one user
where={"user_id": user_id} limits the search to one user's rows. One collection holds every user, and this filter is what keeps them apart.
def recall(user_id, query_vector, k):
hit = store.query(query_embeddings=[query_vector], n_results=k,
where={"user_id": user_id}, include=["documents", "distances"])
return list(zip(hit["distances"][0], hit["documents"][0]))Recalling the first session in a later one
The stored messages are the five user turns of the notebook's first session with Chiru, plus the cricket sentence, plus one message from a second user, Priya, taken from the notebook's isolation test. The question is a turn of the notebook's second session, asked when the conversation buffer is empty and only the store remembers. Put the four pieces above in one file and add these lines.
session_1 = ["Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000.",
"My monthly expenses are ₹60,000 for rent, food, and transport.",
"I have an FD of ₹50,000 maturing in 3 months. I'm conservative with investments.",
"My long-term goal is to retire comfortably by age 55. I'm currently 32.",
"Based on my profile, what should be my investment priority right now?",
"I enjoy playing cricket"]
for n, text in enumerate(session_1):
remember("chiru_001", n, text)
remember("priya_002", 0, "I am Priya. My salary is ₹80,000 per month.")
question = "How much monthly savings do I have available for new investments?"
qv = embed(question) # embedded once, reused for every search below
print("All memories of chiru_001, closest first:")
for dist, text in recall("chiru_001", qv, 10):
print(f" {dist:.4f} {text}")
print("All memories of priya_002:", [text for dist, text in recall("priya_002", qv, 10)])
MAX_DISTANCE = 0.40 # chosen from the distances printed above
kept = [text for dist, text in recall("chiru_001", qv, 3) if dist < MAX_DISTANCE]
print("\nInjected into the prompt (k=3):")
for text in kept:
print(" -", text)
reply = client.chat.completions.create(model=MODEL, temperature=0, max_tokens=600, messages=[
{"role": "system", "content": "You are FinCoach, a personal finance assistant for users in India. "
"Answer in at most two sentences. Use only the memories you are given. "
"If a number you need is missing, say which one."},
{"role": "system", "content": "Memories from past sessions:\n" + "\n".join("- " + t for t in kept)},
{"role": "user", "content": question}])
print("\nFinCoach:", reply.choices[0].message.content)All memories of chiru_001, closest first: 0.2374 Based on my profile, what should be my investment priority right now? 0.2877 Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000. 0.3125 I have an FD of ₹50,000 maturing in 3 months. I'm conservative with investments. 0.3137 My long-term goal is to retire comfortably by age 55. I'm currently 32. 0.3510 My monthly expenses are ₹60,000 for rent, food, and transport. 0.4763 I enjoy playing cricket All memories of priya_002: ['I am Priya. My salary is ₹80,000 per month.'] Injected into the prompt (k=3): - Based on my profile, what should be my investment priority right now? - Hi! I'm Chiru. My monthly take-home salary is ₹1,20,000. - I have an FD of ₹50,000 maturing in 3 months. I'm conservative with investments. FinCoach: I’d need to know your regular monthly expenses (rent, bills, food, transport, etc.) to subtract from your ₹1,20,000 take‑home salary and determine the amount you can save each month for new investments.
What the store returned and what the model did with it
- Priya's message never appears in Chiru's list, and none of Chiru's six appears in hers. All seven rows are in one collection; the
wherefilter is the only thing that keeps the two users apart. - The five messages about money sit between 0.2374 and 0.3510, and the cricket sentence is at 0.4763. The cut-off of 0.40 is chosen from that gap. The notebook keeps memories with
distance < 0.7, a value set for OpenAI's embeddings; with these vectors it would let the cricket sentence through. A cut-off belongs to one embedding model and is measured again when the model changes. - The closest memory, at 0.2374, is a question from session 1. It holds no fact, yet it takes one of the three slots, because one question about investing resembles another.
- The expenses message is fifth, at 0.3510, and k = 3 leaves it out. The model received the salary but not the expenses.
- The reply shows the gap. The model did not invent a number. It asked for the monthly expenses, which were in the store all along. The saved output of the notebook's run has the same kind of miss on this question: there the salary message was not retrieved, and the assistant asked for the monthly income.
Stale facts in a vector store
Vector store memory only ever adds rows. When a fact changes, the old row and the new row are both in the store, and both are close to the same question. The notebook shows this with two messages about the user's job, stored a year and a half apart.
import os
import numpy as np
from google import genai
gem = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
def embed(text):
return np.array(gem.models.embed_content(model="gemini-embedding-2", contents=text).embeddings[0].values)
memories = [("2024-12-01", "I work at Infosys as a software engineer and earn ₹90,000 per month."),
("2026-06-12", "I recently changed jobs. I now work at TCS and my salary is ₹1,20,000 per month.")]
q = embed("Where do I work and what is my current salary?")
for date, text in memories:
print(f"stored {date} distance {1 - float(q @ embed(text)):.4f} {text}")stored 2024-12-01 distance 0.2312 I work at Infosys as a software engineer and earn ₹90,000 per month. stored 2026-06-12 distance 0.2254 I recently changed jobs. I now work at TCS and my salary is ₹1,20,000 per month.
Both messages come back, at distances 0.2312 and 0.2254, a difference of 0.0058. Each has a date, but a similarity search does not look at dates, so the store cannot tell which fact is current. In this run the newer message is ahead by a hair; nothing guarantees that. A model given both would read two employers and two salaries, ₹90,000 and ₹1,20,000, and have to guess. Keeping one current value per fact is the job of Entity memory.
Vector store memory vs buffer memory
| Buffer techniques | Vector store memory | |
|---|---|---|
| Where memory lives | A Python list in RAM | A vector database |
| After the session ends | Gone | Still stored |
| How context is chosen | By position: the most recent turns | By meaning: the closest stored turns |
| Users | One user, one buffer | Many users in one store, kept apart by a metadata filter |
| Extra cost per turn | None | One embedding call to write, one to search |
| A changed fact | The old turn falls out of the buffer in time | Old and new are both retrieved |
Where you use vector store memory
- Returning users. An assistant that should recall what a user said in an earlier session without being told again.
- Long histories. A fact from turn 3 is found at turn 47 because the lookup is by meaning, not by position.
- Next to a short-term buffer. The video's suggestion for a hybrid: when a sliding window or a token buffer evicts a message, embed it and store it, so the buffer gives recency and the store gives recall.
add is called with documents= but no embeddings=, or query with query_texts=, ChromaDB downloads a small local model and embeds the text itself, with a different model from yours. Pass embeddings= on every add and query_embeddings= on every query.Related
- Previous: Token buffer memory
- Next: Entity memory
- Reference: ChromaDB: configuring a collection, Gemini API embeddings
- In the three-vector example, change
queryto[0.0, 0.0, 0.1]: the cosine collection ranks c first at distance 0.0, and the default collection ranks it last at 3.61. - In the session example, set
MAX_DISTANCE = 0.30: only the first two memories are injected, since the FD message sits at 0.3125. - Change the last
recall("chiru_001", qv, 3)torecall("chiru_001", qv, 5): the expenses message at 0.3510 is injected too, so the model has both numbers and can answer that ₹60,000 is left each month.
You understood something today that you didn't yesterday.