A persistent vector store with Chroma
Chroma is a vector store that runs inside your own process and writes its index to a folder. The policies are embedded once and read back on every later start instead of rebuilt.
Last updated: 27 Sep, 2026 · LangChain 1.4
A store that survives a restart
An in-memory store is rebuilt every time the program starts, which means embedding every chunk again. The RAG crash course's VectorStore class uses ChromaDB with a persistent directory instead, so whatever the store holds is saved on the hard disk. Its constructor takes a collection name, pdf_documents, and a persist directory in the data folder, ../data/vector_store, and creates that folder if it does not exist. It then opens chromadb.PersistentClient on the folder and calls get_or_create_collection with the collection name and some metadata, so a later run opens the same collection. Its add_documents method prepares what Chroma needs for every chunk (an id made with uuid, the chunk's metadata plus its index and content length, the text, and the embedding as a list) and adds them all to the collection. Created fresh, the store reports 0 existing documents.
Persistence brings one new mistake. The video's add_documents adds whatever it is given, so running the adding step again puts the chunks in a second time: the duplicate problem from Embeddings and a vector store. The code below adds the policies only when the collection is empty, so running it twice is safe.
In the shop, this replaces the top of search.py; the tool below it keeps its shape. desk.py is back on DeskModel, as the end of the hosted-model lesson asked, so the output below and the five desk tests stay exact and only the store changes.
Creating a Chroma store
from langchain_chroma import Chroma
store = Chroma(collection_name="policies", embedding_function=WordEmbeddings(),
persist_directory="./policy_store", # the folder it writes to
collection_metadata={"hnsw:space": "cosine"}) # score as cosine, higher is closer- written in Documents and splitting
- written in Embeddings and a vector store
- written in Several tool calls at once
- written in Three tools and a model that picks
- written in Three tools and a model that picks
- written in The desk's guardrails
- written in The desk on a hosted model
- written in The desk's guardrails
View the code here
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter
POLICIES = {
"refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
"\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
"shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
"\n\nExpress shipping arrives the next working day and costs 9 euros.",
"accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}
docs = [Document(page_content=text, metadata={"source": name}) for name, text in POLICIES.items()]
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs)
import re
import zlib
from langchain_core.embeddings import Embeddings
COMMON = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i",
"if", "is", "it", "my", "of", "on", "the", "to", "what", "with", "you", "your"}
class WordEmbeddings(Embeddings):
def embed_query(self, text):
vector = [0.0] * 256
for word in re.findall(r"[a-z]+", text.lower()):
if word not in COMMON:
vector[zlib.crc32(word.rstrip("s").encode()) % 256] += 1.0
return vector
def embed_documents(self, texts):
return [self.embed_query(text) for text in texts]
import re
from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult
class ShopModel(BaseChatModel):
tools: list = []
@property
def _llm_type(self):
return "shop"
def bind_tools(self, tools, **kwargs):
return self.model_copy(update={"tools": tools}) # a copy holding the tools
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
message = self.decide(messages) # the reply comes from decide
return ChatResult(generations=[ChatGeneration(message=message)])
def decide(self, messages):
results = [] # the tool results at the end
for m in reversed(messages):
if not isinstance(m, ToolMessage):
break
results.insert(0, m.text)
if results: # results are back: answer with them
return AIMessage(" ".join(results))
text = messages[-1].text
orders = re.findall(r"\b[A-Z]\d+\b", text)
tool = "refund_order" if "refund" in text.lower() else "lookup_order"
if orders and tool in [t.name for t in self.tools]: # one call per order id
calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
for o in orders]
return AIMessage("", tool_calls=calls)
if orders: # that tool is not bound
return AIMessage(f"I have no way to look up {orders[0]} yet.")
return AIMessage("Hello. Which order is this about?")
from dataclasses import dataclass
from langchain.tools import ToolRuntime, tool
ORDERS = {"A17": ("ravi", "shipped on 3 March"), "C40": ("mei", "waiting for stock")}
@dataclass
class Customer:
name: str
@tool
def lookup_order(order_id: str, runtime: ToolRuntime[Customer]) -> str:
"""Look up one of the customer's orders by its id, such as A17."""
owner, status = ORDERS.get(order_id, (None, None))
if owner != runtime.context.name:
return f"{order_id} is not one of your orders."
return f"{order_id} {status}."
@tool
def refund_order(order_id: str, runtime: ToolRuntime[Customer]) -> str:
"""Refund one of the customer's orders in full. This cannot be undone."""
owner, _ = ORDERS.get(order_id, (None, None))
if owner != runtime.context.name:
return f"{order_id} is not one of your orders, so it cannot be refunded."
return f"Refunded {order_id}."
import re
from langchain.messages import AIMessage
from shop_model import ShopModel
class DeskModel(ShopModel):
def decide(self, messages):
last = messages[-1]
if last.type == "tool" and last.text == "No policy covers this.":
return AIMessage("Our policies do not cover that. A person will reply.")
if last.type == "tool" or re.findall(r"\b[A-Z]\d+\b", last.text):
return super().decide(messages)
query = {"name": "search_policies", "args": {"query": last.text}, "id": "call_p"}
return AIMessage("", tool_calls=[query])
from langchain.agents.middleware import before_agent
from langchain.messages import AIMessage
@before_agent(can_jump_to=["end"])
def no_passwords(state, runtime):
if "password" in state["messages"][-1].text.lower():
answer = AIMessage("I cannot help with passwords. Please use the reset link.")
return {"messages": [answer], "jump_to": "end"}
from langchain.agents import create_agent
from langchain.agents.middleware import HumanInTheLoopMiddleware, ModelCallLimitMiddleware, PIIMiddleware
from langgraph.checkpoint.memory import InMemorySaver
from desk_model import DeskModel
from password_check import no_passwords
from search import search_policies
from desk_tools import Customer, lookup_order, refund_order
agent = create_agent(
DeskModel(),
system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned. If a tool says an order is not the customer's, say exactly that. Add nothing the tools did not say.",
tools=[lookup_order, refund_order, search_policies],
context_schema=Customer,
middleware=[
no_passwords,
PIIMiddleware("credit_card", strategy="mask"),
ModelCallLimitMiddleware(run_limit=6),
HumanInTheLoopMiddleware(interrupt_on={"refund_order": True}),
],
checkpointer=InMemorySaver(),
)
from langgraph.types import Command
from desk import Customer, agent
def say(who, text, thread):
config = {"configurable": {"thread_id": thread}}
result = agent.invoke({"messages": [{"role": "user", "content": text}]}, config,
context=Customer(who), version="v2")
if result.interrupts:
print(f"{who}: {text}\n paused for approval: {result.interrupts[0].value['action_requests'][0]['args']}")
result = agent.invoke(Command(resume={"decisions": [{"type": "approve"}]}), config,
context=Customer(who), version="v2")
text = "(approved)"
print(f"{who}: {text}\n desk: {result.value['messages'][-1].text}")
Swapping in the Chroma store
The top of search.py changes: the store becomes a Chroma store, and the chunks are added only when the folder is empty.
from langchain.tools import tool
from langchain_chroma import Chroma
from policies import chunks
from word_embeddings import WordEmbeddings
store = Chroma(collection_name="policies", embedding_function=WordEmbeddings(),
persist_directory="./policy_store",
collection_metadata={"hnsw:space": "cosine"})
if not store.get()["ids"]:
store.add_documents(chunks)persist_directory is the folder it writes to, and get() returns what is already in there, so the chunks are added once rather than on every start.
The tool, reading relevance scores
The tool below it is untouched, except that it now reads relevance scores from the new store.
@tool
def search_policies(query: str) -> str:
"""Search the shop's policies on refunds, shipping and accounts.
Pass the customer's question, word for word, as the query."""
found = [doc for doc, score in store.similarity_search_with_relevance_scores(query, k=2)
if score >= 0.3]
if not found:
return "No policy covers this."
return "\n".join(f"[{doc.metadata['source']}] {doc.page_content}" for doc in found)The score is not the same number
A store also decides what its scores mean. Chroma's default answers with a distance, where lower is closer, and a plain similarity_search_with_score shows it.
from langchain_chroma import Chroma
from policies import chunks
from word_embeddings import WordEmbeddings
plain = Chroma(collection_name="plain-store", embedding_function=WordEmbeddings())
plain.add_documents(chunks)
for doc, score in plain.similarity_search_with_score("do you sell gift cards", k=2):
print(round(score, 2), doc.metadata["source"])11.0 shipping.md 12.0 refunds.md
Those are distances: 11 and 12, for a question no policy covers. A score >= 0.3 cut would have let both through and the desk would have quoted the shipping policy at a customer asking about gift cards. Two changes keep the old behaviour: hnsw:space asks for cosine, and similarity_search_with_relevance_scores returns 0 to 1 with higher meaning closer.
The desk on the real store
from chat import say
from search import store
say("ravi", "How long does a refund take?", "ravi-9")
say("ravi", "Do you sell gift cards?", "ravi-9")
print(len(store.get()["ids"]), "chunks kept in ./policy_store")ravi: How long does a refund take? desk: [refunds.md] Refunds go back to the card you paid with. They take up to 5 working days to arrive. ravi: Do you sell gift cards? desk: Our policies do not cover that. A person will reply. 5 chunks kept in ./policy_store
The refund policy is found and the gift card question is still refused, so the cut survived the move. The five chunks now sit in ./policy_store, and the next start reads them instead of splitting the documents again.
What the real store changed
- Chroma's default score is a distance, where lower is closer, so 11 and 12 mean the gift-card question matched nothing well.
hnsw:spaceasks for cosine, andsimilarity_search_with_relevance_scoresreturns 0 to 1 with higher meaning closer, soscore >= 0.3means the same thing again.- The index survives the exit: five chunks stay in
./policy_store, and the next start reads them instead of splitting the documents again.
The embedder is the same swap
WordEmbeddings is two methods, and a hosted embedder is the same two. Google's Gemini embeddings are free with the Gemini key from the setup lesson, GOOGLE_API_KEY. Set the key the same way as GROQ_API_KEY: export GOOGLE_API_KEY=..., $env:GOOGLE_API_KEY = "..." in PowerShell, or a Colab secret. Then install the package and change the argument.
pip install "langchain-google-genai==4.4.0"from langchain_google_genai import GoogleGenerativeAIEmbeddings
store = Chroma(collection_name="policies",
embedding_function=GoogleGenerativeAIEmbeddings(model="gemini-embedding-001"),
persist_directory="./policy_store",
collection_metadata={"hnsw:space": "cosine"})It reads GOOGLE_API_KEY and sends each chunk to Google to embed. Two things change with it. Vectors from one embedder mean nothing to another, so delete ./policy_store and let the chunks be added again. And the 0.3 cut has to be checked again, because real vectors sit much closer together than word counts do.
Here are those scores from a real run, with the same chunks embedded by Gemini in a fresh store:
from langchain_chroma import Chroma
from langchain_google_genai import GoogleGenerativeAIEmbeddings
from policies import chunks
gemini = Chroma(collection_name="gemini-store",
embedding_function=GoogleGenerativeAIEmbeddings(model="gemini-embedding-001"), # uses your GOOGLE_API_KEY
collection_metadata={"hnsw:space": "cosine"})
gemini.add_documents(chunks)
for question in ["How long does a refund take?", "Do you sell gift cards?"]:
print(question)
for doc, score in gemini.similarity_search_with_relevance_scores(question, k=2):
print(" ", round(score, 2), doc.metadata["source"])How long does a refund take? 0.76 refunds.md 0.69 refunds.md Do you sell gift cards? 0.6 refunds.md 0.6 shipping.md
The refund question finds the refund policy at 0.76. The gift-card question, which no policy answers, still scores 0.6, twice the 0.3 cut, so with this embedder the desk would answer it instead of refusing.
Raise the cut in search_policies to score >= 0.7 and check both questions again:
from langchain_chroma import Chroma
from langchain_google_genai import GoogleGenerativeAIEmbeddings
from policies import chunks
gemini = Chroma(collection_name="gemini-cut",
embedding_function=GoogleGenerativeAIEmbeddings(model="gemini-embedding-001"), # uses your GOOGLE_API_KEY
collection_metadata={"hnsw:space": "cosine"})
gemini.add_documents(chunks)
for question in ["How long does a refund take?", "Do you sell gift cards?"]:
found = [doc for doc, score in gemini.similarity_search_with_relevance_scores(question, k=2) if score >= 0.7]
print(question, "->", [doc.metadata["source"] for doc in found] or "No policy covers this.")How long does a refund take? -> ['refunds.md'] Do you sell gift cards? -> No policy covers this.
At 0.7 the refund question still finds refunds.md, which scored 0.76, and the gift-card question, at 0.6, is refused again. The cut is tied to the embedder: change the model and it has to be checked again. These were checks: leave search.py on WordEmbeddings() with the 0.3 cut, so the desk and its tests stay key-free.
When policies live on disk
- Keeping an embedded policy set on disk so a restart does not rebuild it.
- Moving from the in-memory store to a database that other processes can read.
score >= 0.3 cut written for similarities lets everything through. Ask for cosine and read relevance scores, then re-check the cut.Related
- Previous: The desk on a hosted model
- Next: Saving threads with a SqliteSaver checkpointer
- Reference: Retrieval and vector stores
- Delete
./policy_storeand run it again. - Take out
collection_metadataand see what the refusal does. - Swap
WordEmbeddings()forGoogleGenerativeAIEmbeddings(model="gemini-embedding-001"), delete./policy_store, and print the relevance scores for the gift-card question.
This is what real progress feels like.