LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

A persistent vector store with Chroma

Chroma is a vector store that runs inside your own process and writes its index to a folder. The policies are embedded once and read back on every later start instead of rebuilt.

Last updated: 27 Sep, 2026 · LangChain 1.4

A persistent Chroma vector store · from the Complete RAG Crash Course With LangChain · 69:30 to 74:22

A store that survives a restart

An in-memory store is rebuilt every time the program starts, which means embedding every chunk again. The RAG crash course's VectorStore class uses ChromaDB with a persistent directory instead, so whatever the store holds is saved on the hard disk. Its constructor takes a collection name, pdf_documents, and a persist directory in the data folder, ../data/vector_store, and creates that folder if it does not exist. It then opens chromadb.PersistentClient on the folder and calls get_or_create_collection with the collection name and some metadata, so a later run opens the same collection. Its add_documents method prepares what Chroma needs for every chunk (an id made with uuid, the chunk's metadata plus its index and content length, the text, and the embedding as a list) and adds them all to the collection. Created fresh, the store reports 0 existing documents.

Persistence brings one new mistake. The video's add_documents adds whatever it is given, so running the adding step again puts the chunks in a second time: the duplicate problem from Embeddings and a vector store. The code below adds the policies only when the collection is empty, so running it twice is safe.

In the shop, this replaces the top of search.py; the tool below it keeps its shape. desk.py is back on DeskModel, as the end of the hosted-model lesson asked, so the output below and the five desk tests stay exact and only the store changes.

Creating a Chroma store

python
from langchain_chroma import Chroma

store = Chroma(collection_name="policies", embedding_function=WordEmbeddings(),
               persist_directory="./policy_store",           # the folder it writes to
               collection_metadata={"hnsw:space": "cosine"})  # score as cosine, higher is closer
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
policies.py
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter

POLICIES = {
    "refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
                  "\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
    "shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
                   "\n\nExpress shipping arrives the next working day and costs 9 euros.",
    "accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}

docs = [Document(page_content=text, metadata={"source": name}) for name, text in POLICIES.items()]
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs)
word_embeddings.py
import re
import zlib

from langchain_core.embeddings import Embeddings

COMMON = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i",
          "if", "is", "it", "my", "of", "on", "the", "to", "what", "with", "you", "your"}


class WordEmbeddings(Embeddings):
    def embed_query(self, text):
        vector = [0.0] * 256
        for word in re.findall(r"[a-z]+", text.lower()):
            if word not in COMMON:
                vector[zlib.crc32(word.rstrip("s").encode()) % 256] += 1.0
        return vector

    def embed_documents(self, texts):
        return [self.embed_query(text) for text in texts]
shop_model.py
import re

from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult


class ShopModel(BaseChatModel):
    tools: list = []

    @property
    def _llm_type(self):
        return "shop"

    def bind_tools(self, tools, **kwargs):
        return self.model_copy(update={"tools": tools})   # a copy holding the tools

    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        message = self.decide(messages)                   # the reply comes from decide
        return ChatResult(generations=[ChatGeneration(message=message)])

    def decide(self, messages):
        results = []                                # the tool results at the end
        for m in reversed(messages):
            if not isinstance(m, ToolMessage):
                break
            results.insert(0, m.text)
        if results:                                 # results are back: answer with them
            return AIMessage(" ".join(results))
        text = messages[-1].text
        orders = re.findall(r"\b[A-Z]\d+\b", text)
        tool = "refund_order" if "refund" in text.lower() else "lookup_order"
        if orders and tool in [t.name for t in self.tools]:   # one call per order id
            calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
                     for o in orders]
            return AIMessage("", tool_calls=calls)
        if orders:                                  # that tool is not bound
            return AIMessage(f"I have no way to look up {orders[0]} yet.")
        return AIMessage("Hello. Which order is this about?")
desk_tools.py
from dataclasses import dataclass

from langchain.tools import ToolRuntime, tool

ORDERS = {"A17": ("ravi", "shipped on 3 March"), "C40": ("mei", "waiting for stock")}


@dataclass
class Customer:
    name: str


@tool
def lookup_order(order_id: str, runtime: ToolRuntime[Customer]) -> str:
    """Look up one of the customer's orders by its id, such as A17."""
    owner, status = ORDERS.get(order_id, (None, None))
    if owner != runtime.context.name:
        return f"{order_id} is not one of your orders."
    return f"{order_id} {status}."


@tool
def refund_order(order_id: str, runtime: ToolRuntime[Customer]) -> str:
    """Refund one of the customer's orders in full. This cannot be undone."""
    owner, _ = ORDERS.get(order_id, (None, None))
    if owner != runtime.context.name:
        return f"{order_id} is not one of your orders, so it cannot be refunded."
    return f"Refunded {order_id}."
desk_model.py
import re

from langchain.messages import AIMessage
from shop_model import ShopModel


class DeskModel(ShopModel):
    def decide(self, messages):
        last = messages[-1]
        if last.type == "tool" and last.text == "No policy covers this.":
            return AIMessage("Our policies do not cover that. A person will reply.")
        if last.type == "tool" or re.findall(r"\b[A-Z]\d+\b", last.text):
            return super().decide(messages)
        query = {"name": "search_policies", "args": {"query": last.text}, "id": "call_p"}
        return AIMessage("", tool_calls=[query])
password_check.py
from langchain.agents.middleware import before_agent
from langchain.messages import AIMessage


@before_agent(can_jump_to=["end"])
def no_passwords(state, runtime):
    if "password" in state["messages"][-1].text.lower():
        answer = AIMessage("I cannot help with passwords. Please use the reset link.")
        return {"messages": [answer], "jump_to": "end"}
desk.py
from langchain.agents import create_agent
from langchain.agents.middleware import HumanInTheLoopMiddleware, ModelCallLimitMiddleware, PIIMiddleware
from langgraph.checkpoint.memory import InMemorySaver
from desk_model import DeskModel
from password_check import no_passwords
from search import search_policies
from desk_tools import Customer, lookup_order, refund_order

agent = create_agent(
    DeskModel(),
    system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned. If a tool says an order is not the customer's, say exactly that. Add nothing the tools did not say.",
    tools=[lookup_order, refund_order, search_policies],
    context_schema=Customer,
    middleware=[
        no_passwords,
        PIIMiddleware("credit_card", strategy="mask"),
        ModelCallLimitMiddleware(run_limit=6),
        HumanInTheLoopMiddleware(interrupt_on={"refund_order": True}),
    ],
    checkpointer=InMemorySaver(),
)
chat.py
from langgraph.types import Command
from desk import Customer, agent


def say(who, text, thread):
    config = {"configurable": {"thread_id": thread}}
    result = agent.invoke({"messages": [{"role": "user", "content": text}]}, config,
                          context=Customer(who), version="v2")
    if result.interrupts:
        print(f"{who}: {text}\n  paused for approval: {result.interrupts[0].value['action_requests'][0]['args']}")
        result = agent.invoke(Command(resume={"decisions": [{"type": "approve"}]}), config,
                              context=Customer(who), version="v2")
        text = "(approved)"
    print(f"{who}: {text}\n  desk: {result.value['messages'][-1].text}")

Swapping in the Chroma store

The top of search.py changes: the store becomes a Chroma store, and the chunks are added only when the folder is empty.

python
from langchain.tools import tool
from langchain_chroma import Chroma
from policies import chunks
from word_embeddings import WordEmbeddings

store = Chroma(collection_name="policies", embedding_function=WordEmbeddings(),
               persist_directory="./policy_store",
               collection_metadata={"hnsw:space": "cosine"})
if not store.get()["ids"]:
    store.add_documents(chunks)

persist_directory is the folder it writes to, and get() returns what is already in there, so the chunks are added once rather than on every start.

The tool, reading relevance scores

The tool below it is untouched, except that it now reads relevance scores from the new store.

python
@tool
def search_policies(query: str) -> str:
    """Search the shop's policies on refunds, shipping and accounts.
    Pass the customer's question, word for word, as the query."""
    found = [doc for doc, score in store.similarity_search_with_relevance_scores(query, k=2)
             if score >= 0.3]
    if not found:
        return "No policy covers this."
    return "\n".join(f"[{doc.metadata['source']}] {doc.page_content}" for doc in found)

The score is not the same number

A store also decides what its scores mean. Chroma's default answers with a distance, where lower is closer, and a plain similarity_search_with_score shows it.

Example
from langchain_chroma import Chroma
from policies import chunks
from word_embeddings import WordEmbeddings

plain = Chroma(collection_name="plain-store", embedding_function=WordEmbeddings())
plain.add_documents(chunks)
for doc, score in plain.similarity_search_with_score("do you sell gift cards", k=2):
    print(round(score, 2), doc.metadata["source"])

Those are distances: 11 and 12, for a question no policy covers. A score >= 0.3 cut would have let both through and the desk would have quoted the shipping policy at a customer asking about gift cards. Two changes keep the old behaviour: hnsw:space asks for cosine, and similarity_search_with_relevance_scores returns 0 to 1 with higher meaning closer.

The desk on the real store

Example
from chat import say
from search import store

say("ravi", "How long does a refund take?", "ravi-9")
say("ravi", "Do you sell gift cards?", "ravi-9")
print(len(store.get()["ids"]), "chunks kept in ./policy_store")

The refund policy is found and the gift card question is still refused, so the cut survived the move. The five chunks now sit in ./policy_store, and the next start reads them instead of splitting the documents again.

What the real store changed

  • Chroma's default score is a distance, where lower is closer, so 11 and 12 mean the gift-card question matched nothing well.
  • hnsw:space asks for cosine, and similarity_search_with_relevance_scores returns 0 to 1 with higher meaning closer, so score >= 0.3 means the same thing again.
  • The index survives the exit: five chunks stay in ./policy_store, and the next start reads them instead of splitting the documents again.

The embedder is the same swap

WordEmbeddings is two methods, and a hosted embedder is the same two. Google's Gemini embeddings are free with the Gemini key from the setup lesson, GOOGLE_API_KEY. Set the key the same way as GROQ_API_KEY: export GOOGLE_API_KEY=..., $env:GOOGLE_API_KEY = "..." in PowerShell, or a Colab secret. Then install the package and change the argument.

pip install "langchain-google-genai==4.4.0"
python
from langchain_google_genai import GoogleGenerativeAIEmbeddings

store = Chroma(collection_name="policies",
               embedding_function=GoogleGenerativeAIEmbeddings(model="gemini-embedding-001"),
               persist_directory="./policy_store",
               collection_metadata={"hnsw:space": "cosine"})

It reads GOOGLE_API_KEY and sends each chunk to Google to embed. Two things change with it. Vectors from one embedder mean nothing to another, so delete ./policy_store and let the chunks be added again. And the 0.3 cut has to be checked again, because real vectors sit much closer together than word counts do.

Here are those scores from a real run, with the same chunks embedded by Gemini in a fresh store:

ExampleAPI key
from langchain_chroma import Chroma
from langchain_google_genai import GoogleGenerativeAIEmbeddings
from policies import chunks

gemini = Chroma(collection_name="gemini-store",
                embedding_function=GoogleGenerativeAIEmbeddings(model="gemini-embedding-001"),  # uses your GOOGLE_API_KEY
                collection_metadata={"hnsw:space": "cosine"})
gemini.add_documents(chunks)

for question in ["How long does a refund take?", "Do you sell gift cards?"]:
    print(question)
    for doc, score in gemini.similarity_search_with_relevance_scores(question, k=2):
        print(" ", round(score, 2), doc.metadata["source"])

The refund question finds the refund policy at 0.76. The gift-card question, which no policy answers, still scores 0.6, twice the 0.3 cut, so with this embedder the desk would answer it instead of refusing.

Raise the cut in search_policies to score >= 0.7 and check both questions again:

ExampleAPI key
from langchain_chroma import Chroma
from langchain_google_genai import GoogleGenerativeAIEmbeddings
from policies import chunks

gemini = Chroma(collection_name="gemini-cut",
                embedding_function=GoogleGenerativeAIEmbeddings(model="gemini-embedding-001"),  # uses your GOOGLE_API_KEY
                collection_metadata={"hnsw:space": "cosine"})
gemini.add_documents(chunks)

for question in ["How long does a refund take?", "Do you sell gift cards?"]:
    found = [doc for doc, score in gemini.similarity_search_with_relevance_scores(question, k=2) if score >= 0.7]
    print(question, "->", [doc.metadata["source"] for doc in found] or "No policy covers this.")

At 0.7 the refund question still finds refunds.md, which scored 0.76, and the gift-card question, at 0.6, is refused again. The cut is tied to the embedder: change the model and it has to be checked again. These were checks: leave search.py on WordEmbeddings() with the 0.3 cut, so the desk and its tests stay key-free.

When policies live on disk

  • Keeping an embedded policy set on disk so a restart does not rebuild it.
  • Moving from the in-memory store to a database that other processes can read.
Watch out. A new store can change what its score means. Chroma's default is a distance, where lower is closer, so a score >= 0.3 cut written for similarities lets everything through. Ask for cosine and read relevance scores, then re-check the cut.
Try it yourself
  • Delete ./policy_store and run it again.
  • Take out collection_metadata and see what the refusal does.
  • Swap WordEmbeddings() for GoogleGenerativeAIEmbeddings(model="gemini-embedding-001"), delete ./policy_store, and print the relevance scores for the gift-card question.

This is what real progress feels like.