Embeddings and a vector store
An embedding is a list of numbers standing for a text, and a vector store is an object that keeps chunks with their embeddings and returns the ones closest to a question.
Last updated: 27 Sep, 2026 · LangChain 1.4
From chunks to vectors
An embedding model turns a piece of text into a vector, a list of numbers, so that texts with similar meaning get vectors that point in similar directions. The RAG crash course uses sentence-transformers with all-MiniLM-L6-v2 from Hugging Face, which converts every text into 384 numbers, and wraps it in its own EmbeddingManager class. The constructor takes the model name, _load_model loads the SentenceTransformer and reads its size with get_sentence_embedding_dimension, and generate_embeddings calls model.encode on a list of texts, with a progress bar, and returns a numpy array. Creating embedding_manager = EmbeddingManager() loads the model and reports a dimension of 384.
A vector store keeps those vectors next to their chunks and, given a new vector, finds the closest ones. The video's VectorStore is its own wrapper around a ChromaDB collection, built in A persistent vector store with Chroma; it is not one of LangChain's vector stores, and the shop version below uses LangChain's InMemoryVectorStore. Created fresh, the store reports 0 documents in its collection. The video then takes the page_content of every chunk into texts, embeds them with embedding_manager.generate_embeddings, and passes the chunks and their embeddings to vectorstore.add_documents:
### Convert the text to embeddings
texts=[doc.page_content for doc in chunks]
## Generate the Embeddings
embeddings=embedding_manager.generate_embeddings(texts)
##store int he vector dtaabase
vectorstore.add_documents(chunks,embeddings)Generating embeddings for 359 texts... Batches: 100%|██████████| 12/12 [00:06<00:00, 1.78it/s] Generated embeddings with shape: (359, 384) Adding 359 documents to vector store... Successfully added 359 documents to vector store Total documents in collection: 1077
Shown as the video's notebook printed it, not run here. It relies on the video's own EmbeddingManager and VectorStore classes and on sentence-transformers, a large download; the shop version below builds the same two pieces small enough to run.
all-MiniLM-L6-v2 gives every chunk a vector of 384 numbers, so the 359 chunks became a (359, 384) array, embedded in 12 batches. In the video the collection then held 359 documents, saved in the vector_store folder on disk. The last line above says 1,077, three times 359: this output was saved after the cell had been run three times, and each run added every chunk again. A persistent store keeps what you add, so adding the same chunks again duplicates them, and a search then returns the same passage more than once. A persistent vector store with Chroma checks for an existing collection before adding for this reason.
The video's model is the large-download kind: it pulls in PyTorch. The shop version builds both pieces small enough to see every number. Its embedding only counts words, so it runs on your machine with no key and no download. The last part of the course swaps in a hosted one.
InMemoryVectorStore and a retriever
store = InMemoryVectorStore(embedding_model) # keeps chunks and their vectors
store.add_documents(chunks) # embed each chunk and store it
store.similarity_search_with_score(query, k=1) # -> [(Document, score), ...]
retriever = store.as_retriever(search_kwargs={"k": 1}) # same invoke/batch as a modelThe common words
An Embeddings class needs two methods: one for a question, one for a list of documents. The code in this section goes in one file, word_embeddings.py. Start with the imports and the words to skip.
import re
import zlib
from langchain_core.embeddings import Embeddings
# words that appear everywhere and would make every text look alike
COMMON = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i",
"if", "is", "it", "my", "of", "on", "the", "to", "what", "with", "you", "your"}A word-counting embedding
Each word is hashed into one of 256 slots and counted. Hashing turns a word into a fixed number (zlib.crc32 does it here), and % 256 folds that number into a slot from 0 to 255, so the same word always lands in the same slot. Stripping a final s makes "refund" and "refunds" count as one word.
class WordEmbeddings(Embeddings):
def embed_query(self, text):
vector = [0.0] * 256
for word in re.findall(r"[a-z]+", text.lower()):
if word not in COMMON: # skip common words
vector[zlib.crc32(word.rstrip("s").encode()) % 256] += 1.0 # hash into a slot
return vector
def embed_documents(self, texts): # one vector per document
return [self.embed_query(text) for text in texts]Each word hashed to a slot
Embed a few words and print which slots fill.
model = WordEmbeddings()
for text in ["refund", "Refunds", "shipping"]:
vector = model.embed_query(text)
print(f"{text:<9} slots {[i for i, v in enumerate(vector) if v]}")refund slots [88] Refunds slots [88] shipping slots [36]
"refund" and "Refunds" land in the same slot, and "shipping" lands somewhere else. Texts about the same thing fill the same slots, which is what a search has to work with.
- written in Documents and splitting
View the code here
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter
POLICIES = {
"refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
"\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
"shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
"\n\nExpress shipping arrives the next working day and costs 9 euros.",
"accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}
docs = [Document(page_content=text, metadata={"source": name}) for name, text in POLICIES.items()]
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs)
A vector store
Build the store from the chunks and the embedding model.
from langchain_core.vectorstores import InMemoryVectorStore
from policies import chunks # the chunks from the documents lesson
from word_embeddings import WordEmbeddings
store = InMemoryVectorStore(WordEmbeddings()) # embeds each chunk on add
store.add_documents(chunks)InMemoryVectorStore embeds each chunk as it is added and keeps it in a list. policies.py is the documents lesson's file.
for question in ["Is shipping free?", "How do I reset my password?"]:
doc, score = store.similarity_search_with_score(question, k=1)[0]
print(f"{score:.2f} {doc.metadata['source']:<12} {question}")0.61 shipping.md Is shipping free? 0.69 accounts.md How do I reset my password?
The score is the cosine similarity of the two vectors: 1 for the same words in the same proportions, 0 for nothing in common. Both questions found the policy that answers them.
Where counting words fails
Ask two questions the policies do not cover. The store still returns its closest chunk, which is the limit to see before trusting it.
for question in ["Can I pay with bitcoin?", "Do you sell gift cards?"]:
doc, score = store.similarity_search_with_score(question, k=1)[0]
print(f"{score:.2f} {doc.metadata['source']:<12} {doc.page_content[:40]}")0.25 shipping.md Express shipping arrives the next workin 0.17 refunds.md You can ask for a refund within 30 days
Neither question is covered, but the store always returns its closest chunk. Bitcoin landed on the express shipping chunk because "pay" in the question and "next" in the chunk both hash to slot 60, a collision that hashing cannot avoid. Gift cards landed on the refund window chunk the same way: "gift" in the question and "be" in "can be refunded" both hash to slot 13. Both matches are collisions, not meaning. Low scores are the signal: the retrieval lesson treats anything under 0.3 as no answer.
A retriever
Wrap the store as a retriever, an object that takes a question and returns documents.
retriever = store.as_retriever(search_kwargs={"k": 1})
print(retriever.invoke("Is shipping free?")[0].page_content)Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros.
It has the same invoke and batch methods as a chat model. There is also a similarity_score_threshold search type for filtering by score; with InMemoryVectorStore in this version it raises NotImplementedError, so the retrieval lesson filters by score itself.
What the scores show
- "refund" and "Refunds" landed in the same slot, because the text is lowercased and a final s is stripped, so a search treats them as one word.
- similarity_search_with_score returns the cosine similarity, 1 for the same words in the same proportions and 0 for nothing in common.
- Both covered questions found the right policy, while the two uncovered ones still returned their closest chunk at a low score, both through hash collisions.
- A low score is the signal there is no real answer, which the retrieval lesson uses by dropping anything under 0.3.
- as_retriever wraps the store as a retriever, with the same invoke and batch methods as a chat model.
Word-count embedding vs hosted embedding
| This word-count model | A hosted embedding model | |
|---|---|---|
| Matches on | Shared words | Meaning |
| "bitcoin" vs "payment" | No match unless the word repeats | Related, so it can match |
| Runs | On your machine | As an API call |
Where embeddings fit
- A policy or docs search that returns the closest chunks to a question.
- Any place you need a score to tell a real match from a guess.
Related
- Previous: Documents and splitting
- Next: Retrieval as a tool
- Reference: Retrieval
- Remove
"you"fromCOMMONand rerun the two good questions. - Change 256 to 16 in both places and look for more collisions.
- Print the top two results with
k=2for "Is shipping free?".
Little by little, you're building something great.