Embeddings and a vector store
An embedding is a list of numbers standing for a text, and a vector store is an object that keeps chunks with their embeddings and returns the ones closest to a question.
Last updated: 27 Sep, 2026 · LangChain 1.4
Search needs each text as numbers so it can measure closeness. An embedding model turns text into a vector where similar texts get similar vectors. The one built here counts words, which is enough to see how the rest works, and it runs on your machine.
InMemoryVectorStore and a retriever
store = InMemoryVectorStore(embedding_model) # keeps chunks and their vectors
store.add_documents(chunks) # embed each chunk and store it
store.similarity_search_with_score(query, k=1) # -> [(Document, score), ...]
retriever = store.as_retriever(search_kwargs={"k": 1}) # same invoke/batch as a modelThe common words
An Embeddings class needs two methods: one for a question, one for a list of documents. Start with the imports and the words to skip.
import re
import zlib
from langchain_core.embeddings import Embeddings
# words that appear everywhere and would make every text look alike
COMMON = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i",
"if", "is", "it", "my", "of", "on", "the", "to", "what", "with", "you", "your"}A word-counting embedding
Each word is hashed into one of 256 slots and counted. Stripping a final s makes "refund" and "refunds" count as one word.
class WordEmbeddings(Embeddings):
def embed_query(self, text):
vector = [0.0] * 256
for word in re.findall(r"[a-z]+", text.lower()):
if word not in COMMON: # skip common words
vector[zlib.crc32(word.rstrip("s").encode()) % 256] += 1.0 # hash into a slot
return vector
def embed_documents(self, texts): # one vector per document
return [self.embed_query(text) for text in texts]Each word hashed to a slot
Embed a few words and print which slots fill.
model = WordEmbeddings()
for text in ["refund", "Refunds", "shipping"]:
vector = model.embed_query(text)
print(f"{text:<9} slots {[i for i, v in enumerate(vector) if v]}")"refund" and "Refunds" land in the same slot, and "shipping" lands somewhere else. Texts about the same thing fill the same slots, which is what a search has to work with.
A vector store
Build the store from the chunks and the embedding model.
from langchain_core.vectorstores import InMemoryVectorStore
from policies import chunks # the chunks from the documents lesson
from word_embeddings import WordEmbeddings
store = InMemoryVectorStore(WordEmbeddings()) # embeds each chunk on add
store.add_documents(chunks)InMemoryVectorStore embeds each chunk as it is added and keeps it in a list. policies.py is lesson 26's file.
for question in ["Is shipping free?", "How do I reset my password?"]:
doc, score = store.similarity_search_with_score(question, k=1)[0]
print(f"{score:.2f} {doc.metadata['source']:<12} {question}")The score is the cosine similarity of the two vectors: 1 for the same words in the same proportions, 0 for nothing in common. Both questions found the policy that answers them.
Where counting words fails
Ask two questions the policies do not cover. The store still returns its closest chunk, which is the limit to see before trusting it.
for question in ["Can I pay with bitcoin?", "Do you sell gift cards?"]:
doc, score = store.similarity_search_with_score(question, k=1)[0]
print(f"{score:.2f} {doc.metadata['source']:<12} {doc.page_content[:40]}")Neither question is covered, but the store always returns its closest chunk. Bitcoin landed on the express shipping chunk because two different words fell into the same slot, a collision that hashing cannot avoid. Low scores are the signal: lesson 28 treats anything under 0.3 as no answer.
A retriever
Wrap the store as a retriever, an object that takes a question and returns documents.
retriever = store.as_retriever(search_kwargs={"k": 1})
print(retriever.invoke("Is shipping free?")[0].page_content)as_retriever wraps the store as a retriever: an object that takes a question and returns documents, with the same invoke and batch methods as a chat model. There is also a similarity_score_threshold search type for filtering by score; with InMemoryVectorStore in this version it raises NotImplementedError, so lesson 28 filters by score itself.
What the scores show
- "refund" and "Refunds" landed in the same slot, because the text is lowercased and a final s is stripped, so a search treats them as one word.
- similarity_search_with_score returns the cosine similarity, 1 for the same words in the same proportions and 0 for nothing in common.
- Both covered questions found the right policy, while the two uncovered ones still returned their closest chunk at a low score, one from a hash collision.
- A low score is the signal there is no real answer, which the retrieval lesson uses by dropping anything under 0.3.
- as_retriever wraps the store as a retriever, with the same invoke and batch methods as a chat model.
Word-count embedding vs hosted embedding
| This word-count model | A hosted embedding model | |
|---|---|---|
| Matches on | Shared words | Meaning |
| "bitcoin" vs "payment" | No match unless the word repeats | Related, so it can match |
| Runs | On your machine | As an API call |
Where embeddings fit
- A policy or docs search that returns the closest chunks to a question.
- Any place you need a score to tell a real match from a guess.
Related
- Previous: Documents and splitting
- Next: Retrieval as a tool
- Reference: Retrieval
- Remove
"you"fromCOMMONand rerun the two good questions. - Change 256 to 16 in both places and look for more collisions.
- Print the top two results with
k=2for "Is shipping free?".
Little by little, you're building something great.