Retrieval: finding the right documents
Retrieval is finding the few documents that answer a question, so the model can read them instead of guessing. Here you build a tiny keyword retriever to make the idea concrete.
Last updated: 27 Sep, 2026 · LangGraph 1.2
A model does not know your policies or your data. Retrieval fetches the relevant lines first; the next lesson has the model answer from them.
The documents
Start with the documents to search. Here they are three short support lines.
docs = [
"Refunds are processed within 5 working days.",
"Reset your password from the settings page.",
"Orders ship within 2 working days.",
]The stop words
Some words carry no meaning for matching, such as "how" and "the". Put those in a stop set so the search ignores them.
STOP = {"how", "long", "for", "a", "the", "to", "do", "you", "is"} # words to ignoreThe search function
The search splits the question into words, drops the stop words, and keeps any document that contains one of the remaining words.
def search(query, docs):
terms = [w for w in query.lower().split() if w not in STOP] # content words only
return [d for d in docs if any(t in d.lower() for t in terms)] # docs that share oneAsking a question
Ask a question. After the stop words are dropped, "refund" is the word that matters, and it appears in the first document.
print(search("how long for a refund", docs))The retriever in a run
The same pieces in one file, ready to run.
docs = [
"Refunds are processed within 5 working days.",
"Reset your password from the settings page.",
"Orders ship within 2 working days.",
]
STOP = {"how", "long", "for", "a", "the", "to", "do", "you", "is"}
def search(query, docs):
terms = [w for w in query.lower().split() if w not in STOP]
return [d for d in docs if any(t in d.lower() for t in terms)]
print(search("how long for a refund", docs))What the search returned
- The retriever keeps any document that shares a content word with the question; common words like "how" and "for" are ignored.
- "refund" matched the first line, so only that one came back.
- A real retriever uses embeddings to match by meaning, not exact words, but the shape is the same: question in, a short list of documents out.
Keyword vs embeddings
| Keyword match | Embeddings | |
|---|---|---|
| Matches | Shared words | Similar meaning |
| "refund" vs "money back" | Misses it | Finds it |
| Setup | None | An embedding model and a vector store |
When to retrieve documents
- Answering from a company's own documents, policies or a knowledge base.
- The retrieve step of a RAG pipeline, before the model reads the results.
Related
- Previous: Capstone: the support agent
- Next: RAG: answering from what you found
- Add a document and a query that should match it.
- Return only the single best match instead of all matches.
Slow is fine. Stopping is the only problem.