Why retrieval: a model does not know your documents
Retrieval is the step that finds the passage of your documents that answers a question, and it is harder than matching words because people ask in words the documents never use.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
The overview described RAG as find, then answer. This lesson looks at the find step on its own, using the obvious approach first, matching words, so you can see where it breaks and why the rest of the course exists.
Sending every document with every question would cost tokens and bury the answer, so the assistant has to pick the relevant passage. Start with the plainest way to pick it: keep the words of the question and return any file that shares one.
A tiny help centre and a word matcher
Three one-line documents stand in for the shop's pages, and keyword_search keeps words longer than three letters and returns every file that shares one:
DOCS = {
"refunds.md": "You can get a full refund within 30 days of delivery.",
"delivery.md": "Standard delivery takes 3 to 5 working days.",
"lamps.md": "The LMP-204 desk lamp has a known cable fault.",
}
def keyword_search(question):
words = {w for w in question.lower().replace("?", "").split() if len(w) > 3}
return [name for name, text in DOCS.items() if words & set(text.lower().replace(".", "").split())]Running keyword search over four questions
for question in ["How long does delivery take?", "Can I get a refund?", "How do I get my money back?", "My parcel is late"]:
print(f"{question:30} -> {keyword_search(question)}")How long does delivery take? -> ['refunds.md', 'delivery.md'] Can I get a refund? -> ['refunds.md'] How do I get my money back? -> [] My parcel is late -> []
Reading what keyword search missed
- The first two questions find their file, and the delivery question also turns up refunds, whose text happens to mention delivery.
- "How do I get my money back?" returns an empty list: it means a refund and never says the word.
- "My parcel is late" returns an empty list too, though it is a delivery question asked in other words.
People describe the same need in different words, so search has to match meaning, not spelling. That is what embeddings do, in the next lesson.
Keyword search vs search by meaning
| Keyword search | Search by meaning | |
|---|---|---|
| Matches on | Shared words | Similar meaning |
| "money back" finds refund | No | Yes |
| Needs a model or key | No | No, with a local model |
| Good for | Exact codes and names | Questions in the reader's own words |
The pieces of a RAG system
- Documents: the source material, split into chunks.
- An index: the chunks stored in a form that can be searched by meaning.
- A retriever: finds the chunks that best match a question.
- A model: answers from those chunks, and says so when they do not hold the answer.
When keyword search is still enough
- Part numbers and error codes, where the exact string is what the reader typed.
- Small, controlled vocabularies where every reader uses the same words.
- As one half of hybrid search later in the course, alongside search by meaning.
Related
- Previous: Installation and setup
- Next: Embeddings: text as numbers that carry meaning
- Reference: High-level concepts: RAG
- Add the word
moneyto the refunds text. Does that fix the search, and would it fix"cash back"? - Search for
"the"and read the result. - Write three questions about delivery that share no word with its text.
You understood something today that you didn't yesterday.