RAG: answering from what you found
RAG, retrieval-augmented generation, means: retrieve the relevant documents, then have the model answer using them. It grounds answers in your own data and lets the agent refuse when it finds nothing.
Last updated: 27 Sep, 2026 · LangGraph 1.2
Retrieval found the documents. RAG puts them in front of the model and asks it to answer from them, not from memory, which is what keeps the answer grounded and checkable.
The documents and stop words
Start with the documents and the same stop set from the retrieval lesson.
docs = [
"Refunds are processed within 5 working days.",
"Orders ship within 2 working days.",
]
STOP = {"how", "long", "for", "a", "the", "to", "do", "you", "is"} # words to ignoreThe search function
Reuse the keyword search from the previous lesson. It returns the documents that share a content word with the question.
def search(query, docs):
terms = [w for w in query.lower().split() if w not in STOP]
return [d for d in docs if any(t in d.lower() for t in terms)]The answer step
Now the answer step. It searches first; if nothing came back it refuses, and if a document matched it answers from that document. A real graph puts a model call where the last line is.
def answer(query, docs):
found = search(query, docs)
if not found:
return "I do not have information on that." # grounded refusal
return "According to the docs: " + found[0] # a real model would write thisAsking two questions
Ask two questions: one the documents cover, and one they do not.
print(answer("how long for a refund", docs)) # a document matched
print(answer("do you deliver to Mars", docs)) # nothing matchedRAG end to end
The same pieces in one file, ready to run.
docs = [
"Refunds are processed within 5 working days.",
"Orders ship within 2 working days.",
]
STOP = {"how", "long", "for", "a", "the", "to", "do", "you", "is"}
def search(query, docs):
terms = [w for w in query.lower().split() if w not in STOP]
return [d for d in docs if any(t in d.lower() for t in terms)]
def answer(query, docs):
found = search(query, docs)
if not found:
return "I do not have information on that." # grounded refusal
return "According to the docs: " + found[0] # a real model would write this
print(answer("how long for a refund", docs))
print(answer("do you deliver to Mars", docs))What the two answers showed
- The first question retrieved a matching document, so the answer is built from it and cites it.
- The second retrieved nothing, so the agent refuses instead of inventing an answer.
- A real graph replaces the last line with a model call whose prompt includes the retrieved documents and the instruction to answer only from them.
Plain model vs RAG
| Plain model | RAG | |
|---|---|---|
| Answers from | Its training | Your documents |
| Your data | Not known | Retrieved and read each time |
| When it lacks an answer | May invent one | Can refuse, with nothing to cite |
When to ground answers in data
- A documentation or policy assistant that must answer from a fixed source.
- Any answer that has to be grounded in your data and checkable, not guessed.
Related
- Previous: Retrieval: finding the right documents
- Next: Integrations: swapping in real backends
- Reference: Agentic RAG
- Add a document and ask a question it answers.
- Return the matched document as a citation alongside the answer.
This is what real progress feels like.