LangGraphLangGraph 1.2 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
38 small wins to finish your pathNext lesson →

RAG: answering from what you found

RAG, retrieval-augmented generation, means: retrieve the relevant documents, then have the model answer using them. It grounds answers in your own data and lets the agent refuse when it finds nothing.

Last updated: 27 Sep, 2026 · LangGraph 1.2

Retrieval found the documents. RAG puts them in front of the model and asks it to answer from them, not from memory, which is what keeps the answer grounded and checkable.

The documents and stop words

Start with the documents and the same stop set from the retrieval lesson.

python
docs = [
    "Refunds are processed within 5 working days.",
    "Orders ship within 2 working days.",
]

STOP = {"how", "long", "for", "a", "the", "to", "do", "you", "is"}   # words to ignore

The search function

Reuse the keyword search from the previous lesson. It returns the documents that share a content word with the question.

python
def search(query, docs):
    terms = [w for w in query.lower().split() if w not in STOP]
    return [d for d in docs if any(t in d.lower() for t in terms)]

The answer step

Now the answer step. It searches first; if nothing came back it refuses, and if a document matched it answers from that document. A real graph puts a model call where the last line is.

python
def answer(query, docs):
    found = search(query, docs)
    if not found:
        return "I do not have information on that."     # grounded refusal
    return "According to the docs: " + found[0]           # a real model would write this

Asking two questions

Ask two questions: one the documents cover, and one they do not.

python
print(answer("how long for a refund", docs))    # a document matched
print(answer("do you deliver to Mars", docs))    # nothing matched

RAG end to end

The same pieces in one file, ready to run.

Example
docs = [
    "Refunds are processed within 5 working days.",
    "Orders ship within 2 working days.",
]

STOP = {"how", "long", "for", "a", "the", "to", "do", "you", "is"}

def search(query, docs):
    terms = [w for w in query.lower().split() if w not in STOP]
    return [d for d in docs if any(t in d.lower() for t in terms)]

def answer(query, docs):
    found = search(query, docs)
    if not found:
        return "I do not have information on that."     # grounded refusal
    return "According to the docs: " + found[0]           # a real model would write this

print(answer("how long for a refund", docs))
print(answer("do you deliver to Mars", docs))

What the two answers showed

  • The first question retrieved a matching document, so the answer is built from it and cites it.
  • The second retrieved nothing, so the agent refuses instead of inventing an answer.
  • A real graph replaces the last line with a model call whose prompt includes the retrieved documents and the instruction to answer only from them.

Plain model vs RAG

Plain modelRAG
Answers fromIts trainingYour documents
Your dataNot knownRetrieved and read each time
When it lacks an answerMay invent oneCan refuse, with nothing to cite

When to ground answers in data

  • A documentation or policy assistant that must answer from a fixed source.
  • Any answer that has to be grounded in your data and checkable, not guessed.
Watch out. Tell the model to answer only from the retrieved documents and to say when it cannot. Without that instruction it falls back on training and the grounding is lost.
Try it yourself
  • Add a document and ask a question it answers.
  • Return the matched document as a citation alongside the answer.

This is what real progress feels like.