RAG
RAG, retrieval-augmented generation, means: retrieve the relevant documents, then have the model answer using them. It grounds answers in your own data and lets the agent refuse when it has nothing to answer from.
Last updated: 29 Sep, 2026 · LangGraph 1.2
Retrieval found the documents. RAG puts them in front of the model and asks it to answer from them, not from memory, which keeps the answer grounded and checkable. The agentic RAG video builds it as a graph with one more step in front: a node that decides whether the question needs the documents at all.
The agentic RAG graph from the video
The video's state has four keys: question, documents, answer and needs_retrieval. decide_retrieval is plain Python: if the question contains "what", "how", "explain", "describe" or "tell me", it sets needs_retrieval to true. An LLM with a prompt could make that decision instead. retrieve_documents asks a FAISS retriever (a vector store searched by meaning) for the question's documents, and generate_answer builds a prompt from them, or asks the LLM directly when there are none.
should_retrieve reads needs_retrieval and returns "retrieve" or "generate", and add_conditional_edges("decide", should_retrieve, {...}) maps that to the nodes; retrieve then goes to generate, and generate to END.

The video's retriever uses OpenAI embeddings and gpt-4.1, so it needs an OpenAI key to run. Its as_retriever(k=3) ignores k; as_retriever(search_kwargs={"k": 3}) is the form that sets it. It starts the graph with set_entry_point("decide"), the same as add_edge(START, "decide"), the form the docs use.
The shop's version below keeps the video's three nodes and its keyword rule, retrieves with the keyword search from Retrieval, and differs in one decision: with nothing retrieved it refuses instead of letting the model answer from memory.
The shop's decide, retrieve and generate nodes
Build it one piece at a time.
The documents and the search
Start with the two support lines and the keyword search from the retrieval lesson.
docs = [
"Refunds are processed within 5 working days.",
"Orders ship within 2 working days.",
]
STOP = {"how", "long", "for", "a", "the", "to", "do", "you", "is"}
def search(query, docs):
terms = [w for w in query.lower().split() if w not in STOP]
return [d for d in docs if any(t in d.lower() for t in terms)]The state
The state carries the question, the documents found, the answer and the decision, as in the video.
from typing_extensions import TypedDict
class RagState(TypedDict):
question: str
documents: list[str]
answer: str
needs_retrieval: boolThe decide and retrieve nodes
decide uses the video's keywords. retrieve runs the search and stores what it found.
def decide(state):
# the video's rule: these words mean the question needs the documents
retrieval_keywords = ["what", "how", "explain", "describe", "tell me"]
question = state["question"].lower()
return {"needs_retrieval": any(k in question for k in retrieval_keywords)}
def retrieve(state):
return {"documents": search(state["question"], docs)}The generate node
generate refuses when nothing was retrieved. Otherwise it gives the model the documents and asks it to answer only from them.
from langchain.chat_models import init_chat_model
from langchain.messages import SystemMessage, HumanMessage
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEY
def generate(state):
if not state["documents"]:
return {"answer": "I do not have information on that."} # nothing retrieved: refuse
context = "\n".join(state["documents"])
reply = model.invoke([
SystemMessage("Answer in one short sentence, using only these documents:\n" + context),
HumanMessage(state["question"]),
])
return {"answer": reply.content}Wiring the graph
should_retrieve is the router, and the path map names the node for each answer.
from langgraph.graph import StateGraph, START, END
def should_retrieve(state):
return "retrieve" if state["needs_retrieval"] else "generate"
b = StateGraph(RagState)
b.add_node("decide", decide)
b.add_node("retrieve", retrieve)
b.add_node("generate", generate)
b.add_edge(START, "decide")
b.add_conditional_edges("decide", should_retrieve, {"retrieve": "retrieve", "generate": "generate"})
b.add_edge("retrieve", "generate")
b.add_edge("generate", END)
app = b.compile()Asking two questions
Ask one question the documents cover and one they do not.
for q in ["how long for a refund", "do you deliver to Mars"]:
result = app.invoke({"question": q, "documents": [], "answer": "", "needs_retrieval": False})
print(q, "| retrieve:", result["needs_retrieval"], "| docs:", len(result["documents"]))
print(" ", result["answer"])Agentic RAG end to end
The pieces in one file:
docs = [
"Refunds are processed within 5 working days.",
"Orders ship within 2 working days.",
]
STOP = {"how", "long", "for", "a", "the", "to", "do", "you", "is"}
def search(query, docs):
terms = [w for w in query.lower().split() if w not in STOP]
return [d for d in docs if any(t in d.lower() for t in terms)]
from typing_extensions import TypedDict
class RagState(TypedDict):
question: str
documents: list[str]
answer: str
needs_retrieval: bool
def decide(state):
# the video's rule: these words mean the question needs the documents
retrieval_keywords = ["what", "how", "explain", "describe", "tell me"]
question = state["question"].lower()
return {"needs_retrieval": any(k in question for k in retrieval_keywords)}
def retrieve(state):
return {"documents": search(state["question"], docs)}
from langchain.chat_models import init_chat_model
from langchain.messages import SystemMessage, HumanMessage
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEY
def generate(state):
if not state["documents"]:
return {"answer": "I do not have information on that."} # nothing retrieved: refuse
context = "\n".join(state["documents"])
reply = model.invoke([
SystemMessage("Answer in one short sentence, using only these documents:\n" + context),
HumanMessage(state["question"]),
])
return {"answer": reply.content}
from langgraph.graph import StateGraph, START, END
def should_retrieve(state):
return "retrieve" if state["needs_retrieval"] else "generate"
b = StateGraph(RagState)
b.add_node("decide", decide)
b.add_node("retrieve", retrieve)
b.add_node("generate", generate)
b.add_edge(START, "decide")
b.add_conditional_edges("decide", should_retrieve, {"retrieve": "retrieve", "generate": "generate"})
b.add_edge("retrieve", "generate")
b.add_edge("generate", END)
app = b.compile()
for q in ["how long for a refund", "do you deliver to Mars"]:
result = app.invoke({"question": q, "documents": [], "answer": "", "needs_retrieval": False})
print(q, "| retrieve:", result["needs_retrieval"], "| docs:", len(result["documents"]))
print(" ", result["answer"])how long for a refund | retrieve: True | docs: 1 Refunds are processed within 5 working days. do you deliver to Mars | retrieve: False | docs: 0 I do not have information on that.
What the two runs showed
- The first question contains "how", so
decidesent it toretrieve, the search matched the refunds line, andgenerateanswered from that one document. - The second question has none of the keywords, so the run went from
decidestraight togeneratewith no documents, andgeneraterefused without calling the model.
Plain model vs RAG
| Plain model | RAG | |
|---|---|---|
| Answers from | Its training | Your documents |
| Your data | Not known | Retrieved and read each time |
| When it lacks an answer | May invent one | Can refuse, with nothing to cite |
When to ground answers in data
- A documentation or policy assistant that must answer from a fixed source.
- Any answer that has to be grounded in your data and checkable, not guessed.
Related
- Previous: Retrieval
- Next: Integrations
- Reference: Agentic RAG
- Ask "tell me about shipping times" and predict the path before you run it.
- Remove "how" from
retrieval_keywordsand run the refund question again. - Print
result["documents"]for each question to see what reached the model.
This is what real progress feels like.