Retrieval as a tool
Retrieval as a tool is a pattern where the search sits behind a tool, so the agent decides when to look something up and then answers only from what it found.
Last updated: 27 Sep, 2026 · LangChain 1.4
The two RAG pipelines
Retrieval-augmented generation, RAG, means finding the passages that answer a question, then answering from them. It runs as two pipelines. The ingestion pipeline runs once, ahead of time: the last two lessons built it. The retrieval pipeline runs on every question.
The retriever: query in, context out
Retrieval is the second half of RAG. The user's query is converted into an embedding with the same model as the chunks, the vector store is searched with it, and the closest chunks come back as the context. The RAG crash course wraps this in a RAGRetriever class. Its constructor takes the vector store and the embedding manager from the last two lessons, because a retriever is an interface built on top of a vector store: a query goes in and a response comes back. Its retrieve(query, top_k, score_threshold) method embeds the query with generate_embeddings, calls query on the Chroma collection with that embedding and top_k, and turns each distance into a similarity score as 1 minus the distance. Results whose score is above score_threshold, 0.0 by default, go into a list of dictionaries holding each document's content and metadata.
rag_retriever = RAGRetriever(vectorstore, embedding_manager) creates the retriever, and the video asks it "What is attention is all you need". The query is embedded into a (1, 384) array, because all-MiniLM-L6-v2 gives 384 numbers, and one document comes back, with its content and metadata: a passage from the attention paper headed "3.2 Attention" that begins "An attention function can be described as mapping a query and a set of key-value pairs to an output": that is the context. How high a good match scores depends on the embedding model, so score_threshold has to be chosen for each model. The shop's word-count model scores good matches around 0.6, so this lesson cuts at 0.3; A persistent vector store with Chroma shows Gemini needing about 0.7.
The video's pipeline always searches before answering. In the shop version the search is a tool, so the agent decides whether to search at all, and when nothing matches it says so instead of guessing. A desk that also answers order questions needs that choice: "Where is A17?" has nothing to find in the policies.
The policy search as a tool
- written in Documents and splitting
- written in Embeddings and a vector store
View the code here
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter
POLICIES = {
"refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
"\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
"shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
"\n\nExpress shipping arrives the next working day and costs 9 euros.",
"accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}
docs = [Document(page_content=text, metadata={"source": name}) for name, text in POLICIES.items()]
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs)
import re
import zlib
from langchain_core.embeddings import Embeddings
COMMON = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i",
"if", "is", "it", "my", "of", "on", "the", "to", "what", "with", "you", "your"}
class WordEmbeddings(Embeddings):
def embed_query(self, text):
vector = [0.0] * 256
for word in re.findall(r"[a-z]+", text.lower()):
if word not in COMMON:
vector[zlib.crc32(word.rstrip("s").encode()) % 256] += 1.0
return vector
def embed_documents(self, texts):
return [self.embed_query(text) for text in texts]
Building the policy store
Build the store from the previous lesson's chunks and embedding model. The store and the tool below go in one file, search.py; later lessons import search_policies from it.
from langchain.tools import tool
from langchain_core.vectorstores import InMemoryVectorStore
from policies import chunks
from word_embeddings import WordEmbeddings
store = InMemoryVectorStore(WordEmbeddings())
store.add_documents(chunks)The search tool
Put the search behind a tool. It keeps the chunks scoring 0.3 or more and returns them with their sources; below that it returns a fixed sentence, so there is nothing to invent from. The docstring also tells the model to pass the customer's question word for word: left to itself, the model rewrites "How long does a refund take?" as "refund time", and those two words score only 0.21 against the refund chunk, under the cut.
@tool
def search_policies(query: str) -> str:
"""Search the shop's policies on refunds, shipping and accounts.
Pass the customer's question, word for word, as the query."""
found = [doc for doc, score in store.similarity_search_with_score(query, k=2) if score >= 0.3]
if not found:
return "No policy covers this." # nothing to answer from
return "\n".join(f"[{doc.metadata['source']}] {doc.page_content}" for doc in found)Giving the agent the tool
Give a real model the tool and a system prompt that says to answer only from what the tool returned, name the source file, and admit when nothing was found. Then let the agent loop run.
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from search import search_policies
agent = create_agent(init_chat_model("groq:openai/gpt-oss-120b", temperature=0), # uses your GROQ_API_KEY
tools=[search_policies],
system_prompt="Answer in one or two short sentences of plain text, using only what the search_policies tool returned. Name the source file in square brackets, like [refunds.md]. If the tool finds nothing, say the policies do not cover it.")Running a covered and an uncovered question
Ask one question the policies cover and one they do not.
for question in ["How long does a refund take?", "Can I pay with bitcoin?"]:
result = agent.invoke({"messages": [{"role": "user", "content": question}]})
print(result["messages"][-1].text, end="\n\n")Refunds are processed back to the original card and can take up to 5 working days to arrive. [refunds.md] The policies do not cover it.
The refund question found the refund policy, and the answer names its file. The bitcoin question scored under the cut, so the tool returned its fixed sentence and the reply says the policies do not cover it. That second answer is the one that keeps a support desk trustworthy.
How the score decides the answer
- The refund question scored above the cut, so the tool returned the refund chunk and the answer names its file.
- The bitcoin question scored under 0.3, so the tool returned its fixed sentence and the reply admits the gap instead of guessing.
- The model answers only from what the tool returned, which is what keeps the answer grounded and checkable.
Plain model vs retrieval tool
| Plain model | Retrieval as a tool | |
|---|---|---|
| Answers from | Its training | The chunks the search returned |
| When nothing matches | May invent an answer | Says the policies do not cover it |
| Cites a source | No | Yes, the file each chunk came from |
When to use retrieval as a tool
- A support desk that must answer from fixed policies and refuse the rest.
- Any answer that has to be grounded in your data, with a source, not guessed.
Related
- Previous: Embeddings and a vector store
- Next: MCPAdapter: tools from another program
- Reference: Retrieval
- Lower the cut to 0.2 and ask about bitcoin again.
- Ask "Is express shipping free?" and read which chunks come back.
- Return the score with each chunk and print what the model receives.
You understood something today that you didn't yesterday.