Documents and splitting
A Document is an object that holds text and metadata, and a text splitter is a tool that cuts long text into chunks small enough to search.
Last updated: 27 Sep, 2026 · LangChain 1.4
The agent knows orders. Customers also ask about the rules: how long a refund takes, when shipping is free. Those answers live in three short policy files, and turning them into searchable pieces starts here. Two packages are needed: the text splitters, and numpy, which the in-memory search in the next lesson uses but does not install.
pip install "langchain-text-splitters==1.1.2" "numpy==2.5.3"Document and RecursiveCharacterTextSplitter
Document(page_content="...", metadata={"source": "file.md"}) # text + where it came from
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs) # -> a list of smaller DocumentsThe policy text
Start with the three policy files as plain strings, keyed by file name.
POLICIES = {
"refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
"\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
"shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
"\n\nExpress shipping arrives the next working day and costs 9 euros.",
"accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}One Document per file
Wrap each file in a Document, putting the file name in metadata so an answer can say where it came from.
from langchain_core.documents import Document
# one Document per file: the text, plus its source
docs = [Document(page_content=text, metadata={"source": name})
for name, text in POLICIES.items()]The text splitter
Cut the documents into chunks. The splitter breaks on paragraphs first and only cuts smaller when a piece is still over chunk_size.
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs) # add_start_index records where each chunk beganThe three documents
print(len(docs))
print(docs[0].metadata)
print(docs[0].page_content[:60])A Document has page_content, the text, and metadata, a dictionary for anything else. The file name goes in the metadata so every answer can say where it came from.
Chunks
Split the documents and print each chunk's source, start index, and opening text.
for chunk in chunks:
print(chunk.metadata["source"], chunk.metadata["start_index"], chunk.page_content[:45])A search should find the paragraph that answers the question, not a whole file. RecursiveCharacterTextSplitter splits on paragraph breaks first and only cuts smaller when a piece is still over chunk_size characters. Every paragraph here fits in 120, so each became one chunk. add_start_index records where each chunk began in its document.
Chunks that overlap
At a smaller size the splitter has to cut inside paragraphs. chunk_overlap repeats a little text across each cut.
small = RecursiveCharacterTextSplitter(chunk_size=60, chunk_overlap=20)
for piece in small.split_text(POLICIES["refunds.md"])[:4]:
print(repr(piece))At 60 characters the splitter had to cut inside paragraphs, at spaces. chunk_overlap repeats up to 20 characters from the end of one chunk at the start of the next, so a sentence cut in two still appears whole somewhere. The documentation's tutorial uses chunks of 1,000 characters with 200 of overlap for a long PDF.
What the chunks show
- A Document holds page_content and metadata; the file name in metadata lets every later answer name its source.
- The splitter broke on paragraph breaks first, so each policy paragraph became one chunk while it fit inside chunk_size.
- add_start_index records where each chunk began in its document, which lets an answer point to an exact spot.
- At size 60 the splitter cut inside paragraphs, and chunk_overlap repeated up to 20 characters across each cut, so a split sentence still appears whole somewhere.
No overlap vs overlap
| chunk_overlap=0 | chunk_overlap=20 | |
|---|---|---|
| Chunk edges | Clean, no repeats | Each chunk repeats up to 20 characters from the last |
| A sentence cut in two | Split across two chunks | Appears whole in at least one chunk |
| Characters stored | Fewer | More |
Where splitting fits
- Turning policy files, docs, or a long PDF into pieces small enough to search.
- Keeping a citation to the source file and offset with every chunk.
chunk_size to the answers you expect, and add chunk_overlap so a split sentence is not lost.Related
- Previous: Guardrails
- Next: Embeddings and a vector store
- Reference: Retrieval
- Set
chunk_size=50and count the chunks for all three documents. - Add a fourth policy,
"payments.md", and check its chunks. - Split with
chunk_overlap=0at size 60 and compare the pieces.
Every expert started right here.