Documents and splitting
A Document is an object that holds text and metadata, and a text splitter is a tool that cuts long text into chunks small enough to search.
Last updated: 27 Sep, 2026 · LangChain 1.4
The Document structure
A LangChain Document is a data structure with two parts. page_content holds the content read from a file, and metadata holds extra information about that file: its name, how many pages it has, when it was created. Every loader hands back this structure: a PDF loader, a CSV loader, a web-based loader or a directory loader each read their source and return Documents. The crash course then imports Document from langchain_core.documents and builds one by hand, with a line of text as page_content and a source, a page count, an author and a creation date as metadata. The metadata matters once the chunks are embedded and stored in a vector database: a similarity search can then apply filters on it, such as only the documents by one author.
from langchain_core.documents import Document
doc=Document(
page_content="this is the main text content I am using to create RAG",
metadata={
"source":"exmaple.txt",
"pages":1,
"author":"Krish Naik",
"date_created":"2025-01-01"
}
)
print(repr(doc))Document(metadata={'source': 'exmaple.txt', 'pages': 1, 'author': 'Krish Naik', 'date_created': '2025-01-01'}, page_content='this is the main text content I am using to create RAG')Loaders turn files into Documents
You rarely build Documents by hand. A loader reads a source and returns them with the metadata filled in. TextLoader reads one text file, the video's python_intro.txt with encoding="utf-8", and load() returns a list of Documents whose metadata already holds the source path. DirectoryLoader reads a whole folder: glob is the pattern that picks the files and loader_cls is the loader used for each one. Over the video's folder of two text files it returned two Documents. The code below is the PDF version of the same call:
from langchain_community.document_loaders import TextLoader, DirectoryLoader, PyMuPDFLoader
loader = TextLoader("../data/text_files/python_intro.txt", encoding="utf-8")
document = loader.load() # [Document(metadata={'source': ...}, page_content=...)]
dir_loader = DirectoryLoader(
"../data/pdf",
glob="**/*.pdf", # which files to pick up
loader_cls=PyMuPDFLoader, # how to read each one
)
pdf_documents = dir_loader.load() # one Document per PDF pageFor the PDFs, DirectoryLoader points at the pdf folder with PyMuPDFLoader as loader_cls. The community package has both PyPDFLoader and PyMuPDFLoader, and the video picks PyMuPDF. It does not take an encoding argument, so the video drops it. The PDF loader returns one Document per page, and its metadata is richer: the file path, the creation date, the total pages (15, 27 and 21 for the three research papers) and, where the PDF has one, the author. Whatever the loader, every item in the list is still a Document. The video's loaders need pip install langchain-community pymupdf; the shop version below needs neither.
Chunking a folder of PDFs
Before chunking, the video checks all_pdf_documents, the pages its own process_all_pdfs function read from the pdf folder. Each page's metadata holds the PDF's own fields (author, dates, total pages, source) plus source_file and file_type, which that function adds. split_documents then cuts them into chunks with chunk_size=1000 and chunk_overlap=200. RecursiveCharacterTextSplitter, from langchain_text_splitters, splits recursively on the separators in order: paragraph breaks, new lines, spaces, then single characters, so a cut lands at a natural break whenever it can. chunk_overlap means some text is repeated between two neighbouring chunks, so a sentence split across two chunks is still whole in one of them. The function prints the count, the first 200 characters of the first chunk and its metadata:
### Text splitting get into chunks
def split_documents(documents,chunk_size=1000,chunk_overlap=200):
"""Split documents into smaller chunks for better RAG performance"""
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=chunk_size,
chunk_overlap=chunk_overlap,
length_function=len,
separators=["\n\n", "\n", " ", ""]
)
split_docs = text_splitter.split_documents(documents)
print(f"Split {len(documents)} documents into {len(split_docs)} chunks")
# Show example of a chunk
if split_docs:
print(f"\nExample chunk:")
print(f"Content: {split_docs[0].page_content[:200]}...")
print(f"Metadata: {split_docs[0].metadata}")
return split_docs
chunks = split_documents(all_pdf_documents)Split 64 documents into 359 chunks
Example chunk:
Content: Provided proper attribution is provided, Google hereby grants permission to
reproduce the tables and figures in this paper solely for use in journalistic or
scholarly works.
Attention Is All You Need
...
Metadata: {'producer': 'pdfTeX-1.40.25', 'creator': 'LaTeX with hyperref', 'creationdate': '2024-04-10T21:11:43+00:00', 'author': '', 'keywords': '', 'moddate': '2024-04-10T21:11:43+00:00', 'ptex.fullbanner': 'This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5', 'subject': '', 'title': '', 'trapped': '/False', 'source': '..\\data\\pdf\\attention.pdf', 'total_pages': 15, 'page': 0, 'page_label': '1', 'source_file': 'attention.pdf', 'file_type': 'pdf'}Shown as it ran in the video, not run here. It uses the video's PDFs and loaders (pip install langchain-community pymupdf), so it is here for the numbers; the shop version below runs with what you have.
There was one Document per page, 64 pages from four PDFs in the video's folder: the "Attention Is All You Need" paper with 15 pages, an embedding paper with 27, an object detection paper with 21 and a one-page proposal. They became 359 chunks of up to 1,000 characters. Each chunk keeps its page's metadata, source_file and file_type included, so any chunk found later can still say which file and page it came from.
The video's pipeline starts from research papers. The shop's version is smaller: three short policy files, so every chunk fits on your screen. Customers ask the desk how long a refund takes or when shipping is free, and those answers live in these files. Two packages are needed: the text splitters, and numpy, which the in-memory search in the next lesson uses but does not install.
pip install "langchain-text-splitters==1.1.2" "numpy==2.5.3"Document and RecursiveCharacterTextSplitter
Document(page_content="...", metadata={"source": "file.md"}) # text + where it came from
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs) # -> a list of smaller DocumentsThe policy text
Everything in this section goes in one file, policies.py; the next lessons import chunks from it. Start with the three policy files as plain strings, keyed by file name.
POLICIES = {
"refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
"\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
"shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
"\n\nExpress shipping arrives the next working day and costs 9 euros.",
"accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}One Document per file
Wrap each file in a Document, putting the file name in metadata so an answer can say where it came from.
from langchain_core.documents import Document
# one Document per file: the text, plus its source
docs = [Document(page_content=text, metadata={"source": name})
for name, text in POLICIES.items()]The text splitter
Cut the documents into chunks. The splitter breaks on paragraphs first and only cuts smaller when a piece is still over chunk_size. This is the last part of policies.py.
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs) # add_start_index records where each chunk beganThe three documents
Run these checks in a separate script that imports from policies.py, so policies.py itself stays quiet when later lessons import it.
print(len(docs))
print(docs[0].metadata)
print(docs[0].page_content[:60])3
{'source': 'refunds.md'}
Refunds go back to the card you paid with. They take up to 5Chunks
Split the documents and print each chunk's source, start index, and opening text.
for chunk in chunks:
print(chunk.metadata["source"], chunk.metadata["start_index"], chunk.page_content[:45])refunds.md 0 Refunds go back to the card you paid with. Th refunds.md 86 You can ask for a refund within 30 days of de shipping.md 0 Standard shipping takes 3 to 5 working days. shipping.md 88 Express shipping arrives the next working day accounts.md 0 To reset your password, use the reset link on
Every paragraph here fits in 120 characters, so each became one chunk. add_start_index records where each chunk began in its document.
Chunks that overlap
Try a smaller size, with a 20-character overlap.
small = RecursiveCharacterTextSplitter(chunk_size=60, chunk_overlap=20)
for piece in small.split_text(POLICIES["refunds.md"])[:4]:
print(repr(piece))'Refunds go back to the card you paid with. They take up to 5' 'They take up to 5 working days to arrive.' 'You can ask for a refund within 30 days of delivery. Opened' 'of delivery. Opened items can be refunded if they are'
At 60 characters the splitter had to cut inside paragraphs, at spaces. chunk_overlap repeats up to 20 characters from the end of one chunk at the start of the next, so a sentence cut in two is more likely to appear whole in one of them.
What the chunks show
- A Document holds page_content and metadata; the file name in metadata lets every later answer name its source.
- The splitter broke on paragraph breaks first, so each policy paragraph became one chunk while it fit inside chunk_size.
- add_start_index records where each chunk began in its document, which lets an answer point to an exact spot.
- At size 60 the splitter cut inside paragraphs, and chunk_overlap repeated up to 20 characters across each cut, so a split sentence still appears whole somewhere.
No overlap vs overlap
| chunk_overlap=0 | chunk_overlap=20 | |
|---|---|---|
| Chunk edges | Clean, no repeats | Each chunk repeats up to 20 characters from the last |
| A sentence cut in two | Split across two chunks | More likely to appear whole in one chunk |
| Characters stored | Fewer | More |
Where splitting fits
- Turning policy files, docs, or a long PDF into pieces small enough to search.
- Keeping a citation to the source file and offset with every chunk.
chunk_size to the answers you expect, and add chunk_overlap so a split sentence is not lost.Related
- Previous: Guardrails
- Next: Embeddings and a vector store
- Reference: Retrieval
- Set
chunk_size=50and count the chunks for all three documents. - Add a fourth policy,
"payments.md", and check its chunks. - Split with
chunk_overlap=0at size 60 and compare the pieces.
Every expert started right here.