LangChainLangChain 1.4 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Documents and splitting

A Document is an object that holds text and metadata, and a text splitter is a tool that cuts long text into chunks small enough to search.

Last updated: 27 Sep, 2026 · LangChain 1.4

The agent knows orders. Customers also ask about the rules: how long a refund takes, when shipping is free. Those answers live in three short policy files, and turning them into searchable pieces starts here. Two packages are needed: the text splitters, and numpy, which the in-memory search in the next lesson uses but does not install.

pip install "langchain-text-splitters==1.1.2" "numpy==2.5.3"

Document and RecursiveCharacterTextSplitter

python
Document(page_content="...", metadata={"source": "file.md"})   # text + where it came from

splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs)   # -> a list of smaller Documents

The policy text

Start with the three policy files as plain strings, keyed by file name.

python
POLICIES = {
    "refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
                  "\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
    "shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
                   "\n\nExpress shipping arrives the next working day and costs 9 euros.",
    "accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}

One Document per file

Wrap each file in a Document, putting the file name in metadata so an answer can say where it came from.

python
from langchain_core.documents import Document

# one Document per file: the text, plus its source
docs = [Document(page_content=text, metadata={"source": name})
        for name, text in POLICIES.items()]

The text splitter

Cut the documents into chunks. The splitter breaks on paragraphs first and only cuts smaller when a piece is still over chunk_size.

python
from langchain_text_splitters import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs)   # add_start_index records where each chunk began

The three documents

Example
print(len(docs))
print(docs[0].metadata)
print(docs[0].page_content[:60])

A Document has page_content, the text, and metadata, a dictionary for anything else. The file name goes in the metadata so every answer can say where it came from.

Chunks

Split the documents and print each chunk's source, start index, and opening text.

Example
for chunk in chunks:
    print(chunk.metadata["source"], chunk.metadata["start_index"], chunk.page_content[:45])

A search should find the paragraph that answers the question, not a whole file. RecursiveCharacterTextSplitter splits on paragraph breaks first and only cuts smaller when a piece is still over chunk_size characters. Every paragraph here fits in 120, so each became one chunk. add_start_index records where each chunk began in its document.

Chunks that overlap

At a smaller size the splitter has to cut inside paragraphs. chunk_overlap repeats a little text across each cut.

Example
small = RecursiveCharacterTextSplitter(chunk_size=60, chunk_overlap=20)
for piece in small.split_text(POLICIES["refunds.md"])[:4]:
    print(repr(piece))

At 60 characters the splitter had to cut inside paragraphs, at spaces. chunk_overlap repeats up to 20 characters from the end of one chunk at the start of the next, so a sentence cut in two still appears whole somewhere. The documentation's tutorial uses chunks of 1,000 characters with 200 of overlap for a long PDF.

What the chunks show

  • A Document holds page_content and metadata; the file name in metadata lets every later answer name its source.
  • The splitter broke on paragraph breaks first, so each policy paragraph became one chunk while it fit inside chunk_size.
  • add_start_index records where each chunk began in its document, which lets an answer point to an exact spot.
  • At size 60 the splitter cut inside paragraphs, and chunk_overlap repeated up to 20 characters across each cut, so a split sentence still appears whole somewhere.

No overlap vs overlap

chunk_overlap=0chunk_overlap=20
Chunk edgesClean, no repeatsEach chunk repeats up to 20 characters from the last
A sentence cut in twoSplit across two chunksAppears whole in at least one chunk
Characters storedFewerMore

Where splitting fits

  • Turning policy files, docs, or a long PDF into pieces small enough to search.
  • Keeping a citation to the source file and offset with every chunk.
Watch out. A chunk larger than the passage you want makes search return whole files; too small cuts sentences apart. Match chunk_size to the answers you expect, and add chunk_overlap so a split sentence is not lost.
Try it yourself
  • Set chunk_size=50 and count the chunks for all three documents.
  • Add a fourth policy, "payments.md", and check its chunks.
  • Split with chunk_overlap=0 at size 60 and compare the pieces.
PreviousGuardrails

Every expert started right here.