1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →
Chunking documents with SentenceSplitter
A node is one chunk of a document, and SentenceSplitter is the parser that decides where each chunk ends, which sets whether a retrieved chunk holds the full answer.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
The reader from the last lesson gave whole files. Retrieval works on chunks, not files, so the next choice is how to cut each file, and that choice decides what a question can find.
Splitting documents into nodes
SentenceSplitter cuts at sentence boundaries where it can, keeping each chunk under chunk_size tokens. chunk_overlap repeats the end of one chunk at the start of the next:
from llama_index.core.node_parser import SentenceSplitter
splitter = SentenceSplitter(chunk_size=80, chunk_overlap=0)
nodes = splitter.get_nodes_from_documents(documents)Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
help/lamps.md
# Lamps
The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.
All lamps come with a two year guarantee against electrical faults.
Bulbs are not covered by the refund policy once they have been used.
The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
help/refunds.md
# Refunds
You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.
Items bought in a sale can be refunded too, but the delivery charge is not returned.
To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.
Personalised items cannot be refunded unless they arrive damaged.
help/delivery.md
# Delivery
Standard delivery takes 3 to 5 working days and is free on orders over 40.
Express delivery arrives the next working day if you order before 2pm. It costs 6.
We deliver to the mainland only. Parcels to islands take 2 extra working days.
If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
Splitting the help centre into nodes
from llama_index.core.node_parser import SentenceSplitter
nodes = SentenceSplitter(chunk_size=80, chunk_overlap=0).get_nodes_from_documents(documents)
print(len(documents), "documents ->", len(nodes), "nodes")
for node in nodes:
print(node.metadata["file_name"], "|", node.get_content().replace("\n", " ")[:60])Output
3 documents -> 6 nodes delivery.md | # Delivery Standard delivery takes 3 to 5 working days and delivery.md | If a parcel has not arrived after 10 working days, contact u lamps.md | # Lamps The LMP-204 desk lamp has a known cable fault. Stop lamps.md | The LMP-310 floor lamp needs a bulb with an E27 fitting, whi refunds.md | # Refunds You can get a full refund within 30 days of deliv refunds.md | Print the label and drop the parcel at any post office. Per
Reading the nodes
- Three documents become six nodes: each file was long enough to split into two chunks at this size.
- A node carries its file name, copied from the document, plus a link back to it.
- Each chunk starts on a sentence, so no chunk begins in the middle of one.
Trying a small and a large size
The same file at 60 tokens and at 400 tokens shows the trade-off. Watch the refund file:
for size in (60, 400):
chunks = SentenceSplitter(chunk_size=size, chunk_overlap=0).get_nodes_from_documents(documents)
# count and print the refund chunksComparing a small and a large chunk size
for size in (60, 400):
chunks = SentenceSplitter(chunk_size=size, chunk_overlap=0).get_nodes_from_documents(documents)
refund = [n.get_content().replace("\n", " ") for n in chunks if n.metadata["file_name"] == "refunds.md"]
print(size, "tokens:", len(chunks), "nodes")
for text in refund:
print(" ", text[:90])Output
60 tokens: 6 nodes
# Refunds You can get a full refund within 30 days of delivery. The money goes back to th
To start a refund, open the order in your account and choose Return an item. Print the lab
400 tokens: 3 nodes
# Refunds You can get a full refund within 30 days of delivery. The money goes back to thReading the two sizes
- At 60 tokens the refund file is two chunks, so the steps for returning an item split across the end of one chunk and the start of the next; a question about returning an item can retrieve only half of them.
- At 400 tokens each file is a single chunk, fine for three short files and wasteful for a long manual, where every retrieved chunk would drag pages of unrelated text into the prompt.
- There is no correct size, only a trade-off: small chunks match precisely and lose context, large ones keep context and match loosely.
Small chunks vs large chunks
| Small chunk_size | Large chunk_size | |
|---|---|---|
| Match precision | High | Lower |
| Context in a chunk | Little | More |
| Risk | Splits an answer | Drags in unrelated text |
| Tokens per retrieved chunk | Few | Many |
When to use a smaller or larger size
- Smaller for dense reference text, where a precise match matters and answers are short.
- Larger for flowing prose, where an answer spans several sentences.
- Measure it on real questions rather than guessing; a later lesson counts hit rate for exactly this.
A warning you may see
Watch out. Metadata counts towards the chunk size. With very small chunks, LlamaIndex warns that the metadata leaves little room for text, which is one more reason to keep metadata short.
Related
- Previous: SimpleDirectoryReader: loading a folder of documents
- Next: Metadata: what gets embedded and what the model sees
- Reference: Node parsers
Try it yourself
- Try
chunk_size=128withchunk_overlap=20and compare the refund chunks. - Print
nodes[0].ref_doc_idand compare it with a document'sdoc_id. - Count the nodes at sizes 64, 128 and 256.
Little by little, you're building something great.