LlamaIndexllama-index-core 0.14 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Chunking documents with SentenceSplitter

A node is one chunk of a document, and SentenceSplitter is the parser that decides where each chunk ends, which sets whether a retrieved chunk holds the full answer.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

The reader from the last lesson gave whole files. Retrieval works on chunks, not files, so the next choice is how to cut each file, and that choice decides what a question can find.

Splitting documents into nodes

SentenceSplitter cuts at sentence boundaries where it can, keeping each chunk under chunk_size tokens. chunk_overlap repeats the end of one chunk at the start of the next:

python
from llama_index.core.node_parser import SentenceSplitter

splitter = SentenceSplitter(chunk_size=80, chunk_overlap=0)
nodes = splitter.get_nodes_from_documents(documents)
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
help/lamps.md
# Lamps

The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.

All lamps come with a two year guarantee against electrical faults.

Bulbs are not covered by the refund policy once they have been used.

The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
help/refunds.md
# Refunds

You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.

Items bought in a sale can be refunded too, but the delivery charge is not returned.

To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.

Personalised items cannot be refunded unless they arrive damaged.
help/delivery.md
# Delivery

Standard delivery takes 3 to 5 working days and is free on orders over 40.

Express delivery arrives the next working day if you order before 2pm. It costs 6.

We deliver to the mainland only. Parcels to islands take 2 extra working days.

If a parcel has not arrived after 10 working days, contact us and we will send a replacement.

Splitting the help centre into nodes

Example
from llama_index.core.node_parser import SentenceSplitter

nodes = SentenceSplitter(chunk_size=80, chunk_overlap=0).get_nodes_from_documents(documents)

print(len(documents), "documents ->", len(nodes), "nodes")
for node in nodes:
    print(node.metadata["file_name"], "|", node.get_content().replace("\n", " ")[:60])

Reading the nodes

  • Three documents become six nodes: each file was long enough to split into two chunks at this size.
  • A node carries its file name, copied from the document, plus a link back to it.
  • Each chunk starts on a sentence, so no chunk begins in the middle of one.

Trying a small and a large size

The same file at 60 tokens and at 400 tokens shows the trade-off. Watch the refund file:

python
for size in (60, 400):
    chunks = SentenceSplitter(chunk_size=size, chunk_overlap=0).get_nodes_from_documents(documents)
    # count and print the refund chunks

Comparing a small and a large chunk size

Example
for size in (60, 400):
    chunks = SentenceSplitter(chunk_size=size, chunk_overlap=0).get_nodes_from_documents(documents)
    refund = [n.get_content().replace("\n", " ") for n in chunks if n.metadata["file_name"] == "refunds.md"]
    print(size, "tokens:", len(chunks), "nodes")
    for text in refund:
        print("   ", text[:90])

Reading the two sizes

  • At 60 tokens the refund file is two chunks, so the steps for returning an item split across the end of one chunk and the start of the next; a question about returning an item can retrieve only half of them.
  • At 400 tokens each file is a single chunk, fine for three short files and wasteful for a long manual, where every retrieved chunk would drag pages of unrelated text into the prompt.
  • There is no correct size, only a trade-off: small chunks match precisely and lose context, large ones keep context and match loosely.

Small chunks vs large chunks

Small chunk_sizeLarge chunk_size
Match precisionHighLower
Context in a chunkLittleMore
RiskSplits an answerDrags in unrelated text
Tokens per retrieved chunkFewMany

When to use a smaller or larger size

  • Smaller for dense reference text, where a precise match matters and answers are short.
  • Larger for flowing prose, where an answer spans several sentences.
  • Measure it on real questions rather than guessing; a later lesson counts hit rate for exactly this.
A warning you may see
Watch out. Metadata counts towards the chunk size. With very small chunks, LlamaIndex warns that the metadata leaves little room for text, which is one more reason to keep metadata short.
Try it yourself
  • Try chunk_size=128 with chunk_overlap=20 and compare the refund chunks.
  • Print nodes[0].ref_doc_id and compare it with a document's doc_id.
  • Count the nodes at sizes 64, 128 and 256.

Little by little, you're building something great.