1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Why do we chunk documents, and what are the common chunking strategies?
30-second answerSay your answer out loud first, then reveal.
Why chunk
- Precision: a vector for a whole document averages all its topics. A chunk vector represents one idea, so it matches specific questions better.
- Context budget: you want 5–10 relevant passages in the prompt, not 5 entire documents.
- Model limits: embedding models truncate long inputs.
Strategies
| Strategy | How | Pros | Cons |
|---|---|---|---|
| Fixed-size | Every N tokens, with overlap | Simple, predictable | Cuts sentences and ideas mid-way |
| Recursive | Try splitting on paragraphs, then sentences, then words until under N | Respects natural boundaries; good default | Still structure-blind |
| Structure-aware | Split on headings, sections, HTML/Markdown elements, pages | Chunks map to real sections; great metadata | Needs good parsing; uneven sizes |
| Semantic | Embed sentences; split where similarity between neighbours drops | Topic-coherent chunks | Slower, more compute; gains are often modest |
| Document-specific | Code by function/class; tables kept whole; Q&A pairs kept together | Best quality for that type | Custom work per format |
| LLM-based ("agentic") | LLM decides boundaries or writes propositions | Highest coherence | Expensive at scale |
Good default answer. "Start with structure-aware recursive splitting at roughly 300–800 tokens with 10–20% overlap, keep headings as metadata, keep tables and code blocks intact, then tune with retrieval evals."
Common mistakes
- Choosing chunk size by gut feel without measuring retrieval recall.
- Splitting a table row away from its header row, which makes the numbers meaningless.
Related
This is what real progress feels like.