Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q4EasyConcept

Why do we chunk documents, and what are the common chunking strategies?

30-second answerSay your answer out loud first, then reveal.

Why chunk

  1. Precision: a vector for a whole document averages all its topics. A chunk vector represents one idea, so it matches specific questions better.
  2. Context budget: you want 5–10 relevant passages in the prompt, not 5 entire documents.
  3. Model limits: embedding models truncate long inputs.

Strategies

StrategyHowProsCons
Fixed-sizeEvery N tokens, with overlapSimple, predictableCuts sentences and ideas mid-way
RecursiveTry splitting on paragraphs, then sentences, then words until under NRespects natural boundaries; good defaultStill structure-blind
Structure-awareSplit on headings, sections, HTML/Markdown elements, pagesChunks map to real sections; great metadataNeeds good parsing; uneven sizes
SemanticEmbed sentences; split where similarity between neighbours dropsTopic-coherent chunksSlower, more compute; gains are often modest
Document-specificCode by function/class; tables kept whole; Q&A pairs kept togetherBest quality for that typeCustom work per format
LLM-based ("agentic")LLM decides boundaries or writes propositionsHighest coherenceExpensive at scale
Good default answer. "Start with structure-aware recursive splitting at roughly 300–800 tokens with 10–20% overlap, keep headings as metadata, keep tables and code blocks intact, then tune with retrieval evals."

Common mistakes

  • Choosing chunk size by gut feel without measuring retrieval recall.
  • Splitting a table row away from its header row, which makes the numbers meaningless.

This is what real progress feels like.