1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you choose chunk size and overlap?
30-second answerSay your answer out loud first, then reveal.
Factors
- Question type:
• Factoid ("What's the notice period?") → smaller chunks (100–300 tokens) retrieve precisely.
• Explanatory or analytical ("Explain the escalation process") → larger chunks (500–1000) keep the reasoning together. - Document structure: FAQs are naturally small units. Legal clauses vary. Narrative docs need larger windows.
- Embedding model: its context length and how well it handles longer text.
- LLM context budget: k × chunk size must fit comfortably, along with instructions and chat history.
Overlap: repeating 10–20% of tokens between neighbouring chunks reduces the chance that an answer is split across a boundary. Too much overlap means duplicate content in results and a bigger index.
Decouple retrieval units from generation units: you don't have to choose one size. Retrieve on small chunks for precision, then pass the larger parent section to the LLM for context. This is the parent-child or small-to-big pattern (Q21) and often beats any single size.
How to tune (what interviewers want to hear)
for size in [256, 512, 1024]:
for overlap in [0, 64, 128]:
rebuild index → run eval set → measure recall@5, MRR, faithfulness
pick the best quality/cost trade-offCommon mistakes
- "512 tokens is the best size." There is no universal best; it depends on data and questions.
Related
Every expert started right here.