Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q5EasyConcept

How do you choose chunk size and overlap?

30-second answerSay your answer out loud first, then reveal.

Factors

  1. Question type:
    • Factoid ("What's the notice period?") → smaller chunks (100–300 tokens) retrieve precisely.
    • Explanatory or analytical ("Explain the escalation process") → larger chunks (500–1000) keep the reasoning together.
  2. Document structure: FAQs are naturally small units. Legal clauses vary. Narrative docs need larger windows.
  3. Embedding model: its context length and how well it handles longer text.
  4. LLM context budget: k × chunk size must fit comfortably, along with instructions and chat history.

Overlap: repeating 10–20% of tokens between neighbouring chunks reduces the chance that an answer is split across a boundary. Too much overlap means duplicate content in results and a bigger index.

Decouple retrieval units from generation units: you don't have to choose one size. Retrieve on small chunks for precision, then pass the larger parent section to the LLM for context. This is the parent-child or small-to-big pattern (Q21) and often beats any single size.

How to tune (what interviewers want to hear)

python
for size in [256, 512, 1024]:
    for overlap in [0, 64, 128]:
        rebuild index → run eval set → measure recall@5, MRR, faithfulness
pick the best quality/cost trade-off

Common mistakes

  • "512 tokens is the best size." There is no universal best; it depends on data and questions.

Every expert started right here.