1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How would you scale RAG to 100 million+ documents (billions of chunks)?
30-second answerSay your answer out loud first, then reveal.

Design areas
- Storage math: 2B chunks × 768 dims × 4 bytes ≈ 6 TB of raw vectors. In RAM, that's infeasible economically, so you need quantization (int8 → 1.5 TB, PQ or binary → much smaller) plus disk-based ANN or tiered memory.
- Sharding strategies:
• By tenant / domain: queries hit one shard (best when filters are natural).
• By hash: even load, but every query fans out to all shards (scatter-gather), so tail latency matters.
• By time: recent data hot, old data cold. - Two-tier retrieval: a cheap first stage (BM25 or binary vectors) over everything → full-precision rescoring → cross-encoder on the top 100.
- Indexing throughput: distributed embedding (GPU batch jobs), checkpointed pipelines, back-pressure. Re-embedding the whole corpus takes days and costs real money, so plan model migrations carefully (Q45).
- Freshness: a small, fast "delta" index for recent documents merged with the large static index; periodic compaction.
- Replication for QPS and availability; consistent snapshots.
- Document-level vs chunk-level retrieval: first retrieve documents (using summaries), then chunks within the top documents. This reduces the search space.
Common mistakes
- Designing one giant HNSW index in RAM without doing the memory math.
Related
Slow is fine. Stopping is the only problem.