1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
With 1M+ token context windows, is RAG dead? Argue both sides.
30-second answerSay your answer out loud first, then reveal.
Arguments that long context reduces the need for RAG
- Small corpora (one handbook, one codebase module, one contract) fit entirely, and no retrieval errors are possible.
- Holistic reasoning across a whole document is better than across fragmented chunks.
- Simpler architecture: no chunking or index tuning.
- Prompt caching makes repeatedly sending the same large context cheaper.
Arguments that RAG remains essential
- Scale: enterprise corpora are billions of tokens, orders of magnitude beyond any window.
- Cost and latency: sending 1M tokens per query is slow and expensive compared with 5K well-chosen tokens, even with caching.
- Freshness: RAG indexes update continuously; there's no need to rebuild a giant context.
- Access control: each user must see a different subset, which retrieval filters handle naturally.
- Quality: models still miss or misuse information in very long contexts, especially with many similar distractors ("needle in a haystack" tests are easier than real multi-fact reasoning).
- Citations and auditability are cleaner with retrieved passages.
Synthesis (what senior candidates say)
- Corpus fits and changes rarely → long context (+ caching).
- Large or dynamic corpus → RAG, but with larger chunks or whole-document retrieval feeding a long-context model.
- Decide with an eval comparing quality, latency and cost on your workload.
Related
Slow is fine. Stopping is the only problem.