Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q43HardConcept

With 1M+ token context windows, is RAG dead? Argue both sides.

30-second answerSay your answer out loud first, then reveal.

Arguments that long context reduces the need for RAG

  • Small corpora (one handbook, one codebase module, one contract) fit entirely, and no retrieval errors are possible.
  • Holistic reasoning across a whole document is better than across fragmented chunks.
  • Simpler architecture: no chunking or index tuning.
  • Prompt caching makes repeatedly sending the same large context cheaper.

Arguments that RAG remains essential

  1. Scale: enterprise corpora are billions of tokens, orders of magnitude beyond any window.
  2. Cost and latency: sending 1M tokens per query is slow and expensive compared with 5K well-chosen tokens, even with caching.
  3. Freshness: RAG indexes update continuously; there's no need to rebuild a giant context.
  4. Access control: each user must see a different subset, which retrieval filters handle naturally.
  5. Quality: models still miss or misuse information in very long contexts, especially with many similar distractors ("needle in a haystack" tests are easier than real multi-fact reasoning).
  6. Citations and auditability are cleaner with retrieved passages.

Synthesis (what senior candidates say)

  • Corpus fits and changes rarely → long context (+ caching).
  • Large or dynamic corpus → RAG, but with larger chunks or whole-document retrieval feeding a long-context model.
  • Decide with an eval comparing quality, latency and cost on your workload.

Slow is fine. Stopping is the only problem.