Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q30IntermediateConcept

What drives the cost of a RAG system, and how do you reduce it?

30-second answerSay your answer out loud first, then reveal.

Cost buckets and levers

BucketDriverLevers
IndexingEmbedding tokens; LLM enrichment (contextual retrieval, summaries)Incremental updates only; batch APIs; prompt caching for enrichment; cheaper or self-hosted embedder
StorageNumber of vectors × dimensions × bytes; HNSW RAMQuantization (int8/binary); Matryoshka truncation; disk-based indexes; dedupe
Query: retrievalEmbedding + search + rerankSelf-host embedder; limit rerank candidates
Query: generationContext tokens × requestsFewer chunks; compression; smaller model; caching; routing
OpsInfra, monitoring, evalsManaged services vs self-hosting trade-off

Back-of-envelope (interviewers love this)

  • 10M chunks × 1024 dims × 4 bytes (float32) ≈ 41 GB of raw vectors, plus HNSW graph overhead.
  • int8 quantization → ~10 GB; binary → ~1.3 GB (with rescoring to recover accuracy).
  • Query side: 1M queries/month × 4K context tokens = 4B input tokens/month. Halving the context halves this cost.

Key insight: better retrieval precision reduces cost and improves quality, because you send fewer, more relevant tokens.

Every expert started right here.