1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What drives the cost of a RAG system, and how do you reduce it?
30-second answerSay your answer out loud first, then reveal.
Cost buckets and levers
| Bucket | Driver | Levers |
|---|---|---|
| Indexing | Embedding tokens; LLM enrichment (contextual retrieval, summaries) | Incremental updates only; batch APIs; prompt caching for enrichment; cheaper or self-hosted embedder |
| Storage | Number of vectors × dimensions × bytes; HNSW RAM | Quantization (int8/binary); Matryoshka truncation; disk-based indexes; dedupe |
| Query: retrieval | Embedding + search + rerank | Self-host embedder; limit rerank candidates |
| Query: generation | Context tokens × requests | Fewer chunks; compression; smaller model; caching; routing |
| Ops | Infra, monitoring, evals | Managed services vs self-hosting trade-off |
Back-of-envelope (interviewers love this)
- 10M chunks × 1024 dims × 4 bytes (float32) ≈ 41 GB of raw vectors, plus HNSW graph overhead.
- int8 quantization → ~10 GB; binary → ~1.3 GB (with rescoring to recover accuracy).
- Query side: 1M queries/month × 4K context tokens = 4B input tokens/month. Halving the context halves this cost.
Key insight: better retrieval precision reduces cost and improves quality, because you send fewer, more relevant tokens.
Related
Every expert started right here.