1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Break down the latency of a RAG request. How do you make it faster?
30-second answerSay your answer out loud first, then reveal.
Example latency budget (illustrative)
| Stage | Typical | Optimisations |
|---|---|---|
| Query rewrite (LLM) | 200–800ms | Small fast model; skip for clear queries |
| Query embedding | 20–100ms | Self-host near the app; cache frequent queries |
| Vector + BM25 search | 10–100ms | Run in parallel; tune ANN params (ef_search) |
| Rerank 50 candidates | 50–300ms | Smaller reranker; fewer candidates; GPU |
| LLM time-to-first-token | 200–1000ms+ | Fewer context tokens; prompt caching; faster model |
| Generation | 1–5s | Streaming; concise answers; smaller model if quality allows |
Key optimisations
- Streaming cuts perceived latency the most: users start reading within a second.
- Fewer context tokens: better reranking means fewer chunks, which lowers time-to-first-token and cost.
- Parallelism: run BM25, vector and multiple sub-queries concurrently.
- Caching:
Exact-match answer cache for frequent questions.
Semantic cache (similar questions → cached answer). Be careful with personalised, permissioned or time-sensitive answers.
Prompt caching for the static system prompt. - Routing: FAQs to cached answers; simple questions without rewriting; complex questions to the full pipeline.
Measure p50 and p95 per stage with tracing before optimising.
Related
This is what real progress feels like.