Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q29IntermediateConcept

Break down the latency of a RAG request. How do you make it faster?

30-second answerSay your answer out loud first, then reveal.

Example latency budget (illustrative)

StageTypicalOptimisations
Query rewrite (LLM)200–800msSmall fast model; skip for clear queries
Query embedding20–100msSelf-host near the app; cache frequent queries
Vector + BM25 search10–100msRun in parallel; tune ANN params (ef_search)
Rerank 50 candidates50–300msSmaller reranker; fewer candidates; GPU
LLM time-to-first-token200–1000ms+Fewer context tokens; prompt caching; faster model
Generation1–5sStreaming; concise answers; smaller model if quality allows

Key optimisations

  1. Streaming cuts perceived latency the most: users start reading within a second.
  2. Fewer context tokens: better reranking means fewer chunks, which lowers time-to-first-token and cost.
  3. Parallelism: run BM25, vector and multiple sub-queries concurrently.
  4. Caching:
    Exact-match answer cache for frequent questions.
    Semantic cache (similar questions → cached answer). Be careful with personalised, permissioned or time-sensitive answers.
    Prompt caching for the static system prompt.
  5. Routing: FAQs to cached answers; simple questions without rewriting; complex questions to the full pipeline.

Measure p50 and p95 per stage with tracing before optimising.

This is what real progress feels like.