Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q16IntermediateSystem design

Design a semantic cache for an LLM-powered FAQ and support assistant.

30-second answerSay your answer out loud first, then reveal.
Semantic cache flow: an embedded query is matched against a scoped vector index of cached queries; a hit above the similarity threshold (optionally verified for same intent) returns the cached answer, otherwise the full RAG and LLM pipeline runs and the new answer is stored, and updated docs invalidate the entries that cite them.

Key design decisions

  1. What to cache: only non-personalised, non-sensitive answers (FAQ-like). Skip queries with user-specific data ("my order").
  2. Threshold tuning: label 500 query pairs as same intent or different intent. Plot precision vs threshold, and choose θ for ≥ 99% precision. A wrong cached answer costs more than a miss.
  3. Near-duplicate traps: "How do I cancel my order?" vs "How do I cancel my subscription?" are similar embeddings but different answers. A cheap reranker or LLM verifier on candidate hits catches these.
  4. Scope keys: tenant, language, product, prompt version, model version.
  5. Invalidation: store source doc IDs with each answer; when a doc changes, purge the entries that cite it. Plus TTLs.
  6. Storage: vector DB or Redis with vector search; LRU eviction.

Metrics: hit rate, latency saved, cost saved, cached-answer error rate (sampled review), staleness incidents.

Expected gains: for FAQ-heavy traffic with many repeated questions, hit rates can be meaningful. Measure real traffic overlap before promising savings.

Little by little, you're building something great.