1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Design a semantic cache for an LLM-powered FAQ and support assistant.
30-second answerSay your answer out loud first, then reveal.

Key design decisions
- What to cache: only non-personalised, non-sensitive answers (FAQ-like). Skip queries with user-specific data ("my order").
- Threshold tuning: label 500 query pairs as same intent or different intent. Plot precision vs threshold, and choose θ for ≥ 99% precision. A wrong cached answer costs more than a miss.
- Near-duplicate traps: "How do I cancel my order?" vs "How do I cancel my subscription?" are similar embeddings but different answers. A cheap reranker or LLM verifier on candidate hits catches these.
- Scope keys: tenant, language, product, prompt version, model version.
- Invalidation: store source doc IDs with each answer; when a doc changes, purge the entries that cite it. Plus TTLs.
- Storage: vector DB or Redis with vector search; LRU eviction.
Metrics: hit rate, latency saved, cost saved, cached-answer error rate (sampled review), staleness incidents.
Expected gains: for FAQ-heavy traffic with many repeated questions, hit rates can be meaningful. Measure real traffic overlap before promising savings.
Related
Little by little, you're building something great.