Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q6EasyConcept

What kinds of caching can you use in LLM systems?

30-second answerSay your answer out loud first, then reveal.
A request checks the exact cache, then the semantic cache within the same scope, and only on two misses goes through the retrieval cache and an LLM call with prefix caching before results are stored with TTL and scope keys.

Cache types and their caveats

CacheSavesRisk / caveat
Exact responseFull LLM callLow hit rate for free-text; good for FAQs and repeated workflows
SemanticFull LLM callWrong answer if "similar" isn't equivalent ("cancel order" vs "cancel subscription"); tune threshold
Prompt / prefix (provider or vLLM)Input token cost + TTFTPrefix must be byte-identical; put static content first
EmbeddingEmbedding callsInvalidate when the model changes
Retrieval / tool resultsDB/API latencyStaleness; TTLs

Cache key scope (critical): include tenant, user permissions, locale, model version and prompt version. Never serve user A's personalised or permissioned answer to user B.

When not to cache: personalised content, time-sensitive answers ("today's balance"), high-stakes outputs without revalidation.

Little by little, you're building something great.