Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q18IntermediateConcept

What is the KV cache? Why is it needed, and how do you calculate its memory?

30-second answerSay your answer out loud first, then reveal.
At step t only the new token's q, k and v are computed; k and v are appended to the KV cache, and attention compares q with all cached keys and sums cached values to produce the next-token logits.

Memory formula

text
KV bytes per token = 2 (K and V) × n_layers × n_kv_heads × head_dim × bytes_per_value
KV total = per_token × context_length × batch_size (concurrent sequences)

Worked example: a 70B Llama-style model with 80 layers, 8 KV heads (GQA), head_dim 128, FP16 (2 bytes):

  • Per token: 2 × 80 × 8 × 128 × 2 = 327,680 bytes ≈ 0.33 MB
  • One 8K-token sequence: ≈ 2.7 GB
  • 32 concurrent 8K sequences: ≈ 86 GB, on top of ~140 GB of FP16 weights.
  • Without GQA (64 KV heads) it would be 8x larger, about 690 GB, which is impractical.

Ways to reduce KV memory

  • GQA / MQA / MLA (fewer or compressed KV heads).
  • KV cache quantization (FP8 / INT8).
  • PagedAttention: allocate in blocks to avoid fragmentation and waste (Q33).
  • Sliding-window attention for some layers (bounded cache).
  • Prefix caching: share the cache for common prompt prefixes across requests.
Interview tip. Being able to do this calculation quickly is a strong signal for LLM infra and FDE roles.

Slow is fine. Stopping is the only problem.