Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q20IntermediateConcept

Explain the prefill and decode phases of LLM inference. What are TTFT and TPOT, and which phase is compute-bound vs memory-bound?

30-second answerSay your answer out loud first, then reveal.
Prefill processes the 2,000-token prompt in parallel and builds the KV cache, the user sees the first token after TTFT, then decode generates one token per step until the full answer after TPOT × output length.

Metrics

  • TTFT (time to first token): queueing + prefill. Grows with prompt length.
  • TPOT / ITL (time per output token / inter-token latency): decode step time.
  • End-to-end latency ≈ TTFT + TPOT × (output tokens − 1).
  • Throughput: total tokens/sec across all users (what you pay for).

Why decode is memory-bound: with batch size 1, generating each token requires reading all ~140 GB of a 70B FP16 model's weights from HBM for one token's worth of maths. GPU compute sits mostly idle. Batching many sequences shares each weight read across many tokens, which raises throughput.

Optimisation map

PhaseBottleneckLevers
PrefillComputeFlashAttention, prompt/prefix caching, chunked prefill, shorter prompts
DecodeMemory bandwidthBatching (continuous batching), quantization (fewer bytes to read), GQA, speculative decoding, higher-bandwidth GPUs

Advanced: disaggregated serving runs prefill and decode on separate GPU pools so long prompts don't stall others' decoding.

Every expert started right here.