1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Explain the prefill and decode phases of LLM inference. What are TTFT and TPOT, and which phase is compute-bound vs memory-bound?
30-second answerSay your answer out loud first, then reveal.

Metrics
- TTFT (time to first token): queueing + prefill. Grows with prompt length.
- TPOT / ITL (time per output token / inter-token latency): decode step time.
- End-to-end latency ≈ TTFT + TPOT × (output tokens − 1).
- Throughput: total tokens/sec across all users (what you pay for).
Why decode is memory-bound: with batch size 1, generating each token requires reading all ~140 GB of a 70B FP16 model's weights from HBM for one token's worth of maths. GPU compute sits mostly idle. Batching many sequences shares each weight read across many tokens, which raises throughput.
Optimisation map
| Phase | Bottleneck | Levers |
|---|---|---|
| Prefill | Compute | FlashAttention, prompt/prefix caching, chunked prefill, shorter prompts |
| Decode | Memory bandwidth | Batching (continuous batching), quantization (fewer bytes to read), GQA, speculative decoding, higher-bandwidth GPUs |
Advanced: disaggregated serving runs prefill and decode on separate GPU pools so long prompts don't stall others' decoding.
Related
Every expert started right here.