1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is speculative decoding, and why does it speed up inference without changing outputs?
30-second answerSay your answer out loud first, then reveal.

Why it works: decoding is limited by reading the weights, not by arithmetic (Q20). Verifying k tokens in parallel reuses the same weight reads, so it's nearly free compared with k sequential steps.
Speedup depends on
- Acceptance rate: how often the draft agrees with the target. High for predictable text (code, boilerplate, structured output), lower for creative text.
- Draft cost: the draft must be much faster than the target.
- Typical speedups are around 2–3x in favourable settings.
Variants
- Separate draft model (same tokenizer required).
- Medusa / EAGLE: extra lightweight heads on the target model predict future tokens.
- Prompt lookup / n-gram decoding: copy candidate spans from the prompt, which is great for editing and RAG.
Trade-off: extra memory for the draft, and less benefit at very high batch sizes (where the GPU becomes compute-bound).
Related
You understood something today that you didn't yesterday.