Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q32IntermediateConcept

What is speculative decoding, and why does it speed up inference without changing outputs?

30-second answerSay your answer out loud first, then reveal.
Speculative decoding: a draft model proposes four tokens, the target model verifies all four in one pass, the first three are kept, the fourth is replaced by the target's token, giving four tokens for about one target step.

Why it works: decoding is limited by reading the weights, not by arithmetic (Q20). Verifying k tokens in parallel reuses the same weight reads, so it's nearly free compared with k sequential steps.

Speedup depends on

  • Acceptance rate: how often the draft agrees with the target. High for predictable text (code, boilerplate, structured output), lower for creative text.
  • Draft cost: the draft must be much faster than the target.
  • Typical speedups are around 2–3x in favourable settings.

Variants

  • Separate draft model (same tokenizer required).
  • Medusa / EAGLE: extra lightweight heads on the target model predict future tokens.
  • Prompt lookup / n-gram decoding: copy candidate spans from the prompt, which is great for editing and RAG.

Trade-off: extra memory for the draft, and less benefit at very high batch sizes (where the GPU becomes compute-bound).

You understood something today that you didn't yesterday.