Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q33IntermediateConcept

What are continuous batching and PagedAttention (vLLM), and why do they matter for serving?

30-second answerSay your answer out loud first, then reveal.

Static batching problem

text
Req A: ████████████████████ (500 tokens)
Req B: ███ (50 tokens) ... idle until A finishes
Req C: waiting in queue until the whole batch finishes

Continuous batching: at every step the scheduler checks which sequences finished, frees their slots, and inserts waiting requests (including their prefill). The result is higher utilisation, lower queueing delay and higher throughput.

PagedAttention (Kwon et al. 2023, vLLM)

  • Naive serving pre-allocates KV memory for the maximum sequence length per request, which wastes most of it (internal fragmentation) and leaves gaps (external fragmentation).
  • PagedAttention stores the KV cache in small blocks (e.g. 16 tokens) via a block table, allocated as the sequence grows.
  • Enables sharing blocks across sequences: common prefixes (system prompts), parallel samples, beam search (copy-on-write).
  • Prefix caching: reuse the KV cache for repeated prompt prefixes (system prompts, few-shot examples, RAG documents).
  • Chunked prefill: split long prompts so they don't stall ongoing decodes.
  • Multi-LoRA serving: many adapters on one base model.

Engines: vLLM, SGLang, TensorRT-LLM, TGI, plus llama.cpp/Ollama for local use.

Slow is fine. Stopping is the only problem.