Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q24IntermediateScenario

Your self-hosted vLLM servers start throwing out-of-memory errors or preempting requests under load. What do you do?

30-second answerSay your answer out loud first, then reveal.

Diagnosis

  • Logs and metrics: KV-cache utilisation near 100%, preemption counts, OOM stack traces, the request size distribution at failure time.
  • Traffic pattern: a few very long prompts (e.g. 100K tokens) can spike memory during prefill.

Levers

LeverEffectTrade-off
Lower max concurrent sequencesLess KV memory pressureLower throughput
Lower max model lengthCaps per-sequence KVRejects long inputs
Chunked prefillSmooths prefill memory and latencySlight overhead
KV-cache FP8 quantization~2x more cache capacitySmall quality risk; evaluate
Weight quantizationFrees memory for cacheQuality evals needed
Tensor parallelism / bigger GPUsMore memory per replicaCost
Gateway limits on prompt sizePrevents pathological requestsNeed a separate path for long docs
Route long-context requests to a dedicated poolIsolationMore pools to manage

Prevention: load tests with realistic token-length distributions (including long-tail requests), alerts on KV-cache utilisation and preemptions, and capacity headroom.

This is what real progress feels like.