1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your self-hosted vLLM servers start throwing out-of-memory errors or preempting requests under load. What do you do?
30-second answerSay your answer out loud first, then reveal.
Diagnosis
- Logs and metrics: KV-cache utilisation near 100%, preemption counts, OOM stack traces, the request size distribution at failure time.
- Traffic pattern: a few very long prompts (e.g. 100K tokens) can spike memory during prefill.
Levers
| Lever | Effect | Trade-off |
|---|---|---|
| Lower max concurrent sequences | Less KV memory pressure | Lower throughput |
| Lower max model length | Caps per-sequence KV | Rejects long inputs |
| Chunked prefill | Smooths prefill memory and latency | Slight overhead |
| KV-cache FP8 quantization | ~2x more cache capacity | Small quality risk; evaluate |
| Weight quantization | Frees memory for cache | Quality evals needed |
| Tensor parallelism / bigger GPUs | More memory per replica | Cost |
| Gateway limits on prompt size | Prevents pathological requests | Need a separate path for long docs |
| Route long-context requests to a dedicated pool | Isolation | More pools to manage |
Prevention: load tests with realistic token-length distributions (including long-tail requests), alerts on KV-cache utilisation and preemptions, and capacity headroom.
Related
This is what real progress feels like.