Design an LLM inference platform serving 1,000 requests per second across several models.

1. Requirements and traffic model:
- 1,000 RPS × average (input, output) tokens, e.g. 1,500 in / 300 out, gives 1.5M prefill tokens/s and 300K decode tokens/s.
- SLAs: p95 TTFT < 1s, TPOT < 50ms. Peak-to-average ratio, multi-region needs.
2. Capacity planning: benchmark each model/GPU combination to find sustainable throughput at the SLA → replicas = peak load ÷ per-replica capacity + headroom.
3. Serving engine choices: continuous batching, PagedAttention, prefix caching, chunked prefill, quantization (FP8), speculative decoding for latency-sensitive routes. Possibly disaggregated prefill/decode pools for long-prompt traffic.
4. Routing:
- Model routing: a cheap model for easy requests, a large one for hard ones (classifier or rules).
- Load-aware: least queue depth / KV utilisation.
- Cache-aware: send requests with the same long prefix to the same replica (prefix cache hits).
- Fairness: per-tenant quotas, priority classes (interactive vs batch).
5. Autoscaling: scale on queue length, KV-cache utilisation and TTFT, not CPU. GPU cold starts (loading 100+ GB of weights) take minutes, so keep warm capacity, use fast weight loading, and predictive scaling.
6. Reliability: health checks, retries on another replica, circuit breakers, graceful degradation (fall back to a smaller model), multi-region.
7. Observability and cost: per-request token counts, TTFT/TPOT histograms, GPU utilisation, cost per 1M tokens per model, per-tenant usage.
Related
You understood something today that you didn't yesterday.