Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q37HardSystem design

Design an LLM inference platform serving 1,000 requests per second across several models.

30-second answerSay your answer out loud first, then reveal.
LLM inference platform: clients go through an API gateway (which logs usage and billing) to a router that sends requests to three GPU pools or an external API fallback, while pool metrics feed an autoscaler that scales the pools.

1. Requirements and traffic model:

  • 1,000 RPS × average (input, output) tokens, e.g. 1,500 in / 300 out, gives 1.5M prefill tokens/s and 300K decode tokens/s.
  • SLAs: p95 TTFT < 1s, TPOT < 50ms. Peak-to-average ratio, multi-region needs.

2. Capacity planning: benchmark each model/GPU combination to find sustainable throughput at the SLA → replicas = peak load ÷ per-replica capacity + headroom.

3. Serving engine choices: continuous batching, PagedAttention, prefix caching, chunked prefill, quantization (FP8), speculative decoding for latency-sensitive routes. Possibly disaggregated prefill/decode pools for long-prompt traffic.

4. Routing:

  • Model routing: a cheap model for easy requests, a large one for hard ones (classifier or rules).
  • Load-aware: least queue depth / KV utilisation.
  • Cache-aware: send requests with the same long prefix to the same replica (prefix cache hits).
  • Fairness: per-tenant quotas, priority classes (interactive vs batch).

5. Autoscaling: scale on queue length, KV-cache utilisation and TTFT, not CPU. GPU cold starts (loading 100+ GB of weights) take minutes, so keep warm capacity, use fast weight loading, and predictive scaling.

6. Reliability: health checks, retries on another replica, circuit breakers, graceful degradation (fall back to a smaller model), multi-region.

7. Observability and cost: per-request token counts, TTFT/TPOT histograms, GPU utilisation, cost per 1M tokens per model, per-tenant usage.

You understood something today that you didn't yesterday.