1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you do capacity planning for a self-hosted LLM deployment?
30-second answerSay your answer out loud first, then reveal.
Worked approach
Peak load: 40 req/s
Avg tokens per request: 1,500 input + 300 output
Benchmarked capacity (1 replica = 2 GPUs, FP8, at p95 TTFT < 1s): 12 req/s
Replicas at peak: 40 / 12 = 3.3 → 4
Headroom (30%) + N+1: 4 × 1.3 = 5.2 → 6, plus 1 for failure = 7 replicas
GPUs: 7 × 2 = 14 GPUs
Total capacity: 7 replicas × 12 req/s = 84 req/s
Utilisation check: average load 15 req/s → 15 / 84 ≈ 18% average utilisationThat last line shows a key insight: sizing for peak leaves GPUs idle off-peak. Options: autoscaling (if GPU capacity is obtainable quickly), sharing GPUs with batch jobs off-peak, or overflowing peaks to an API provider.
Inputs to track: traffic growth forecasts, new features, model changes (a bigger model means fewer req/s per GPU), engine improvements (new versions often raise throughput), and GPU procurement lead times.
Related
This is what real progress feels like.