Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q29IntermediateConcept

How do you do capacity planning for a self-hosted LLM deployment?

30-second answerSay your answer out loud first, then reveal.

Worked approach

text
Peak load:                 40 req/s
Avg tokens per request:    1,500 input + 300 output
Benchmarked capacity (1 replica = 2 GPUs, FP8, at p95 TTFT < 1s): 12 req/s
Replicas at peak:          40 / 12 = 3.3 → 4
Headroom (30%) + N+1:      4 × 1.3 = 5.2 → 6, plus 1 for failure = 7 replicas
GPUs:                      7 × 2 = 14 GPUs
Total capacity:            7 replicas × 12 req/s = 84 req/s
Utilisation check:         average load 15 req/s → 15 / 84 ≈ 18% average utilisation

That last line shows a key insight: sizing for peak leaves GPUs idle off-peak. Options: autoscaling (if GPU capacity is obtainable quickly), sharing GPUs with batch jobs off-peak, or overflowing peaks to an API provider.

Inputs to track: traffic growth forecasts, new features, model changes (a bigger model means fewer req/s per GPU), engine improvements (new versions often raise throughput), and GPU procurement lead times.

This is what real progress feels like.