1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What metrics should you autoscale self-hosted LLM servers on, and why not CPU?
30-second answerSay your answer out loud first, then reveal.
Signals
| Metric | Why |
|---|---|
| Waiting requests (queue depth) | Direct indicator that capacity is insufficient |
| KV-cache utilisation % | Near 100% means preemptions and evictions, so latency spikes |
| Running requests per replica | Concurrency vs tested capacity |
| p95 TTFT vs SLO | User-facing latency; scale before breaching |
| Tokens/s per replica vs benchmarked max | Throughput saturation |
Scaling policies
- Scale up early and fast (thresholds well below saturation), and scale down slowly (avoid thrashing; GPUs are expensive to re-acquire).
- Predictive / scheduled scaling for known daily peaks (e.g. 10am IST spikes).
- Minimum replicas sized for baseline load + N+1 redundancy.
- Overflow routing to an API provider or smaller model while new replicas start.
Caveat: GPU availability itself can be the constraint (cloud capacity shortages), so reserved capacity or committed-use contracts may be necessary for critical services.
Related
Every expert started right here.