Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q20IntermediateConcept

What metrics should you autoscale self-hosted LLM servers on, and why not CPU?

30-second answerSay your answer out loud first, then reveal.

Signals

MetricWhy
Waiting requests (queue depth)Direct indicator that capacity is insufficient
KV-cache utilisation %Near 100% means preemptions and evictions, so latency spikes
Running requests per replicaConcurrency vs tested capacity
p95 TTFT vs SLOUser-facing latency; scale before breaching
Tokens/s per replica vs benchmarked maxThroughput saturation

Scaling policies

  • Scale up early and fast (thresholds well below saturation), and scale down slowly (avoid thrashing; GPUs are expensive to re-acquire).
  • Predictive / scheduled scaling for known daily peaks (e.g. 10am IST spikes).
  • Minimum replicas sized for baseline load + N+1 redundancy.
  • Overflow routing to an API provider or smaller model while new replicas start.

Caveat: GPU availability itself can be the constraint (cloud capacity shortages), so reserved capacity or committed-use contracts may be necessary for critical services.

Every expert started right here.