Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q37HardSystem design

Design a self-hosted inference cluster on Kubernetes serving five open models (chat, code, embeddings, reranker, small classifier) with SLOs.

30-second answerSay your answer out loud first, then reveal.
Inference cluster: internal apps call an inference gateway that routes to classifier, reranker, embeddings, code and chat model pools (fed by a model registry and weights cache) or to an external API fallback on overflow; pools report to Prometheus and DCGM, which drive per-pool autoscalers.

SLO-driven sizing

Examples:

ModelSLOKey metric for scaling
Chat (large)TTFT p95 < 1.5s, 99.5% availabilityQueue depth, KV-cache %, TTFT
Code completion (small, fast)TTFT p95 < 300msIn-flight requests, TTFT
Embeddingsp95 < 100ms for a batch of 32Batch queue length
Rerankerp95 < 150ms for 50 docsQueue length
Classifierp95 < 50msRPS per replica

Operational design

  • Node pools per GPU type, with taints; a separate pool for batch jobs (spot) that can lend capacity off-peak.
  • Release management: blue-green for engine upgrades; canary for model version changes; per-model eval gates.
  • Multi-zone replicas with N+1 redundancy for chat and code.
  • Security: network policies, mTLS, no internet egress from model pods, signed images and weights.
  • Cost reporting: GPU-hours per model and per consuming team.

You understood something today that you didn't yesterday.