Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q19IntermediateConcept

How do you run GPU inference workloads on Kubernetes?

30-second answerSay your answer out loud first, then reveal.
GPU inference on Kubernetes: an ingress feeds a load-aware model service that routes to vLLM pods loading weights from object storage into a node NVMe cache; pod metrics (queue depth, KV-cache %, TTFT, tokens/s) drive a KEDA or custom HPA autoscaler that scales the pods and triggers a node autoscaler for the GPU node pool.

Key considerations

  1. Node pools: separate pools per GPU type (e.g. 80 GB vs 24 GB cards); taints so only GPU workloads land there.
  2. Pod sizing: one model replica may span multiple GPUs (tensor parallelism); request nvidia.com/gpu: N; set /dev/shm size.
  3. GPU sharing: MIG partitions or time-slicing for small models (with caution for latency-sensitive workloads).
  4. Routing: least-outstanding-requests or KV-cache-aware routing beats round-robin for LLMs; prefix-cache affinity improves hit rates.
  5. Cold starts: node provisioning + image pull + weight loading can take 5–15 minutes, so keep minimum replicas, use predictive scaling, pre-pull images and cache weights.
  6. Disruption: PodDisruptionBudgets; graceful shutdown that drains in-flight streams.
  7. Observability: DCGM exporter for GPU metrics and inference engine metrics in Prometheus/Grafana.

This is what real progress feels like.