1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you run GPU inference workloads on Kubernetes?
30-second answerSay your answer out loud first, then reveal.

Key considerations
- Node pools: separate pools per GPU type (e.g. 80 GB vs 24 GB cards); taints so only GPU workloads land there.
- Pod sizing: one model replica may span multiple GPUs (tensor parallelism); request
nvidia.com/gpu: N; set/dev/shmsize. - GPU sharing: MIG partitions or time-slicing for small models (with caution for latency-sensitive workloads).
- Routing: least-outstanding-requests or KV-cache-aware routing beats round-robin for LLMs; prefix-cache affinity improves hit rates.
- Cold starts: node provisioning + image pull + weight loading can take 5–15 minutes, so keep minimum replicas, use predictive scaling, pre-pull images and cache weights.
- Disruption: PodDisruptionBudgets; graceful shutdown that drains in-flight streams.
- Observability: DCGM exporter for GPU metrics and inference engine metrics in Prometheus/Grafana.
Related
This is what real progress feels like.