1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What's different about containerising LLM applications and model servers?
30-second answerSay your answer out loud first, then reveal.
Practices
| Concern | Approach |
|---|---|
| Image size | Slim base images; inference engine + dependencies only; weights downloaded or mounted at startup |
| Weights | Model registry / object storage → local NVMe cache or a pre-populated volume; checksum verification |
| Startup time | Readiness probe only after the model is loaded and warmed up; pre-pulled images on GPU nodes; fast weight loading (safetensors, parallel download) |
| GPU compatibility | Match CUDA / driver versions; use GPU operator-managed nodes; pin engine versions |
| Resources | Request whole GPUs (or MIG slices); set shared memory size for tensor parallelism |
| Health | Liveness (process alive) vs readiness (model loaded, can serve); a test generation in the readiness check |
| Security | Non-root, minimal packages, signed images, no secrets baked in |
Why it matters: slow startup affects autoscaling and recovery. A pod that takes 8 minutes to become ready can't absorb a sudden spike.
Related
Every expert started right here.