1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you optimise GPU costs for self-hosted LLM inference?
30-second answerSay your answer out loud first, then reveal.
Levers
| Lever | How it saves |
|---|---|
| Smaller / distilled models | Fewer GPUs per token |
| Quantization (FP8/INT8/INT4) | Fewer GPUs; higher throughput (memory-bound decode) |
| Engine efficiency | Continuous batching, paged KV cache, prefix caching, speculative decoding |
| Right GPU type | Inference often benefits more from memory bandwidth and capacity than raw FLOPs; benchmark cost per token per GPU type |
| Consolidation | MIG slices or shared pools for small models; multi-LoRA instead of one deployment per fine-tune |
| Utilisation scheduling | Batch jobs fill off-peak capacity |
| Purchasing | Reserved/committed for base load, on-demand for growth, spot for batch |
| Hybrid overflow | API for rare peaks instead of idle GPUs |
| Prompt and output diet | Fewer tokens processed |
Key metric: cost per 1M tokens served (or per request) per model, including idle capacity, tracked over time. Track utilisation honestly. A fleet at 25% average utilisation is usually the biggest waste.
Related
This is what real progress feels like.