Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q44HardConcept

How do you optimise GPU costs for self-hosted LLM inference?

30-second answerSay your answer out loud first, then reveal.

Levers

LeverHow it saves
Smaller / distilled modelsFewer GPUs per token
Quantization (FP8/INT8/INT4)Fewer GPUs; higher throughput (memory-bound decode)
Engine efficiencyContinuous batching, paged KV cache, prefix caching, speculative decoding
Right GPU typeInference often benefits more from memory bandwidth and capacity than raw FLOPs; benchmark cost per token per GPU type
ConsolidationMIG slices or shared pools for small models; multi-LoRA instead of one deployment per fine-tune
Utilisation schedulingBatch jobs fill off-peak capacity
PurchasingReserved/committed for base load, on-demand for growth, spot for batch
Hybrid overflowAPI for rare peaks instead of idle GPUs
Prompt and output dietFewer tokens processed

Key metric: cost per 1M tokens served (or per request) per model, including idle capacity, tracked over time. Track utilisation honestly. A fleet at 25% average utilisation is usually the biggest waste.

This is what real progress feels like.