1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What does "7B" or "70B parameters" mean, and how much memory does a model need?
30-second answerSay your answer out loud first, then reveal.
Inference memory (weights only)
| Model | FP16/BF16 | INT8 | INT4 |
|---|---|---|---|
| 7–8B | ~14–16 GB | ~7–8 GB | ~4 GB |
| 13B | ~26 GB | ~13 GB | ~7 GB |
| 70B | ~140 GB | ~70 GB | ~35–40 GB |
Plus at inference:
- KV cache: grows with context length × concurrent requests and can exceed the weights at high concurrency (Q18, Q36).
- Activations, framework overhead, CUDA graphs: budget an extra 10–20%.
Training memory (full fine-tuning with Adam, mixed precision) ≈ 16 bytes/param
- BF16 weights (2) + BF16 gradients (2) + FP32 master weights (4) + Adam moments m and v (4 + 4).
- 7B → ~112 GB before activations, which is why full fine-tuning of even 7B models needs multiple GPUs or memory-saving tricks (ZeRO, gradient checkpointing). LoRA/QLoRA reduce this dramatically (Q22).
More parameters generally means more capability and knowledge, but also more cost and latency. A well-trained smaller model can beat an older larger one (training data and recipe matter, Q26).
Related
Slow is fine. Stopping is the only problem.