1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What are scaling laws? What did the Chinchilla paper show?
30-second answerSay your answer out loud first, then reveal.
Key ideas
- Loss ≈ a power law in N (params), D (tokens) and C (compute), so performance improves smoothly and predictably, which allows planning big training runs from small experiments.
- Training compute ≈ 6 × N × D FLOPs (forward + backward pass).
Chinchilla result
- Chinchilla (70B params, 1.4T tokens) outperformed Gopher (280B params, 300B tokens) with similar compute.
- Compute-optimal: D ≈ 20 × N.
Why modern models "over-train"
- Chinchilla optimises training compute, but a deployed model is run billions of times, so inference cost dominates.
- A smaller model trained on far more tokens (e.g. 8B params on ~15T tokens, almost 2,000 tokens per parameter) is cheaper to serve with good quality.
Caveats
- Loss improvements don't map linearly onto specific capabilities. Some abilities appear to emerge more abruptly (though this partly depends on the metric).
- Data quality and mixture matter as much as quantity. High-quality data is becoming a bottleneck (synthetic data, filtering).
- New scaling axes: test-time compute (Q15) and RL compute.
Related
Little by little, you're building something great.