Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q26IntermediateConcept

What are scaling laws? What did the Chinchilla paper show?

30-second answerSay your answer out loud first, then reveal.

Key ideas

  • Loss ≈ a power law in N (params), D (tokens) and C (compute), so performance improves smoothly and predictably, which allows planning big training runs from small experiments.
  • Training compute ≈ 6 × N × D FLOPs (forward + backward pass).

Chinchilla result

  • Chinchilla (70B params, 1.4T tokens) outperformed Gopher (280B params, 300B tokens) with similar compute.
  • Compute-optimal: D ≈ 20 × N.

Why modern models "over-train"

  • Chinchilla optimises training compute, but a deployed model is run billions of times, so inference cost dominates.
  • A smaller model trained on far more tokens (e.g. 8B params on ~15T tokens, almost 2,000 tokens per parameter) is cheaper to serve with good quality.

Caveats

  • Loss improvements don't map linearly onto specific capabilities. Some abilities appear to emerge more abruptly (though this partly depends on the metric).
  • Data quality and mixture matter as much as quantity. High-quality data is becoming a bottleneck (synthetic data, filtering).
  • New scaling axes: test-time compute (Q15) and RL compute.

Little by little, you're building something great.