1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Outline how you'd pretrain an LLM from scratch. What are the main stages and decisions?
30-second answerSay your answer out loud first, then reveal.

Key decisions
- Compute budget → size and tokens: use scaling laws. Decide whether to optimise for training (Chinchilla) or inference (smaller model, more tokens).
- Data (the biggest quality lever):
• Quality filtering (classifier-based, e.g. "educational value" scorers), language ID, heuristic rules.
• Deduplication (improves quality and reduces memorisation).
• Mixture weights: web vs code vs maths vs multilingual; upsample high-quality sources. - Tokenizer: vocabulary size, multilingual coverage (Q3), number handling, special tokens.
- Architecture: decoder-only with RMSNorm, SwiGLU, RoPE, GQA (Q47); dense vs MoE.
- Training setup: AdamW, warmup + cosine or WSD (warmup-stable-decay) learning rate schedule, BF16 mixed precision, large batch sizes, gradient clipping.
- Stability: watch for loss spikes; mitigations include lower LR, QK-norm, z-loss, skipping bad batches, restarting from a checkpoint.
- Infrastructure: parallelism strategy (Q39), fast checkpointing, failure handling (at thousands of GPUs, hardware failures happen frequently), throughput monitoring (MFU, model FLOPs utilisation).
- Evaluation during training: small benchmark suites at checkpoints to catch problems early.
Compute estimate: FLOPs ≈ 6 × N × D. For example, 8B params × 15T tokens ≈ 7.2 × 10²³ FLOPs.
Related
Slow is fine. Stopping is the only problem.