Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q38HardSystem design

Outline how you'd pretrain an LLM from scratch. What are the main stages and decisions?

30-second answerSay your answer out loud first, then reveal.
LLM pretraining pipeline: raw data is filtered, deduplicated, decontaminated, mixed and tokenized, then goes through distributed pretraining, annealing, post-training (SFT then preference/RL), and finally evals, red-teaming and release.

Key decisions

  1. Compute budget → size and tokens: use scaling laws. Decide whether to optimise for training (Chinchilla) or inference (smaller model, more tokens).
  2. Data (the biggest quality lever):
    • Quality filtering (classifier-based, e.g. "educational value" scorers), language ID, heuristic rules.
    • Deduplication (improves quality and reduces memorisation).
    • Mixture weights: web vs code vs maths vs multilingual; upsample high-quality sources.
  3. Tokenizer: vocabulary size, multilingual coverage (Q3), number handling, special tokens.
  4. Architecture: decoder-only with RMSNorm, SwiGLU, RoPE, GQA (Q47); dense vs MoE.
  5. Training setup: AdamW, warmup + cosine or WSD (warmup-stable-decay) learning rate schedule, BF16 mixed precision, large batch sizes, gradient clipping.
  6. Stability: watch for loss spikes; mitigations include lower LR, QK-norm, z-loss, skipping bad batches, restarting from a checkpoint.
  7. Infrastructure: parallelism strategy (Q39), fast checkpointing, failure handling (at thousands of GPUs, hardware failures happen frequently), throughput monitoring (MFU, model FLOPs utilisation).
  8. Evaluation during training: small benchmark suites at checkpoints to catch problems early.

Compute estimate: FLOPs ≈ 6 × N × D. For example, 8B params × 15T tokens ≈ 7.2 × 10²³ FLOPs.

Slow is fine. Stopping is the only problem.