1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Explain the architecture choices of modern LLMs: pre-norm vs post-norm, RMSNorm, SwiGLU, RoPE, and GQA. Why did they win?
30-second answerSay your answer out loud first, then reveal.
| Component | Original Transformer (2017) | Modern recipe | Why |
|---|---|---|---|
| Norm placement | Post-norm (after residual add) | Pre-norm | Stable gradients in deep stacks; less warmup sensitivity |
| Norm type | LayerNorm (mean + variance) | RMSNorm (variance only, no mean-centring) | Faster, similar quality |
| FFN activation | ReLU | SwiGLU (Swish-gated linear unit) | Better quality; uses 3 matrices, with hidden size ~8/3·d to match params |
| Positions | Sinusoidal absolute | RoPE | Relative positions; context extension (Q40) |
| Attention heads | MHA | GQA | Smaller KV cache, faster decoding |
| Biases | Yes | Often removed | Stability, simplicity |
| Precision | FP32 | BF16 / FP8 mixed | Speed and memory |
Pre-norm vs post-norm
Post-norm: x = Norm(x + Sublayer(x))
Pre-norm: x = x + Sublayer(Norm(x))Pre-norm keeps a clean residual path (identity), so gradients flow directly through many layers.
SwiGLU: FFN(x) = W₂ · (Swish(W₁x) ⊙ W₃x). The gate lets the network modulate information multiplicatively.
Other modern additions to mention: QK-norm for stability, sliding-window/global attention hybrids, MoE FFNs, multi-token prediction objectives, MLA attention.
Related
You understood something today that you didn't yesterday.