Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q47HardConcept

Explain the architecture choices of modern LLMs: pre-norm vs post-norm, RMSNorm, SwiGLU, RoPE, and GQA. Why did they win?

30-second answerSay your answer out loud first, then reveal.
ComponentOriginal Transformer (2017)Modern recipeWhy
Norm placementPost-norm (after residual add)Pre-normStable gradients in deep stacks; less warmup sensitivity
Norm typeLayerNorm (mean + variance)RMSNorm (variance only, no mean-centring)Faster, similar quality
FFN activationReLUSwiGLU (Swish-gated linear unit)Better quality; uses 3 matrices, with hidden size ~8/3·d to match params
PositionsSinusoidal absoluteRoPERelative positions; context extension (Q40)
Attention headsMHAGQASmaller KV cache, faster decoding
BiasesYesOften removedStability, simplicity
PrecisionFP32BF16 / FP8 mixedSpeed and memory

Pre-norm vs post-norm

text
Post-norm:  x = Norm(x + Sublayer(x))
Pre-norm:   x = x + Sublayer(Norm(x))

Pre-norm keeps a clean residual path (identity), so gradients flow directly through many layers.

SwiGLU: FFN(x) = W₂ · (Swish(W₁x) ⊙ W₃x). The gate lets the network modulate information multiplicatively.

Other modern additions to mention: QK-norm for stability, sliding-window/global attention hybrids, MoE FFNs, multi-token prediction objectives, MLA attention.

You understood something today that you didn't yesterday.