1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What are the components of a transformer block in a modern LLM?
30-second answerSay your answer out loud first, then reveal.

Roles
- Attention: mixes information across positions (context).
- FFN / MLP: transforms each position independently, with an expansion then a projection back down. The FFN holds roughly two-thirds of a dense model's parameters.
- Residual connections: each sub-layer adds to the input. This makes training deep networks stable (gradient flow) and creates a "residual stream" that layers read from and write to.
- Normalisation: keeps activations in a stable range. Pre-norm (normalise before each sub-layer) trains more stably than the original post-norm (Q47).
After the last block: a final norm → a linear projection to vocabulary size → logits.
Parameter intuition: for hidden size d, attention has ~4d² parameters (Q, K, V, O) and the FFN ~8d² (more with SwiGLU), per layer.
Related
Little by little, you're building something great.