Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q6EasyConcept

What are the components of a transformer block in a modern LLM?

30-second answerSay your answer out loud first, then reveal.
A pre-norm transformer block: input hidden states go through RMSNorm and multi-head self-attention, added back to the input by a residual connection, then through RMSNorm and a SwiGLU feed-forward MLP with a second residual add, and on to the next block.

Roles

  • Attention: mixes information across positions (context).
  • FFN / MLP: transforms each position independently, with an expansion then a projection back down. The FFN holds roughly two-thirds of a dense model's parameters.
  • Residual connections: each sub-layer adds to the input. This makes training deep networks stable (gradient flow) and creates a "residual stream" that layers read from and write to.
  • Normalisation: keeps activations in a stable range. Pre-norm (normalise before each sub-layer) trains more stably than the original post-norm (Q47).

After the last block: a final norm → a linear projection to vocabulary size → logits.

Parameter intuition: for hidden size d, attention has ~4d² parameters (Q, K, V, O) and the FFN ~8d² (more with SwiGLU), per layer.

Little by little, you're building something great.