Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q16IntermediateConcept

Write the scaled dot-product attention formula. Why divide by √d_k, and what is the causal mask?

30-second answerSay your answer out loud first, then reveal.

Shapes (one head, sequence length n, head dim d_k):

  • Q, K, V: (n × d_k)
  • QKᵀ: (n × n), the score of every token pair
  • softmax row-wise, then × V → (n × d_k)

Why √d_k: if the components of q and k are independent with mean 0 and variance 1, then q·k has variance d_k. With d_k = 128, scores have a standard deviation of about 11, and softmax over such values is extremely peaked. Dividing by √d_k restores unit variance and healthy gradients.

Causal mask

text
        t1   t2   t3   t4
t1  [   0   -∞   -∞   -∞ ]
t2  [   0    0   -∞   -∞ ]
t3  [   0    0    0   -∞ ]
t4  [   0    0    0    0 ]
  • Training on a whole sequence in parallel predicts every next token at once. The mask stops position t from "cheating" by looking at t+1.
  • At inference, the same property enables the KV cache (Q18).

Padding mask: in batched inputs, also mask pad tokens so they get no attention.

Follow-ups to expect

  • Complexity of attention? O(n² · d) compute and O(n²) score memory per head without optimisations (Q19).

Little by little, you're building something great.