1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Write the scaled dot-product attention formula. Why divide by √d_k, and what is the causal mask?
30-second answerSay your answer out loud first, then reveal.
Shapes (one head, sequence length n, head dim d_k):
- Q, K, V: (n × d_k)
- QKᵀ: (n × n), the score of every token pair
- softmax row-wise, then × V → (n × d_k)
Why √d_k: if the components of q and k are independent with mean 0 and variance 1, then q·k has variance d_k. With d_k = 128, scores have a standard deviation of about 11, and softmax over such values is extremely peaked. Dividing by √d_k restores unit variance and healthy gradients.
Causal mask
t1 t2 t3 t4
t1 [ 0 -∞ -∞ -∞ ]
t2 [ 0 0 -∞ -∞ ]
t3 [ 0 0 0 -∞ ]
t4 [ 0 0 0 0 ]- Training on a whole sequence in parallel predicts every next token at once. The mask stops position t from "cheating" by looking at t+1.
- At inference, the same property enables the KV cache (Q18).
Padding mask: in batched inputs, also mask pad tokens so they get no attention.
Follow-ups to expect
- Complexity of attention? O(n² · d) compute and O(n²) score memory per head without optimisations (Q19).
Related
Little by little, you're building something great.