1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Explain self-attention intuitively. What are queries, keys and values?
30-second answerSay your answer out loud first, then reveal.
Example: in "The animal didn't cross the street because it was too tired", the token "it" should attend strongly to "animal". Its query matches the key of "animal" more than the key of "street", so the representation of "it" absorbs information from "animal".

Steps
- Project each token vector into Q, K and V with learned matrices.
- Score every pair: q_i · k_j (how relevant token j is to token i).
- Scale by √d_k and apply the causal mask (in decoders a token can't see future tokens).
- Softmax each row into weights that sum to 1.
- Output_i = Σ_j weight_ij × v_j.
Why it was a breakthrough: unlike RNNs, every token connects directly to every other token (no long chains), and the computation is parallel across the sequence during training. That's what made scaling to huge datasets practical.
Follow-ups to expect
- Why divide by √d_k? (Q16.)
- What is multi-head attention? (Q17.)
Related
Every expert started right here.