Masked self-attention
Masked self-attention is self-attention with a mask applied to the scores before the softmax, so that each position attends only to the positions it is allowed to see: earlier tokens and itself, and never padding tokens.
Last updated: 07 Oct, 2026 · NumPy
It is the first sub-layer of every decoder layer. Its steps are those of the encoder's Multi-head attention (input embedding and positional encoding, linear projections for Q, K and V, scaled dot-product attention, multi-head attention, concatenation and the final linear projection, residual and layer normalization) with one step added: mask application.
Ignoring padding tokens with a padding mask
Masking helps manage the structure of the sequences being processed and makes the model behave correctly during training and inference. The first reason is variable-length sequences. Sentences in a batch have different lengths, so the shorter ones are padded with a padding token, 0, until they are all equally long. The video's example has the sequences [1, 2, 3] and [4, 5, 0].
Without a mask, the padding takes part in attention like a real word. One zero looks harmless, but with a maximum length of 100 and an output of two words, y1 and y2, 98 padding tokens would influence the attention and lead to incorrect or biased predictions. The padding mask marks real tokens with 1 and padding with 0: [1, 1, 1] for the first sequence and [1, 1, 0] for the second.
Attention scores form a matrix with one row per query (the token that is looking) and one column per key (the token being looked at). The padding mask blocks the padding column on every row, so for [4, 5, 0] the 2-D padding mask is [[1, 1, 0], [1, 1, 0], [1, 1, 0]]. Padding masks are used in the encoder's self-attention and in cross-attention as well, wherever padded sentences are attended to.
Hiding future tokens with a look-ahead mask
The second reason is the auto-regressive property of the decoder. The decoder produces its output one word at a time: for “How are you” it writes the Hindi words one by one. To predict a word it may use the words before it, never the words after it, because at inference those do not exist yet. The look-ahead mask (also called the causal mask) makes sure each position in the decoder's output sequence attends only to earlier positions and to itself, never to future positions. It matters for language modelling and translation, where the order of the sequence is the point.
For three positions the look-ahead mask is lower triangular, [[1, 0, 0], [1, 1, 0], [1, 1, 1]]: row 1 sees token 1, row 2 sees tokens 1 and 2, row 3 sees all three. The decoder uses both masks at once by multiplying them element by element. For [4, 5, 0] that gives [[1, 0, 0], [1, 1, 0], [1, 1, 0]].
Turning the mask into −∞ before the softmax
Wherever the combined mask is 0, −∞ is added to the score (adding −∞ or replacing the score with −∞ gives the same result). The reason is the next step, the softmax: e−∞ = 0, so a blocked position gets an attention weight of exactly 0. Padded tokens and, in the look-ahead case, future tokens then have no influence on the attention weights, and the weighted sum of the values ignores them.
Working the board's 4 × 4 example
The video's toy decoder input is the target [1, 2, 3] padded to [1, 2, 3, 0], each token a 4-dimensional embedding: [0.1, 0.2, 0.3, 0.4], [0.5, 0.6, 0.7, 0.8], [0.9, 1.0, 1.1, 1.2] and [0.0, 0.0, 0.0, 0.0]. To keep the arithmetic visible, the positional encoding is taken as zero and WQ = WK = WV = I, so Q = K = V = the embedding matrix. With dk = 4 the scores are QKᵀ/√4 = QKᵀ/2.
- QKᵀ row 1 is 0.1·0.1 + 0.2·0.2 + 0.3·0.3 + 0.4·0.4 = 0.3, then 0.7, 1.1 and 0. The full matrix is [[0.3, 0.7, 1.1, 0], [0.7, 1.74, 2.78, 0], [1.1, 2.78, 4.46, 0], [0, 0, 0, 0]].
- Scaled scores (÷ 2): [[0.15, 0.35, 0.55, 0], [0.35, 0.87, 1.39, 0], [0.55, 1.39, 2.23, 0], [0, 0, 0, 0]].
- Combined mask = look-ahead × key padding [1, 1, 1, 0] on every row: [[1, 0, 0, 0], [1, 1, 0, 0], [1, 1, 1, 0], [1, 1, 1, 0]].
- Softmax of the masked rows: [1, 0, 0, 0]; [0.3729, 0.6271, 0, 0]; [0.1152, 0.2668, 0.6180, 0]; [⅓, ⅓, ⅓, 0].
- Attention output = weights × V, one 4-value vector per position.
Masking the scores in NumPy
np.tril builds the look-ahead mask, broadcasting multiplies it by the key-padding row, and np.where puts −∞ wherever the mask is 0 before the softmax.
import numpy as np
# output embeddings of the target [1, 2, 3] padded to [1, 2, 3, 0]; W_Q = W_K = W_V = I, PE = 0
E = np.array([[0.1, 0.2, 0.3, 0.4],
[0.5, 0.6, 0.7, 0.8],
[0.9, 1.0, 1.1, 1.2],
[0.0, 0.0, 0.0, 0.0]])
Q = K = V = E
d_k = 4
scores = Q @ K.T / np.sqrt(d_k) # scaled scores
look_ahead = np.tril(np.ones((4, 4))) # 1 = may attend (earlier positions and itself)
key_padding = np.array([1, 1, 1, 0]) # the 4th token is padding
mask = look_ahead * key_padding # same key mask on every row
masked = np.where(mask == 1, scores, -np.inf) # blocked positions become -inf
weights = np.exp(masked - masked.max(axis=-1, keepdims=True))
weights = weights / weights.sum(axis=-1, keepdims=True)
output = weights @ V
np.set_printoptions(suppress=True)
print("Q K^T:\n", np.round(Q @ K.T, 2))
print("scaled scores:\n", np.round(scores, 4))
print("combined mask:\n", mask.astype(int))
print("attention weights:\n", np.round(weights, 4))
print("attention output:\n", np.round(output, 4))Q K^T: [[0.3 0.7 1.1 0. ] [0.7 1.74 2.78 0. ] [1.1 2.78 4.46 0. ] [0. 0. 0. 0. ]] scaled scores: [[0.15 0.35 0.55 0. ] [0.35 0.87 1.39 0. ] [0.55 1.39 2.23 0. ] [0. 0. 0. 0. ]] combined mask: [[1 0 0 0] [1 1 0 0] [1 1 1 0] [1 1 1 0]] attention weights: [[1. 0. 0. 0. ] [0.3729 0.6271 0. 0. ] [0.1152 0.2668 0.618 0. ] [0.3333 0.3333 0.3333 0. ]] attention output: [[0.1 0.2 0.3 0.4 ] [0.3509 0.4509 0.5509 0.6509] [0.7011 0.8011 0.9011 1.0011] [0.5 0.6 0.7 0.8 ]]
What the masked weights show
- Row 1 is [1, 0, 0, 0]: the first token may only look at itself.
- Row 2 is [0.3729, 0.6271, 0, 0], the softmax of 0.35 and 0.87; tokens 3 and 4 get exactly 0.
- Row 3 is [0.1152, 0.2668, 0.618, 0]; its output [0.7011, 0.8011, 0.9011, 1.0011] is mostly token 3's own value vector.
- The padding column is 0 on every row, so the padding vector never enters an output.
- Row 4, the padding query, averages tokens 1 to 3. Its output is never used: the loss ignores padded positions.
Reading a fully masked row and a multiplied mask
Two ways to get masking wrong, run on row 2's scores [0.35, 0.87, 1.39]:
import numpy as np
def masked_softmax(scores, mask, fill):
s = np.where(mask == 1, scores, fill)
e = np.exp(s - s.max(axis=-1, keepdims=True))
return e / e.sum(axis=-1, keepdims=True)
scores = np.array([[0.35, 0.87, 1.39]])
all_blocked = np.array([[0, 0, 0]]) # a row where nothing may be attended
with np.errstate(invalid="ignore"):
print("fill -inf: ", masked_softmax(scores, all_blocked, -np.inf))
print("fill -1e9: ", np.round(masked_softmax(scores, all_blocked, -1e9), 4))
print("multiply by mask instead:", np.round(masked_softmax(scores * np.array([[1, 1, 0]]), np.ones((1, 3)), 0), 4))
print("additive -inf mask: ", np.round(masked_softmax(scores, np.array([[1, 1, 0]]), -np.inf), 4))fill -inf: [[nan nan nan]] fill -1e9: [[0.3333 0.3333 0.3333]] multiply by mask instead: [[0.2953 0.4967 0.2081]] additive -inf mask: [[0.3729 0.6271 0. ]]
- A row with every position blocked becomes softmax(−∞, −∞, −∞) = NaN, and the NaN spreads through the next layers. This happens if a padding query's whole row is masked. Masking only the padding keys avoids it, and many libraries fill with a large negative number such as −1e9, which gives an even split instead of NaN.
- Multiplying the scores by the mask turns a blocked score into 0, and e⁰ = 1 still gets weight: the blocked third position keeps 0.2081 of the attention.
- The additive −∞ mask gives [0.3729, 0.6271, 0], the right answer.
Padding mask vs look-ahead mask
| Padding mask | Look-ahead mask | |
|---|---|---|
| Purpose | ignore padding tokens | keep generation auto-regressive |
| Shape | one row [1, 1, 0], repeated for every query | lower-triangular n × n |
| Depends on | the actual lengths in the batch | only the sequence length |
| Used in | encoder, decoder and cross-attention | decoder self-attention only |
| Blocks | columns of padding tokens | everything above the diagonal |
Where you use masked self-attention
- Every GPT-style model: each layer is masked self-attention plus a feed-forward network.
- Training a translation decoder on a whole target sentence at once.
- Batched inference with sentences of different lengths, where padding masks keep results identical to running each sentence alone.
Related
- Previous: Transformer encoder
- Next: Transformer decoder
- See also: Scaled dot-product attention
- Change
key_paddingto[1, 1, 0, 0]and check which weights become 0. - Replace
np.trilwithnp.ones((4, 4))(no look-ahead mask) and compare row 1 with the masked run. - Remove the division by
np.sqrt(d_k)and compare row 2 with [0.3729, 0.6271].
Every expert started right here.