Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Masked self-attention

Masked self-attention is self-attention with a mask applied to the scores before the softmax, so that each position attends only to the positions it is allowed to see: earlier tokens and itself, and never padding tokens.

Last updated: 07 Oct, 2026 · NumPy

It is the first sub-layer of every decoder layer. Its steps are those of the encoder's Multi-head attention (input embedding and positional encoding, linear projections for Q, K and V, scaled dot-product attention, multi-head attention, concatenation and the final linear projection, residual and layer normalization) with one step added: mask application.

Padding mask · from the Complete Transformers for NLP One Shot video · 4:03:07 to 4:07:50

Ignoring padding tokens with a padding mask

Masking helps manage the structure of the sequences being processed and makes the model behave correctly during training and inference. The first reason is variable-length sequences. Sentences in a batch have different lengths, so the shorter ones are padded with a padding token, 0, until they are all equally long. The video's example has the sequences [1, 2, 3] and [4, 5, 0].

Without a mask, the padding takes part in attention like a real word. One zero looks harmless, but with a maximum length of 100 and an output of two words, y1 and y2, 98 padding tokens would influence the attention and lead to incorrect or biased predictions. The padding mask marks real tokens with 1 and padding with 0: [1, 1, 1] for the first sequence and [1, 1, 0] for the second.

Attention scores form a matrix with one row per query (the token that is looking) and one column per key (the token being looked at). The padding mask blocks the padding column on every row, so for [4, 5, 0] the 2-D padding mask is [[1, 1, 0], [1, 1, 0], [1, 1, 0]]. Padding masks are used in the encoder's self-attention and in cross-attention as well, wherever padded sentences are attended to.

Look-ahead mask · from the Complete Transformers for NLP One Shot video · 4:09:44 to 4:13:28

Hiding future tokens with a look-ahead mask

The second reason is the auto-regressive property of the decoder. The decoder produces its output one word at a time: for “How are you” it writes the Hindi words one by one. To predict a word it may use the words before it, never the words after it, because at inference those do not exist yet. The look-ahead mask (also called the causal mask) makes sure each position in the decoder's output sequence attends only to earlier positions and to itself, never to future positions. It matters for language modelling and translation, where the order of the sequence is the point.

For three positions the look-ahead mask is lower triangular, [[1, 0, 0], [1, 1, 0], [1, 1, 1]]: row 1 sees token 1, row 2 sees tokens 1 and 2, row 3 sees all three. The decoder uses both masks at once by multiplying them element by element. For [4, 5, 0] that gives [[1, 0, 0], [1, 1, 0], [1, 1, 0]].

Why minus infinity · from the Complete Transformers for NLP One Shot video · 4:28:43 to 4:30:21

Turning the mask into −∞ before the softmax

Wherever the combined mask is 0, −∞ is added to the score (adding −∞ or replacing the score with −∞ gives the same result). The reason is the next step, the softmax: e−∞ = 0, so a blocked position gets an attention weight of exactly 0. Padded tokens and, in the look-ahead case, future tokens then have no influence on the attention weights, and the weighted sum of the values ignores them.

Masked scaled dot-product attention: M holds 0 where attention is allowed and −∞ where it is blocked

Working the board's 4 × 4 example

The video's toy decoder input is the target [1, 2, 3] padded to [1, 2, 3, 0], each token a 4-dimensional embedding: [0.1, 0.2, 0.3, 0.4], [0.5, 0.6, 0.7, 0.8], [0.9, 1.0, 1.1, 1.2] and [0.0, 0.0, 0.0, 0.0]. To keep the arithmetic visible, the positional encoding is taken as zero and WQ = WK = WV = I, so Q = K = V = the embedding matrix. With dk = 4 the scores are QKᵀ/√4 = QKᵀ/2.

  • QKᵀ row 1 is 0.1·0.1 + 0.2·0.2 + 0.3·0.3 + 0.4·0.4 = 0.3, then 0.7, 1.1 and 0. The full matrix is [[0.3, 0.7, 1.1, 0], [0.7, 1.74, 2.78, 0], [1.1, 2.78, 4.46, 0], [0, 0, 0, 0]].
  • Scaled scores (÷ 2): [[0.15, 0.35, 0.55, 0], [0.35, 0.87, 1.39, 0], [0.55, 1.39, 2.23, 0], [0, 0, 0, 0]].
  • Combined mask = look-ahead × key padding [1, 1, 1, 0] on every row: [[1, 0, 0, 0], [1, 1, 0, 0], [1, 1, 1, 0], [1, 1, 1, 0]].
  • Softmax of the masked rows: [1, 0, 0, 0]; [0.3729, 0.6271, 0, 0]; [0.1152, 0.2668, 0.6180, 0]; [⅓, ⅓, ⅓, 0].
  • Attention output = weights × V, one 4-value vector per position.
The look-ahead mask is lower triangular, the key-padding mask zeroes the padding column on every row, their product is the combined mask, and after minus infinity and softmax the attention weight rows are 1; 0.37 and 0.63; 0.12, 0.27 and 0.62; and one third three times.

Masking the scores in NumPy

np.tril builds the look-ahead mask, broadcasting multiplies it by the key-padding row, and np.where puts −∞ wherever the mask is 0 before the softmax.

ExampleFrom the video, run with NumPy
import numpy as np

# output embeddings of the target [1, 2, 3] padded to [1, 2, 3, 0]; W_Q = W_K = W_V = I, PE = 0
E = np.array([[0.1, 0.2, 0.3, 0.4],
              [0.5, 0.6, 0.7, 0.8],
              [0.9, 1.0, 1.1, 1.2],
              [0.0, 0.0, 0.0, 0.0]])
Q = K = V = E
d_k = 4

scores = Q @ K.T / np.sqrt(d_k)                       # scaled scores
look_ahead = np.tril(np.ones((4, 4)))                 # 1 = may attend (earlier positions and itself)
key_padding = np.array([1, 1, 1, 0])                  # the 4th token is padding
mask = look_ahead * key_padding                       # same key mask on every row

masked = np.where(mask == 1, scores, -np.inf)         # blocked positions become -inf
weights = np.exp(masked - masked.max(axis=-1, keepdims=True))
weights = weights / weights.sum(axis=-1, keepdims=True)
output = weights @ V

np.set_printoptions(suppress=True)
print("Q K^T:\n", np.round(Q @ K.T, 2))
print("scaled scores:\n", np.round(scores, 4))
print("combined mask:\n", mask.astype(int))
print("attention weights:\n", np.round(weights, 4))
print("attention output:\n", np.round(output, 4))

What the masked weights show

  • Row 1 is [1, 0, 0, 0]: the first token may only look at itself.
  • Row 2 is [0.3729, 0.6271, 0, 0], the softmax of 0.35 and 0.87; tokens 3 and 4 get exactly 0.
  • Row 3 is [0.1152, 0.2668, 0.618, 0]; its output [0.7011, 0.8011, 0.9011, 1.0011] is mostly token 3's own value vector.
  • The padding column is 0 on every row, so the padding vector never enters an output.
  • Row 4, the padding query, averages tokens 1 to 3. Its output is never used: the loss ignores padded positions.

Reading a fully masked row and a multiplied mask

Two ways to get masking wrong, run on row 2's scores [0.35, 0.87, 1.39]:

ExampleRun with NumPy
import numpy as np

def masked_softmax(scores, mask, fill):
    s = np.where(mask == 1, scores, fill)
    e = np.exp(s - s.max(axis=-1, keepdims=True))
    return e / e.sum(axis=-1, keepdims=True)

scores = np.array([[0.35, 0.87, 1.39]])
all_blocked = np.array([[0, 0, 0]])                    # a row where nothing may be attended

with np.errstate(invalid="ignore"):
    print("fill -inf: ", masked_softmax(scores, all_blocked, -np.inf))
print("fill -1e9: ", np.round(masked_softmax(scores, all_blocked, -1e9), 4))
print("multiply by mask instead:", np.round(masked_softmax(scores * np.array([[1, 1, 0]]), np.ones((1, 3)), 0), 4))
print("additive -inf mask:      ", np.round(masked_softmax(scores, np.array([[1, 1, 0]]), -np.inf), 4))
  • A row with every position blocked becomes softmax(−∞, −∞, −∞) = NaN, and the NaN spreads through the next layers. This happens if a padding query's whole row is masked. Masking only the padding keys avoids it, and many libraries fill with a large negative number such as −1e9, which gives an even split instead of NaN.
  • Multiplying the scores by the mask turns a blocked score into 0, and e⁰ = 1 still gets weight: the blocked third position keeps 0.2081 of the attention.
  • The additive −∞ mask gives [0.3729, 0.6271, 0], the right answer.

Padding mask vs look-ahead mask

Padding maskLook-ahead mask
Purposeignore padding tokenskeep generation auto-regressive
Shapeone row [1, 1, 0], repeated for every querylower-triangular n × n
Depends onthe actual lengths in the batchonly the sequence length
Used inencoder, decoder and cross-attentiondecoder self-attention only
Blockscolumns of padding tokenseverything above the diagonal

Where you use masked self-attention

  • Every GPT-style model: each layer is masked self-attention plus a feed-forward network.
  • Training a translation decoder on a whole target sentence at once.
  • Batched inference with sentences of different lengths, where padding masks keep results identical to running each sentence alone.
Watch out. Apply the mask to the scores before the softmax and with −∞ (or a large negative number), never by multiplying the scores by 0 and never after the softmax. Zeroing weights after the softmax leaves rows that no longer add up to 1.
Try it yourself
  • Change key_padding to [1, 1, 0, 0] and check which weights become 0.
  • Replace np.tril with np.ones((4, 4)) (no look-ahead mask) and compare row 1 with the masked run.
  • Remove the division by np.sqrt(d_k) and compare row 2 with [0.3729, 0.6271].

Every expert started right here.