Attention mechanism (Bahdanau and Luong)
The attention mechanism is a technique that lets a decoder build a new context vector at every output step, as a weighted sum of all the encoder's hidden states, with weights that score how well each source word matches what the decoder is about to write.
Last updated: 07 Oct, 2026 · NumPy
The Encoder-decoder (seq2seq) models lesson handed the decoder one context vector C for the whole sentence. Attention keeps every encoder state and lets each output word choose which ones to read.
Squeezing a long sentence into one vector
The video starts from the encoder-decoder diagram of the research paper: the encoder produces a vector w, the context vector, and passes it to the decoder, which writes the translation. Researchers tried this neural machine translation on sentences of different lengths. Short sentences were translated well; long ones, a hundred words for example, were not.
The plot of BLEU score against sentence length shows it. Up to about 10 words the score is high, and it keeps falling as sentences get longer. The reason is the vector w. The encoder reads the words one by one, as Word2Vec or embedding-layer vectors, but only its final state after the end of the sentence becomes w; its outputs at the other steps are not used. A short sentence fits in that one vector and a long one does not. The video compares it to asking a person to memorise a hundred English words and then translate them into French from memory.
Cho et al. (2014) measured this drop. Bahdanau, Cho and Bengio (2015) removed it with attention: their attention model, trained on sentences of up to 50 words, kept its BLEU score on sentences of 50 words and more, where the plain encoder-decoder's score fell.
Building a context vector for every output word
Bahdanau attention keeps all the encoder's states. The encoder is a bidirectional RNN, so each source word j gets an annotation hj = [→hj ; ←hj] that joins its forward and backward states. The diagram, redrawn from the attention figure in the Complete Transformers for NLP One Shot video, uses the sentence "Hello, What's Up". Before it writes output word i, the decoder:
- scores every annotation hj against its previous state si−1 with a small feed-forward network, the alignment model, giving the scores ei,j;
- turns the scores into weights αi,j with a softmax, so they are positive and sum to 1;
- takes the weighted sum ci = Σj αi,j hj, a context vector made for this word only;
- computes its new state si from si−1, the previous word yi−1 and ci, and predicts the word through a softmax.
The alignment model is trained together with the encoder and the decoder by backpropagation; nobody labels which source word belongs to which target word. Because the score adds two projected vectors inside the tanh, this form is called additive attention.
Scoring with Luong's dot, general and concat
Luong, Pham and Manning (2015) simplified the recipe. They score the decoder's current state ht, where Bahdanau uses the previous one, against each encoder state h̄s, and offer three score functions:
The weights and the context vector follow as before. The context is then joined with ht to give the attentional state h̃t = tanh(Wc[ct ; ht]), which predicts the word. Luong et al. also tested local attention, which looks only at a window of 2D + 1 source positions around an aligned position pt, against global attention over every position. The dot score needs no weights of its own, and it is the score the transformer later scales and uses everywhere.
Computing attention weights in NumPy
The example scores three encoder states with all three functions. For the states it borrows the board's vectors for The, cat and sat, [1, 0, 1, 0], [0, 1, 0, 1] and [1, 1, 1, 1]; the decoder state is s = [0.5, 1, 0, 0.5]. The additive and general weights are random with a fixed seed, so only the dot scores can be checked by hand: 0.5, 1.5 and 2.
Additive scores
# Bahdanau: a feed-forward net scores s against each encoder state h
e_additive = np.array([v_a @ np.tanh(W_a @ s + U_a @ h) for h in H])Dot and general scores
e_dot = H @ s # Luong dot: one dot product per h_j
e_general = np.array([s @ W_g @ h for h in H]) # Luong general: s^T W h_j
alpha = softmax(e_dot) # weights over The, cat, sat
c = alpha @ H # the context vectorimport numpy as np
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
H = np.array([[1, 0, 1, 0], # h1: the encoder state at "The"
[0, 1, 0, 1], # h2: at "cat"
[1, 1, 1, 1]], float) # h3: at "sat"
s = np.array([0.5, 1.0, 0.0, 0.5]) # the decoder state
rng = np.random.default_rng(42)
W_a = rng.normal(0, 0.5, (3, 4)) # Bahdanau: decoder state -> 3 hidden units
U_a = rng.normal(0, 0.5, (3, 4)) # Bahdanau: encoder state -> 3 hidden units
v_a = rng.normal(0, 0.5, 3)
W_g = rng.normal(0, 0.5, (4, 4)) # Luong general
scores = {
"additive (Bahdanau)": np.array([v_a @ np.tanh(W_a @ s + U_a @ h) for h in H]),
"dot (Luong)": H @ s,
"general (Luong)": np.array([s @ W_g @ h for h in H]),
}
for name, e in scores.items():
alpha = softmax(e) # attention weights over The, cat, sat
c = alpha @ H # context vector: weighted sum of h1, h2, h3
print(name)
print(" scores e =", np.round(e, 4))
print(" weights alpha =", np.round(alpha, 4), " sum =", round(alpha.sum(), 4))
print(" context c =", np.round(c, 4))additive (Bahdanau) scores e = [0.1547 0.0283 0.0645] weights alpha = [0.3578 0.3153 0.3269] sum = 1.0 context c = [0.6847 0.6422 0.6847 0.6422] dot (Luong) scores e = [0.5 1.5 2. ] weights alpha = [0.122 0.3315 0.5465] sum = 1.0 context c = [0.6685 0.878 0.6685 0.878 ] general (Luong) scores e = [-0.1126 0.7099 0.5973] weights alpha = [0.1883 0.4287 0.383 ] sum = 1.0 context c = [0.5713 0.8117 0.5713 0.8117]
What the three sets of weights show
- The dot scores 0.5, 1.5 and 2 become the weights 0.122, 0.3315 and 0.5465: sat matches s best, so it gets more than half of the weight.
- Every set of weights sums to 1.0, so each context vector is a weighted average of h1, h2 and h3 and has four numbers, like each of them.
- The additive weights 0.3578, 0.3153 and 0.3269 are nearly even, because the random Wa, Ua and va have learned nothing yet. Training is what makes the weights sharp.
- The context depends on s: a different decoder state makes a different source word win, which is how each output word reads its own part of the sentence.
Bahdanau vs Luong attention
| Bahdanau (2015) | Luong (2015) | |
|---|---|---|
| Decoder state scored | the previous state sᵢ₋₁ | the current state hₜ |
| Score function | additive: vₐᵀ tanh(Wₐs + Uₐh) | dot, general or concat |
| Encoder | bidirectional RNN | stacked LSTM |
| Where the context goes | into the RNN step that computes sᵢ | joined with hₜ after the RNN step |
| Source positions attended | all of them | all (global) or a window (local) |
Where you use attention
- Neural machine translation with RNNs, where it first lifted BLEU on long sentences.
- Reading the alignment: plotting α as a grid of target words against source words shows which source words each output word used.
- Image captioning and speech recognition, where the decoder attends over image regions or audio frames instead of words.
Related
- Previous: Encoder-decoder (seq2seq) models
- Next: Transformers
- Reference: Bahdanau, Cho and Bengio, Neural Machine Translation by Jointly Learning to Align and Translate
- Reference: Luong, Pham and Manning, Effective Approaches to Attention-based Neural Machine Translation
- Change the decoder state to
s = np.array([1.0, 0.0, 1.0, 0.0])and see which word gets the most dot-product weight. - Multiply
sby 10 and rerun: the dot-product weights become much sharper. Scaled dot-product attention is about this effect. - Set
W_g = np.eye(4)and check that the general scores now equal the dot scores.
Every expert started right here.