Transformer architecture
The transformer architecture is an encoder-decoder design in which a stack of N = 6 identical encoder layers turns the input tokens into contextual vectors and a stack of 6 identical decoder layers generates the output from them, every layer built from attention and feed-forward sub-layers wrapped in residual connections and layer normalization.
Last updated: 07 Oct, 2026 · NumPy
The Transformers lesson gave the two reasons for the design, parallel processing and contextual vectors. This lesson opens the block diagram.
Stacking six encoders and six decoders
Seen from outside, the transformer is one block that takes an English sentence and returns its French translation, a sequence-to-sequence task. Inside, it follows the encoder-decoder architecture. The encoder side is not one encoder but several stacked ones: the input goes from one encoder to the next, bottom to top, and the result is passed to the decoder side, which is a stack of decoders. The paper uses six encoders and six decoders; six is the paper's choice, not a requirement. The example is "How are you" in and "Comment vas-tu ?" out.
Inside one encoder there are two layers: self-attention, then a feed-forward neural network. One decoder has a self-attention and a feed-forward network too, with an extra layer between them, encoder-decoder attention, where the decoder reads the encoder's output.
Adding residual connections and layer normalization
The drawing in the video keeps only the sub-layers. In the paper each sub-layer, attention or feed-forward, also has a residual connection around it followed by layer normalization, drawn as Add & Norm: the sub-layer's input is added to its output and the sum is normalized.
For the addition to work, every sub-layer and the embeddings produce vectors of one size, dmodel = 512. Normalizing after the addition is called post-norm. GPT-2 and most large language models since use pre-norm, x + Sublayer(LayerNorm(x)), which trains more stably in deep stacks. Both are covered in Residual connections and layer normalization.
Naming the three attention layers
The full diagram has three kinds of attention, all computed with Scaled dot-product attention:
- Encoder self-attention: queries, keys and values all come from the encoder's own states, so every input token looks at every other.
- Masked decoder self-attention: each output token looks only at the tokens before it, so the decoder cannot read the word it is about to predict (Masked self-attention).
- Encoder-decoder (cross) attention: the queries come from the decoder, the keys and values from the final encoder output (Cross-attention (encoder-decoder attention)).
At the bottom, the input tokens become embeddings and a positional encoding is added. At the top of the decoder, a linear layer and a softmax turn the last vector into probabilities over the vocabulary (Linear and softmax output layer).
Tracing shapes through the encoder stack
The example builds the encoder side at the paper's size: dmodel = 512, h = 8 attention heads of dk = 64, a feed-forward width dff = 2048 and N = 6 layers, each with its own random weights. Three random vectors stand in for "How are you" after embedding and positional encoding.
One encoder layer
def encoder_layer(x, p):
x = layer_norm(x + multi_head_attention(x, *p["attn"])) # sub-layer 1 + Add & Norm
ffn = np.maximum(0, x @ p["W1"]) @ p["W2"] # sub-layer 2: position-wise FFN
return layer_norm(x + ffn) # Add & NormSix layers in a row
for i in range(6):
x = encoder_layer(x, new_layer()) # (3, 512) in, (3, 512) outimport numpy as np
d_model, h, d_ff, N = 512, 8, 2048, 6 # the paper's base model
d_k = d_model // h
rng = np.random.default_rng(42)
def layer_norm(x, eps=1e-6):
return (x - x.mean(-1, keepdims=True)) / np.sqrt(x.var(-1, keepdims=True) + eps)
def multi_head_attention(x, Wq, Wk, Wv, Wo):
n = x.shape[0]
q, k, v = [(x @ W).reshape(n, h, d_k).transpose(1, 0, 2) for W in (Wq, Wk, Wv)]
s = q @ k.transpose(0, 2, 1) / np.sqrt(d_k) # (h, n, n)
w = np.exp(s - s.max(-1, keepdims=True))
w /= w.sum(-1, keepdims=True)
z = (w @ v).transpose(1, 0, 2).reshape(n, d_model) # concat the 8 heads
return z @ Wo
def encoder_layer(x, p):
x = layer_norm(x + multi_head_attention(x, *p["attn"])) # sub-layer 1 + Add & Norm
ffn = np.maximum(0, x @ p["W1"]) @ p["W2"] # sub-layer 2: position-wise FFN
return layer_norm(x + ffn) # Add & Norm
def new_layer():
m = lambda a, b: rng.normal(0, 1 / np.sqrt(a), (a, b))
return {"attn": [m(d_model, d_model) for _ in range(4)],
"W1": m(d_model, d_ff), "W2": m(d_ff, d_model)}
x = rng.normal(size=(3, d_model)) # "How are you": 3 token vectors (embedding + position)
print("input:", x.shape, "How starts", np.round(x[0, :3], 4))
for i in range(1, N + 1):
x = encoder_layer(x, new_layer()) # each of the 6 layers has its own weights
print(f"encoder {i}: {x.shape}, How starts {np.round(x[0, :3], 4)}, "
f"row mean {abs(x.mean(1)).max():.4f}, row std {x.std(1).mean():.4f}")input: (3, 512) How starts [ 0.3047 -1.04 0.7505] encoder 1: (3, 512), How starts [-0.5038 -1.2813 0.7583], row mean 0.0000, row std 1.0000 encoder 2: (3, 512), How starts [-1.0977 -1.1858 0.8107], row mean 0.0000, row std 1.0000 encoder 3: (3, 512), How starts [-0.785 0.4721 0.232 ], row mean 0.0000, row std 1.0000 encoder 4: (3, 512), How starts [-0.5324 1.3154 -0.4582], row mean 0.0000, row std 1.0000 encoder 5: (3, 512), How starts [-0.855 0.0403 0.0367], row mean 0.0000, row std 1.0000 encoder 6: (3, 512), How starts [ 1.0884 -0.0659 -1.2513], row mean 0.0000, row std 1.0000
What the six layers did to the vectors
- The shape stays (3, 512) from the input to encoder 6: one 512-number vector per token, which is what lets the layers stack and the residual additions work.
- The vector of How changes at every layer: its first three numbers go from [0.3047, −1.04, 0.7505] to [1.0884, −0.0659, −1.2513].
- Every row has mean 0.0000 and standard deviation 1.0000 after each layer, the work of the last layer normalization. In the paper the layer norm also has a learned gain and bias, which stay at 1 and 0 here.
Counting the weights of each layer
The same sizes give the number of trained weights in each part of a layer.
d_model, d_ff, N = 512, 2048, 6
mha = 4 * (d_model * d_model + d_model) # W_Q, W_K, W_V, W_O and their biases
ffn = d_model * d_ff + d_ff + d_ff * d_model + d_model
norm = 2 * d_model # layer norm gain and bias
encoder_layer = mha + ffn + 2 * norm # self-attention, FFN, 2 Add & Norm
decoder_layer = 2 * mha + ffn + 3 * norm # masked self-attn, cross-attn, FFN, 3 Add & Norm
print("multi-head attention:", f"{mha:,}")
print("feed-forward network:", f"{ffn:,}")
print("one encoder layer: ", f"{encoder_layer:,}")
print("one decoder layer: ", f"{decoder_layer:,}")
print("6 encoders + 6 decoders:", f"{N * (encoder_layer + decoder_layer):,}")multi-head attention: 1,050,624 feed-forward network: 2,099,712 one encoder layer: 3,152,384 one decoder layer: 4,204,032 6 encoders + 6 decoders: 44,138,496
- One encoder layer holds 3,152,384 weights: 1,050,624 in multi-head attention, 2,099,712 in the feed-forward network and 2,048 in two layer norms.
- A decoder layer holds 4,204,032, because it has a second attention sub-layer and a third layer norm.
- The 12 layers hold 44,138,496 weights. The paper reports 65 million for its base model; most of the rest sit in the token embeddings.
Encoder layer vs decoder layer
| Encoder layer | Decoder layer | |
|---|---|---|
| Sub-layers | self-attention, feed-forward | masked self-attention, cross-attention, feed-forward |
| Attends to | all input tokens | earlier output tokens, then all encoder outputs |
| Add & Norm blocks | 2 | 3 |
| Weights at the paper's size | 3,152,384 | 4,204,032 |
| Runs during generation | once per input | once per generated token |
Where you use the architecture
- The full encoder-decoder for translation and summarisation: the original transformer, T5, BART.
- The encoder stack alone for understanding tasks such as classification and NER: BERT.
- The decoder stack alone, without cross-attention, for text generation: GPT.
Related
- Previous: Transformers
- Next: Self-attention
- See also: Transformer encoder, Transformer decoder
- Reference: Vaswani et al., Attention Is All You Need (2017)
- Set
N = 2and check that the output shape is still (3, 512). - Change
d_ff = 2048tod_ff = 1024in the weight count and see how much smaller the feed-forward network gets. - Use
x = rng.normal(size=(10, d_model))for a 10-token sentence. Does the number of weights change?
You understood something today that you didn't yesterday.