Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Transformer architecture

The transformer architecture is an encoder-decoder design in which a stack of N = 6 identical encoder layers turns the input tokens into contextual vectors and a stack of 6 identical decoder layers generates the output from them, every layer built from attention and feed-forward sub-layers wrapped in residual connections and layer normalization.

Last updated: 07 Oct, 2026 · NumPy

The Transformers lesson gave the two reasons for the design, parallel processing and contextual vectors. This lesson opens the block diagram.

Encoder and decoder stacks · from the Complete Transformers for NLP One Shot video · 24:17 to 28:56

Stacking six encoders and six decoders

Seen from outside, the transformer is one block that takes an English sentence and returns its French translation, a sequence-to-sequence task. Inside, it follows the encoder-decoder architecture. The encoder side is not one encoder but several stacked ones: the input goes from one encoder to the next, bottom to top, and the result is passed to the decoder side, which is a stack of decoders. The paper uses six encoders and six decoders; six is the paper's choice, not a requirement. The example is "How are you" in and "Comment vas-tu ?" out.

Inside one encoder there are two layers: self-attention, then a feed-forward neural network. One decoder has a self-attention and a feed-forward network too, with an extra layer between them, encoder-decoder attention, where the decoder reads the encoder's output.

How are you ? enters a stack of 6 encoders; the output of encoder 6 goes to every one of the 6 decoders, which read the French so far and produce Comment vas-tu ?. One encoder holds self-attention and a feed-forward network; one decoder holds masked self-attention, encoder-decoder attention and a feed-forward network.

Adding residual connections and layer normalization

The drawing in the video keeps only the sub-layers. In the paper each sub-layer, attention or feed-forward, also has a residual connection around it followed by layer normalization, drawn as Add & Norm: the sub-layer's input is added to its output and the sum is normalized.

Add & Norm around every sub-layer (post-norm, as in the paper)

For the addition to work, every sub-layer and the embeddings produce vectors of one size, dmodel = 512. Normalizing after the addition is called post-norm. GPT-2 and most large language models since use pre-norm, x + Sublayer(LayerNorm(x)), which trains more stably in deep stacks. Both are covered in Residual connections and layer normalization.

The full transformer: input embedding plus positional encoding into N encoder layers of multi-head self-attention and feed-forward, each with Add and Norm and a residual path; output embedding of the target shifted right plus positional encoding into N decoder layers of masked self-attention, cross-attention with keys and values from the encoder, and feed-forward, each with Add and Norm; then a linear layer and a softmax.

Naming the three attention layers

The full diagram has three kinds of attention, all computed with Scaled dot-product attention:

  • Encoder self-attention: queries, keys and values all come from the encoder's own states, so every input token looks at every other.
  • Masked decoder self-attention: each output token looks only at the tokens before it, so the decoder cannot read the word it is about to predict (Masked self-attention).
  • Encoder-decoder (cross) attention: the queries come from the decoder, the keys and values from the final encoder output (Cross-attention (encoder-decoder attention)).

At the bottom, the input tokens become embeddings and a positional encoding is added. At the top of the decoder, a linear layer and a softmax turn the last vector into probabilities over the vocabulary (Linear and softmax output layer).

Tracing shapes through the encoder stack

The example builds the encoder side at the paper's size: dmodel = 512, h = 8 attention heads of dk = 64, a feed-forward width dff = 2048 and N = 6 layers, each with its own random weights. Three random vectors stand in for "How are you" after embedding and positional encoding.

One encoder layer

python
def encoder_layer(x, p):
    x = layer_norm(x + multi_head_attention(x, *p["attn"]))   # sub-layer 1 + Add & Norm
    ffn = np.maximum(0, x @ p["W1"]) @ p["W2"]                # sub-layer 2: position-wise FFN
    return layer_norm(x + ffn)                                # Add & Norm

Six layers in a row

python
for i in range(6):
    x = encoder_layer(x, new_layer())   # (3, 512) in, (3, 512) out
ExampleRun on NumPy 2.5
import numpy as np

d_model, h, d_ff, N = 512, 8, 2048, 6      # the paper's base model
d_k = d_model // h
rng = np.random.default_rng(42)

def layer_norm(x, eps=1e-6):
    return (x - x.mean(-1, keepdims=True)) / np.sqrt(x.var(-1, keepdims=True) + eps)

def multi_head_attention(x, Wq, Wk, Wv, Wo):
    n = x.shape[0]
    q, k, v = [(x @ W).reshape(n, h, d_k).transpose(1, 0, 2) for W in (Wq, Wk, Wv)]
    s = q @ k.transpose(0, 2, 1) / np.sqrt(d_k)              # (h, n, n)
    w = np.exp(s - s.max(-1, keepdims=True))
    w /= w.sum(-1, keepdims=True)
    z = (w @ v).transpose(1, 0, 2).reshape(n, d_model)       # concat the 8 heads
    return z @ Wo

def encoder_layer(x, p):
    x = layer_norm(x + multi_head_attention(x, *p["attn"]))  # sub-layer 1 + Add & Norm
    ffn = np.maximum(0, x @ p["W1"]) @ p["W2"]               # sub-layer 2: position-wise FFN
    return layer_norm(x + ffn)                               # Add & Norm

def new_layer():
    m = lambda a, b: rng.normal(0, 1 / np.sqrt(a), (a, b))
    return {"attn": [m(d_model, d_model) for _ in range(4)],
            "W1": m(d_model, d_ff), "W2": m(d_ff, d_model)}

x = rng.normal(size=(3, d_model))          # "How are you": 3 token vectors (embedding + position)
print("input:", x.shape, "How starts", np.round(x[0, :3], 4))
for i in range(1, N + 1):
    x = encoder_layer(x, new_layer())      # each of the 6 layers has its own weights
    print(f"encoder {i}: {x.shape}, How starts {np.round(x[0, :3], 4)}, "
          f"row mean {abs(x.mean(1)).max():.4f}, row std {x.std(1).mean():.4f}")

What the six layers did to the vectors

  • The shape stays (3, 512) from the input to encoder 6: one 512-number vector per token, which is what lets the layers stack and the residual additions work.
  • The vector of How changes at every layer: its first three numbers go from [0.3047, −1.04, 0.7505] to [1.0884, −0.0659, −1.2513].
  • Every row has mean 0.0000 and standard deviation 1.0000 after each layer, the work of the last layer normalization. In the paper the layer norm also has a learned gain and bias, which stay at 1 and 0 here.

Counting the weights of each layer

The same sizes give the number of trained weights in each part of a layer.

ExampleRun on Python 3.12
d_model, d_ff, N = 512, 2048, 6

mha = 4 * (d_model * d_model + d_model)          # W_Q, W_K, W_V, W_O and their biases
ffn = d_model * d_ff + d_ff + d_ff * d_model + d_model
norm = 2 * d_model                               # layer norm gain and bias

encoder_layer = mha + ffn + 2 * norm             # self-attention, FFN, 2 Add & Norm
decoder_layer = 2 * mha + ffn + 3 * norm         # masked self-attn, cross-attn, FFN, 3 Add & Norm
print("multi-head attention:", f"{mha:,}")
print("feed-forward network:", f"{ffn:,}")
print("one encoder layer:   ", f"{encoder_layer:,}")
print("one decoder layer:   ", f"{decoder_layer:,}")
print("6 encoders + 6 decoders:", f"{N * (encoder_layer + decoder_layer):,}")
  • One encoder layer holds 3,152,384 weights: 1,050,624 in multi-head attention, 2,099,712 in the feed-forward network and 2,048 in two layer norms.
  • A decoder layer holds 4,204,032, because it has a second attention sub-layer and a third layer norm.
  • The 12 layers hold 44,138,496 weights. The paper reports 65 million for its base model; most of the rest sit in the token embeddings.

Encoder layer vs decoder layer

Encoder layerDecoder layer
Sub-layersself-attention, feed-forwardmasked self-attention, cross-attention, feed-forward
Attends toall input tokensearlier output tokens, then all encoder outputs
Add & Norm blocks23
Weights at the paper's size3,152,3844,204,032
Runs during generationonce per inputonce per generated token

Where you use the architecture

  • The full encoder-decoder for translation and summarisation: the original transformer, T5, BART.
  • The encoder stack alone for understanding tasks such as classification and NER: BERT.
  • The decoder stack alone, without cross-attention, for text generation: GPT.
Watch out. "Six encoders" means six layers stacked one on top of another, each with its own weights, not six encoders reading the input side by side. Only the output of the last encoder goes to the decoders, and every decoder layer reads that same output.
Try it yourself
  • Set N = 2 and check that the output shape is still (3, 512).
  • Change d_ff = 2048 to d_ff = 1024 in the weight count and see how much smaller the feed-forward network gets.
  • Use x = rng.normal(size=(10, d_model)) for a 10-token sentence. Does the number of weights change?
PreviousTransformers

You understood something today that you didn't yesterday.