Position-wise feed-forward network
The position-wise feed-forward network is the second sub-layer of every transformer layer: two linear layers with a ReLU between them, FFN(x) = max(0, xW₁ + b₁)W₂ + b₂, applied to each token's vector separately and with the same weights at every position.
Last updated: 07 Oct, 2026 · NumPy
After Multi-head attention and its Add & Norm, every token's vector goes through a small neural network of its own. Attention mixes information between tokens; this layer works on each token alone.
Passing each contextual vector through the network
The video follows "How are you" through one encoder. An embedding layer turns each word into a vector. Self-attention converts these vectors into new ones, Z1, Z2 and Z3, called contextual vectors because each one now depends on the other words. Then each vector is passed to the feed-forward neural network, and its output goes on to encoder 2, which again has a self-attention and a feed-forward network.
The handwritten notes draw this step as three separate network ovals inside one shared box: one network, applied at each position.
Expanding to 2048 and back to 512
The first linear layer widens each 512-number vector to dff = 2048, the ReLU sets the negative values to 0 (ReLU and its variants), and the second layer brings it back to dmodel = 512, ready for the residual addition. The paper notes that the same weights are used at every position but differ from layer to layer, and that the network is the same as two convolutions with kernel size 1.
"Position-wise" is the key word. The vector of How never meets the vector of you inside this network. Mixing across tokens happens only in attention; the feed-forward network then processes what each token has gathered.
Running the FFN on The cat sat
The example feeds the three contextual vectors from the Self-attention lesson into a small FFN with dmodel = 4 and dff = 16, the same 4× widening as 512 to 2048. The weights are random with a fixed seed.
The network
def ffn(x):
return np.maximum(0, x @ W1 + b1) @ W2 + b2 # (n, 4) -> (n, 16) -> (n, 4)One position at a time or all at once
one_by_one = np.array([ffn(z) for z in Z]) # a loop over the three tokens
all_at_once = ffn(Z) # one matrix product for all threeimport numpy as np
Z = np.array([[0.8446, 0.5777, 0.8446, 0.5777], # The, from the self-attention lesson
[0.5777, 0.8446, 0.5777, 0.8446], # cat
[0.7881, 0.7881, 0.7881, 0.7881]]) # sat
d_model, d_ff = 4, 16 # 4x wider inside, like 512 -> 2048
rng = np.random.default_rng(42)
W1, b1 = rng.normal(0, 0.5, (d_model, d_ff)), np.zeros(d_ff)
W2, b2 = rng.normal(0, 0.5, (d_ff, d_model)), np.zeros(d_model)
def ffn(x):
return np.maximum(0, x @ W1 + b1) @ W2 + b2 # linear -> ReLU -> linear
one_by_one = np.array([ffn(z) for z in Z]) # each position on its own
all_at_once = ffn(Z) # the whole (3, 4) matrix in one go
print("FFN output:\n", np.round(all_at_once, 4))
print("same as one position at a time:", np.allclose(one_by_one, all_at_once))
print("hidden units switched off by ReLU per word:", (Z @ W1 + b1 <= 0).sum(axis=1))
swapped = ffn(Z[[2, 1, 0]]) # feed the rows as sat, cat, The
print("swapping positions only swaps the outputs:", np.allclose(swapped, all_at_once[[2, 1, 0]]))FFN output: [[-0.8922 -0.1113 1.1279 0.9931] [-0.7854 0.1088 0.9081 1.2013] [-0.9078 0.0132 1.1043 1.2399]] same as one position at a time: True hidden units switched off by ReLU per word: [8 8 7] swapping positions only swaps the outputs: True
What the per-position run shows
- The loop over tokens and the single matrix product give the same output (True), so the network runs for all tokens in parallel.
- Swapping the rows only swaps the outputs (True): a token's output depends on its own vector alone.
- ReLU switches off 8, 8 and 7 of the 16 hidden units for The, cat and sat. The three sets overlap, because the three inputs are similar, but they are not identical.
Counting the weights of the FFN
At the paper's size the feed-forward network is the largest part of a layer.
d_model, d_ff = 512, 2048
ffn = d_model * d_ff + d_ff + d_ff * d_model + d_model # W1, b1, W2, b2
mha = 4 * (d_model * d_model + d_model) # W_Q, W_K, W_V, W_O with biases
norm = 2 * 2 * d_model # two layer norms (gain and bias)
layer = mha + ffn + norm
print(f"FFN weights per layer: {ffn:,}")
print(f"encoder layer total: {layer:,}")
print(f"share in the FFN: {ffn / layer:.1%}")FFN weights per layer: 2,099,712 encoder layer total: 3,152,384 share in the FFN: 66.6%
- The FFN holds 2,099,712 weights: 512 × 2048 + 2048 + 2048 × 512 + 512.
- That is 66.6% of an encoder layer's 3,152,384 weights, twice as many as multi-head attention.
Attention vs feed-forward sub-layer
| Multi-head attention | Position-wise FFN | |
|---|---|---|
| Mixes tokens | yes, every token with every other | no, each token alone |
| Weights per layer (paper) | 1,050,624 | 2,099,712 |
| Cost as the sentence length n grows | grows as n² (the score matrix) | grows as n |
| Nonlinearity | softmax over the scores | ReLU between two linear layers |
Where you use the position-wise FFN
- Every encoder and decoder layer of the transformer, after attention.
- Newer models change the activation: BERT and GPT-2 use GELU, and LLaMA uses a gated variant, SwiGLU.
- Mixture-of-experts models replace the one FFN with several and send each token to only a few of them.
Related
- Previous: Multi-head attention
- Next: Positional encoding
- Reference: Vaswani et al., Attention Is All You Need (2017)
- Set
d_ff = 4(no widening) and count how many hidden units ReLU switches off. - Replace
np.maximum(0, ...)with the plain value and check that the two linear layers collapse into one:np.allclose(ffn(Z), Z @ (W1 @ W2)). - In the weight count, use
d_model, d_ff = 768, 3072(BERT-base) and compare the FFN share.
Little by little, you're building something great.