Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Position-wise feed-forward network

The position-wise feed-forward network is the second sub-layer of every transformer layer: two linear layers with a ReLU between them, FFN(x) = max(0, xW₁ + b₁)W₂ + b₂, applied to each token's vector separately and with the same weights at every position.

Last updated: 07 Oct, 2026 · NumPy

After Multi-head attention and its Add & Norm, every token's vector goes through a small neural network of its own. Attention mixes information between tokens; this layer works on each token alone.

The encoder from embeddings to the feed-forward network · from the Complete Transformers for NLP One Shot video · 28:56 to 32:48

Passing each contextual vector through the network

The video follows "How are you" through one encoder. An embedding layer turns each word into a vector. Self-attention converts these vectors into new ones, Z1, Z2 and Z3, called contextual vectors because each one now depends on the other words. Then each vector is passed to the feed-forward neural network, and its output goes on to encoder 2, which again has a self-attention and a feed-forward network.

The handwritten notes draw this step as three separate network ovals inside one shared box: one network, applied at each position.

How, are and you pass through self-attention and its Add and Norm, which mix the positions, and come out as z1, z2 and z3; each z goes through its own copy of the feed-forward network, 512 to 2048 with ReLU and back to 512, with the same weights W1, b1, W2 and b2 at every position and no connection between positions, before Add and Norm and the next encoder layer.
The position-wise feed-forward network (Vaswani et al., 2017, section 3.3)

Expanding to 2048 and back to 512

The first linear layer widens each 512-number vector to dff = 2048, the ReLU sets the negative values to 0 (ReLU and its variants), and the second layer brings it back to dmodel = 512, ready for the residual addition. The paper notes that the same weights are used at every position but differ from layer to layer, and that the network is the same as two convolutions with kernel size 1.

"Position-wise" is the key word. The vector of How never meets the vector of you inside this network. Mixing across tokens happens only in attention; the feed-forward network then processes what each token has gathered.

Running the FFN on The cat sat

The example feeds the three contextual vectors from the Self-attention lesson into a small FFN with dmodel = 4 and dff = 16, the same 4× widening as 512 to 2048. The weights are random with a fixed seed.

The network

python
def ffn(x):
    return np.maximum(0, x @ W1 + b1) @ W2 + b2    # (n, 4) -> (n, 16) -> (n, 4)

One position at a time or all at once

python
one_by_one = np.array([ffn(z) for z in Z])   # a loop over the three tokens
all_at_once = ffn(Z)                         # one matrix product for all three
ExampleRun on NumPy 2.5
import numpy as np

Z = np.array([[0.8446, 0.5777, 0.8446, 0.5777],    # The, from the self-attention lesson
              [0.5777, 0.8446, 0.5777, 0.8446],    # cat
              [0.7881, 0.7881, 0.7881, 0.7881]])   # sat
d_model, d_ff = 4, 16                              # 4x wider inside, like 512 -> 2048
rng = np.random.default_rng(42)
W1, b1 = rng.normal(0, 0.5, (d_model, d_ff)), np.zeros(d_ff)
W2, b2 = rng.normal(0, 0.5, (d_ff, d_model)), np.zeros(d_model)

def ffn(x):
    return np.maximum(0, x @ W1 + b1) @ W2 + b2    # linear -> ReLU -> linear

one_by_one = np.array([ffn(z) for z in Z])         # each position on its own
all_at_once = ffn(Z)                               # the whole (3, 4) matrix in one go
print("FFN output:\n", np.round(all_at_once, 4))
print("same as one position at a time:", np.allclose(one_by_one, all_at_once))
print("hidden units switched off by ReLU per word:", (Z @ W1 + b1 <= 0).sum(axis=1))

swapped = ffn(Z[[2, 1, 0]])                        # feed the rows as sat, cat, The
print("swapping positions only swaps the outputs:", np.allclose(swapped, all_at_once[[2, 1, 0]]))

What the per-position run shows

  • The loop over tokens and the single matrix product give the same output (True), so the network runs for all tokens in parallel.
  • Swapping the rows only swaps the outputs (True): a token's output depends on its own vector alone.
  • ReLU switches off 8, 8 and 7 of the 16 hidden units for The, cat and sat. The three sets overlap, because the three inputs are similar, but they are not identical.

Counting the weights of the FFN

At the paper's size the feed-forward network is the largest part of a layer.

ExampleRun on Python 3.12
d_model, d_ff = 512, 2048

ffn = d_model * d_ff + d_ff + d_ff * d_model + d_model     # W1, b1, W2, b2
mha = 4 * (d_model * d_model + d_model)                    # W_Q, W_K, W_V, W_O with biases
norm = 2 * 2 * d_model                                     # two layer norms (gain and bias)
layer = mha + ffn + norm
print(f"FFN weights per layer: {ffn:,}")
print(f"encoder layer total:   {layer:,}")
print(f"share in the FFN:      {ffn / layer:.1%}")
  • The FFN holds 2,099,712 weights: 512 × 2048 + 2048 + 2048 × 512 + 512.
  • That is 66.6% of an encoder layer's 3,152,384 weights, twice as many as multi-head attention.

Attention vs feed-forward sub-layer

Multi-head attentionPosition-wise FFN
Mixes tokensyes, every token with every otherno, each token alone
Weights per layer (paper)1,050,6242,099,712
Cost as the sentence length n growsgrows as n² (the score matrix)grows as n
Nonlinearitysoftmax over the scoresReLU between two linear layers

Where you use the position-wise FFN

  • Every encoder and decoder layer of the transformer, after attention.
  • Newer models change the activation: BERT and GPT-2 use GELU, and LLaMA uses a gated variant, SwiGLU.
  • Mixture-of-experts models replace the one FFN with several and send each token to only a few of them.
Watch out. The output projection WO of multi-head attention is not this network. WO is a single 512 × 512 matrix inside the attention sub-layer; the FFN is a separate two-layer network with a 2048-wide ReLU layer, after Add & Norm.
Try it yourself
  • Set d_ff = 4 (no widening) and count how many hidden units ReLU switches off.
  • Replace np.maximum(0, ...) with the plain value and check that the two linear layers collapse into one: np.allclose(ffn(Z), Z @ (W1 @ W2)).
  • In the weight count, use d_model, d_ff = 768, 3072 (BERT-base) and compare the FFN share.

Little by little, you're building something great.