Transformers
A transformer is a neural network architecture that processes all the tokens of a sequence at once with self-attention instead of recurrence, so each token's vector is computed from the whole sentence in parallel.
Last updated: 07 Oct, 2026 · NumPy
The Attention mechanism (Bahdanau and Luong) lesson ended on a limit: attention removed the context bottleneck but still sat on top of RNNs. "Attention Is All You Need" (Vaswani et al., 2017) dropped the recurrence and built the whole model from attention.
Processing every word in parallel
The attention model the video revises uses a bidirectional LSTM as the encoder and an LSTM as the decoder. In both, the words go in by time step: one word at t = 1, the next at t = 2, and so on. The words of a sentence cannot be sent in parallel, so training cannot be spread across them, and with a huge dataset the encoder-decoder with attention is not scalable with respect to training.
Transformers use no LSTM or RNN in the encoder or the decoder. They use a self-attention module, and all the words of the sentence are sent to the encoder together. That is also why transformers need positional encoding: with no time steps, the order of the words has to be added to the input another way (Positional encoding). Because training scales with data, bigger datasets gave state-of-the-art NLP models, and through transfer learning the same architecture moved into multimodal tasks that mix text and images.
The exact cause is the recurrence ht = f(ht−1, xt): step t cannot start until step t − 1 is done. Self-attention has no such chain. It scores every word against every other word with one matrix product, so the n positions of a layer are computed together. This holds for the encoder and for training the decoder; a trained decoder still writes its output one token at a time.
Getting contextual embeddings instead of fixed vectors
The second problem is about the vectors themselves. Take the sentence "My name is Krish and I play cricket". An embedding layer based on Word2Vec gives every word a fixed vector, the same vector wherever the word appears. But the words of a sentence relate to each other: "I" is the person named before it, and "cricket" is what that person plays. A contextual embedding gives "cricket" a vector that changes with this relationship instead of a fixed one. Self-attention produces such vectors, which is a large part of why transformers are more accurate.
Seeing a word's vector change with its sentence
A runnable check on the board's three-word example. The fixed vector of "cat" is [0, 1, 0, 1] in any sentence. A small self-attention step, worked through number by number in the Self-attention lesson, gives it a different vector in "The cat sat" and in "The cat".
The self-attention step
scores = X @ X.T / np.sqrt(X.shape[1]) # all word pairs in one product
w = np.exp(scores) / np.exp(scores).sum(axis=1, keepdims=True)
Z = w @ X # each row mixes all the wordsimport numpy as np
E = {"The": [1, 0, 1, 0], "cat": [0, 1, 0, 1], "sat": [1, 1, 1, 1]} # fixed, like Word2Vec
def self_attention(X):
scores = X @ X.T / np.sqrt(X.shape[1]) # every word scored against every word at once
w = np.exp(scores) / np.exp(scores).sum(axis=1, keepdims=True)
return w @ X # each row mixes all the words
for sentence in (["The", "cat", "sat"], ["The", "cat"]):
X = np.array([E[word] for word in sentence], float)
Z = self_attention(X)
i = sentence.index("cat")
print(" ".join(sentence))
print(" fixed vector of cat: ", X[i])
print(" contextual vector of cat:", np.round(Z[i], 4))The cat sat fixed vector of cat: [0. 1. 0. 1.] contextual vector of cat: [0.5777 0.8446 0.5777 0.8446] The cat fixed vector of cat: [0. 1. 0. 1.] contextual vector of cat: [0.2689 0.7311 0.2689 0.7311]
What the two contexts do to cat
- The fixed vector stays [0, 1, 0, 1] in both sentences, as a Word2Vec or embedding-layer lookup would give it.
- In "The cat sat" cat becomes [0.5777, 0.8446, 0.5777, 0.8446]: it has mixed in The and sat.
- In "The cat" it becomes [0.2689, 0.7311, 0.2689, 0.7311]: without sat the mix changes, so the same word gets a different vector in a different sentence.
Transformers vs RNN encoder-decoders
| RNN or LSTM with attention | Transformer | |
|---|---|---|
| Reads the input | one word per time step | all words at once |
| Sequential operations per layer | O(n) | O(1) |
| Work per layer | O(n · d²) | O(n² · d) |
| Path between two distant words | O(n) steps | 1 step |
| Word vectors | fixed embeddings fed into a recurrence | contextual, from self-attention |
| Word order | built into the time steps | added with positional encoding |
Here n is the sentence length and d the vector size. The complexity rows are those of Table 1 in the paper: self-attention does more work per layer when n is larger than d, but none of it has to wait.
Where you use transformers
- Translation, the task of the original paper, with an encoder and a decoder.
- Language models: BERT keeps only the encoder and GPT only the decoder, both pretrained on huge text and then fine-tuned (BERT, GPT and T5).
- Images and mixed inputs: vision transformers split an image into patches, and DALL-E (2021) generated images from a text prompt with a transformer.
Related
- Previous: Attention mechanism (Bahdanau and Luong)
- Next: Transformer architecture
- Reference: Vaswani et al., Attention Is All You Need (2017)
- Add
"mat": [0, 0, 1, 1]toEand run["The", "cat", "sat", "mat"]. Does cat's contextual vector change again? - Print
Zfor all three words of "The cat sat" and check that every entry lies between 0 and 1. - Run
["sat", "cat", "The"]: cat's contextual vector does not change. Positional encoding explains why, and how transformers fix it.
Little by little, you're building something great.