Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Transformers

A transformer is a neural network architecture that processes all the tokens of a sequence at once with self-attention instead of recurrence, so each token's vector is computed from the whole sentence in parallel.

Last updated: 07 Oct, 2026 · NumPy

The Attention mechanism (Bahdanau and Luong) lesson ended on a limit: attention removed the context bottleneck but still sat on top of RNNs. "Attention Is All You Need" (Vaswani et al., 2017) dropped the recurrence and built the whole model from attention.

Why RNN attention does not scale · from the Complete Transformers for NLP One Shot video · 12:15 to 15:46

Processing every word in parallel

The attention model the video revises uses a bidirectional LSTM as the encoder and an LSTM as the decoder. In both, the words go in by time step: one word at t = 1, the next at t = 2, and so on. The words of a sentence cannot be sent in parallel, so training cannot be spread across them, and with a huge dataset the encoder-decoder with attention is not scalable with respect to training.

Transformers use no LSTM or RNN in the encoder or the decoder. They use a self-attention module, and all the words of the sentence are sent to the encoder together. That is also why transformers need positional encoding: with no time steps, the order of the words has to be added to the input another way (Positional encoding). Because training scales with data, bigger datasets gave state-of-the-art NLP models, and through transfer learning the same architecture moved into multimodal tasks that mix text and images.

The exact cause is the recurrence ht = f(ht−1, xt): step t cannot start until step t − 1 is done. Self-attention has no such chain. It scores every word against every other word with one matrix product, so the n positions of a layer are computed together. This holds for the encoder and for training the decoder; a trained decoder still writes its output one token at a time.

Top: an RNN or LSTM encoder reads x1 to x4 in four time steps, each hidden state waiting for the one before. Bottom: self-attention takes x1 to x4 together and produces z1 to z4 in one step.
Contextual embeddings · from the Complete Transformers for NLP One Shot video · 18:42 to 22:22

Getting contextual embeddings instead of fixed vectors

The second problem is about the vectors themselves. Take the sentence "My name is Krish and I play cricket". An embedding layer based on Word2Vec gives every word a fixed vector, the same vector wherever the word appears. But the words of a sentence relate to each other: "I" is the person named before it, and "cricket" is what that person plays. A contextual embedding gives "cricket" a vector that changes with this relationship instead of a fixed one. Self-attention produces such vectors, which is a large part of why transformers are more accurate.

The sentence My name is Krish and I play cricket through two layers: a Word2Vec lookup gives 8 fixed vectors, the same in every sentence; a self-attention layer gives 8 contextual vectors, where I and cricket take in information from the name.

Seeing a word's vector change with its sentence

A runnable check on the board's three-word example. The fixed vector of "cat" is [0, 1, 0, 1] in any sentence. A small self-attention step, worked through number by number in the Self-attention lesson, gives it a different vector in "The cat sat" and in "The cat".

The self-attention step

python
scores = X @ X.T / np.sqrt(X.shape[1])       # all word pairs in one product
w = np.exp(scores) / np.exp(scores).sum(axis=1, keepdims=True)
Z = w @ X                                    # each row mixes all the words
ExampleRun on NumPy 2.5
import numpy as np

E = {"The": [1, 0, 1, 0], "cat": [0, 1, 0, 1], "sat": [1, 1, 1, 1]}   # fixed, like Word2Vec

def self_attention(X):
    scores = X @ X.T / np.sqrt(X.shape[1])      # every word scored against every word at once
    w = np.exp(scores) / np.exp(scores).sum(axis=1, keepdims=True)
    return w @ X                                # each row mixes all the words

for sentence in (["The", "cat", "sat"], ["The", "cat"]):
    X = np.array([E[word] for word in sentence], float)
    Z = self_attention(X)
    i = sentence.index("cat")
    print(" ".join(sentence))
    print("  fixed vector of cat:     ", X[i])
    print("  contextual vector of cat:", np.round(Z[i], 4))

What the two contexts do to cat

  • The fixed vector stays [0, 1, 0, 1] in both sentences, as a Word2Vec or embedding-layer lookup would give it.
  • In "The cat sat" cat becomes [0.5777, 0.8446, 0.5777, 0.8446]: it has mixed in The and sat.
  • In "The cat" it becomes [0.2689, 0.7311, 0.2689, 0.7311]: without sat the mix changes, so the same word gets a different vector in a different sentence.

Transformers vs RNN encoder-decoders

RNN or LSTM with attentionTransformer
Reads the inputone word per time stepall words at once
Sequential operations per layerO(n)O(1)
Work per layerO(n · d²)O(n² · d)
Path between two distant wordsO(n) steps1 step
Word vectorsfixed embeddings fed into a recurrencecontextual, from self-attention
Word orderbuilt into the time stepsadded with positional encoding

Here n is the sentence length and d the vector size. The complexity rows are those of Table 1 in the paper: self-attention does more work per layer when n is larger than d, but none of it has to wait.

Where you use transformers

  • Translation, the task of the original paper, with an encoder and a decoder.
  • Language models: BERT keeps only the encoder and GPT only the decoder, both pretrained on huge text and then fine-tuned (BERT, GPT and T5).
  • Images and mixed inputs: vision transformers split an image into patches, and DALL-E (2021) generated images from a text prompt with a transformer.
Watch out. "Transformers process words in parallel" is true for the encoder and for training. A decoder that generates text still produces one token, feeds it back and produces the next, so generation time grows with the length of the output.
Try it yourself
  • Add "mat": [0, 0, 1, 1] to E and run ["The", "cat", "sat", "mat"]. Does cat's contextual vector change again?
  • Print Z for all three words of "The cat sat" and check that every entry lies between 0 and 1.
  • Run ["sat", "cat", "The"]: cat's contextual vector does not change. Positional encoding explains why, and how transformers fix it.

Little by little, you're building something great.