Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

BERT, GPT and T5

BERT, GPT and T5 are three families of pretrained transformers that each keep a different part of the original architecture: BERT is encoder-only, GPT is decoder-only and T5 is encoder-decoder, and each is pretrained on unlabelled text with its own objective.

Last updated: 07 Oct, 2026 · NumPy

The 2017 Transformer was trained from scratch for one task, translation. Its successors are pretrained once on huge amounts of plain text and then adapted to many tasks, the same idea as Transfer learning with VGG16 for images. Which half of the Transformer architecture a model keeps decides what it is good at.

Reading text in both directions: BERT

BERT (Devlin et al., 2018) is a stack of transformer encoders. Every token attends to every other token, left and right, so its vectors are built from the whole sentence. It is pretrained with two objectives:

  • Masked language modelling (MLM). 15% of the tokens are picked; of those, 80% are replaced by [MASK], 10% by a random token and 10% left unchanged, and the model predicts the original tokens.
  • Next sentence prediction. Given two sentences, predict whether the second follows the first in the text. Later models such as RoBERTa dropped this objective.

Each input starts with a [CLS] token, whose final vector is used for classification, and sentences are separated by [SEP]. BERT-base has 12 layers, a hidden size of 768 and 12 heads, about 110 million parameters; BERT-large has 24 layers, 1024 and 16 heads, about 340 million. It reads, it does not write: classification, named entity recognition, extractive question answering and search embeddings.

Predicting the next token: GPT

GPT (Radford et al., 2018) is a stack of transformer decoders without the cross-attention sub-layer, since there is no encoder: each layer is Masked self-attention plus a feed-forward network. It is pretrained on one objective, predicting the next token from all the tokens before it. GPT-1 had 12 layers; GPT-2 (2019) went up to 1.5 billion parameters and moved layer normalization to the start of each sub-layer (pre-norm); GPT-3 (2020) has 175 billion and can follow a task described in the prompt with a few examples. The same next-token objective, followed by instruction tuning, underlies today's chat models.

Turning every task into text: T5

T5 (Raffel et al., 2020) keeps the full encoder-decoder. Every task is written as text in and text out, with a prefix that names it: “translate English to German: That is good.”, “summarize: ...”, “cola sentence: ...”. It is pretrained by span corruption: spans of the input are replaced with sentinel tokens and the decoder writes the missing spans. Sizes go from 60 million to 11 billion parameters. BART is a similar encoder-decoder model trained to rebuild corrupted text.

BERT is encoder-only and predicts masked words from both sides; GPT is decoder-only with a causal mask and predicts the next word; T5 is encoder-decoder, reading a prompt such as translate English to German and writing the answer as text.

Counting BERT-base's parameters from its configuration

The “110 million” can be checked from BERT-base's published configuration: a 30,522-piece WordPiece vocabulary, 512 positions, 2 segment types, hidden size 768, 12 layers and a feed-forward size of 3072. Each layer is the encoder layer from the transformer encoder lesson at these sizes.

ExampleRun with plain Python
# BERT-base as published: 12 layers, hidden size 768, 12 heads, FFN 3072, WordPiece vocabulary 30,522
V, P, T, H, L, F = 30522, 512, 2, 768, 12, 3072

embeddings = V * H + P * H + T * H + 2 * H          # token, position, segment, layer norm
attention = 4 * (H * H + H) + 2 * H                 # Q, K, V, output projection + layer norm
ffn = H * F + F + F * H + H + 2 * H                 # 768 -> 3072 -> 768 + layer norm
pooler = H * H + H                                  # the [CLS] pooler on top

total = embeddings + L * (attention + ffn) + pooler
print("embeddings:  ", f"{embeddings:,}")
print("one layer:   ", f"{attention + ffn:,}")
print("12 layers:   ", f"{L * (attention + ffn):,}")
print("total:       ", f"{total:,}")
  • The embeddings hold 23,837,184 weights, almost all of them the 30,522 × 768 token table.
  • One layer holds 7,087,872; twelve hold 85,054,464.
  • The total is 109,482,240, the “110M” of the BERT paper.

Masking 15% of a sentence the BERT way

The 80/10/10 rule on the decoder definition from the video, “the transformer decoder is responsible for generating the output sequence one token at a time ...”, split on spaces for simplicity (BERT would split it into WordPiece tokens).

ExampleRun with NumPy
import numpy as np

rng = np.random.default_rng(1)
sentence = ("the transformer decoder is responsible for generating the output sequence "
            "one token at a time using the encoder output and the previously generated tokens")
tokens = sentence.split()
vocab = sorted(set(tokens))

n_pick = round(0.15 * len(tokens))                       # BERT picks 15% of the tokens
picked = rng.choice(len(tokens), size=n_pick, replace=False)
inputs, labels = list(tokens), {}
for i in sorted(picked):
    labels[i] = tokens[i]                                 # the model must predict the original
    r = rng.random()
    if r < 0.8:
        inputs[i] = "[MASK]"                              # 80%: replace with [MASK]
    elif r < 0.9:
        inputs[i] = str(rng.choice(vocab))                # 10%: a random word
    # else 10%: keep the word unchanged

print(len(tokens), "tokens,", n_pick, "picked:", sorted(int(i) for i in picked))
print("input :", " ".join(inputs))
print("labels:", {int(i): w for i, w in labels.items()})
  • 4 of the 24 tokens are picked: round(0.15 × 24) = 4.
  • Three become [MASK]: sequence, token and generated.
  • Position 17, “encoder”, became the random word “transformer”; the model still has to predict “encoder” there, so it cannot trust any input token blindly.
  • The labels hold only the four picked positions; the loss is computed on those, not on the other 20.

Encoder-only vs decoder-only vs encoder-decoder

BERT (encoder-only)GPT (decoder-only)T5 (encoder-decoder)
Attentionbidirectionalcausal (earlier tokens)bidirectional encoder, causal decoder + cross
Pretrainingmasked language model (+ next sentence)next-token predictionspan corruption
Outputa vector per tokenthe next token, repeatedlya new text sequence
Best atclassification, NER, searchgeneration, chattranslation, summarization
Base size110M (12 layers, 768)117M for GPT-1 (12 layers, 768)220M for T5-base

Where you use BERT, GPT and T5

  • BERT-style encoders for labelling text and for embeddings in semantic search and retrieval.
  • GPT-style decoders for anything that writes: chat, code, drafting, answering from a prompt.
  • T5-style encoder-decoders when the output is a rewrite of the input: translation, summaries.
Watch out. An encoder-only model such as BERT has no causal mask and was never trained to continue text, so it cannot generate a paragraph left to right. Pick the family that matches the task before picking a size.
Try it yourself
  • Change L, H, F to 24, 1024, 4096 and compare the total with BERT-large's 340 million.
  • Change the seed in the masking example and count how many picked tokens become [MASK] over a few runs.
  • Set V = 50257 and P = 1024 (GPT-2's vocabulary and context) and see how much the embedding count grows.

Every expert started right here.