Subword tokenization (BPE and WordPiece)
Subword tokenization is a way of splitting text into pieces that are smaller than words and larger than characters, learned from a corpus, so that a fixed vocabulary of a few tens of thousands of pieces can spell any word.
Last updated: 07 Oct, 2026 · Python 3.12
The Linear and softmax output layer has one cell per vocabulary entry, and the embedding layer one row per entry, so the vocabulary has to be fixed before training. The Transformer paper's vocabularies were not words: about 37,000 byte-pair tokens for English-German and 32,000 word-pieces for English-French. This is how such a vocabulary is built.
Splitting text into words, characters or subwords
Tokenization with NLTK splits text into words. A word vocabulary has two problems for a neural model: it gets very large (every form of every word: token, tokens, tokenize, tokenizer), and any word not seen in training, a name, a typo, a new term, becomes an unknown token. Characters avoid unknown tokens, but a sentence becomes several times longer, and attention's cost grows with the square of the length.
Subwords sit in between. Frequent words stay whole, rare words are spelt from frequent pieces: “tokenizers” can be token + izer + s. Every word can still be written with the vocabulary, because the single characters (or bytes) are always in it.
Learning merges with byte-pair encoding (BPE)
Byte-pair encoding, adapted for neural translation by Sennrich, Haddow and Birch (2016), learns its vocabulary bottom-up:
- Write every word of the corpus as characters, with an end-of-word mark
</w>, and count how often each word occurs. - Count every pair of adjacent symbols, weighted by the word counts.
- Merge the most frequent pair into one new symbol everywhere, and record the merge.
- Repeat until the vocabulary reaches the size you want (a set number of merges).
To tokenize a new word, split it into characters and replay the recorded merges in the same order. The result depends only on the merge list, so the same word is always split the same way.
Scoring pairs the WordPiece way
WordPiece, the tokenizer of BERT, builds its vocabulary the same way but picks the pair to merge by a score instead of a raw count. A pair is merged when the two pieces occur together much more often than their own frequencies would suggest. Pieces inside a word carry a ## prefix, so “stemming” starts as s ##t ##e ##m ##m ##i ##n ##g.
Choosing a tokenizer: byte-level BPE, WordPiece and SentencePiece
- Byte-level BPE (GPT-2 and later GPT models) runs BPE on the bytes of UTF-8 text, so any string, any language or emoji, can be encoded without an unknown token. GPT-2's vocabulary has 50,257 entries.
- WordPiece (BERT): the uncased English BERT models use a vocabulary of 30,522 pieces.
- SentencePiece (T5 and many multilingual models) treats the text as a raw stream with spaces included, so it needs no word splitting first; T5 uses a 32,000-piece SentencePiece vocabulary.
A BPE merge loop in plain Python
The corpus is seven words from the earlier lessons with made-up counts: token 6, tokens 4, tokenize 3, tokenizer 2, stem 4, stems 3, stemming 2.
Counting adjacent pairs
from collections import Counter
def pair_counts(words):
pairs = Counter()
for symbols, c in words.items(): # symbols = one word as a tuple of pieces
for a, b in zip(symbols, symbols[1:]):
pairs[(a, b)] += c # weighted by how often the word occurs
return pairsMerging the best pair
def merge(words, pair):
out = {}
for symbols, c in words.items():
s, i = [], 0
while i < len(symbols):
if i < len(symbols) - 1 and (symbols[i], symbols[i + 1]) == pair:
s.append(symbols[i] + symbols[i + 1]); i += 2 # glue the pair
else:
s.append(symbols[i]); i += 1
out[tuple(s)] = c
return outcounts = {"token": 6, "tokens": 4, "tokenize": 3, "tokenizer": 2,
"stem": 4, "stems": 3, "stemming": 2}
words = {tuple(w) + ("</w>",): c for w, c in counts.items()} # characters + end-of-word mark
merges = []
for step in range(1, 11):
pairs = pair_counts(words)
best = max(pairs, key=pairs.get) # most frequent pair (the first one wins a tie)
words = merge(words, best)
merges.append(best)
print(f"merge {step:2}: {best[0]} + {best[1]} -> {best[0] + best[1]} (count {pairs[best]})")
print()
for symbols in words:
print(" ".join(symbols))
def tokenize(word, merges):
symbols = tuple(word) + ("</w>",)
for pair in merges: # replay the merges in the order they were learned
symbols = next(iter(merge({symbols: 1}, pair)))
return list(symbols)
print()
for w in ["tokens", "stemmed", "tokenizers"]:
print(w, "->", tokenize(w, merges))merge 1: t + o -> to (count 15) merge 2: to + k -> tok (count 15) merge 3: tok + e -> toke (count 15) merge 4: toke + n -> token (count 15) merge 5: s + t -> st (count 9) merge 6: st + e -> ste (count 9) merge 7: ste + m -> stem (count 9) merge 8: s + </w> -> s</w> (count 7) merge 9: token + </w> -> token</w> (count 6) merge 10: token + i -> tokeni (count 5) token</w> token s</w> tokeni z e </w> tokeni z e r </w> stem </w> stem s</w> stem m i n g </w> tokens -> ['token', 's</w>'] stemmed -> ['stem', 'm', 'e', 'd', '</w>'] tokenizers -> ['tokeni', 'z', 'e', 'r', 's</w>']
What the ten merges learned
- The first four merges build “token”: t + o, to + k, tok + e, toke + n, each with count 15, because 6 + 4 + 3 + 2 = 15 words contain those letters in a row.
- Merges 5 to 7 build “stem” (count 4 + 3 + 2 = 9).
- Merge 8 is s + </w> (count 7, from tokens and stems): a plural ending becomes its own piece.
- “tokens” is split into token + s</w>, two pieces both learned from the corpus.
- “stemmed” was never in the corpus, and it still tokenizes: stem + m + e + d + </w>. Nothing is unknown, the rare part falls back to characters.
- “tokenizers” becomes tokeni + z + e + r + s</w>; with ten more merges this small corpus glues it into “tokenize” and then “tokenizer”, whole words, because no other word shares those letters.
Comparing the BPE and WordPiece choice
from collections import Counter
counts = {"token": 6, "tokens": 4, "tokenize": 3, "tokenizer": 2,
"stem": 4, "stems": 3, "stemming": 2}
# WordPiece marks every piece after the first with ##
words = {(w[0],) + tuple("##" + ch for ch in w[1:]): c for w, c in counts.items()}
unit, pairs = Counter(), Counter()
for symbols, c in words.items():
for s in symbols:
unit[s] += c
for a, b in zip(symbols, symbols[1:]):
pairs[(a, b)] += c
by_count = max(pairs, key=pairs.get)
score = {p: n / (unit[p[0]] * unit[p[1]]) for p, n in pairs.items()}
by_score = max(score, key=score.get)
print("most frequent pair:", by_count, "count", pairs[by_count], "score", round(score[by_count], 4))
print("best WordPiece pair:", by_score, "count", pairs[by_score], "score", round(score[by_score], 4))most frequent pair: ('t', '##o') count 15 score 0.0667
best WordPiece pair: ('##i', '##z') count 5 score 0.1429- BPE's first choice is t + ##o, the most frequent pair (15 times), but its score is only 0.0667 because t and ##o are each common on their own.
- WordPiece's first choice is ##i + ##z: only 5 occurrences, but ##z never appears without ##i before it, so its score is 0.1429.
Using a pretrained tokenizer from Hugging Face
A pretrained model must be used with the exact tokenizer it was trained with. The Hugging Face transformers library downloads it by the model's name (not run here: it downloads the tokenizer files from the Hugging Face Hub).
from transformers import AutoTokenizer
bert = AutoTokenizer.from_pretrained("google-bert/bert-base-uncased") # WordPiece
gpt2 = AutoTokenizer.from_pretrained("openai-community/gpt2") # byte-level BPE
bert.tokenize("Tokenizers split stemming") # list of WordPiece strings, "##" marks a continuation
gpt2.tokenize("Tokenizers split stemming") # list of BPE strings, "Ġ" marks a leading space
bert("Tokenizers split stemming") # input_ids (with [CLS] ... [SEP]), token_type_ids, attention_masktokenize() returns a list of piece strings: BERT's WordPiece marks a piece that continues a word with ##, GPT-2's byte-level BPE marks a piece that starts after a space with Ġ. Calling the tokenizer itself returns a dictionary-like BatchEncoding with input_ids (the vocabulary indices, with BERT's [CLS] at the start and [SEP] at the end) token_type_ids (0 for the first sentence, 1 for a second one) and an attention_mask of 1s for real tokens.
BPE vs WordPiece vs SentencePiece
| BPE (byte-level) | WordPiece | SentencePiece | |
|---|---|---|---|
| Merge rule | most frequent pair | highest count(ab) / (count(a)·count(b)) | BPE or a unigram language model |
| Marks | Ġ for a leading space | ## for a continuation | ▁ for a leading space |
| Unknown tokens | none (bytes) | [UNK] for unseen characters | rare (character fallback) |
| Used by | GPT-2, GPT-3, RoBERTa | BERT, DistilBERT | T5, ALBERT, many multilingual models |
| Vocabulary size | 50,257 (GPT-2) | 30,522 (BERT uncased) | 32,000 (T5) |
Where you use subword tokenization
- Every pretrained transformer: the tokenizer is saved with the model and loaded with it.
- Counting tokens for an API's context window or price, which is set in tokens, not words.
- Training a model on a new language or domain, where a tokenizer trained on that text gives shorter sequences than an English one.
Related
- Previous: Linear and softmax output layer
- Next: BERT, GPT and T5
- See also: Tokenization
- Reference: Byte-pair encoding tokenization in the Hugging Face LLM course
- Run 20 merges instead of 10 and find the merge that makes “tokenizer” a single piece.
- Add
"stemmed": 5tocountsand check how “stemmed” is tokenized now. - Change the tie rule to
min(pairs, key=lambda p: (-pairs[p], p))(alphabetical on ties) and compare the merge list.
This is what real progress feels like.