Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Subword tokenization (BPE and WordPiece)

Subword tokenization is a way of splitting text into pieces that are smaller than words and larger than characters, learned from a corpus, so that a fixed vocabulary of a few tens of thousands of pieces can spell any word.

Last updated: 07 Oct, 2026 · Python 3.12

The Linear and softmax output layer has one cell per vocabulary entry, and the embedding layer one row per entry, so the vocabulary has to be fixed before training. The Transformer paper's vocabularies were not words: about 37,000 byte-pair tokens for English-German and 32,000 word-pieces for English-French. This is how such a vocabulary is built.

Splitting text into words, characters or subwords

Tokenization with NLTK splits text into words. A word vocabulary has two problems for a neural model: it gets very large (every form of every word: token, tokens, tokenize, tokenizer), and any word not seen in training, a name, a typo, a new term, becomes an unknown token. Characters avoid unknown tokens, but a sentence becomes several times longer, and attention's cost grows with the square of the length.

Subwords sit in between. Frequent words stay whole, rare words are spelt from frequent pieces: “tokenizers” can be token + izer + s. Every word can still be written with the vocabulary, because the single characters (or bytes) are always in it.

Learning merges with byte-pair encoding (BPE)

Byte-pair encoding, adapted for neural translation by Sennrich, Haddow and Birch (2016), learns its vocabulary bottom-up:

  1. Write every word of the corpus as characters, with an end-of-word mark </w>, and count how often each word occurs.
  2. Count every pair of adjacent symbols, weighted by the word counts.
  3. Merge the most frequent pair into one new symbol everywhere, and record the merge.
  4. Repeat until the vocabulary reaches the size you want (a set number of merges).

To tokenize a new word, split it into characters and replay the recorded merges in the same order. The result depends only on the merge list, so the same word is always split the same way.

BPE on the word tokenizer: it starts as single characters with an end-of-word mark, the first four merges build the piece token, merge 10 adds token plus i, and the word stemmed, which is not in the corpus, is split into stem, m, e, d and the end-of-word mark by replaying the merges.

Scoring pairs the WordPiece way

WordPiece, the tokenizer of BERT, builds its vocabulary the same way but picks the pair to merge by a score instead of a raw count. A pair is merged when the two pieces occur together much more often than their own frequencies would suggest. Pieces inside a word carry a ## prefix, so “stemming” starts as s ##t ##e ##m ##m ##i ##n ##g.

The WordPiece merge score

Choosing a tokenizer: byte-level BPE, WordPiece and SentencePiece

  • Byte-level BPE (GPT-2 and later GPT models) runs BPE on the bytes of UTF-8 text, so any string, any language or emoji, can be encoded without an unknown token. GPT-2's vocabulary has 50,257 entries.
  • WordPiece (BERT): the uncased English BERT models use a vocabulary of 30,522 pieces.
  • SentencePiece (T5 and many multilingual models) treats the text as a raw stream with spaces included, so it needs no word splitting first; T5 uses a 32,000-piece SentencePiece vocabulary.

A BPE merge loop in plain Python

The corpus is seven words from the earlier lessons with made-up counts: token 6, tokens 4, tokenize 3, tokenizer 2, stem 4, stems 3, stemming 2.

Counting adjacent pairs

python
from collections import Counter

def pair_counts(words):
    pairs = Counter()
    for symbols, c in words.items():                 # symbols = one word as a tuple of pieces
        for a, b in zip(symbols, symbols[1:]):
            pairs[(a, b)] += c                       # weighted by how often the word occurs
    return pairs

Merging the best pair

python
def merge(words, pair):
    out = {}
    for symbols, c in words.items():
        s, i = [], 0
        while i < len(symbols):
            if i < len(symbols) - 1 and (symbols[i], symbols[i + 1]) == pair:
                s.append(symbols[i] + symbols[i + 1]); i += 2      # glue the pair
            else:
                s.append(symbols[i]); i += 1
        out[tuple(s)] = c
    return out
ExampleRun with plain Python
counts = {"token": 6, "tokens": 4, "tokenize": 3, "tokenizer": 2,
          "stem": 4, "stems": 3, "stemming": 2}
words = {tuple(w) + ("</w>",): c for w, c in counts.items()}   # characters + end-of-word mark

merges = []
for step in range(1, 11):
    pairs = pair_counts(words)
    best = max(pairs, key=pairs.get)          # most frequent pair (the first one wins a tie)
    words = merge(words, best)
    merges.append(best)
    print(f"merge {step:2}: {best[0]} + {best[1]} -> {best[0] + best[1]}  (count {pairs[best]})")

print()
for symbols in words:
    print(" ".join(symbols))

def tokenize(word, merges):
    symbols = tuple(word) + ("</w>",)
    for pair in merges:                       # replay the merges in the order they were learned
        symbols = next(iter(merge({symbols: 1}, pair)))
    return list(symbols)

print()
for w in ["tokens", "stemmed", "tokenizers"]:
    print(w, "->", tokenize(w, merges))

What the ten merges learned

  • The first four merges build “token”: t + o, to + k, tok + e, toke + n, each with count 15, because 6 + 4 + 3 + 2 = 15 words contain those letters in a row.
  • Merges 5 to 7 build “stem” (count 4 + 3 + 2 = 9).
  • Merge 8 is s + </w> (count 7, from tokens and stems): a plural ending becomes its own piece.
  • “tokens” is split into token + s</w>, two pieces both learned from the corpus.
  • “stemmed” was never in the corpus, and it still tokenizes: stem + m + e + d + </w>. Nothing is unknown, the rare part falls back to characters.
  • “tokenizers” becomes tokeni + z + e + r + s</w>; with ten more merges this small corpus glues it into “tokenize” and then “tokenizer”, whole words, because no other word shares those letters.

Comparing the BPE and WordPiece choice

ExampleRun with plain Python
from collections import Counter

counts = {"token": 6, "tokens": 4, "tokenize": 3, "tokenizer": 2,
          "stem": 4, "stems": 3, "stemming": 2}
# WordPiece marks every piece after the first with ##
words = {(w[0],) + tuple("##" + ch for ch in w[1:]): c for w, c in counts.items()}

unit, pairs = Counter(), Counter()
for symbols, c in words.items():
    for s in symbols:
        unit[s] += c
    for a, b in zip(symbols, symbols[1:]):
        pairs[(a, b)] += c

by_count = max(pairs, key=pairs.get)
score = {p: n / (unit[p[0]] * unit[p[1]]) for p, n in pairs.items()}
by_score = max(score, key=score.get)
print("most frequent pair:", by_count, "count", pairs[by_count], "score", round(score[by_count], 4))
print("best WordPiece pair:", by_score, "count", pairs[by_score], "score", round(score[by_score], 4))
  • BPE's first choice is t + ##o, the most frequent pair (15 times), but its score is only 0.0667 because t and ##o are each common on their own.
  • WordPiece's first choice is ##i + ##z: only 5 occurrences, but ##z never appears without ##i before it, so its score is 0.1429.

Using a pretrained tokenizer from Hugging Face

A pretrained model must be used with the exact tokenizer it was trained with. The Hugging Face transformers library downloads it by the model's name (not run here: it downloads the tokenizer files from the Hugging Face Hub).

python
from transformers import AutoTokenizer

bert = AutoTokenizer.from_pretrained("google-bert/bert-base-uncased")   # WordPiece
gpt2 = AutoTokenizer.from_pretrained("openai-community/gpt2")          # byte-level BPE

bert.tokenize("Tokenizers split stemming")    # list of WordPiece strings, "##" marks a continuation
gpt2.tokenize("Tokenizers split stemming")    # list of BPE strings, "Ġ" marks a leading space
bert("Tokenizers split stemming")             # input_ids (with [CLS] ... [SEP]), token_type_ids, attention_mask

tokenize() returns a list of piece strings: BERT's WordPiece marks a piece that continues a word with ##, GPT-2's byte-level BPE marks a piece that starts after a space with Ġ. Calling the tokenizer itself returns a dictionary-like BatchEncoding with input_ids (the vocabulary indices, with BERT's [CLS] at the start and [SEP] at the end) token_type_ids (0 for the first sentence, 1 for a second one) and an attention_mask of 1s for real tokens.

BPE vs WordPiece vs SentencePiece

BPE (byte-level)WordPieceSentencePiece
Merge rulemost frequent pairhighest count(ab) / (count(a)·count(b))BPE or a unigram language model
MarksĠ for a leading space## for a continuation▁ for a leading space
Unknown tokensnone (bytes)[UNK] for unseen charactersrare (character fallback)
Used byGPT-2, GPT-3, RoBERTaBERT, DistilBERTT5, ALBERT, many multilingual models
Vocabulary size50,257 (GPT-2)30,522 (BERT uncased)32,000 (T5)

Where you use subword tokenization

  • Every pretrained transformer: the tokenizer is saved with the model and loaded with it.
  • Counting tokens for an API's context window or price, which is set in tokens, not words.
  • Training a model on a new language or domain, where a tokenizer trained on that text gives shorter sequences than an English one.
Watch out. A tokenizer and a model only work as a pair. Encoding text with GPT-2's tokenizer and feeding the ids to BERT gives indices that mean different pieces in BERT's vocabulary; the model runs without an error and its output is meaningless.
Try it yourself
  • Run 20 merges instead of 10 and find the merge that makes “tokenizer” a single piece.
  • Add "stemmed": 5 to counts and check how “stemmed” is tokenized now.
  • Change the tie rule to min(pairs, key=lambda p: (-pairs[p], p)) (alphabetical on ties) and compare the merge list.

This is what real progress feels like.