1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is tokenization, and why do LLMs use subword tokens?
30-second answerSay your answer out loud first, then reveal.
Why not characters or words?
| Unit | Vocabulary | Sequence length | Problem |
|---|---|---|---|
| Characters | Tiny | Very long | Expensive attention; slow learning |
| Words | Huge (millions) | Short | Unknown words, typos, inflections, new names |
| Subwords | 32K–200K | Moderate | Good balance; any string can be represented |
BPE in brief
- Start with bytes or characters as the vocabulary.
- Count the most frequent adjacent pair in the training corpus (e.g. "t"+"h").
- Merge it into a new token ("th"), and repeat until the vocabulary reaches its target size.
- At inference, apply the learned merges to new text.
Example (illustrative): "unhappiness" → ["un", "happiness"]; "Bengaluru" → ["Beng", "al", "uru"].
Byte-level BPE (GPT-family tokenizers) starts from raw bytes, so any Unicode text, emoji or code can be encoded with no unknown tokens.
Practical consequences
- Spaces are often part of tokens (" Hello" ≠ "Hello").
- Numbers may be split inconsistently, which affects arithmetic (Q41).
- Each model has its own tokenizer. Token counts differ across providers for the same text.
Common mistakes
- Saying "1 token = 1 word." For English, roughly 1 token ≈ 0.75 words (about 4 characters), and it varies widely by language.
Related
You understood something today that you didn't yesterday.