Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q2EasyConcept

What is tokenization, and why do LLMs use subword tokens?

30-second answerSay your answer out loud first, then reveal.

Why not characters or words?

UnitVocabularySequence lengthProblem
CharactersTinyVery longExpensive attention; slow learning
WordsHuge (millions)ShortUnknown words, typos, inflections, new names
Subwords32K–200KModerateGood balance; any string can be represented

BPE in brief

  1. Start with bytes or characters as the vocabulary.
  2. Count the most frequent adjacent pair in the training corpus (e.g. "t"+"h").
  3. Merge it into a new token ("th"), and repeat until the vocabulary reaches its target size.
  4. At inference, apply the learned merges to new text.
Example (illustrative): "unhappiness" → ["un", "happiness"]; "Bengaluru" → ["Beng", "al", "uru"].

Byte-level BPE (GPT-family tokenizers) starts from raw bytes, so any Unicode text, emoji or code can be encoded with no unknown tokens.

Practical consequences

  • Spaces are often part of tokens (" Hello" ≠ "Hello").
  • Numbers may be split inconsistently, which affects arithmetic (Q41).
  • Each model has its own tokenizer. Token counts differ across providers for the same text.

Common mistakes

  • Saying "1 token = 1 word." For English, roughly 1 token ≈ 0.75 words (about 4 characters), and it varies widely by language.

You understood something today that you didn't yesterday.