Tokens
A token is a piece of text, often a whole word or part of one, that a model reads and writes as a number from a fixed vocabulary.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
A model never sees letters or words. A tokenizer cuts the text into tokens and gives each one its number, and the numbers are what the model reads. Prices, speed and length limits are all counted in tokens.
The video describes a token as a word or about three quarters of a word, then types hi how are you? into OpenAI's online tokenizer page, which shows 5 tokens for 15 characters. The page counts the same text in Python with tiktoken, using o200k_harmony, the tokenizer gpt-oss uses.
Syntax: encoding.encode(text) gives a list of token ids; encoding.decode(ids) turns ids back into text.
Counting the tokens in the video's example
import tiktoken
encoding = tiktoken.get_encoding("o200k_harmony") # the tokenizer gpt-oss uses
ids = encoding.encode("hi how are you?")
print(ids)
print([encoding.decode([i]) for i in ids])
print(len(ids), "tokens")[3686, 1495, 553, 481, 30] ['hi', ' how', ' are', ' you', '?'] 5 tokens
- The first line is the five token ids, the numbers the model reads.
- The second line decodes each id on its own. Three of the five pieces carry the space in front of the word:
' how', not'how'. - 5 tokens, the same count the video's tokenizer page shows.
Splitting rare words and other languages
The next example continues the same file, so encoding is already defined.
for text in ["refund", "unsubscribing", "My parcel has not arrived", "मेरा पार्सल नहीं आया"]:
ids = encoding.encode(text)
print(len(ids), [encoding.decode([i]) for i in ids])1 ['refund'] 4 ['un', 'sub', 'scrib', 'ing'] 5 ['My', ' parcel', ' has', ' not', ' arrived'] 7 ['म', 'ेरा', ' पार', '्स', 'ल', ' नहीं', ' आया']
- refund is common, so it is one token.
- unsubscribing is rarer and splits into four pieces.
- The English sentence is five tokens, one per word.
- The Hindi sentence says the same thing in four words and takes seven tokens. Some pieces are a single letter or a vowel sign, which read correctly only when joined back together. The same message can cost more tokens in one language than in another.
Decoding ids back to text
ids = encoding.encode("The parcel arrived but the box was crushed")
print(ids)
print(encoding.decode(ids))
print(encoding.n_vocab, "tokens in the vocabulary")[976, 37708, 18157, 889, 290, 5506, 673, 52894] The parcel arrived but the box was crushed 201088 tokens in the vocabulary
Decoding the ids gives back exactly the text that was encoded. The last line is the size of the vocabulary: gpt-oss chooses every token it writes from these 201,088 entries.
Tokens vs words vs characters
| Characters | Words | Tokens | |
|---|---|---|---|
hi how are you? | 15 | 4 | 5 |
| What models count | No | No | Yes |
| Same across models | Yes | Yes | No, each tokenizer splits differently |
When token counts matter
- Price. APIs charge per token, sent and received (Cost).
- Limits. A model reads at most a fixed number of tokens per request (Context window).
- Speed. A model writes one token at a time, so a longer answer takes longer (Latency).
Related
- Previous: Choosing a model
- Next: Next-token prediction
- Reference: tiktoken on GitHub
- Count the tokens in your own name and in a long email address.
- Encode
" refund"with a space in front and"refund"without. Are the ids the same? - Encode
"1234567"and print the pieces to see how a number is split.
Slow is fine. Stopping is the only problem.