LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Tokens

A token is a piece of text, often a whole word or part of one, that a model reads and writes as a number from a fixed vocabulary.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

A model never sees letters or words. A tokenizer cuts the text into tokens and gives each one its number, and the numbers are what the model reads. Prices, speed and length limits are all counted in tokens.

Tokens in the OpenAI tokenizer · from the How AI Actually Works + Why Your Prompts Keep Failing session · 16:44 to 17:20

The video describes a token as a word or about three quarters of a word, then types hi how are you? into OpenAI's online tokenizer page, which shows 5 tokens for 15 characters. The page counts the same text in Python with tiktoken, using o200k_harmony, the tokenizer gpt-oss uses.

Syntax: encoding.encode(text) gives a list of token ids; encoding.decode(ids) turns ids back into text.

Counting the tokens in the video's example

ExampleFrom the video, run with tiktoken
import tiktoken

encoding = tiktoken.get_encoding("o200k_harmony")  # the tokenizer gpt-oss uses

ids = encoding.encode("hi how are you?")
print(ids)
print([encoding.decode([i]) for i in ids])
print(len(ids), "tokens")
  • The first line is the five token ids, the numbers the model reads.
  • The second line decodes each id on its own. Three of the five pieces carry the space in front of the word: ' how', not 'how'.
  • 5 tokens, the same count the video's tokenizer page shows.

Splitting rare words and other languages

The next example continues the same file, so encoding is already defined.

Example
for text in ["refund", "unsubscribing", "My parcel has not arrived", "मेरा पार्सल नहीं आया"]:
    ids = encoding.encode(text)
    print(len(ids), [encoding.decode([i]) for i in ids])
  • refund is common, so it is one token.
  • unsubscribing is rarer and splits into four pieces.
  • The English sentence is five tokens, one per word.
  • The Hindi sentence says the same thing in four words and takes seven tokens. Some pieces are a single letter or a vowel sign, which read correctly only when joined back together. The same message can cost more tokens in one language than in another.

Decoding ids back to text

Example
ids = encoding.encode("The parcel arrived but the box was crushed")
print(ids)
print(encoding.decode(ids))
print(encoding.n_vocab, "tokens in the vocabulary")

Decoding the ids gives back exactly the text that was encoded. The last line is the size of the vocabulary: gpt-oss chooses every token it writes from these 201,088 entries.

Tokens vs words vs characters

CharactersWordsTokens
hi how are you?1545
What models countNoNoYes
Same across modelsYesYesNo, each tokenizer splits differently

When token counts matter

  • Price. APIs charge per token, sent and received (Cost).
  • Limits. A model reads at most a fixed number of tokens per request (Context window).
  • Speed. A model writes one token at a time, so a longer answer takes longer (Latency).
Watch out. A token count belongs to one tokenizer. The same text gives different counts under different tokenizers, so count with the tokenizer of the model you call. Cost compares two.
Try it yourself
  • Count the tokens in your own name and in a long email address.
  • Encode " refund" with a space in front and "refund" without. Are the ids the same?
  • Encode "1234567" and print the pieces to see how a number is split.

Slow is fine. Stopping is the only problem.