Tokenization
Tokenization is the process of splitting text into smaller units called tokens, such as sentences or words, so that each unit can be counted, cleaned and later turned into a vector.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
Every text preprocessing step in the Natural language processing (NLP) roadmap works on tokens, so tokenization comes first. Four terms come with it, and the rest of the course uses them all the time: corpus, documents, vocabulary and words.
Naming the corpus, documents, vocabulary and words
The video defines the four terms on its first example:
- Corpus: the whole body of text. In the video's example it is one paragraph.
- Documents: the units the corpus is made of. In the example each sentence is a document.
- Vocabulary: all the unique words in the corpus, the way a dictionary's vocabulary is its count of unique words.
- Words: every word in the corpus, repeats included.
In a real dataset the corpus is the collection of all documents, and a document is whatever one row holds: one e-mail, one review, one tweet. The spam messages in part 5 form a corpus of 5,572 documents. The vocabulary is then the set of unique tokens across that whole collection.
Splitting a paragraph into sentences and words
Tokenization takes a paragraph or a sentence and converts it into tokens. The video's paragraph is "My name is Krish and I have an interest in teaching ML, NLP and DL. I am also a YouTuber."
- Paragraph into sentences. The tokenizer looks for the characters that end a sentence, such as a full stop or an exclamation mark. The paragraph gives two sentence tokens, document 1 and document 2.
- Sentence into words. Applied to a sentence, tokenization gives word tokens, "My", "name", "is" and so on, each one separate.
So words can be tokens and sentences can be tokens. Word tokens matter because each word is later converted into a vector, and each one is cleaned first, which is what the rest of this part does.
Counting words and building the vocabulary
The second example has two sentences: "I like to drink Apple Juice. My friend likes Mango Juice." The full stop splits it into two sentence tokens, and counting gives three numbers:
- 11 words: I, like, to, drink, Apple, Juice, My, friend, likes, Mango, Juice.
- 10 unique words: Juice appears twice and is counted once. Like and likes are two different words. These 10 words are the vocabulary.
- 9 unique words if sentence 2 said "like" instead of "likes", since then like would repeat too.
The interview definition the video gives: tokenization is a process to convert a paragraph or sentences into tokens. A paragraph can become sentences, and a paragraph or a sentence can become words.
Counting the juice example in Python
Plain Python is enough to check the counts. Regular expressions from the re module find the pieces; Tokenization with NLTK does the same job with trained tokenizers.
The sentences and the words
A sentence ends after ., ! or ? followed by a space. A word is a run of letters, so the full stops are left out; a token can also be one punctuation mark.
sentences = re.split(r"(?<=[.!?])\s+", corpus) # split after . ! or ? and a space
words = re.findall(r"[A-Za-z]+", corpus) # runs of letters only
tokens = re.findall(r"\w+|[^\w\s]", corpus) # words and punctuation marksThe vocabulary and its index
A set keeps one copy of each word. Lower-casing first makes Apple and apple the same word, and numbering the sorted vocabulary gives every word an index, the first step toward a vector.
vocabulary = sorted(set(w.lower() for w in words)) # unique words, a to z
index = {w: i for i, w in enumerate(vocabulary)} # word -> positionCounting words, tokens and unique words
import re
corpus = "I like to drink Apple Juice. My friend likes Mango Juice."
sentences = re.split(r"(?<=[.!?])\s+", corpus)
words = re.findall(r"[A-Za-z]+", corpus)
tokens = re.findall(r"\w+|[^\w\s]", corpus)
vocabulary = sorted(set(w.lower() for w in words))
index = {w: i for i, w in enumerate(vocabulary)}
print("sentences:", sentences)
print("words:", len(words), words)
print("tokens with punctuation:", len(tokens))
print("vocabulary:", len(vocabulary), vocabulary)
print("index of juice:", index["juice"])
liked = re.findall(r"[A-Za-z]+", corpus.replace("likes", "like"))
print("vocabulary with like:", len(set(w.lower() for w in liked)))sentences: ['I like to drink Apple Juice.', 'My friend likes Mango Juice.'] words: 11 ['I', 'like', 'to', 'drink', 'Apple', 'Juice', 'My', 'friend', 'likes', 'Mango', 'Juice'] tokens with punctuation: 13 vocabulary: 10 ['apple', 'drink', 'friend', 'i', 'juice', 'like', 'likes', 'mango', 'my', 'to'] index of juice: 4 vocabulary with like: 9
What the counts show
- Two sentences, split after each full stop.
- 11 words, the board's total, and 13 tokens once the two full stops count as tokens too. A tokenizer that keeps punctuation reports 13.
- A vocabulary of 10, a to z, with juice once and like and likes apart. Juice sits at index 4.
- 9 when likes becomes like, as on the board.
Words vs tokens vs vocabulary
| Words | Tokens | Vocabulary | |
|---|---|---|---|
| What it counts | Every word, repeats included | Every unit the tokenizer emits | Each distinct word once |
| Punctuation | Left out | Kept as tokens | Left out |
| Juice example | 11 | 13 | 10 (9 with like) |
| Used for | The length of a document | The input to the next step | The columns of a vector, one per word |
Where you use tokenization
- Every NLP pipeline: stemming, stopword removal, bag of words and Word2Vec all start from word tokens.
- Sizing a model: the vocabulary size sets how many columns a bag-of-words vector has.
- Splitting long text: a document is cut into sentences before it is summarised or translated.
Related
- Previous: NLP use cases
- Next: Tokenization with NLTK
- Reference: re, regular expressions
- Add a third sentence, "Juice is my favourite drink!", and predict the new word count and vocabulary size before running.
- Remove
.lower()from the vocabulary line and add "apple" in lower case to the corpus: how many unique words now? - Print
tokensto see which two tokens are not words.
Slow is fine. Stopping is the only problem.