Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Tokenization

Tokenization is the process of splitting text into smaller units called tokens, such as sentences or words, so that each unit can be counted, cleaned and later turned into a vector.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Every text preprocessing step in the Natural language processing (NLP) roadmap works on tokens, so tokenization comes first. Four terms come with it, and the rest of the course uses them all the time: corpus, documents, vocabulary and words.

Naming the corpus, documents, vocabulary and words

Corpus, documents and sentence tokens · from the Complete NLP Machine Learning in One Shot video · 21:53 to 26:03

The video defines the four terms on its first example:

  • Corpus: the whole body of text. In the video's example it is one paragraph.
  • Documents: the units the corpus is made of. In the example each sentence is a document.
  • Vocabulary: all the unique words in the corpus, the way a dictionary's vocabulary is its count of unique words.
  • Words: every word in the corpus, repeats included.

In a real dataset the corpus is the collection of all documents, and a document is whatever one row holds: one e-mail, one review, one tweet. The spam messages in part 5 form a corpus of 5,572 documents. The vocabulary is then the set of unique tokens across that whole collection.

Splitting a paragraph into sentences and words

Tokenization takes a paragraph or a sentence and converts it into tokens. The video's paragraph is "My name is Krish and I have an interest in teaching ML, NLP and DL. I am also a YouTuber."

  • Paragraph into sentences. The tokenizer looks for the characters that end a sentence, such as a full stop or an exclamation mark. The paragraph gives two sentence tokens, document 1 and document 2.
  • Sentence into words. Applied to a sentence, tokenization gives word tokens, "My", "name", "is" and so on, each one separate.
The paragraph is the corpus; sentence tokenization splits it at the full stop into two documents, and word tokenization splits the first sentence into word tokens, with the comma and the full stop as tokens of their own.

So words can be tokens and sentences can be tokens. Word tokens matter because each word is later converted into a vector, and each one is cleaned first, which is what the rest of this part does.

Counting words and building the vocabulary

Counting words and the vocabulary · from the Complete NLP Machine Learning in One Shot video · 27:49 to 31:31

The second example has two sentences: "I like to drink Apple Juice. My friend likes Mango Juice." The full stop splits it into two sentence tokens, and counting gives three numbers:

  • 11 words: I, like, to, drink, Apple, Juice, My, friend, likes, Mango, Juice.
  • 10 unique words: Juice appears twice and is counted once. Like and likes are two different words. These 10 words are the vocabulary.
  • 9 unique words if sentence 2 said "like" instead of "likes", since then like would repeat too.

The interview definition the video gives: tokenization is a process to convert a paragraph or sentences into tokens. A paragraph can become sentences, and a paragraph or a sentence can become words.

I like to drink Apple Juice. My friend likes Mango Juice. has 13 tokens with the full stops and 11 words; the second Juice repeats, so the vocabulary has 10 unique words, and 9 with like instead of likes.

Counting the juice example in Python

Plain Python is enough to check the counts. Regular expressions from the re module find the pieces; Tokenization with NLTK does the same job with trained tokenizers.

The sentences and the words

A sentence ends after ., ! or ? followed by a space. A word is a run of letters, so the full stops are left out; a token can also be one punctuation mark.

python
sentences = re.split(r"(?<=[.!?])\s+", corpus)      # split after . ! or ? and a space
words = re.findall(r"[A-Za-z]+", corpus)              # runs of letters only
tokens = re.findall(r"\w+|[^\w\s]", corpus)          # words and punctuation marks

The vocabulary and its index

A set keeps one copy of each word. Lower-casing first makes Apple and apple the same word, and numbering the sorted vocabulary gives every word an index, the first step toward a vector.

python
vocabulary = sorted(set(w.lower() for w in words))   # unique words, a to z
index = {w: i for i, w in enumerate(vocabulary)}      # word -> position

Counting words, tokens and unique words

ExampleFrom the video, run on Python 3.12
import re

corpus = "I like to drink Apple Juice. My friend likes Mango Juice."

sentences = re.split(r"(?<=[.!?])\s+", corpus)
words = re.findall(r"[A-Za-z]+", corpus)
tokens = re.findall(r"\w+|[^\w\s]", corpus)
vocabulary = sorted(set(w.lower() for w in words))
index = {w: i for i, w in enumerate(vocabulary)}

print("sentences:", sentences)
print("words:", len(words), words)
print("tokens with punctuation:", len(tokens))
print("vocabulary:", len(vocabulary), vocabulary)
print("index of juice:", index["juice"])

liked = re.findall(r"[A-Za-z]+", corpus.replace("likes", "like"))
print("vocabulary with like:", len(set(w.lower() for w in liked)))

What the counts show

  • Two sentences, split after each full stop.
  • 11 words, the board's total, and 13 tokens once the two full stops count as tokens too. A tokenizer that keeps punctuation reports 13.
  • A vocabulary of 10, a to z, with juice once and like and likes apart. Juice sits at index 4.
  • 9 when likes becomes like, as on the board.

Words vs tokens vs vocabulary

WordsTokensVocabulary
What it countsEvery word, repeats includedEvery unit the tokenizer emitsEach distinct word once
PunctuationLeft outKept as tokensLeft out
Juice example111310 (9 with like)
Used forThe length of a documentThe input to the next stepThe columns of a vector, one per word

Where you use tokenization

  • Every NLP pipeline: stemming, stopword removal, bag of words and Word2Vec all start from word tokens.
  • Sizing a model: the vocabulary size sets how many columns a bag-of-words vector has.
  • Splitting long text: a document is cut into sentences before it is summarised or translated.
Watch out. The vocabulary size depends on every choice you make: lower-casing, keeping punctuation, merging likes into like. The same corpus can give 10 or 9 unique words. Fix the choices before counting, and apply the same ones to new text.
Try it yourself
  • Add a third sentence, "Juice is my favourite drink!", and predict the new word count and vocabulary size before running.
  • Remove .lower() from the vocabulary line and add "apple" in lower case to the corpus: how many unique words now?
  • Print tokens to see which two tokens are not words.
PreviousNLP use cases

Slow is fine. Stopping is the only problem.