Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Tokenization with NLTK

NLTK's tokenizers are functions that split text into sentence tokens (sent_tokenize) or word tokens (word_tokenize, wordpunct_tokenize, TreebankWordTokenizer), each with its own rule for punctuation.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Tokenization split text with hand-written regular expressions. NLTK's tokenizers are tested on real English: they know that "Dr." does not end a sentence and that "don't" is two words.

The corpus is the video's own. Its two lines are written to test the tokenizers: a comma with no space after it, a possessive, and a lower-case word after the exclamation mark:

python
corpus = """Hello Welcome,to Krish Naik's NLP Tutorials.
Please do watch the entire course! to become expert in NLP.
"""

Downloading the sentence model

sent_tokenize and word_tokenize need the Punkt model, downloaded once:

python
import nltk
nltk.download("punkt_tab")    # once per machine: the Punkt sentence model

Splitting a paragraph into sentences with sent_tokenize

sent_tokenize(text) converts a paragraph into a list of sentences, with NLTK's recommended sentence tokenizer for the language (English by default), currently the PunktSentenceTokenizer. Punkt is a trained model: it learned from English text which full stops end a sentence and which belong to abbreviations such as "Dr." and "p.m.". It splits at ., ! and ? when they end a sentence. A line break alone does not end a sentence.

python
from nltk.tokenize import sent_tokenize

documents = sent_tokenize(corpus)    # paragraph -> list of sentences
for sentence in documents:
    print(sentence)
ExampleFrom the video, run on NLTK 3.10.3
corpus = """Hello Welcome,to Krish Naik's NLP Tutorials.
Please do watch the entire course! to become expert in NLP.
"""

from nltk.tokenize import sent_tokenize

documents = sent_tokenize(corpus)
print(type(documents), len(documents))
for sentence in documents:
    print(repr(sentence))

print(sent_tokenize("Hello Welcome to the tutorial\nPlease do watch the course"))
print(sent_tokenize("Dr. Smith went to Washington. He arrived at 3 p.m. on Monday! Was it late? Yes."))

What sent_tokenize did

  • Three sentences from the corpus, in a Python list. The full stop after Tutorials ends the first and the exclamation mark ends the second, so "to become expert in NLP." is a sentence of its own.
  • The line break is kept inside the text: with no full stop before it, "tutorial\nPlease" stays one sentence. The split in the corpus came from the full stop, not from the new line.
  • Abbreviations survive: "Dr." and "p.m." do not end a sentence, while "!" and "?" do.

Splitting words with word_tokenize

word_tokenize, wordpunct_tokenize and Treebank · from the Complete NLP Machine Learning in One Shot video · 41:13 to 45:09

word_tokenize splits a paragraph, or each sentence from sent_tokenize, into words. The comma, the full stops and the exclamation mark become separate tokens. Every word is a separate token because each one has a different importance and gets its own preprocessing.

wordpunct_tokenize treats punctuation as separate words too, and goes one step further: the apostrophe s that word_tokenize keeps as one token, 's, is split into ' and s.

TreebankWordTokenizer differs in one small way, the video's quiz. A full stop in the middle of the text is not treated as a separate word: it stays in the previous word, as in "Tutorials.". Only the last full stop becomes a separate token. Most of the time word_tokenize and sent_tokenize are the ones to use.

Why the three tokenizers differ

  • word_tokenize runs sent_tokenize first, then an improved Treebank tokenizer on each sentence. Every sentence's final full stop is therefore split off, and possessives and contractions are split the Penn Treebank way: Naik's becomes Naik + 's, don't becomes do + n't.
  • wordpunct_tokenize is one regular expression, \w+|[^\w\s]+: every run of letters and digits, and every run of punctuation, is a token.
  • TreebankWordTokenizer expects one sentence at a time, so it only detaches the full stop at the very end of the string it is given.
python
from nltk.tokenize import word_tokenize, wordpunct_tokenize, TreebankWordTokenizer

word_tokenize(corpus)                        # Treebank rules, sentence by sentence
wordpunct_tokenize(corpus)                   # every punctuation run is a token
TreebankWordTokenizer().tokenize(corpus)     # one sentence assumed

Comparing the three word tokenizers on the corpus

ExampleFrom the video, run on NLTK 3.10.3
corpus = """Hello Welcome,to Krish Naik's NLP Tutorials.
Please do watch the entire course! to become expert in NLP.
"""

from nltk.tokenize import sent_tokenize, word_tokenize, wordpunct_tokenize, TreebankWordTokenizer

for name, tokens in [("word_tokenize", word_tokenize(corpus)),
                     ("wordpunct_tokenize", wordpunct_tokenize(corpus)),
                     ("TreebankWordTokenizer", TreebankWordTokenizer().tokenize(corpus))]:
    print(f"{name} ({len(tokens)}):", tokens)

for sentence in sent_tokenize(corpus):
    print(word_tokenize(sentence))

print(word_tokenize("I don't think it's late, isn't it?"))
print(wordpunct_tokenize("I don't think it's late, isn't it?"))
On the same corpus word_tokenize gives 23 tokens with Naik and 's apart, wordpunct_tokenize gives 24 because it splits 's into an apostrophe and s, and TreebankWordTokenizer gives 22 because it keeps Tutorials. with its full stop.

Reading the token lists

  • 23, 24 and 22 tokens, the same counts as in the video's notebook.
  • word_tokenize gives Naik and 's as two tokens and every full stop as its own token.
  • wordpunct_tokenize has one token more: ' and s.
  • TreebankWordTokenizer has one token less: Tutorials. keeps its full stop.
  • Sentence by sentence, word_tokenize gives the three lists the video prints in its loop.
  • Contractions: word_tokenize gives do and n't, is and n't; wordpunct_tokenize gives don, ', t, which loses the negation's shape.

word_tokenize vs wordpunct_tokenize vs TreebankWordTokenizer

word_tokenizewordpunct_tokenizeTreebankWordTokenizer
Methodsent_tokenize, then improved Treebank rulesOne regular expressionTreebank rules on the whole string
Naik'sNaik, 'sNaik, ', sNaik, 's
Full stop inside the textSplit offSplit offKept: Tutorials.
don'tdo, n'tdon, ', tdo, n't
Needs punkt_tabYesNoNo
Tokens on the corpus232422

Where you use NLTK's tokenizers

  • Before stemming, lemmatization and stopword removal: each of them works on word tokens.
  • Sentence-level tasks: summarising, translating or tagging one sentence at a time starts with sent_tokenize.
  • Counting: word frequencies for the bag of words in part 3.
Watch out. On a fresh install, sent_tokenize and word_tokenize stop with LookupError: Resource 'punkt_tab' not found. Run nltk.download('punkt_tab') once. wordpunct_tokenize and TreebankWordTokenizer need no data, which is why they still run when the others fail.
Try it yourself
  • Pass language="german" to sent_tokenize on a German sentence pair such as "Ich bin Dr. Meyer. Wer sind Sie?" and count the sentences.
  • Add a space after the comma in "Welcome,to" and check whether any token list changes.
  • Run TreebankWordTokenizer().tokenize on each sentence from sent_tokenize and compare the result with word_tokenize(corpus).
PreviousTokenization

This is what real progress feels like.