Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Lemmatization

Lemmatization is a text normalisation technique that maps each word to its lemma, the dictionary form found in a lexicon such as WordNet, using the word's part of speech (going → go as a verb, history → history).

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Stemming ended with forms that are not words, histori and goe. Lemmatization fixes that by looking words up instead of cutting them.

Downloading WordNet

python
import nltk
nltk.download("wordnet")    # once per machine: the WordNet dictionary

Finding the lemma with WordNetLemmatizer

WordNetLemmatizer and the POS argument · from the Complete NLP Machine Learning in One Shot video · 1:05:53 to 1:10:28

Lemmatization is like stemming, but its output is a lemma: a valid word that means the same thing, where stemming gives a stem. Eating becomes eat, history stays history, goes becomes go. The aim is the exact, meaningful form of the word, with the meaning unchanged.

NLTK provides the WordNetLemmatizer class, a thin wrapper around the WordNet corpus. It uses the morphy() function of the WordNet corpus reader to find a lemma: a dictionary of words it compares against.

lemmatize takes two parameters, the word and its POS tag (part of speech), with the default pos='n', noun. The codes are n for noun, v for verb, a for adjective and r for adverb. Lemmatizing "going" as a noun returns going; as a verb it returns go; as an adjective or an adverb, going again. For going, the verb is the correct choice.

python
from nltk.stem import WordNetLemmatizer

lemmatizer = WordNetLemmatizer()
lemmatizer.lemmatize("going")             # 'going', read as a noun
lemmatizer.lemmatize("going", pos="v")    # 'go'

Lemmatizing the ten words as nouns, adjectives and verbs

ExampleFrom the video, run on NLTK 3.10.3
words = ["eating", "eats", "eaten", "writing", "writes", "programming", "programs", "history", "finally", "finalized"]

from nltk.stem import WordNetLemmatizer

lemmatizer = WordNetLemmatizer()
print(f"{'word':13}{'noun':13}{'adjective':13}verb")
for word in words:
    print(f"{word:13}{lemmatizer.lemmatize(word):13}{lemmatizer.lemmatize(word, pos='a'):13}"
          f"{lemmatizer.lemmatize(word, pos='v')}")

print(lemmatizer.lemmatize("going"), lemmatizer.lemmatize("going", pos="v"))
print(lemmatizer.lemmatize("goes"), lemmatizer.lemmatize("goes", pos="v"))
print(lemmatizer.lemmatize("fairly", pos="v"), lemmatizer.lemmatize("sportingly"))

Reading the three columns

  • As nouns, only programs changes, to program. The other words are not noun forms in WordNet, so they come back unchanged.
  • As adjectives, nothing changes.
  • As verbs, eating, eats and eaten all become eat, writing and writes write, programming and programs program, and finalized finalize. History and finally stay as they are, real words, where Porter gave histori and final.
  • goes becomes go with either POS, because WordNet lists "go" as a noun too.
  • fairly and sportingly stay unchanged: an adverb's lemma is itself. Lemmatization undoes inflection (goes → go), not derivation (fairly → fair).
Thirteen words as Porter stems and as WordNet lemmas: with the default noun POS only programs and goes change, while with pos v eating, eats and eaten all become eat, history stays history instead of histori, and fairly stays fairly.

Picking the POS for each word

The lemmatizer only knows the word you give it, so it cannot tell a verb from a noun. Passing pos='v' for every word fixes the verbs and misses the plural nouns; the right POS has to come from the sentence. NLTK's tagger, nltk.pos_tag, labels each word with a Penn Treebank tag, covered in Parts of speech (POS) tagging. Its first letter maps to WordNet's codes: J for adjectives, V for verbs, R for adverbs, everything else a noun.

python
def get_wordnet_pos(tag):
    if tag.startswith("J"):
        return "a"    # JJ, JJR, JJS: adjectives
    if tag.startswith("V"):
        return "v"    # VB, VBD, VBG ...: verbs
    if tag.startswith("R"):
        return "r"    # RB, RBR, RBS: adverbs
    return "n"        # nouns and the rest

Downloading the tagger

python
nltk.download("punkt_tab")
nltk.download("averaged_perceptron_tagger_eng")
ExampleRun on NLTK 3.10.3
import nltk
from nltk.stem import WordNetLemmatizer, PorterStemmer

def get_wordnet_pos(tag):
    if tag.startswith("J"):
        return "a"
    if tag.startswith("V"):
        return "v"
    if tag.startswith("R"):
        return "r"
    return "n"

lemmatizer = WordNetLemmatizer()
sentence = "The children were eating mangoes and the cats were sitting on better chairs"
tagged = nltk.pos_tag(nltk.word_tokenize(sentence))
print(tagged)

print("pos from the tagger:", [lemmatizer.lemmatize(w.lower(), get_wordnet_pos(t)) for w, t in tagged])
print("pos='n' for all:   ", [lemmatizer.lemmatize(w.lower()) for w, t in tagged])
print("pos='v' for all:   ", [lemmatizer.lemmatize(w.lower(), "v") for w, t in tagged])
print("Porter stems:      ", [PorterStemmer().stem(w) for w, t in tagged])

What the tagged lemmas show

  • With the tagger's POS every word reaches its lemma: children → child, were → be, eating → eat, mangoes → mango, sitting → sit, and better (JJR, a comparative adjective) → good.
  • All nouns fixes the plurals but leaves were, eating and sitting.
  • All verbs fixes the verbs but leaves children and mangoes.
  • Porter gives eat and sit, but keeps were, children and better, and has no way to reach good.

Comparing the speed of stemming and lemmatization

The video's question: which takes more time, the WordNet lemmatizer or stemming? Lemmatization costs more in a pipeline, for two reasons. The first call loads WordNet, an 11 MB dictionary, into memory, which takes a moment. And a correct lemma needs the POS of every word, so the whole text goes through the tagger first, the slowest step of the three. Once WordNet is loaded, one lookup is a fast dictionary search; a stemmer needs no data and no tags at all.

Stemming vs lemmatization

StemmingLemmatization
OutputA stem, not always a word (histori)A lemma, a dictionary word (history); a word WordNet does not know comes back unchanged
MethodRules that cut suffixesA dictionary lookup (WordNet) plus rules
Needs the POSNoYes: pos='n' by default
Irregular formseaten → eaten, goes → goeeaten → eat, goes → go (as a verb)
CostNo data, very fastWordNet download and load, a tagger for the POS
Use it forClassification, searchOutput people read: keywords, topics, chatbot answers

Where you use lemmatization

  • Question answering and chatbots built on classic pipelines, where the matched words must be real words.
  • Text summarisation and keyword extraction: a keyword list of lemmas reads naturally.
  • Topic labels and word counts shown in a report.
Watch out. lemmatize is case-sensitive: lemmatize("Eating", pos="v") returns Eating, while a stemmer lower-cases first. Lower-case the words before lemmatizing, and pass a POS; with the default noun most verbs come back unchanged.
Try it yourself
  • Lemmatize "mice", "studies" and "was" with the right POS and stem the same words with Porter.
  • Lemmatize "better" with pos='a', pos='r' and the default, and explain the three answers.
  • Run the tagged-lemma example on "The striped bats are hanging on their feet".
PreviousStemming

You understood something today that you didn't yesterday.