Lemmatization
Lemmatization is a text normalisation technique that maps each word to its lemma, the dictionary form found in a lexicon such as WordNet, using the word's part of speech (going → go as a verb, history → history).
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
Stemming ended with forms that are not words, histori and goe. Lemmatization fixes that by looking words up instead of cutting them.
Downloading WordNet
import nltk
nltk.download("wordnet") # once per machine: the WordNet dictionaryFinding the lemma with WordNetLemmatizer
Lemmatization is like stemming, but its output is a lemma: a valid word that means the same thing, where stemming gives a stem. Eating becomes eat, history stays history, goes becomes go. The aim is the exact, meaningful form of the word, with the meaning unchanged.
NLTK provides the WordNetLemmatizer class, a thin wrapper around the WordNet corpus. It uses the morphy() function of the WordNet corpus reader to find a lemma: a dictionary of words it compares against.
lemmatize takes two parameters, the word and its POS tag (part of speech), with the default pos='n', noun. The codes are n for noun, v for verb, a for adjective and r for adverb. Lemmatizing "going" as a noun returns going; as a verb it returns go; as an adjective or an adverb, going again. For going, the verb is the correct choice.
from nltk.stem import WordNetLemmatizer
lemmatizer = WordNetLemmatizer()
lemmatizer.lemmatize("going") # 'going', read as a noun
lemmatizer.lemmatize("going", pos="v") # 'go'Lemmatizing the ten words as nouns, adjectives and verbs
words = ["eating", "eats", "eaten", "writing", "writes", "programming", "programs", "history", "finally", "finalized"]
from nltk.stem import WordNetLemmatizer
lemmatizer = WordNetLemmatizer()
print(f"{'word':13}{'noun':13}{'adjective':13}verb")
for word in words:
print(f"{word:13}{lemmatizer.lemmatize(word):13}{lemmatizer.lemmatize(word, pos='a'):13}"
f"{lemmatizer.lemmatize(word, pos='v')}")
print(lemmatizer.lemmatize("going"), lemmatizer.lemmatize("going", pos="v"))
print(lemmatizer.lemmatize("goes"), lemmatizer.lemmatize("goes", pos="v"))
print(lemmatizer.lemmatize("fairly", pos="v"), lemmatizer.lemmatize("sportingly"))word noun adjective verb eating eating eating eat eats eats eats eat eaten eaten eaten eat writing writing writing write writes writes writes write programming programming programming program programs program programs program history history history history finally finally finally finally finalized finalized finalized finalize going go go go fairly sportingly
Reading the three columns
- As nouns, only programs changes, to program. The other words are not noun forms in WordNet, so they come back unchanged.
- As adjectives, nothing changes.
- As verbs, eating, eats and eaten all become eat, writing and writes write, programming and programs program, and finalized finalize. History and finally stay as they are, real words, where Porter gave histori and final.
- goes becomes go with either POS, because WordNet lists "go" as a noun too.
- fairly and sportingly stay unchanged: an adverb's lemma is itself. Lemmatization undoes inflection (goes → go), not derivation (fairly → fair).
Picking the POS for each word
The lemmatizer only knows the word you give it, so it cannot tell a verb from a noun. Passing pos='v' for every word fixes the verbs and misses the plural nouns; the right POS has to come from the sentence. NLTK's tagger, nltk.pos_tag, labels each word with a Penn Treebank tag, covered in Parts of speech (POS) tagging. Its first letter maps to WordNet's codes: J for adjectives, V for verbs, R for adverbs, everything else a noun.
def get_wordnet_pos(tag):
if tag.startswith("J"):
return "a" # JJ, JJR, JJS: adjectives
if tag.startswith("V"):
return "v" # VB, VBD, VBG ...: verbs
if tag.startswith("R"):
return "r" # RB, RBR, RBS: adverbs
return "n" # nouns and the restDownloading the tagger
nltk.download("punkt_tab")
nltk.download("averaged_perceptron_tagger_eng")import nltk
from nltk.stem import WordNetLemmatizer, PorterStemmer
def get_wordnet_pos(tag):
if tag.startswith("J"):
return "a"
if tag.startswith("V"):
return "v"
if tag.startswith("R"):
return "r"
return "n"
lemmatizer = WordNetLemmatizer()
sentence = "The children were eating mangoes and the cats were sitting on better chairs"
tagged = nltk.pos_tag(nltk.word_tokenize(sentence))
print(tagged)
print("pos from the tagger:", [lemmatizer.lemmatize(w.lower(), get_wordnet_pos(t)) for w, t in tagged])
print("pos='n' for all: ", [lemmatizer.lemmatize(w.lower()) for w, t in tagged])
print("pos='v' for all: ", [lemmatizer.lemmatize(w.lower(), "v") for w, t in tagged])
print("Porter stems: ", [PorterStemmer().stem(w) for w, t in tagged])[('The', 'DT'), ('children', 'NNS'), ('were', 'VBD'), ('eating', 'VBG'), ('mangoes', 'NNS'), ('and', 'CC'), ('the', 'DT'), ('cats', 'NNS'), ('were', 'VBD'), ('sitting', 'VBG'), ('on', 'IN'), ('better', 'JJR'), ('chairs', 'NNS')]
pos from the tagger: ['the', 'child', 'be', 'eat', 'mango', 'and', 'the', 'cat', 'be', 'sit', 'on', 'good', 'chair']
pos='n' for all: ['the', 'child', 'were', 'eating', 'mango', 'and', 'the', 'cat', 'were', 'sitting', 'on', 'better', 'chair']
pos='v' for all: ['the', 'children', 'be', 'eat', 'mangoes', 'and', 'the', 'cat', 'be', 'sit', 'on', 'better', 'chair']
Porter stems: ['the', 'children', 'were', 'eat', 'mango', 'and', 'the', 'cat', 'were', 'sit', 'on', 'better', 'chair']What the tagged lemmas show
- With the tagger's POS every word reaches its lemma: children → child, were → be, eating → eat, mangoes → mango, sitting → sit, and better (JJR, a comparative adjective) → good.
- All nouns fixes the plurals but leaves were, eating and sitting.
- All verbs fixes the verbs but leaves children and mangoes.
- Porter gives eat and sit, but keeps were, children and better, and has no way to reach good.
Comparing the speed of stemming and lemmatization
The video's question: which takes more time, the WordNet lemmatizer or stemming? Lemmatization costs more in a pipeline, for two reasons. The first call loads WordNet, an 11 MB dictionary, into memory, which takes a moment. And a correct lemma needs the POS of every word, so the whole text goes through the tagger first, the slowest step of the three. Once WordNet is loaded, one lookup is a fast dictionary search; a stemmer needs no data and no tags at all.
Stemming vs lemmatization
| Stemming | Lemmatization | |
|---|---|---|
| Output | A stem, not always a word (histori) | A lemma, a dictionary word (history); a word WordNet does not know comes back unchanged |
| Method | Rules that cut suffixes | A dictionary lookup (WordNet) plus rules |
| Needs the POS | No | Yes: pos='n' by default |
| Irregular forms | eaten → eaten, goes → goe | eaten → eat, goes → go (as a verb) |
| Cost | No data, very fast | WordNet download and load, a tagger for the POS |
| Use it for | Classification, search | Output people read: keywords, topics, chatbot answers |
Where you use lemmatization
- Question answering and chatbots built on classic pipelines, where the matched words must be real words.
- Text summarisation and keyword extraction: a keyword list of lemmas reads naturally.
- Topic labels and word counts shown in a report.
lemmatize is case-sensitive: lemmatize("Eating", pos="v") returns Eating, while a stemmer lower-cases first. Lower-case the words before lemmatizing, and pass a POS; with the default noun most verbs come back unchanged.Related
- Previous: Stemming
- Next: Stopwords
- Notebook: 6-Lemmatization
- Reference: nltk.stem.wordnet and WordNet
- Lemmatize "mice", "studies" and "was" with the right POS and stem the same words with Porter.
- Lemmatize "better" with
pos='a',pos='r'and the default, and explain the three answers. - Run the tagged-lemma example on "The striped bats are hanging on their feet".
You understood something today that you didn't yesterday.