Stemming
Stemming is a text normalisation technique that cuts affixes off a word with fixed rules to reach its stem, a shared form that is not always a dictionary word (eating → eat, history → histori).
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
Text cleaning and normalisation made every word one consistent string. Stemming goes one level deeper and merges the different forms of a word into one token.
Reducing word variants to one stem
Take a classification problem: are the comments on a product positive or negative reviews? The reviews are text, and they contain words like eating, eat and eaten, or going, gone and goes. Each group talks about the same thing, and for deciding positive or negative the variety does not matter.
Keeping every variant also increases the number of input features, because each word becomes a column of the vector after preprocessing. So each group should become one shared form, its word stem, such as eat. Stemming is the technique that finds it.
A stemmer works only by cutting endings, so it reaches eat from eating and eats, and go from going. Irregular forms such as eaten, gone and goes need a dictionary of word forms, which is what Lemmatization uses.
Stemming with PorterStemmer
The video's list has ten words: eating, eats, eaten, writing, writes, programming, programs, history, finally and finalized. The first stemmer is the Porter stemmer, the 1980 suffix-stripping algorithm by Martin Porter. In NLTK it is a class: create an object, then call stem on each word.
from nltk.stem import PorterStemmer
stemming = PorterStemmer()
for word in words:
print(word + "---->" + stemming.stem(word))Most results are good: eating and eats become eat, writing and writes write, programming and programs program, finally and finalized final. Eaten stays eaten. History becomes histori, which is not a word, and congratulations becomes congratul, where congratulate was expected. This is the major disadvantage of stemming: for some words the form changes and the meaning is lost. Sitting becomes sit, a good result.
For classification problems such as spam or ham and positive or negative reviews, stemming is a common choice: a stem that is not a word does no harm as long as every form maps to it. Other stemmers improve on Porter's rules.
Stemming the video's ten words
words = ["eating", "eats", "eaten", "writing", "writes", "programming", "programs", "history", "finally", "finalized"]
from nltk.stem import PorterStemmer
stemming = PorterStemmer()
for word in words:
print(word + "---->" + stemming.stem(word))
print(stemming.stem("congratulations"), stemming.stem("sitting"))eating---->eat eats---->eat eaten---->eaten writing---->write writes---->write programming---->program programs---->program history---->histori finally---->final finalized---->final congratul sit
Writing your own rule with RegexpStemmer
RegexpStemmer takes a single regular expression and removes any prefix or suffix that matches it. Created with no arguments it raises an error: the regular expression is required. The second parameter, min, is the minimum length of a word to stem: with min=4, words shorter than four letters are returned unchanged.
The pattern 'ing$|s$|e$|able$' means: at the end of the word, remove ing, s, e or able. $ marks the end. So eating becomes eat, and ingeating becomes ingeat, because only the last ing is at the end. Remove the $ from ing$ and every ing is removed: ingeating becomes eat. To match at the start of a word, the anchor is ^: '^ing' turns ingeating into eating.
from nltk.stem import RegexpStemmer
reg_stemmer = RegexpStemmer("ing$|s$|e$|able$", min=4)
reg_stemmer.stem("eating") # 'eat'
reg_stemmer.stem("ingeating") # 'ingeat'from nltk.stem import RegexpStemmer
words = ["eating", "ingeating", "bringing", "sing", "able", "cars"]
for pattern in ["ing$|s$|e$|able$", "ing|s$|e$|able$", "^ing|s$|e$|able$"]:
reg_stemmer = RegexpStemmer(pattern, min=4)
print(f"{pattern:18}", [reg_stemmer.stem(w) for w in words])ing$|s$|e$|able$ ['eat', 'ingeat', 'bring', 's', '', 'car'] ing|s$|e$|able$ ['eat', 'eat', 'br', 's', '', 'car'] ^ing|s$|e$|able$ ['eating', 'eating', 'bringing', 'sing', '', 'car']
A rule removes every match, so it needs care: ing without an anchor cuts bringing down to br, sing becomes s, and able becomes an empty string, because min checks the word's length before stemming, not the result's.
Improving on Porter with SnowballStemmer
The Snowball stemmer performs better than Porter in the sense of giving a better form of the word. Its English version is Porter2, Porter's own revision of his algorithm, and it is available for other languages too. On fairly and sportingly, Porter gives fairli and sportingli while Snowball gives fair and sport. Both still give goe for goes: however much a stemmer tries, the form of some words changes.
from nltk.stem import SnowballStemmer
snowballsstemmer = SnowballStemmer("english")
snowballsstemmer.stem("fairly") # 'fair'NLTK has a third common stemmer that the video does not show, the Lancaster stemmer, which cuts much harder: writing becomes writ, history hist.
Comparing Porter, Snowball and Lancaster
words = ["eating", "eats", "eaten", "writing", "writes", "programming", "programs", "history", "finally", "finalized"] + ["congratulations", "fairly", "sportingly", "goes"]
from nltk.stem import PorterStemmer, SnowballStemmer, LancasterStemmer
porter, snowball, lancaster = PorterStemmer(), SnowballStemmer("english"), LancasterStemmer()
print(f"{'word':16}{'Porter':12}{'Snowball':12}Lancaster")
for w in words:
print(f"{w:16}{porter.stem(w):12}{snowball.stem(w):12}{lancaster.stem(w)}")
print(SnowballStemmer.languages)
print(porter.stem("university"), porter.stem("universe"))word Porter Snowball Lancaster
eating eat eat eat
eats eat eat eat
eaten eaten eaten eat
writing write write writ
writes write write writ
programming program program program
programs program program program
history histori histori hist
finally final final fin
finalized final final fin
congratulations congratul congratul congrat
fairly fairli fair fair
sportingly sportingli sport sport
goes goe goe goe
('arabic', 'danish', 'dutch', 'english', 'finnish', 'french', 'german', 'hungarian', 'italian', 'norwegian', 'porter', 'portuguese', 'romanian', 'russian', 'spanish', 'swedish')
univers universReading the three stemmers
- Porter and Snowball agree on the ten words; they differ on fairly and sportingly, where Snowball removes the whole -ly.
- Lancaster is the most aggressive: eaten becomes eat, but writing becomes writ and finally fin, real words with other meanings.
- goes → goe in all three: the rules strip the s and leave goe. Reaching go needs a dictionary of word forms, not a suffix rule.
- Snowball supports 16 entries, from Arabic to Swedish, including
porterfor the original algorithm. - university and universe both become univers: over-stemming merges two different words into one token.
Measuring how much stemming shrinks the vocabulary
Fewer variants means fewer columns. The SMS spam dataset, loaded from the course materials, shows by how much:
import re
import pandas as pd
import matplotlib.pyplot as plt
from nltk.stem import PorterStemmer, SnowballStemmer, LancasterStemmer
url = ("https://raw.githubusercontent.com/krishnaik06/The-Grand-Complete-Data-Science-Materials/main/Com"
"plete%20NLP%20For%20ML%20%26%20Deep%20Learning/Practicals/SpamClassifier-master/smsspamcollection/SMSSpamCollection")
messages = pd.read_csv(url, sep="\t", names=["label", "message"])
words = [w for m in messages["message"] for w in re.findall(r"[A-Za-z]+", m)]
lower = {w.lower() for w in words}
counts = {"as written": len(set(words)), "lower case": len(lower)}
for name, s in [("Porter", PorterStemmer()), ("Snowball", SnowballStemmer("english")),
("Lancaster", LancasterStemmer())]:
counts[name] = len({s.stem(w) for w in lower})
print(len(messages), "messages,", len(words), "words")
print(counts)
plt.figure(figsize=(7, 4))
plt.bar(counts.keys(), counts.values(), color="#9370DB")
plt.title("Distinct words in the SMS spam dataset")
plt.ylabel("vocabulary size")
plt.show()5572 messages, 87450 words
{'as written': 9972, 'lower case': 7785, 'Porter': 6422, 'Snowball': 6427, 'Lancaster': 5751}The 5,572 messages hold 87,450 words. Lower-casing alone cuts the vocabulary from 9,972 to 7,785 distinct words, Porter brings it to 6,422 and the aggressive Lancaster to 5,751. Every merged pair is one column fewer for the bag of words, and one more count shared by the forms of a word.
Porter vs Snowball vs Lancaster
| PorterStemmer | SnowballStemmer | LancasterStemmer | |
|---|---|---|---|
| Algorithm | Porter (1980) | Porter2, Porter's revision | Paice-Husk (Lancaster, 1990) |
| How hard it cuts | Moderate | Moderate, a few better rules | Aggressive |
| fairly | fairli | fair | fair |
| writing | write | write | writ |
| Languages | English | 16 entries | English |
| SMS vocabulary | 6,422 | 6,427 | 5,751 |
Where you use stemming
- Bag of words and TF-IDF classifiers for spam or sentiment, where fewer, denser columns help.
- Search: stemming the query and the documents lets "running" find "run".
- Quick baselines: a stemmer needs no dictionary and no part-of-speech tags.
Related
- Previous: Text cleaning and normalisation
- Next: Lemmatization
- Notebook: 5-Stemming And Its Types
- Reference: nltk.stem and the Porter2 algorithm
- Stem "generously" with Porter and with Snowball and compare: which one keeps a word?
- Make a RegexpStemmer with
"ing$"andmin=5and stem "sing" and "singing". - Pass
ignore_stopwords=TruetoSnowballStemmer("english", ...)and stem "having" with and without it.
Little by little, you're building something great.