Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Stemming

Stemming is a text normalisation technique that cuts affixes off a word with fixed rules to reach its stem, a shared form that is not always a dictionary word (eating → eat, history → histori).

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Text cleaning and normalisation made every word one consistent string. Stemming goes one level deeper and merges the different forms of a word into one token.

Reducing word variants to one stem

Take a classification problem: are the comments on a product positive or negative reviews? The reviews are text, and they contain words like eating, eat and eaten, or going, gone and goes. Each group talks about the same thing, and for deciding positive or negative the variety does not matter.

Keeping every variant also increases the number of input features, because each word becomes a column of the vector after preprocessing. So each group should become one shared form, its word stem, such as eat. Stemming is the technique that finds it.

A stemmer works only by cutting endings, so it reaches eat from eating and eats, and go from going. Irregular forms such as eaten, gone and goes need a dictionary of word forms, which is what Lemmatization uses.

Stemming with PorterStemmer

PorterStemmer and its disadvantage · from the Complete NLP Machine Learning in One Shot video · 50:01 to 54:21

The video's list has ten words: eating, eats, eaten, writing, writes, programming, programs, history, finally and finalized. The first stemmer is the Porter stemmer, the 1980 suffix-stripping algorithm by Martin Porter. In NLTK it is a class: create an object, then call stem on each word.

python
from nltk.stem import PorterStemmer

stemming = PorterStemmer()
for word in words:
    print(word + "---->" + stemming.stem(word))

Most results are good: eating and eats become eat, writing and writes write, programming and programs program, finally and finalized final. Eaten stays eaten. History becomes histori, which is not a word, and congratulations becomes congratul, where congratulate was expected. This is the major disadvantage of stemming: for some words the form changes and the meaning is lost. Sitting becomes sit, a good result.

For classification problems such as spam or ham and positive or negative reviews, stemming is a common choice: a stem that is not a word does no harm as long as every form maps to it. Other stemmers improve on Porter's rules.

Stemming the video's ten words

ExampleFrom the video, run on NLTK 3.10.3
words = ["eating", "eats", "eaten", "writing", "writes", "programming", "programs", "history", "finally", "finalized"]

from nltk.stem import PorterStemmer

stemming = PorterStemmer()
for word in words:
    print(word + "---->" + stemming.stem(word))

print(stemming.stem("congratulations"), stemming.stem("sitting"))

Writing your own rule with RegexpStemmer

RegexpStemmer · from the Complete NLP Machine Learning in One Shot video · 54:22 to 57:56

RegexpStemmer takes a single regular expression and removes any prefix or suffix that matches it. Created with no arguments it raises an error: the regular expression is required. The second parameter, min, is the minimum length of a word to stem: with min=4, words shorter than four letters are returned unchanged.

The pattern 'ing$|s$|e$|able$' means: at the end of the word, remove ing, s, e or able. $ marks the end. So eating becomes eat, and ingeating becomes ingeat, because only the last ing is at the end. Remove the $ from ing$ and every ing is removed: ingeating becomes eat. To match at the start of a word, the anchor is ^: '^ing' turns ingeating into eating.

python
from nltk.stem import RegexpStemmer

reg_stemmer = RegexpStemmer("ing$|s$|e$|able$", min=4)
reg_stemmer.stem("eating")       # 'eat'
reg_stemmer.stem("ingeating")    # 'ingeat'
ExampleFrom the video, run on NLTK 3.10.3
from nltk.stem import RegexpStemmer

words = ["eating", "ingeating", "bringing", "sing", "able", "cars"]
for pattern in ["ing$|s$|e$|able$", "ing|s$|e$|able$", "^ing|s$|e$|able$"]:
    reg_stemmer = RegexpStemmer(pattern, min=4)
    print(f"{pattern:18}", [reg_stemmer.stem(w) for w in words])
RegexpStemmer with min 4: the pattern ing$ gives eat, ingeat, bring, s, an empty string and car; ing without the anchor removes ing anywhere, so bringing becomes br; ^ing removes it only at the start.

A rule removes every match, so it needs care: ing without an anchor cuts bringing down to br, sing becomes s, and able becomes an empty string, because min checks the word's length before stemming, not the result's.

Improving on Porter with SnowballStemmer

SnowballStemmer · from the Complete NLP Machine Learning in One Shot video · 1:02:43 to 1:04:17

The Snowball stemmer performs better than Porter in the sense of giving a better form of the word. Its English version is Porter2, Porter's own revision of his algorithm, and it is available for other languages too. On fairly and sportingly, Porter gives fairli and sportingli while Snowball gives fair and sport. Both still give goe for goes: however much a stemmer tries, the form of some words changes.

python
from nltk.stem import SnowballStemmer

snowballsstemmer = SnowballStemmer("english")
snowballsstemmer.stem("fairly")     # 'fair'

NLTK has a third common stemmer that the video does not show, the Lancaster stemmer, which cuts much harder: writing becomes writ, history hist.

Comparing Porter, Snowball and Lancaster

ExampleRun on NLTK 3.10.3
words = ["eating", "eats", "eaten", "writing", "writes", "programming", "programs", "history", "finally", "finalized"] + ["congratulations", "fairly", "sportingly", "goes"]

from nltk.stem import PorterStemmer, SnowballStemmer, LancasterStemmer

porter, snowball, lancaster = PorterStemmer(), SnowballStemmer("english"), LancasterStemmer()
print(f"{'word':16}{'Porter':12}{'Snowball':12}Lancaster")
for w in words:
    print(f"{w:16}{porter.stem(w):12}{snowball.stem(w):12}{lancaster.stem(w)}")

print(SnowballStemmer.languages)
print(porter.stem("university"), porter.stem("universe"))
Fourteen words through Porter, Snowball and Lancaster: eating and eats become eat in all three, eaten stays eaten except in Lancaster, history becomes histori or hist, fairly becomes fairli in Porter but fair in Snowball, and goes becomes goe in all three; stems that are not English words are red.

Reading the three stemmers

  • Porter and Snowball agree on the ten words; they differ on fairly and sportingly, where Snowball removes the whole -ly.
  • Lancaster is the most aggressive: eaten becomes eat, but writing becomes writ and finally fin, real words with other meanings.
  • goes → goe in all three: the rules strip the s and leave goe. Reaching go needs a dictionary of word forms, not a suffix rule.
  • Snowball supports 16 entries, from Arabic to Swedish, including porter for the original algorithm.
  • university and universe both become univers: over-stemming merges two different words into one token.

Measuring how much stemming shrinks the vocabulary

Fewer variants means fewer columns. The SMS spam dataset, loaded from the course materials, shows by how much:

ExampleRun on NLTK 3.10.3 and pandas 3.0.6
import re
import pandas as pd
import matplotlib.pyplot as plt
from nltk.stem import PorterStemmer, SnowballStemmer, LancasterStemmer

url = ("https://raw.githubusercontent.com/krishnaik06/The-Grand-Complete-Data-Science-Materials/main/Com"
       "plete%20NLP%20For%20ML%20%26%20Deep%20Learning/Practicals/SpamClassifier-master/smsspamcollection/SMSSpamCollection")
messages = pd.read_csv(url, sep="\t", names=["label", "message"])
words = [w for m in messages["message"] for w in re.findall(r"[A-Za-z]+", m)]
lower = {w.lower() for w in words}

counts = {"as written": len(set(words)), "lower case": len(lower)}
for name, s in [("Porter", PorterStemmer()), ("Snowball", SnowballStemmer("english")),
                ("Lancaster", LancasterStemmer())]:
    counts[name] = len({s.stem(w) for w in lower})
print(len(messages), "messages,", len(words), "words")
print(counts)

plt.figure(figsize=(7, 4))
plt.bar(counts.keys(), counts.values(), color="#9370DB")
plt.title("Distinct words in the SMS spam dataset")
plt.ylabel("vocabulary size")
plt.show()
A bar chart of distinct words in the 5,572 SMS messages: 9,972 as written, 7,785 in lower case, 6,422 after Porter, 6,427 after Snowball and 5,751 after Lancaster.

The 5,572 messages hold 87,450 words. Lower-casing alone cuts the vocabulary from 9,972 to 7,785 distinct words, Porter brings it to 6,422 and the aggressive Lancaster to 5,751. Every merged pair is one column fewer for the bag of words, and one more count shared by the forms of a word.

Porter vs Snowball vs Lancaster

PorterStemmerSnowballStemmerLancasterStemmer
AlgorithmPorter (1980)Porter2, Porter's revisionPaice-Husk (Lancaster, 1990)
How hard it cutsModerateModerate, a few better rulesAggressive
fairlyfairlifairfair
writingwritewritewrit
LanguagesEnglish16 entriesEnglish
SMS vocabulary6,4226,4275,751

Where you use stemming

  • Bag of words and TF-IDF classifiers for spam or sentiment, where fewer, denser columns help.
  • Search: stemming the query and the documents lets "running" find "run".
  • Quick baselines: a stemmer needs no dictionary and no part-of-speech tags.
Watch out. A stem is a key, not a word: never show histori or congratul to a user, and apply the same stemmer to training and test text. Modern transformer models do not stem at all; their subword tokenizers handle word forms.
Try it yourself
  • Stem "generously" with Porter and with Snowball and compare: which one keeps a word?
  • Make a RegexpStemmer with "ing$" and min=5 and stem "sing" and "singing".
  • Pass ignore_stopwords=True to SnowballStemmer("english", ...) and stem "having" with and without it.

Little by little, you're building something great.