Stopwords
Stopwords are very common words such as the, is, of and I that carry little meaning on their own, and stopword removal is the preprocessing step that filters them out before text becomes vectors.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
Lemmatization and stemming merge the forms of a word. Stopword removal drops whole words, the frequent ones that would otherwise fill every vector.
Downloading the stopword lists
import nltk
nltk.download("stopwords") # once per machine: stopword lists for 33 languagesFinding the words that carry little meaning
Text preprocessing cleans the data and puts it in the right format, because a model trains with mathematical equations and needs numbers. The video's corpus is a speech by Dr APJ Abdul Kalam, former President of India, beginning "I have three visions for India."
Words such as I, the, have, of, to, there and why do not play a big role in tasks like spam or ham classification, or deciding whether a review is positive or negative. Some words do: not plays a very important role. Stopword removal passes the paragraph through a list of such words and removes them.
NLTK ships the lists: from nltk.corpus import stopwords, download them once, then stopwords.words('english') returns the English list. Words in it such as aren't, couldn't and not can change whether a statement is positive or negative, so it is often a good idea to create your own list. There are lists for German, French, Arabic and other languages.
Reading NLTK's stopword lists
from nltk.corpus import stopwords
english = stopwords.words("english")
print(len(english), english[:12])
print([w for w in english if w in ("not", "no", "nor") or w.endswith("n't")])
print(len(stopwords.fileids()), "languages, Hindi included:", "hindi" in stopwords.fileids())
print(len(stopwords.words("german")), stopwords.words("german")[:8])198 ['a', 'about', 'above', 'after', 'again', 'against', 'ain', 'all', 'am', 'an', 'and', 'any'] ["aren't", "couldn't", "didn't", "doesn't", "don't", "hadn't", "hasn't", "haven't", "isn't", "mightn't", "mustn't", "needn't", 'no', 'nor', 'not', "shan't", "shouldn't", "wasn't", "weren't", "won't", "wouldn't"] 33 languages, Hindi included: False 232 ['aber', 'alle', 'allem', 'allen', 'aller', 'alles', 'als', 'also']
What the lists contain
- 198 English stopwords, sorted a to z on today's NLTK.
- The negations are in it: not, no, nor and contractions such as don't, isn't and wasn't. Removing them turns a negative sentence into a positive one.
- 33 languages, with no Hindi list; Hinglish, Bengali, Nepali and Tamil are among them. The German list has 232 words.
Removing stopwords from the Kalam speech
The video combines everything so far on the speech: split it into sentences, split each sentence into words, keep the words that are not stopwords, stem them, and join them back into a sentence.
The stopword set
The list is lower case, so each word is lower-cased before the test; without that, I, We and Because would pass the filter. A set, built once, makes each membership test fast; rebuilding it inside the loop would rebuild it for every word.
stop = set(stopwords.words("english")) # built once, fast lookups
stemmer = PorterStemmer()The filter and stem loop
sentences = nltk.sent_tokenize(paragraph)
for i in range(len(sentences)):
words = nltk.word_tokenize(sentences[i])
words = [stemmer.stem(w) for w in words if w.lower() not in stop]
sentences[i] = " ".join(words) # the list of words back into a sentence# Speech of Dr APJ Abdul Kalam
paragraph = """I have three visions for India. In 3000 years of our history, people from all over
the world have come and invaded us, captured our lands, conquered our minds.
From Alexander onwards, the Greeks, the Turks, the Moguls, the Portuguese, the British,
the French, the Dutch, all of them came and looted us, took over what was ours.
Yet we have not done this to any other nation. We have not conquered anyone.
We have not grabbed their land, their culture,
their history and tried to enforce our way of life on them.
Why? Because we respect the freedom of others.That is why my
first vision is that of freedom. I believe that India got its first vision of
this in 1857, when we started the War of Independence. It is this freedom that
we must protect and nurture and build on. If we are not free, no one will respect us.
My second vision for India’s development. For fifty years we have been a developing nation.
It is time we see ourselves as a developed nation. We are among the top 5 nations of the world
in terms of GDP. We have a 10 percent growth rate in most areas. Our poverty levels are falling.
Our achievements are being globally recognised today. Yet we lack the self-confidence to
see ourselves as a developed nation, self-reliant and self-assured. Isn’t this incorrect?
I have a third vision. India must stand up to the world. Because I believe that unless India
stands up to the world, no one will respect us. Only strength respects strength. We must be
strong not only as a military power but also as an economic power. Both must go hand-in-hand.
My good fortune was to have worked with three great minds. Dr. Vikram Sarabhai of the Dept. of
space, Professor Satish Dhawan, who succeeded him and Dr. Brahm Prakash, father of nuclear material.
I was lucky to have worked with all three of them closely and consider this the great opportunity of my life.
I see four milestones in my career"""
import nltk
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer
stop = set(stopwords.words("english"))
stemmer = PorterStemmer()
sentences = nltk.sent_tokenize(paragraph)
before = after = 0
for i in range(len(sentences)):
words = nltk.word_tokenize(sentences[i])
before += len(words)
words = [stemmer.stem(w) for w in words if w.lower() not in stop]
after += len(words)
sentences[i] = " ".join(words)
print(len(sentences), "sentences,", before, "tokens before,", after, "after")
for s in sentences[:4] + sentences[7:8] + sentences[-1:]:
print(s)31 sentences, 399 tokens before, 218 after three vision india . 3000 year histori , peopl world come invad us , captur land , conquer mind . alexand onward , greek , turk , mogul , portugues , british , french , dutch , came loot us , took . yet done nation . respect freedom others.that first vision freedom . see four mileston career
What the filter removed
- 31 sentences, and 399 tokens shrink to 218: stopwords are almost half the speech.
- "I have three visions for India." became three vision india . : I, have and for are gone, and Porter stemmed visions.
- histori, peopl, invad, captur: the stems from the stemming lesson, now in context.
- others.that survived as one token, because the speech has no space after the full stop. Text cleaning and normalisation fixes such text before tokenizing.
- us stays: it is not in NLTK's list. Every list is a choice.
Keeping not with a custom stopword list
The fifth sentence of the speech, "We have not conquered anyone.", loses its not. A custom list is NLTK's list minus the negations:
import nltk
from nltk.corpus import stopwords
stop = set(stopwords.words("english"))
negations = {"not", "no", "nor"} | {w for w in stop if w.endswith("n't")}
custom = stop - negations
for s in ["The food is good", "The food is not good", "We have not conquered anyone."]:
tokens = nltk.word_tokenize(s.lower())
print([w for w in tokens if w not in stop], [w for w in tokens if w not in custom])
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS
print("scikit-learn's list:", len(ENGLISH_STOP_WORDS), "words, not included:", "not" in ENGLISH_STOP_WORDS)['food', 'good'] ['food', 'good'] ['food', 'good'] ['food', 'not', 'good'] ['conquered', 'anyone', '.'] ['not', 'conquered', 'anyone', '.'] scikit-learn's list: 318 words, not included: True
Reading the custom list
- With NLTK's list, "The food is good" and "The food is not good" both become food good: two opposite reviews give the same tokens.
- With the custom list, the second one keeps not: food not good.
- scikit-learn's built-in English list, used by
stop_words='english'in its vectorizers, has 318 words and contains not too.
NLTK's stopwords vs scikit-learn's stop_words='english'
| NLTK stopwords.words('english') | scikit-learn ENGLISH_STOP_WORDS | |
|---|---|---|
| Size | 198 | 318 |
| Contains not | Yes | Yes |
| Contractions such as don't | Yes | No |
| Other languages | 33 lists | English only |
| Where it is used | Your own filter loop | CountVectorizer(stop_words='english') |
Where you use stopword removal
- Bag of words and TF-IDF: fewer, more meaningful columns for spam and topic classifiers.
- Keyword extraction and word clouds, where the and of would top every list.
- Search indexes, which skip the commonest words to stay small.
Related
- Previous: Lemmatization
- Next: Parts of speech (POS) tagging
- Notebook: 7-Text Preprocessing-Stopwords With NLTK
- Reference: NLTK data
- Replace
PorterStemmerwithWordNetLemmatizerandlemmatize(w.lower(), pos='v')in the loop and compare the first four sentences. - Add "us" to
stopand count the tokens after filtering again. - Print
stopwords.words("hinglish")[:20]and pick out three words.
Slow is fine. Stopping is the only problem.