Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Stopwords

Stopwords are very common words such as the, is, of and I that carry little meaning on their own, and stopword removal is the preprocessing step that filters them out before text becomes vectors.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Lemmatization and stemming merge the forms of a word. Stopword removal drops whole words, the frequent ones that would otherwise fill every vector.

Downloading the stopword lists

python
import nltk
nltk.download("stopwords")    # once per machine: stopword lists for 33 languages

Finding the words that carry little meaning

Stopwords and NLTK's lists · from the Complete NLP Machine Learning in One Shot video · 1:15:34 to 1:20:31

Text preprocessing cleans the data and puts it in the right format, because a model trains with mathematical equations and needs numbers. The video's corpus is a speech by Dr APJ Abdul Kalam, former President of India, beginning "I have three visions for India."

Words such as I, the, have, of, to, there and why do not play a big role in tasks like spam or ham classification, or deciding whether a review is positive or negative. Some words do: not plays a very important role. Stopword removal passes the paragraph through a list of such words and removes them.

NLTK ships the lists: from nltk.corpus import stopwords, download them once, then stopwords.words('english') returns the English list. Words in it such as aren't, couldn't and not can change whether a statement is positive or negative, so it is often a good idea to create your own list. There are lists for German, French, Arabic and other languages.

Reading NLTK's stopword lists

ExampleRun on NLTK 3.10.3
from nltk.corpus import stopwords

english = stopwords.words("english")
print(len(english), english[:12])
print([w for w in english if w in ("not", "no", "nor") or w.endswith("n't")])
print(len(stopwords.fileids()), "languages, Hindi included:", "hindi" in stopwords.fileids())
print(len(stopwords.words("german")), stopwords.words("german")[:8])

What the lists contain

  • 198 English stopwords, sorted a to z on today's NLTK.
  • The negations are in it: not, no, nor and contractions such as don't, isn't and wasn't. Removing them turns a negative sentence into a positive one.
  • 33 languages, with no Hindi list; Hinglish, Bengali, Nepali and Tamil are among them. The German list has 232 words.

Removing stopwords from the Kalam speech

The video combines everything so far on the speech: split it into sentences, split each sentence into words, keep the words that are not stopwords, stem them, and join them back into a sentence.

The stopword set

The list is lower case, so each word is lower-cased before the test; without that, I, We and Because would pass the filter. A set, built once, makes each membership test fast; rebuilding it inside the loop would rebuild it for every word.

python
stop = set(stopwords.words("english"))    # built once, fast lookups
stemmer = PorterStemmer()

The filter and stem loop

python
sentences = nltk.sent_tokenize(paragraph)
for i in range(len(sentences)):
    words = nltk.word_tokenize(sentences[i])
    words = [stemmer.stem(w) for w in words if w.lower() not in stop]
    sentences[i] = " ".join(words)    # the list of words back into a sentence
ExampleFrom the video, run on NLTK 3.10.3
# Speech of Dr APJ Abdul Kalam
paragraph = """I have three visions for India. In 3000 years of our history, people from all over 
               the world have come and invaded us, captured our lands, conquered our minds. 
               From Alexander onwards, the Greeks, the Turks, the Moguls, the Portuguese, the British,
               the French, the Dutch, all of them came and looted us, took over what was ours. 
               Yet we have not done this to any other nation. We have not conquered anyone. 
               We have not grabbed their land, their culture, 
               their history and tried to enforce our way of life on them. 
               Why? Because we respect the freedom of others.That is why my 
               first vision is that of freedom. I believe that India got its first vision of 
               this in 1857, when we started the War of Independence. It is this freedom that
               we must protect and nurture and build on. If we are not free, no one will respect us.
               My second vision for India’s development. For fifty years we have been a developing nation.
               It is time we see ourselves as a developed nation. We are among the top 5 nations of the world
               in terms of GDP. We have a 10 percent growth rate in most areas. Our poverty levels are falling.
               Our achievements are being globally recognised today. Yet we lack the self-confidence to
               see ourselves as a developed nation, self-reliant and self-assured. Isn’t this incorrect?
               I have a third vision. India must stand up to the world. Because I believe that unless India 
               stands up to the world, no one will respect us. Only strength respects strength. We must be 
               strong not only as a military power but also as an economic power. Both must go hand-in-hand. 
               My good fortune was to have worked with three great minds. Dr. Vikram Sarabhai of the Dept. of 
               space, Professor Satish Dhawan, who succeeded him and Dr. Brahm Prakash, father of nuclear material.
               I was lucky to have worked with all three of them closely and consider this the great opportunity of my life. 
               I see four milestones in my career"""

import nltk
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer

stop = set(stopwords.words("english"))
stemmer = PorterStemmer()

sentences = nltk.sent_tokenize(paragraph)
before = after = 0
for i in range(len(sentences)):
    words = nltk.word_tokenize(sentences[i])
    before += len(words)
    words = [stemmer.stem(w) for w in words if w.lower() not in stop]
    after += len(words)
    sentences[i] = " ".join(words)

print(len(sentences), "sentences,", before, "tokens before,", after, "after")
for s in sentences[:4] + sentences[7:8] + sentences[-1:]:
    print(s)

What the filter removed

  • 31 sentences, and 399 tokens shrink to 218: stopwords are almost half the speech.
  • "I have three visions for India." became three vision india . : I, have and for are gone, and Porter stemmed visions.
  • histori, peopl, invad, captur: the stems from the stemming lesson, now in context.
  • others.that survived as one token, because the speech has no space after the full stop. Text cleaning and normalisation fixes such text before tokenizing.
  • us stays: it is not in NLTK's list. Every list is a choice.
We have not conquered anyone. is tokenized and lower-cased, then the NLTK stopwords we, have and not are removed and Porter stems the rest to conquer anyon; the negation is lost, and a custom list without not keeps it.

Keeping not with a custom stopword list

The fifth sentence of the speech, "We have not conquered anyone.", loses its not. A custom list is NLTK's list minus the negations:

ExampleRun on NLTK 3.10.3 and scikit-learn 1.9.1
import nltk
from nltk.corpus import stopwords

stop = set(stopwords.words("english"))
negations = {"not", "no", "nor"} | {w for w in stop if w.endswith("n't")}
custom = stop - negations

for s in ["The food is good", "The food is not good", "We have not conquered anyone."]:
    tokens = nltk.word_tokenize(s.lower())
    print([w for w in tokens if w not in stop], [w for w in tokens if w not in custom])

from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS
print("scikit-learn's list:", len(ENGLISH_STOP_WORDS), "words, not included:", "not" in ENGLISH_STOP_WORDS)

Reading the custom list

  • With NLTK's list, "The food is good" and "The food is not good" both become food good: two opposite reviews give the same tokens.
  • With the custom list, the second one keeps not: food not good.
  • scikit-learn's built-in English list, used by stop_words='english' in its vectorizers, has 318 words and contains not too.

NLTK's stopwords vs scikit-learn's stop_words='english'

NLTK stopwords.words('english')scikit-learn ENGLISH_STOP_WORDS
Size198318
Contains notYesYes
Contractions such as don'tYesNo
Other languages33 listsEnglish only
Where it is usedYour own filter loopCountVectorizer(stop_words='english')

Where you use stopword removal

  • Bag of words and TF-IDF: fewer, more meaningful columns for spam and topic classifiers.
  • Keyword extraction and word clouds, where the and of would top every list.
  • Search indexes, which skip the commonest words to stay small.
Watch out. Do not remove stopwords before part-of-speech tagging or named entity recognition, and not for transformer models: they read the whole sentence, and the small words are context they use. For sentiment, keep the negations.
Try it yourself
  • Replace PorterStemmer with WordNetLemmatizer and lemmatize(w.lower(), pos='v') in the loop and compare the first four sentences.
  • Add "us" to stop and count the tokens after filtering again.
  • Print stopwords.words("hinglish")[:20] and pick out three words.
PreviousLemmatization

Slow is fine. Stopping is the only problem.