Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

N-grams

N-grams are sequences of n consecutive words, bigrams for two and trigrams for three, that bag of words can count as extra features so the vectors keep some of the word order.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Bag of words (BoW) records which words appear and throws away their order. That is why two sentences with opposite meanings can end up with nearly the same vector.

Comparing “The food is good” with “The food is not good”

The video keeps every word, stopwords included. The vocabulary is the, food, is, not and good, five words, so “The food is good” becomes [1 1 1 0 1] and “The food is not good” becomes [1 1 1 1 1]. Only the value for “not” differs.

The food is good vs the food is not good · from the Complete NLP Machine Learning in One Shot video · 2:28:29 to 2:29:45

Plotted, the two vectors point almost the same way, with a small angle between them. Measured with Cosine similarity for documents, they look like near-identical sentences, yet their meanings are opposite:

The cosine of the two bag of words vectors [1 1 1 0 1] and [1 1 1 1 1]

Turning word pairs into features

The board notes in the video's materials add the fix. Remove “The” and “is” but keep “not”: S1 becomes “food good” and S2 “food not good”. With single words, the unigrams, the vectors over food, not, good are [1 0 1] and [1 1 1].

A bigram is a pair of neighbouring words. S1 has the bigram “food good”; S2 has “food not” and “not good”. Adding these three bigrams as columns gives S1 = [1 0 1 1 0 0] and S2 = [1 1 1 0 1 1]. The sentences now share only two of their features, and the cosine drops from 0.816 to 0.516.

N-grams over food not good: windows of one, two and three words; the unigram vectors of food good and food not good are 1 0 1 and 1 1 1 with cosine 0.816, and adding the bigrams food good, food not and not good gives 1 0 1 1 0 0 and 1 1 1 0 1 1 with cosine 0.516.

A trigram is three words in a row, and in general an n-gram is n words in a row. scikit-learn sets the range with ngram_range=(min_n, max_n), listed on the same page:

ngram_rangeFeatures
(1, 1)unigrams (the default)
(1, 2)unigrams and bigrams
(1, 3)unigrams, bigrams and trigrams
(2, 3)bigrams and trigrams only

Counting n-grams with NLTK and CountVectorizer

Listing the n-grams of a sentence

python
from nltk.util import ngrams
list(ngrams("food not good".split(), 2))     # [('food', 'not'), ('not', 'good')]

Adding n-grams to bag of words

python
cv = CountVectorizer(ngram_range=(1, 2))     # single words and word pairs
X = cv.fit_transform(docs)
cv.get_feature_names_out()                   # 'food', 'food good', 'food not', ...
ExampleFrom the video, run on scikit-learn 1.9.1
from nltk.util import ngrams
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.metrics.pairwise import cosine_similarity

tokens = "food not good".split()
for n in (1, 2, 3):
    print(f"n={n}:", [" ".join(g) for g in ngrams(tokens, n)])

docs = ["food good", "food not good"]
for rng in [(1, 1), (1, 2), (2, 2), (1, 3), (2, 3)]:
    cv = CountVectorizer(ngram_range=rng)
    X = cv.fit_transform(docs).toarray()
    print(rng, list(cv.get_feature_names_out()), X.tolist(), f"cosine {cosine_similarity(X)[0, 1]:.4f}")

What the n-gram vectors show

  • The windows over “food not good” give three unigrams, two bigrams and one trigram.
  • (1, 1) reproduces the unigram table: [1 1 0] and [1 1 1] over food, good, not (alphabetical), cosine 0.8165.
  • (1, 2) reproduces the bigram table with the columns in alphabetical order, and the cosine falls to 0.5164.
  • (2, 2) keeps only the pairs, and the two sentences share none of them: cosine 0.0.
  • (1, 3) and (2, 3) add the trigram “food not good”, which only S2 has: the cosine falls to 0.4714 with (1, 3) and to 0.0 with (2, 3).

Keeping “not” when removing stopwords

NLTK's English stopword list contains “not”, “no” and “nor”. Removing it with the full list turns both sentences into “food good”, and no n-gram can separate them any more. A custom list without the negations keeps the difference.

ExampleFrom the video, run on scikit-learn 1.9.1
import nltk
from nltk.corpus import stopwords

nltk.download("stopwords", quiet=True)
stop = set(stopwords.words("english"))
print("'not' in NLTK's list:", "not" in stop)

keep_negations = stop - {"not", "no", "nor"}
for s in ["The food is good", "The food is not good"]:
    words = s.lower().split()
    print([w for w in words if w not in stop], [w for w in words if w not in keep_negations])

With NLTK's full list both sentences become ['food', 'good']. Without the three negations the second keeps ['food', 'not', 'good'].

N-grams on the SMS spam messages

The practical notebook reruns the SMS bag of words with ngram_range=(2, 3): only bigrams and trigrams, and the 100 most frequent of them. corpus is the stemmed list of 5,572 messages built in Bag of words (BoW).

ExampleFrom the video, run on scikit-learn 1.9.1
cv = CountVectorizer(max_features=100, binary=True, ngram_range=(2, 3))
X = cv.fit_transform(corpus)
names = cv.get_feature_names_out()
print(X.shape)
print(names[:8])
print("trigrams kept:", [n for n in names if n.count(" ") == 2][:5])
print("all bigrams and trigrams:", len(CountVectorizer(ngram_range=(2, 3)).fit(corpus).vocabulary_))
  • The features are stemmed word pairs and triples such as “call claim” and “call custom servic”: phrases typical of spam.
  • Without max_features the corpus has 59,253 distinct bigrams and trigrams, so the limit matters much more than with single words.

Unigrams vs n-grams

Unigrams (1, 1)Unigrams and bigrams (1, 2)
food good vs food not goodcosine 0.8165cosine 0.5164
Word orderlostkept for neighbouring words
Vocabulary sizenumber of distinct wordsmany times larger
Catchestopic wordsphrases and negations such as “not good”

Where you use n-grams

  • Sentiment analysis, where “not good” and “not bad” flip the meaning of a single word.
  • Spam filters, where phrases such as “free entry” or “call claim” say more than either word alone.
  • Search and autocomplete, where counts of word pairs and triples estimate which word comes next.
Watch out. Every extra n raises the vocabulary many times over, and most n-grams appear once. Use max_features or min_df to drop the rare ones, or the matrix gets huge and the model overfits.
Try it yourself
  • Add a third document, “food not bad”, to docs and rerun with (1, 2). Which bigram does it share with “food not good”?
  • Change ngrams(tokens, n) to run on “I am not feeling well”. How many trigrams does it have?
  • In the SMS example, set ngram_range=(1, 2) and look at which single words make it into the top 100.

Slow is fine. Stopping is the only problem.