N-grams
N-grams are sequences of n consecutive words, bigrams for two and trigrams for three, that bag of words can count as extra features so the vectors keep some of the word order.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
Bag of words (BoW) records which words appear and throws away their order. That is why two sentences with opposite meanings can end up with nearly the same vector.
Comparing “The food is good” with “The food is not good”
The video keeps every word, stopwords included. The vocabulary is the, food, is, not and good, five words, so “The food is good” becomes [1 1 1 0 1] and “The food is not good” becomes [1 1 1 1 1]. Only the value for “not” differs.
Plotted, the two vectors point almost the same way, with a small angle between them. Measured with Cosine similarity for documents, they look like near-identical sentences, yet their meanings are opposite:
Turning word pairs into features
The board notes in the video's materials add the fix. Remove “The” and “is” but keep “not”: S1 becomes “food good” and S2 “food not good”. With single words, the unigrams, the vectors over food, not, good are [1 0 1] and [1 1 1].
A bigram is a pair of neighbouring words. S1 has the bigram “food good”; S2 has “food not” and “not good”. Adding these three bigrams as columns gives S1 = [1 0 1 1 0 0] and S2 = [1 1 1 0 1 1]. The sentences now share only two of their features, and the cosine drops from 0.816 to 0.516.
A trigram is three words in a row, and in general an n-gram is n words in a row. scikit-learn sets the range with ngram_range=(min_n, max_n), listed on the same page:
| ngram_range | Features |
|---|---|
| (1, 1) | unigrams (the default) |
| (1, 2) | unigrams and bigrams |
| (1, 3) | unigrams, bigrams and trigrams |
| (2, 3) | bigrams and trigrams only |
Counting n-grams with NLTK and CountVectorizer
Listing the n-grams of a sentence
from nltk.util import ngrams
list(ngrams("food not good".split(), 2)) # [('food', 'not'), ('not', 'good')]Adding n-grams to bag of words
cv = CountVectorizer(ngram_range=(1, 2)) # single words and word pairs
X = cv.fit_transform(docs)
cv.get_feature_names_out() # 'food', 'food good', 'food not', ...from nltk.util import ngrams
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.metrics.pairwise import cosine_similarity
tokens = "food not good".split()
for n in (1, 2, 3):
print(f"n={n}:", [" ".join(g) for g in ngrams(tokens, n)])
docs = ["food good", "food not good"]
for rng in [(1, 1), (1, 2), (2, 2), (1, 3), (2, 3)]:
cv = CountVectorizer(ngram_range=rng)
X = cv.fit_transform(docs).toarray()
print(rng, list(cv.get_feature_names_out()), X.tolist(), f"cosine {cosine_similarity(X)[0, 1]:.4f}")n=1: ['food', 'not', 'good'] n=2: ['food not', 'not good'] n=3: ['food not good'] (1, 1) ['food', 'good', 'not'] [[1, 1, 0], [1, 1, 1]] cosine 0.8165 (1, 2) ['food', 'food good', 'food not', 'good', 'not', 'not good'] [[1, 1, 0, 1, 0, 0], [1, 0, 1, 1, 1, 1]] cosine 0.5164 (2, 2) ['food good', 'food not', 'not good'] [[1, 0, 0], [0, 1, 1]] cosine 0.0000 (1, 3) ['food', 'food good', 'food not', 'food not good', 'good', 'not', 'not good'] [[1, 1, 0, 0, 1, 0, 0], [1, 0, 1, 1, 1, 1, 1]] cosine 0.4714 (2, 3) ['food good', 'food not', 'food not good', 'not good'] [[1, 0, 0, 0], [0, 1, 1, 1]] cosine 0.0000
What the n-gram vectors show
- The windows over “food not good” give three unigrams, two bigrams and one trigram.
- (1, 1) reproduces the unigram table: [1 1 0] and [1 1 1] over food, good, not (alphabetical), cosine 0.8165.
- (1, 2) reproduces the bigram table with the columns in alphabetical order, and the cosine falls to 0.5164.
- (2, 2) keeps only the pairs, and the two sentences share none of them: cosine 0.0.
- (1, 3) and (2, 3) add the trigram “food not good”, which only S2 has: the cosine falls to 0.4714 with (1, 3) and to 0.0 with (2, 3).
Keeping “not” when removing stopwords
NLTK's English stopword list contains “not”, “no” and “nor”. Removing it with the full list turns both sentences into “food good”, and no n-gram can separate them any more. A custom list without the negations keeps the difference.
import nltk
from nltk.corpus import stopwords
nltk.download("stopwords", quiet=True)
stop = set(stopwords.words("english"))
print("'not' in NLTK's list:", "not" in stop)
keep_negations = stop - {"not", "no", "nor"}
for s in ["The food is good", "The food is not good"]:
words = s.lower().split()
print([w for w in words if w not in stop], [w for w in words if w not in keep_negations])'not' in NLTK's list: True ['food', 'good'] ['food', 'good'] ['food', 'good'] ['food', 'not', 'good']
With NLTK's full list both sentences become ['food', 'good']. Without the three negations the second keeps ['food', 'not', 'good'].
N-grams on the SMS spam messages
The practical notebook reruns the SMS bag of words with ngram_range=(2, 3): only bigrams and trigrams, and the 100 most frequent of them. corpus is the stemmed list of 5,572 messages built in Bag of words (BoW).
cv = CountVectorizer(max_features=100, binary=True, ngram_range=(2, 3))
X = cv.fit_transform(corpus)
names = cv.get_feature_names_out()
print(X.shape)
print(names[:8])
print("trigrams kept:", [n for n in names if n.count(" ") == 2][:5])
print("all bigrams and trigrams:", len(CountVectorizer(ngram_range=(2, 3)).fit(corpus).vocabulary_))(5572, 100) ['account statement' 'account statement show' 'attempt contact' 'await collect' 'call claim' 'call custom' 'call custom servic' 'call identifi'] trigrams kept: ['account statement show', 'call custom servic', 'call identifi code', 'call land line', 'claim valid hr'] all bigrams and trigrams: 59253
- The features are stemmed word pairs and triples such as “call claim” and “call custom servic”: phrases typical of spam.
- Without
max_featuresthe corpus has 59,253 distinct bigrams and trigrams, so the limit matters much more than with single words.
Unigrams vs n-grams
| Unigrams (1, 1) | Unigrams and bigrams (1, 2) | |
|---|---|---|
| food good vs food not good | cosine 0.8165 | cosine 0.5164 |
| Word order | lost | kept for neighbouring words |
| Vocabulary size | number of distinct words | many times larger |
| Catches | topic words | phrases and negations such as “not good” |
Where you use n-grams
- Sentiment analysis, where “not good” and “not bad” flip the meaning of a single word.
- Spam filters, where phrases such as “free entry” or “call claim” say more than either word alone.
- Search and autocomplete, where counts of word pairs and triples estimate which word comes next.
max_features or min_df to drop the rare ones, or the matrix gets huge and the model overfits.Related
- Previous: Bag of words (BoW)
- Next: TF-IDF
- Reference: Text feature extraction in the scikit-learn user guide
- Add a third document, “food not bad”, to
docsand rerun with(1, 2). Which bigram does it share with “food not good”? - Change
ngrams(tokens, n)to run on “I am not feeling well”. How many trigrams does it have? - In the SMS example, set
ngram_range=(1, 2)and look at which single words make it into the top 100.
Slow is fine. Stopping is the only problem.