Bag of words (BoW)
Bag of words (BoW) is a text representation that turns each sentence into one vector of word counts over a fixed vocabulary, ignoring the order of the words.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
One-hot encoding for text gave every word its own vector, so documents of different lengths got different shapes. Bag of words gives a whole sentence one vector as long as the vocabulary, so every sentence has the same shape and a classifier can train on it.
Lower-casing the sentences and removing stopwords
The video's dataset has three positive sentences, each with output 1: “He is a good boy”, “She is a good girl” and “Boy and girl are good”.
The next step cleans them. Lower-casing comes first, so that “Boy” and “boy” count as one word. Then the Stopwords he, she, is, a, and and are are removed. What remains is S1 “good boy”, S2 “good girl” and S3 “boy girl good”.
Counting the vocabulary and building the vectors
The vocabulary is the unique words of the cleaned sentences, each with its frequency: good appears 3 times, boy 2 times and girl 2 times. On the board they are sorted from the most to the least frequent.
The frequency matters on a large dataset. A vocabulary can hold tens of thousands of words, many of them seen once, so it is common to keep only the top 10 or top 20 most frequent words as features. In scikit-learn that is max_features.
The kept words become the columns: good, boy, girl. Each sentence gets a 1 where a word is present and 0 where it is not: S1 “good boy” is [1 1 0], S2 “good girl” is [1 0 1] and S3 “boy girl good” is [1 1 1]. One-hot encoding made a vector for every word; bag of words makes one vector for the whole sentence.
Choosing binary bag of words or counts
If S2 were “good girl good”, good would appear twice. The board names two versions:
- Bag of words (counts): the value is the number of times the word appears, so good becomes 2.
- Binary bag of words: the value is only 1 or 0, present or absent, so good stays 1.
Building bag of words with CountVectorizer
scikit-learn's CountVectorizer builds the vocabulary and the count matrix in one call.
Cleaning with NLTK stopwords
from nltk.corpus import stopwords
stop = set(stopwords.words("english")) # a set makes each lookup fast
clean = [" ".join(w for w in s.lower().split() if w not in stop) for s in sentences]Counts, binary=True and max_features
cv = CountVectorizer() # counts, every word in the vocabulary
cv = CountVectorizer(binary=True) # 1 if the word is present, else 0
cv = CountVectorizer(max_features=2) # keep the 2 most frequent words
X = cv.fit_transform(clean) # sparse matrix, one row per sentence
cv.get_feature_names_out() # the column wordsimport nltk
from nltk.corpus import stopwords
from sklearn.feature_extraction.text import CountVectorizer
nltk.download("stopwords", quiet=True)
sentences = ["He is a good boy", "She is a good girl", "Boy and girl are good"]
stop = set(stopwords.words("english"))
clean = [" ".join(w for w in s.lower().split() if w not in stop) for s in sentences]
print(clean)
cv = CountVectorizer()
X = cv.fit_transform(clean)
print(cv.get_feature_names_out())
print(X.toarray())
print("vocabulary_:", cv.vocabulary_)
for binary in (False, True):
v = CountVectorizer(binary=binary).fit(clean)
print("binary =", binary, v.transform(["good girl good"]).toarray())
print("max_features=2 keeps:", CountVectorizer(max_features=2).fit(clean).get_feature_names_out())
print("unknown word:", cv.transform(["boy girl good school"]).toarray())['good boy', 'good girl', 'boy girl good']
['boy' 'girl' 'good']
[[1 0 1]
[0 1 1]
[1 1 1]]
vocabulary_: {'good': 2, 'boy': 0, 'girl': 1}
binary = False [[0 1 2]]
binary = True [[0 1 1]]
max_features=2 keeps: ['boy' 'good']
unknown word: [[1 1 1]]What the vectors show
- The cleaned sentences are good boy, good girl and boy girl good, as on the board.
- The columns are alphabetical: boy, girl, good. The rows [1 0 1], [0 1 1] and [1 1 1] are the board's [1 1 0], [1 0 1] and [1 1 1] with the columns in a different order.
vocabulary_maps each word to its column. - Frequency picks which words survive
max_features, not the column order. Withmax_features=2the vectorizer keeps good (3) and one of the two words tied at 2; this run kept boy. Which of two tied words survives is not guaranteed, so do not rely on it. - “good girl good” is [0 1 2] with counts and [0 1 1] with
binary=True. - The unknown word “school” is dropped without a warning: “boy girl good school” gets the same vector as “boy girl good”.
Bag of words on the SMS spam messages
The practical notebook in the video's materials builds bag of words on the SMS Spam Collection: 5,572 text messages, each labelled ham or spam. Each message keeps only its letters, is lower-cased and split into words, loses its stopwords, and every word is reduced by the Porter stemmer. The data loads from the GitHub repository of the materials, so nothing needs downloading by hand.
import re
import nltk
import pandas as pd
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer
from sklearn.feature_extraction.text import CountVectorizer
nltk.download("stopwords", quiet=True)
url = "https://raw.githubusercontent.com/krishnaik06/Avgword2vec-Implementation/main/smsspamcollection/SMSSpamCollection"
messages = pd.read_csv(url, sep="\t", names=["label", "message"])
ps = PorterStemmer()
stop = set(stopwords.words("english"))
corpus = []
for message in messages["message"]:
review = re.sub("[^a-zA-Z]", " ", message).lower().split() # letters only, lower-case
corpus.append(" ".join(ps.stem(word) for word in review if word not in stop))
print(messages.shape)
print(corpus[0])
cv = CountVectorizer(max_features=100, binary=True)
X = cv.fit_transform(corpus)
print(X.shape, "non-zero values:", X.nnz)
names = cv.get_feature_names_out()
print("message 0 has:", [names[j] for j in sorted(X[0].nonzero()[1])])(5572, 2) go jurong point crazi avail bugi n great world la e buffet cine got amor wat (5572, 100) non-zero values: 14746 message 0 has: ['go', 'got', 'great', 'wat']
- The first message becomes “go jurong point crazi avail bugi n great world la e buffet cine got amor wat” after stemming.
- X is 5,572 × 100: one row per message, one column per kept word. Only 14,746 of its 557,200 cells are non-zero, under 3%, so scikit-learn stores it as a sparse matrix.
- Message 0 keeps four features, go, got, great and wat; its other words are not among the 100 most frequent.
Listing the advantages and disadvantages of bag of words
- Simple and intuitive, and quick to compute.
- Fixed-size input. Every sentence, short or long, becomes a vector of the vocabulary's size, which solves one-hot encoding's shape problem.
- Still a sparse matrix. With 50,000 words in the vocabulary, every sentence is 50,000 numbers, nearly all 0.
- Word order is lost. The vector records which words appear, not where, so “dog bites man” and “man bites dog” get identical vectors.
- Out of vocabulary. A test word that was not in the training vocabulary, such as “school”, is ignored, even when it matters to the sentence.
- No semantic meaning. good and great are separate columns with no link between them, and each present word gets the same 1.
from sklearn.feature_extraction.text import CountVectorizer
pair = ["dog bites man", "man bites dog"]
cv = CountVectorizer()
print(cv.fit(pair).get_feature_names_out())
print(cv.transform(pair).toarray())['bites' 'dog' 'man'] [[1 1 1] [1 1 1]]
Both sentences give [1 1 1] over bites, dog and man: bag of words cannot tell who bit whom.
Count bag of words vs binary bag of words
| Count BoW | Binary BoW | |
|---|---|---|
| Value of a word | how many times it appears | 1 if it appears, else 0 |
| “good girl good” over boy, girl, good | [0 1 2] | [0 1 1] |
| CountVectorizer | CountVectorizer() | CountVectorizer(binary=True) |
| Naive Bayes that fits it | MultinomialNB | BernoulliNB |
| Good for | long texts where repetition matters | short texts such as SMS messages |
Where you use bag of words
- Spam and ham filters on short messages, often with Naive Bayes: Spam classifier with BoW and TF-IDF builds one.
- Sentiment classification baselines, before trying TF-IDF or word embeddings.
- Keyword counts for tagging tickets or routing emails by the words they contain.
transform on the test text. Fitting on all the data puts test words into the vocabulary and makes the test score look better than it will be on new messages.Related
- Previous: One-hot encoding for text
- Next: N-grams
- Reference: CountVectorizer in the scikit-learn API
- Change
max_features=2tomax_features=1. Which word is kept, and why that one? - Skip the NLTK cleaning and pass the raw
sentencestoCountVectorizer(). Which words appear, and why is “a” missing? - In the SMS example, change
max_features=100tomax_features=10and print the feature names.
You understood something today that you didn't yesterday.