Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Spam classifier with BoW and TF-IDF

A spam classifier is a text classification model that reads a message and labels it spam or ham (an ordinary message). This one turns 5,572 SMS messages into bag of words and TF-IDF vectors and trains a multinomial naive Bayes model on them.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

The Bag of words (BoW) and TF-IDF lessons turned three short sentences into vectors by hand. This project runs the same encodings on a real dataset, feeds the vectors to a model and measures how well it separates spam from ham.

The plan for the spam project · from Day 5 of the Live NLP series · 9:25 to 12:03

Planning the spam classification project

The plan in the video has three stages. Text preprocessing comes first: tokenize the messages, remove stopwords and reduce each word with stemming or lemmatization. Then the text becomes vectors with bag of words, TF-IDF or Word2Vec. Last, a machine learning algorithm trains on the vectors and its accuracy is measured.

Two steps make the measurement honest. The data is split into training and test messages before any vectorizer learns a vocabulary, and the accuracy is compared with a model that always predicts the most common class. This lesson builds the BoW and TF-IDF versions; the Word2Vec version is the next lesson.

The spam project as a flow: 5,572 SMS messages are cleaned with a letters-only regex, lower-casing, stopword removal and the Porter stemmer, then split 80/20 with stratify; the 4,457 training messages fit the vectorizer and MultinomialNB inside a Pipeline, while the 1,115 test messages are only transformed and predicted, and the result is compared with the 0.866 majority-class baseline.

Loading the SMS Spam Collection

The SMS Spam Collection comes from the UCI Machine Learning Repository. It is one tab-separated text file: each line holds a label, ham or spam, a tab and the message. read_csv with sep="\t" and two column names reads it straight from the materials repo on GitHub, so nothing needs downloading by hand.

ExampleThe dataset from the materials repo, run on pandas 3.0
import pandas as pd

URL = ("https://raw.githubusercontent.com/krishnaik06/The-Grand-Complete-Data-Science-Materials/main/"
       "Complete%20NLP%20For%20ML%20%26%20Deep%20Learning/Practicals/SpamClassifier-master/smsspamcollection/SMSSpamCollection")
messages = pd.read_csv(URL, sep="\t", names=["label", "message"])

print(messages.shape)
print(messages["label"].value_counts().to_dict())
print(messages.loc[2, "message"])
print("share of ham:", round((messages["label"] == "ham").mean(), 4))
  • 5,572 messages in two columns, label and message.
  • 4,825 are ham and 747 are spam, so 86.59% of all messages are ham. A model that always answers ham is already right 86.59% of the time, and every accuracy in this project has to be read against that number.
  • Message 2 is a typical spam text: a prize, a short code to text and a price per message.

Cleaning the messages with stemming

Each message goes through the steps of the Stopwords and Stemming lessons. A regular expression replaces every character that is not a letter, [^a-zA-Z], with a space. The text is lower-cased and split into words, the NLTK English stopwords are dropped, and the Porter stemmer cuts each word to its stem. The stopword list is turned into a set once, because checking a word against a set is fast and the check runs for every word of every message.

The stopword list is an NLTK resource. Download it once:

python
import nltk
nltk.download("stopwords")
ExampleContinues from the loading code above; run on NLTK 3.10
import re
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer

ps = PorterStemmer()
stop = set(stopwords.words("english"))     # a set, built once, for fast lookups

def clean(text):
    words = re.sub("[^a-zA-Z]", " ", text).lower().split()   # letters only, lower case
    return " ".join(ps.stem(w) for w in words if w not in stop)

corpus = messages["message"].apply(clean)
print(corpus[0])

The first message, "Go until jurong point, crazy.. Available only in bugis n great world la e buffet... Cine there got amore wat...", keeps 16 stems. "until", "only", "in" and "there" are stopwords, and "crazy" and "available" become the stems crazi and avail.

Coding the labels as 0 and 1

A model needs numbers for the labels. Here spam is 1 and ham is 0, so row 1 of every report below is the spam class, the one the filter exists for. The project notebooks build the label with pd.get_dummies(messages['label']).iloc[:, 0], which picks the ham column, so in their reports 1 means ham.

python
y = (messages["label"] == "spam").astype(int)     # spam = 1, ham = 0
Bag of words features for the messages · from Day 5 of the Live NLP series · 21:01 to 24:07

Choosing the bag of words settings

CountVectorizer builds the bag of words. max_features=2500 keeps the 2,500 terms that occur most often across the messages and drops the rest, so every message becomes a vector of length 2,500. With binary=True a word that appears twice in a message still counts 1; without it the cell holds the count. ngram_range adds groups of neighbouring words as extra terms (see N-grams).

The live session uses binary=True and single words. The project notebook for this lesson (27.2 in the materials repo) keeps the counts and sets ngram_range=(1, 2), so pairs such as free entri and claim call become features of their own. The code here follows notebook 27.2.

Splitting before fitting the vectorizer

Data leakage is information from the test data reaching the model during training. A vectorizer fitted on all 5,572 messages has chosen its 2,500 terms, and for TF-IDF its idf weights, partly from the test messages. The test score then measures a model that has already seen the test vocabulary. The live session fits the vectorizer on every message to keep the demo short and recommends splitting first and fitting on the training part only; notebook 27.2 does that, and so does the code here.

train_test_split keeps 20% of the messages for testing (see Train and test split). stratify=y keeps the spam share the same, about 13.4%, in both parts, and random_state=42 makes the split the same on every run.

python
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    corpus, y, test_size=0.20, random_state=42, stratify=y)   # same spam share in both parts

Putting the vectorizer and the model in a Pipeline

A Pipeline chains the steps into one estimator. fit() fits the vectorizer on X_train and the model on the vectorizer's output; predict() only transforms new messages with the fitted vocabulary. The leak cannot happen by accident. The model is MultinomialNB, the Naive Bayes variant that models how often each term appears in each class, which is what count vectors hold. scikit-learn's MultinomialNB takes the sparse matrix the vectorizer returns, so .toarray() is not needed.

python
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB

bow_model = Pipeline([
    ("bow", CountVectorizer(max_features=2500, ngram_range=(1, 2))),
    ("nb", MultinomialNB()),
])
bow_model.fit(X_train, y_train)       # the vocabulary comes from the training messages only
y_pred = bow_model.predict(X_test)    # the test messages are only transformed

Measuring against a majority-class baseline

DummyClassifier(strategy="most_frequent") always predicts the class it saw most in training, here ham. Its accuracy is the floor a real model must beat. classification_report takes the true labels first and the predictions second.

python
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score, classification_report

baseline = DummyClassifier(strategy="most_frequent").fit(X_train, y_train)
print(accuracy_score(y_test, baseline.predict(X_test)))     # always predicts ham
print(classification_report(y_test, y_pred, target_names=["ham", "spam"]))   # true labels first

Training the BoW and TF-IDF spam models

The full run trains both encodings with the same split, the same 2,500 terms of one and two words, and the same model, so the only difference is what each cell of the vector holds.

ExampleContinues from the cleaning code above; run on scikit-learn 1.9.1
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix

y = (messages["label"] == "spam").astype(int)     # spam = 1, ham = 0
X_train, X_test, y_train, y_test = train_test_split(
    corpus, y, test_size=0.20, random_state=42, stratify=y)
print("train:", len(X_train), "test:", len(X_test), "spam in test:", y_test.sum())

baseline = DummyClassifier(strategy="most_frequent").fit(X_train, y_train)
print("majority-class baseline:", round(accuracy_score(y_test, baseline.predict(X_test)), 4))

models = {}
for name, vectorizer in [("BoW", CountVectorizer(max_features=2500, ngram_range=(1, 2))),
                         ("TF-IDF", TfidfVectorizer(max_features=2500, ngram_range=(1, 2)))]:
    models[name] = Pipeline([("vec", vectorizer), ("nb", MultinomialNB())]).fit(X_train, y_train)
    y_pred = models[name].predict(X_test)
    print(f"\n{name} + MultinomialNB accuracy:", round(accuracy_score(y_test, y_pred), 4))
    print(confusion_matrix(y_test, y_pred))
    print(classification_report(y_test, y_pred, target_names=["ham", "spam"], digits=3))

What the two reports show

  • The baseline scores 0.8664: the test part holds 966 ham and 149 spam messages, and 966 / 1,115 = 0.8664.
  • BoW + MultinomialNB reaches 0.9839, 11.75 points above the baseline. Its confusion matrix shows 135 of the 149 spam messages caught (recall 0.906) and 4 ham messages flagged as spam (spam precision 0.971).
  • TF-IDF + MultinomialNB reaches 0.9758. It flags only 1 ham message, so its spam precision is 0.992, but it misses 26 spam messages, so its spam recall drops to 0.826.
  • On this data, BoW is the better input for MultinomialNB. The two models differ in the spam they miss, 14 against 26, far more than in the ham they flag.
  • The ham rows look perfect in both (recall 0.996 and 0.999) because ham is the easy majority class. The spam row is where the two models differ.
ExampleContinues from the run above
import matplotlib.pyplot as plt
from sklearn.metrics import ConfusionMatrixDisplay, confusion_matrix

fig, axes = plt.subplots(1, 2, figsize=(9, 4))
for ax, (name, model) in zip(axes, models.items()):
    ConfusionMatrixDisplay.from_estimator(model, X_test, y_test, display_labels=["ham", "spam"],
                                          ax=ax, colorbar=False, cmap="Purples")
    ax.set_title(name + " + MultinomialNB")
    print(name, "[tn, fp, fn, tp]:", confusion_matrix(y_test, model.predict(X_test)).ravel().tolist())
plt.tight_layout()
plt.show()
Two confusion matrices for the 1,115 test messages: BoW with MultinomialNB has 962 ham correct, 4 ham flagged as spam, 14 spam missed and 135 spam caught; TF-IDF with MultinomialNB has 965 ham correct, 1 ham flagged, 26 spam missed and 123 spam caught.

Reading precision and recall the right way round

classification_report(y_true, y_pred) groups the messages by their true class. Spam precision is the share of messages predicted spam that are spam; spam recall is the share of true spam messages the model caught; support is the number of true messages in each class (see Precision, recall and F-beta). With the arguments swapped, scikit-learn treats the predictions as the truth: precision and recall trade places and support counts the predictions.

ExampleContinues from the run above
y_pred = models["BoW"].predict(X_test)
print("classification_report(y_test, y_pred):")
print(classification_report(y_test, y_pred, target_names=["ham", "spam"], digits=3))
print("classification_report(y_pred, y_test), arguments swapped:")
print(classification_report(y_pred, y_test, target_names=["ham", "spam"], digits=3))
  • The correct report gives spam precision 0.971, recall 0.906 and support 149, the number of spam messages in the test part.
  • The swapped report gives spam precision 0.906 and recall 0.971, and its support of 139 is the number of messages the model predicted as spam (135 + 4).
  • The accuracy, 0.984, is the same in both, which is why the swap goes unnoticed. For a spam filter the difference matters: 0.906 recall means about 1 spam in 11 reaches the inbox.

BoW vs TF-IDF for spam

BoW (CountVectorizer)TF-IDF (TfidfVectorizer)
Cell valuehow often the term appears in the messagecount × smoothed idf, each row scaled to length 1
Terms kept here2,500 most frequent unigrams and bigramsthe same 2,500 terms
Accuracy here (MultinomialNB)0.98390.9758
Spam caught (recall)135 of 149 (0.906)123 of 149 (0.826)
Ham flagged as spam41
Helps most whenfrequent words carry the signalcommon words drown out rarer, telling ones

Where you use a spam classifier

  • SMS and email filtering, where a message predicted spam goes to a separate folder instead of the inbox.
  • Comment and review moderation, to hold back promotional or scam posts for a human to check.
  • A first text-classification baseline: BoW or TF-IDF with naive Bayes trains in under a second and sets the score that a heavier model has to beat.
Watch out. Fit the vectorizer on the training messages only, which a Pipeline does for you. A vectorizer fitted before the split chooses its terms and idf weights partly from the test messages. And read every accuracy next to the 0.866 baseline: with 13% spam, the spam recall says more about the filter than the overall accuracy.
Try it yourself
  • Add binary=True to the CountVectorizer, as in the live session, and compare the spam recall with the count version.
  • Change ngram_range=(1, 2) to (1, 1) in both vectorizers and see whether the bigrams were helping.
  • Replace MultinomialNB() with LogisticRegression(max_iter=1000) from sklearn.linear_model and read how the spam precision and recall move.

You understood something today that you didn't yesterday.