Spam classifier with Average Word2Vec
An Average Word2Vec spam classifier is a text classification model that represents each message as the mean of its word vectors and trains a classifier on those dense vectors. Here Word2Vec learns the word vectors from the SMS messages themselves.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
The Spam classifier with BoW and TF-IDF lesson gave every message a sparse vector of 2,500 counts. Average Word2Vec gives every message one dense vector instead: the element-wise mean of the Word2Vec vectors of its words. This project trains those word vectors from scratch on the same SMS Spam Collection and classifies the averaged vectors.
Cleaning with lemmatization and keeping the stopwords
This version cleans with the WordNet lemmatizer from the Lemmatization lesson instead of a stemmer, so tokens stay real words such as joking. Stopwords stay in: Word2Vec learns a word from its neighbours, and words such as "not" and "you" are part of that context. The letters-only regex and lower-casing are the same as before. The lemmatizer needs two NLTK resources:
import nltk
nltk.download("wordnet")
nltk.download("omw-1.4")import re
import numpy as np
import pandas as pd
from nltk.stem import WordNetLemmatizer
from gensim.utils import simple_preprocess
URL = ("https://raw.githubusercontent.com/krishnaik06/The-Grand-Complete-Data-Science-Materials/main/"
"Complete%20NLP%20For%20ML%20%26%20Deep%20Learning/Practicals/SpamClassifier-master/smsspamcollection/SMSSpamCollection")
messages = pd.read_csv(URL, sep="\t", names=["label", "message"])lemmatizer = WordNetLemmatizer()
def clean(text):
words = re.sub("[^a-zA-Z]", " ", text).lower().split() # stopwords stay in
return " ".join(lemmatizer.lemmatize(w) for w in words)
corpus = messages["message"].apply(clean)Tokenizing with simple_preprocess
gensim's simple_preprocess turns a text into a list of lower-case tokens and drops tokens shorter than 2 or longer than 15 characters (its min_len=2 and max_len=15 defaults). One-letter words such as u and n disappear here. Word2Vec trains on lists of tokens, one list per message.
Finding the messages that clean to nothing
The video's feature table has 5,569 rows while the dataset has 5,572 messages. Three messages contain no letters at all, "645", ":) " and ":-) :-)", so the letters-only regex leaves an empty string. The fix is a mask, corpus.str.len() > 0, applied to the documents and to the labels alike, so that row i of the features and row i of the labels still describe the same message.
empty = corpus.str.len() == 0
print("messages:", len(messages), "| empty after cleaning:", empty.sum())
print(messages.loc[empty, "message"].tolist())
keep = ~empty
docs = [simple_preprocess(text) for text in corpus[keep]] # lists of lower-case tokens
y = (messages.loc[keep, "label"] == "spam").astype(int).to_numpy() # spam = 1, ham = 0
print("documents:", len(docs), "| spam:", y.sum())
print(corpus[1], "->", docs[1])messages: 5572 | empty after cleaning: 3 ['645', ':) ', ':-) :-)'] documents: 5569 | spam: 747 ok lar joking wif u oni -> ['ok', 'lar', 'joking', 'wif', 'oni']
- 3 messages clean to nothing, and they are the three the video finds. 5,569 documents remain, 747 of them spam: the three empty ones were ham.
- Message 1, "Ok lar... Joking wif u oni...", becomes five tokens.
simple_preprocessdrops the one-letteru.
Training Word2Vec inside a scikit-learn transformer
The word vectors are learned from text, so they belong to training like the vectorizer did in the BoW lesson. The video trains Word2Vec on all 5,569 messages and splits afterwards; here a small transformer class trains it inside fit(), and a Pipeline then calls fit() on the training messages only. The settings are gensim's defaults, which the project notebook uses: vector_size=100, window=5, min_count=5 (words seen fewer than 5 times get no vector), CBOW (sg=0) and epochs=5. seed=42 with workers=1 makes training repeat exactly; with several worker threads two runs differ slightly even with the same seed. Training Word2Vec with gensim explains each setting.
import numpy as np
from gensim.models import Word2Vec
from sklearn.base import BaseEstimator, TransformerMixin
class AvgWord2Vec(BaseEstimator, TransformerMixin):
def __init__(self, vector_size=100, window=5, min_count=5, epochs=5):
self.vector_size, self.window = vector_size, window
self.min_count, self.epochs = min_count, epochs
def fit(self, docs, y=None): # learns word vectors from the training docs
self.w2v_ = Word2Vec(docs, vector_size=self.vector_size, window=self.window,
min_count=self.min_count, epochs=self.epochs,
seed=42, workers=1) # seed + one worker: the same vectors every run
return selfAveraging the word vectors in transform
transform(), inside the same class, looks up each token, skips the ones without a vector and takes the element-wise mean with np.mean(..., axis=0). Every message comes out as one row of 100 numbers, however many words it has.
def transform(self, docs):
wv = self.w2v_.wv
out = np.zeros((len(docs), self.vector_size), dtype=np.float32)
for i, doc in enumerate(docs):
vecs = [wv[w] for w in doc if w in wv.key_to_index] # skip unknown words
if vecs: # no known word: stays zeros
out[i] = np.mean(vecs, axis=0)
return outHandling messages with no known word
The averaged vectors are dense, all 100 numbers carry a value, so the video picks a random forest, which handles dense numeric features well. Fitting then fails with "Input contains NaN": 12 messages have no word in the vocabulary, and the mean of an empty list is NaN. The video drops those rows. The transformer here gives such a message a vector of zeros instead, so every test message still gets a prediction, as a spam filter must. The label stays a separate array y and never becomes a column of the feature table.
Training the Average Word2Vec spam model
The run splits the 5,569 documents with stratify=y, fits the pipeline on the training part and reports on the test part, next to the majority-class baseline. The classifier is a Random forest with random_state=42, so its trees are the same on every run.
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix
X_train, X_test, y_train, y_test = train_test_split(
docs, y, test_size=0.20, random_state=42, stratify=y)
print("train:", len(X_train), "test:", len(X_test), "| majority-class baseline:", round(1 - y_test.mean(), 4))
model = Pipeline([("avg_w2v", AvgWord2Vec()), ("forest", RandomForestClassifier(random_state=42))])
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
print("accuracy:", round(accuracy_score(y_test, y_pred), 4))
print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred, target_names=["ham", "spam"], digits=3))
wv = model.named_steps["avg_w2v"].w2v_.wv
X_test_vec = model.named_steps["avg_w2v"].transform(X_test)
print("vocabulary:", len(wv.key_to_index), "words | test messages with no known word:",
int((np.abs(X_test_vec).sum(axis=1) == 0).sum()))
print("unknown words in", docs[1], ":", [w for w in docs[1] if w not in wv.key_to_index])
print("closest to 'good':", [(w, round(s, 3)) for w, s in wv.most_similar("good", topn=5)])train: 4455 test: 1114 | majority-class baseline: 0.8662
accuracy: 0.9695
[[954 11]
[ 23 126]]
precision recall f1-score support
ham 0.976 0.989 0.982 965
spam 0.920 0.846 0.881 149
accuracy 0.969 1114
macro avg 0.948 0.917 0.932 1114
weighted avg 0.969 0.969 0.969 1114
vocabulary: 1507 words | test messages with no known word: 5
unknown words in ['ok', 'lar', 'joking', 'wif', 'oni'] : ['oni']
closest to 'good': [('wa', 0.999), ('night', 0.999), ('not', 0.999), ('happy', 0.999), ('amp', 0.999)]What the averaged vectors scored
- The baseline is 0.8662 on 1,114 test messages, and the model reaches 0.9695.
- Spam precision 0.920 and recall 0.846: the confusion matrix shows 126 of the 149 spam messages caught and 11 ham messages flagged.
- The vocabulary has 1,507 words, the training words seen at least 5 times. 5 test messages have none of them and get the zero vector.
onifrom message 1 is one of the unknown words. - Every neighbour of "good" scores 0.999. After 5 epochs over about 4,500 short messages the vectors still point in nearly the same direction, so the averages of different messages are close together. The forest still separates them, but the vectors carry little meaning yet.
Training the vectors for more epochs
More passes over the training messages give the vectors time to move apart. The run below trains the same pipeline with 5 and with 30 epochs, once with the random forest and once with Logistic regression, a linear model that also suits dense features.
from sklearn.linear_model import LogisticRegression
for name, clf in [("random forest", RandomForestClassifier(random_state=42)),
("logistic regression", LogisticRegression(max_iter=2000))]:
for epochs in (5, 30):
m = Pipeline([("avg_w2v", AvgWord2Vec(epochs=epochs)), ("clf", clf)]).fit(X_train, y_train)
print(f"{name:19s} epochs={epochs:2d} accuracy {accuracy_score(y_test, m.predict(X_test)):.4f}")
wv30 = m.named_steps["avg_w2v"].w2v_.wv
print("closest to 'good' after 30 epochs:", [(w, round(s, 3)) for w, s in wv30.most_similar("good", topn=5)])random forest epochs= 5 accuracy 0.9695
random forest epochs=30 accuracy 0.9794
logistic regression epochs= 5 accuracy 0.8743
logistic regression epochs=30 accuracy 0.9722
closest to 'good' after 30 epochs: [('evening', 0.786), ('brings', 0.77), ('nice', 0.752), ('sweet', 0.732), ('wonderful', 0.722)]What longer training changed
- Logistic regression jumps from 0.8743 to 0.9722. With 5 epochs the averaged vectors are so alike that a straight boundary between spam and ham barely exists, and the model is hardly better than the 0.8662 baseline.
- The random forest rises from 0.9695 to 0.9794, the best score of this lesson.
- The neighbours of "good" now make sense: evening, brings, nice, sweet and wonderful, with similarities between 0.72 and 0.79 instead of 0.999 for everything.
- BoW + MultinomialNB scored 0.9839 in the BoW lesson. Word vectors trained on 4,500 short texts do not beat word counts on this task; they help more when the vectors come from a large corpus.
Average Word2Vec vs BoW for spam
| Average Word2Vec | Bag of words | |
|---|---|---|
| Vector per message | 100 dense numbers, the mean of the word vectors | 2,500 sparse counts, mostly zeros |
| Learned from | the training messages (Word2Vec, 30 epochs) | the training messages (vocabulary only) |
| Similar words | close vectors, so "good" and "nice" overlap | separate columns, no overlap |
| Word order | lost in the average | lost, except inside bigrams |
| Unknown words | skipped; no known word gives zeros | dropped |
| Best accuracy here | 0.9794 (random forest) | 0.9839 (MultinomialNB) |
Where you use Average Word2Vec features
- Short-text classification (spam, intent, support-ticket routing) when a fixed-length dense input is wanted for a classical model.
- Domain text that pretrained vectors do not cover, such as SMS slang ("lar", "wif") or product codes, where vectors trained on your own messages know the words.
- A fast baseline before a sequence model, such as the Recurrent neural network (RNN) of the next part, which reads the words in order instead of averaging them.
y is added to the vectors before X is taken from that table, the random forest reads the answer from that column and scores close to 1.0 (0.997 with this split), a number that says nothing about spam. Build X from the vector columns only, and train Word2Vec on the training messages, as the pipeline does.Related
- Previous: Spam classifier with BoW and TF-IDF
- Next: Sentiment analysis of Kindle reviews
- See also: Word2Vec, Random forest
- Reference: gensim Word2Vec documentation
- Set
min_count=2inAvgWord2Vec(...)and check how the vocabulary size and the number of zero-vector test messages change. - Train skip-gram vectors by adding
sg=1to theWord2Vec(...)call infit(), withepochs=30, and compare the accuracy. - Print
wv.most_similar("free", topn=5)after 5 and after 30 epochs: the 5-epoch neighbours all score about 0.999, the 30-epoch ones are spread out and read like spam offers.
Slow is fine. Stopping is the only problem.