Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Spam classifier with Average Word2Vec

An Average Word2Vec spam classifier is a text classification model that represents each message as the mean of its word vectors and trains a classifier on those dense vectors. Here Word2Vec learns the word vectors from the SMS messages themselves.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

The Spam classifier with BoW and TF-IDF lesson gave every message a sparse vector of 2,500 counts. Average Word2Vec gives every message one dense vector instead: the element-wise mean of the Word2Vec vectors of its words. This project trains those word vectors from scratch on the same SMS Spam Collection and classifies the averaged vectors.

Cleaning with lemmatization and keeping the stopwords

This version cleans with the WordNet lemmatizer from the Lemmatization lesson instead of a stemmer, so tokens stay real words such as joking. Stopwords stay in: Word2Vec learns a word from its neighbours, and words such as "not" and "you" are part of that context. The letters-only regex and lower-casing are the same as before. The lemmatizer needs two NLTK resources:

python
import nltk
nltk.download("wordnet")
nltk.download("omw-1.4")
python
import re
import numpy as np
import pandas as pd
from nltk.stem import WordNetLemmatizer
from gensim.utils import simple_preprocess

URL = ("https://raw.githubusercontent.com/krishnaik06/The-Grand-Complete-Data-Science-Materials/main/"
       "Complete%20NLP%20For%20ML%20%26%20Deep%20Learning/Practicals/SpamClassifier-master/smsspamcollection/SMSSpamCollection")
messages = pd.read_csv(URL, sep="\t", names=["label", "message"])
python
lemmatizer = WordNetLemmatizer()
def clean(text):
    words = re.sub("[^a-zA-Z]", " ", text).lower().split()   # stopwords stay in
    return " ".join(lemmatizer.lemmatize(w) for w in words)

corpus = messages["message"].apply(clean)

Tokenizing with simple_preprocess

gensim's simple_preprocess turns a text into a list of lower-case tokens and drops tokens shorter than 2 or longer than 15 characters (its min_len=2 and max_len=15 defaults). One-letter words such as u and n disappear here. Word2Vec trains on lists of tokens, one list per message.

Finding the three empty messages · from the AvgWord2vec Indepth Intuition and Practical Implementation video · 24:35 to 29:21

Finding the messages that clean to nothing

The video's feature table has 5,569 rows while the dataset has 5,572 messages. Three messages contain no letters at all, "645", ":) " and ":-) :-)", so the letters-only regex leaves an empty string. The fix is a mask, corpus.str.len() > 0, applied to the documents and to the labels alike, so that row i of the features and row i of the labels still describe the same message.

ExampleContinues from the cleaning code above; run on gensim 4.4.0
empty = corpus.str.len() == 0
print("messages:", len(messages), "| empty after cleaning:", empty.sum())
print(messages.loc[empty, "message"].tolist())

keep = ~empty
docs = [simple_preprocess(text) for text in corpus[keep]]          # lists of lower-case tokens
y = (messages.loc[keep, "label"] == "spam").astype(int).to_numpy()   # spam = 1, ham = 0
print("documents:", len(docs), "| spam:", y.sum())
print(corpus[1], "->", docs[1])
  • 3 messages clean to nothing, and they are the three the video finds. 5,569 documents remain, 747 of them spam: the three empty ones were ham.
  • Message 1, "Ok lar... Joking wif u oni...", becomes five tokens. simple_preprocess drops the one-letter u.

Training Word2Vec inside a scikit-learn transformer

The word vectors are learned from text, so they belong to training like the vectorizer did in the BoW lesson. The video trains Word2Vec on all 5,569 messages and splits afterwards; here a small transformer class trains it inside fit(), and a Pipeline then calls fit() on the training messages only. The settings are gensim's defaults, which the project notebook uses: vector_size=100, window=5, min_count=5 (words seen fewer than 5 times get no vector), CBOW (sg=0) and epochs=5. seed=42 with workers=1 makes training repeat exactly; with several worker threads two runs differ slightly even with the same seed. Training Word2Vec with gensim explains each setting.

python
import numpy as np
from gensim.models import Word2Vec
from sklearn.base import BaseEstimator, TransformerMixin

class AvgWord2Vec(BaseEstimator, TransformerMixin):
    def __init__(self, vector_size=100, window=5, min_count=5, epochs=5):
        self.vector_size, self.window = vector_size, window
        self.min_count, self.epochs = min_count, epochs

    def fit(self, docs, y=None):                  # learns word vectors from the training docs
        self.w2v_ = Word2Vec(docs, vector_size=self.vector_size, window=self.window,
                             min_count=self.min_count, epochs=self.epochs,
                             seed=42, workers=1)  # seed + one worker: the same vectors every run
        return self

Averaging the word vectors in transform

transform(), inside the same class, looks up each token, skips the ones without a vector and takes the element-wise mean with np.mean(..., axis=0). Every message comes out as one row of 100 numbers, however many words it has.

python
    def transform(self, docs):
        wv = self.w2v_.wv
        out = np.zeros((len(docs), self.vector_size), dtype=np.float32)
        for i, doc in enumerate(docs):
            vecs = [wv[w] for w in doc if w in wv.key_to_index]   # skip unknown words
            if vecs:                                              # no known word: stays zeros
                out[i] = np.mean(vecs, axis=0)
        return out
Average Word2Vec on message 1: Ok lar... Joking wif u oni becomes the tokens ok, lar, joking, wif and oni after simple_preprocess drops u; oni has no vector because it is rarer than min_count, so the four 100-dimensional vectors of ok, lar, joking and wif are averaged element by element into one 100-dimensional row, which a random forest classifies as ham or spam.
Dense vectors, empty rows and the random forest · from the AvgWord2vec Indepth Intuition and Practical Implementation video · 33:21 to 36:00

Handling messages with no known word

The averaged vectors are dense, all 100 numbers carry a value, so the video picks a random forest, which handles dense numeric features well. Fitting then fails with "Input contains NaN": 12 messages have no word in the vocabulary, and the mean of an empty list is NaN. The video drops those rows. The transformer here gives such a message a vector of zeros instead, so every test message still gets a prediction, as a spam filter must. The label stays a separate array y and never becomes a column of the feature table.

Training the Average Word2Vec spam model

The run splits the 5,569 documents with stratify=y, fits the pipeline on the training part and reports on the test part, next to the majority-class baseline. The classifier is a Random forest with random_state=42, so its trees are the same on every run.

ExampleContinues from the code above; run on gensim 4.4.0 and scikit-learn 1.9.1
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix

X_train, X_test, y_train, y_test = train_test_split(
    docs, y, test_size=0.20, random_state=42, stratify=y)
print("train:", len(X_train), "test:", len(X_test), "| majority-class baseline:", round(1 - y_test.mean(), 4))

model = Pipeline([("avg_w2v", AvgWord2Vec()), ("forest", RandomForestClassifier(random_state=42))])
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
print("accuracy:", round(accuracy_score(y_test, y_pred), 4))
print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred, target_names=["ham", "spam"], digits=3))

wv = model.named_steps["avg_w2v"].w2v_.wv
X_test_vec = model.named_steps["avg_w2v"].transform(X_test)
print("vocabulary:", len(wv.key_to_index), "words | test messages with no known word:",
      int((np.abs(X_test_vec).sum(axis=1) == 0).sum()))
print("unknown words in", docs[1], ":", [w for w in docs[1] if w not in wv.key_to_index])
print("closest to 'good':", [(w, round(s, 3)) for w, s in wv.most_similar("good", topn=5)])

What the averaged vectors scored

  • The baseline is 0.8662 on 1,114 test messages, and the model reaches 0.9695.
  • Spam precision 0.920 and recall 0.846: the confusion matrix shows 126 of the 149 spam messages caught and 11 ham messages flagged.
  • The vocabulary has 1,507 words, the training words seen at least 5 times. 5 test messages have none of them and get the zero vector. oni from message 1 is one of the unknown words.
  • Every neighbour of "good" scores 0.999. After 5 epochs over about 4,500 short messages the vectors still point in nearly the same direction, so the averages of different messages are close together. The forest still separates them, but the vectors carry little meaning yet.

Training the vectors for more epochs

More passes over the training messages give the vectors time to move apart. The run below trains the same pipeline with 5 and with 30 epochs, once with the random forest and once with Logistic regression, a linear model that also suits dense features.

ExampleContinues from the run above
from sklearn.linear_model import LogisticRegression

for name, clf in [("random forest", RandomForestClassifier(random_state=42)),
                  ("logistic regression", LogisticRegression(max_iter=2000))]:
    for epochs in (5, 30):
        m = Pipeline([("avg_w2v", AvgWord2Vec(epochs=epochs)), ("clf", clf)]).fit(X_train, y_train)
        print(f"{name:19s} epochs={epochs:2d}  accuracy {accuracy_score(y_test, m.predict(X_test)):.4f}")

wv30 = m.named_steps["avg_w2v"].w2v_.wv
print("closest to 'good' after 30 epochs:", [(w, round(s, 3)) for w, s in wv30.most_similar("good", topn=5)])

What longer training changed

  • Logistic regression jumps from 0.8743 to 0.9722. With 5 epochs the averaged vectors are so alike that a straight boundary between spam and ham barely exists, and the model is hardly better than the 0.8662 baseline.
  • The random forest rises from 0.9695 to 0.9794, the best score of this lesson.
  • The neighbours of "good" now make sense: evening, brings, nice, sweet and wonderful, with similarities between 0.72 and 0.79 instead of 0.999 for everything.
  • BoW + MultinomialNB scored 0.9839 in the BoW lesson. Word vectors trained on 4,500 short texts do not beat word counts on this task; they help more when the vectors come from a large corpus.

Average Word2Vec vs BoW for spam

Average Word2VecBag of words
Vector per message100 dense numbers, the mean of the word vectors2,500 sparse counts, mostly zeros
Learned fromthe training messages (Word2Vec, 30 epochs)the training messages (vocabulary only)
Similar wordsclose vectors, so "good" and "nice" overlapseparate columns, no overlap
Word orderlost in the averagelost, except inside bigrams
Unknown wordsskipped; no known word gives zerosdropped
Best accuracy here0.9794 (random forest)0.9839 (MultinomialNB)

Where you use Average Word2Vec features

  • Short-text classification (spam, intent, support-ticket routing) when a fixed-length dense input is wanted for a classical model.
  • Domain text that pretrained vectors do not cover, such as SMS slang ("lar", "wif") or product codes, where vectors trained on your own messages know the words.
  • A fast baseline before a sequence model, such as the Recurrent neural network (RNN) of the next part, which reads the words in order instead of averaging them.
Watch out. Keep the label out of the feature table. If a column holding y is added to the vectors before X is taken from that table, the random forest reads the answer from that column and scores close to 1.0 (0.997 with this split), a number that says nothing about spam. Build X from the vector columns only, and train Word2Vec on the training messages, as the pipeline does.
Try it yourself
  • Set min_count=2 in AvgWord2Vec(...) and check how the vocabulary size and the number of zero-vector test messages change.
  • Train skip-gram vectors by adding sg=1 to the Word2Vec(...) call in fit(), with epochs=30, and compare the accuracy.
  • Print wv.most_similar("free", topn=5) after 5 and after 30 epochs: the 5-epoch neighbours all score about 0.999, the 30-epoch ones are spread out and read like spam offers.

Slow is fine. Stopping is the only problem.