Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Training Word2Vec with gensim

gensim's Word2Vec class is a library implementation that trains CBOW or skip-gram word vectors on your own tokenized sentences.

Last updated: 07 Oct, 2026 · gensim 4.4 · NLTK 3.10

Google's pretrained vectors in Word2Vec come from news text. They were learned from news articles, where SMS spellings such as ur, lor or txt are rare. Training from scratch gives vectors for the words of your own corpus, and the video's materials train Word2Vec on the SMS spam messages.

Preparing the SMS messages as token lists

gensim's Word2Vec takes a list of sentences, each a list of words. The notebook in the video's materials keeps only the letters of each message, lower-cases it and lemmatizes every word with WordNet; stopwords stay, because they are context for the words around them.

Cleaning and lemmatizing each message

python
lemmatizer = WordNetLemmatizer()
corpus = []
for message in messages["message"]:
    review = re.sub("[^a-zA-Z]", " ", message).lower().split()
    corpus.append(" ".join(lemmatizer.lemmatize(word) for word in review))

Tokenizing with simple_preprocess

simple_preprocess lower-cases and splits the text, and drops tokens shorter than 2 characters, such as n, e and u. A message that is empty after cleaning gives no token list at all.

python
words = []
for text in corpus:
    for sent in sent_tokenize(text):
        words.append(simple_preprocess(sent))   # ['go', 'jurong', 'point', ...]

Setting Word2Vec's main parameters

Every setting of a Word2Vec model is a keyword argument. The example reads the defaults from the installed gensim:

ExampleRun on gensim 4.4.0
import inspect
from gensim.models import Word2Vec

defaults = inspect.signature(Word2Vec).parameters
for name in ["vector_size", "window", "min_count", "sg", "hs", "negative", "sample", "epochs", "workers", "seed"]:
    print(f"{name:12} {defaults[name].default}")

What each of them controls:

ParameterWhat it controls
vector_sizehow many numbers each word vector has
windowthe most words taken on each side of the centre word
min_countwords seen fewer times than this are left out of the vocabulary
sg0 for CBOW, 1 for skip-gram
hs, negativehierarchical softmax (hs=1) or negative sampling with that many noise words (negative=5)
samplehow strongly very frequent words are randomly skipped
epochshow many passes over the corpus
seed, workersthe random seed and the number of threads; workers=1 makes a run repeat exactly

Replacing the softmax with negative sampling

The networks in CBOW (continuous bag of words) and Skip-gram end in a softmax over the whole vocabulary, which means one score per word for every training pair. With 3 million words that is far too slow. gensim offers two cheaper objectives:

  • Negative sampling (the default, negative=5): for each true (word, context) pair, push their vectors together and push the word away from 5 random noise words, drawn with probability proportional to their count to the power 0.75 (ns_exponent).
  • Hierarchical softmax (hs=1, negative=0): arranges the vocabulary as a binary tree, so a probability needs about log₂ V steps instead of V.
The negative sampling objective for word w, true context c and K noise words n_k (maximised)

Subsampling (sample=0.001) randomly drops very frequent words such as the and to from the windows. They add little information, and dropping them both speeds up training and widens the effective window of the rarer words.

Training the model on the SMS messages

With seed=42 and workers=1 the run repeats exactly. The other settings are the defaults: CBOW, 100 numbers per word, 5 epochs.

ExampleFrom the video's materials, run on gensim 4.4.0
import re
import nltk
import numpy as np
import pandas as pd
from gensim.models import Word2Vec
from gensim.utils import simple_preprocess
from nltk import sent_tokenize
from nltk.stem import WordNetLemmatizer

nltk.download("wordnet", quiet=True)
nltk.download("punkt_tab", quiet=True)
url = "https://raw.githubusercontent.com/krishnaik06/Avgword2vec-Implementation/main/smsspamcollection/SMSSpamCollection"
messages = pd.read_csv(url, sep="\t", names=["label", "message"])

lemmatizer = WordNetLemmatizer()
corpus = []
for message in messages["message"]:
    review = re.sub("[^a-zA-Z]", " ", message).lower().split()
    corpus.append(" ".join(lemmatizer.lemmatize(word) for word in review))

words = []
for text in corpus:
    for sent in sent_tokenize(text):
        words.append(simple_preprocess(sent))

print(len(messages), "messages ->", len(words), "token lists; first:", words[0])

def mean_cosine(model, n=500):
    V = model.wv.vectors[:n]                       # the n most frequent words
    V = V / np.linalg.norm(V, axis=1, keepdims=True)
    C = V @ V.T
    return C[np.triu_indices(n, 1)].mean()

model = Word2Vec(words, seed=42, workers=1)        # gensim defaults: CBOW, 5 epochs
print("corpus_count:", model.corpus_count, " epochs:", model.epochs, " vocabulary:", len(model.wv))
print("vector for 'good':", model.wv["good"].shape)
print("similar to 'good':", [(w, round(s, 3)) for w, s in model.wv.similar_by_word("good", topn=5)])
print("mean cosine of the 500 most frequent words:", round(mean_cosine(model), 4))

What the default model shows

  • 5,572 messages give 5,569 token lists: three messages contain no letters at all.
  • The vocabulary has 1,721 words, those seen at least 5 times (min_count=5), and each word gets 100 numbers.
  • Every neighbour of good scores 0.999, and the 500 most frequent words have a mean pairwise cosine of 0.9926: nearly all the vectors point the same way.
  • The model is undertrained. Five passes over 5,569 short messages are not enough to pull the words apart, so the similarity lists mean nothing yet.

Training for more epochs

More passes over the same data spread the vectors out. The example retrains CBOW for 50 epochs and skip-gram for 20, with mean_cosine and the token lists from above.

ExampleFrom the video's materials, run on gensim 4.4.0
cbow50 = Word2Vec(words, epochs=50, seed=42, workers=1)
skip20 = Word2Vec(words, sg=1, epochs=20, seed=42, workers=1)
for name, m in [("CBOW, 50 epochs", cbow50), ("skip-gram, 20 epochs", skip20)]:
    print(name, " mean cosine:", round(mean_cosine(m), 4))
    print("   similar to 'good':", [(w, round(s, 3)) for w, s in m.wv.similar_by_word("good", topn=5)])
    print("   similar to 'prize':", [w for w, s in m.wv.similar_by_word("prize", topn=5)])
  • CBOW with 50 epochs drops the mean cosine from 0.9926 to 0.037, and the neighbours of good become great, brings, blessing, wonderful and lovely.
  • Skip-gram with 20 epochs lands at a mean cosine of 0.2406, with blessing, brings, lovely and great near good.
  • prize finds the vocabulary of spam messages in both models, learned without any labels.

Plotting the trained vectors with PCA

A 100-number vector cannot be drawn, but Principal component analysis (PCA) reduces a set of them to the two directions of largest spread. The plot shows 14 words from the 50-epoch CBOW model.

ExampleFrom the video's materials, run on gensim 4.4.0
import matplotlib.pyplot as plt
from sklearn.decomposition import PCA

model = Word2Vec(words, epochs=50, seed=42, workers=1)
show = ["good", "great", "happy", "love", "free", "prize", "claim", "win", "cash",
        "call", "txt", "home", "night", "morning"]
pca = PCA(n_components=2)
points = pca.fit_transform(model.wv[show])
print("share of the variance kept by 2 components:", round(float(pca.explained_variance_ratio_.sum()), 3))

plt.figure(figsize=(7, 5))
plt.scatter(points[:, 0], points[:, 1], color="tab:blue")
for w, (x, y) in zip(show, points):
    plt.annotate(w, (x, y), textcoords="offset points", xytext=(5, -14) if w == "happy" else (5, 4))
plt.title("Word2Vec trained on the SMS messages, reduced to 2-D with PCA")
plt.xlabel("PC 1")
plt.ylabel("PC 2")
plt.show()
Fourteen words from the Word2Vec model trained on the SMS messages, reduced to two dimensions with PCA: spam words such as free, prize, claim, win and cash form one group, away from everyday words such as good, happy, home, night and morning.

The spam words free, prize, claim, win and cash sit on the right with call and txt, away from everyday words such as good, happy, home and night on the left. The two components keep 41.6% of the variance of these 14 vectors, so neighbours in the plot are a rough guide.

Handling words the model has never seen

The vocabulary is fixed when the model is trained. Asking for an unknown word is an error:

ExampleFrom the video's materials, run on gensim 4.4.0
model = Word2Vec(words, seed=42, workers=1)
print("ineuron" in model.wv, "good" in model.wv)
model.wv["ineuron"]

Check with word in model.wv before the lookup and skip the words that are missing. To keep a model, model.save("sms_word2vec.model") writes it to disk and Word2Vec.load("sms_word2vec.model") reads it back.

CBOW vs skip-gram on the SMS messages

CBOW, 50 epochsSkip-gram, 20 epochs
Settingsg=0, epochs=50sg=1, epochs=20
Mean cosine of the top 500 words0.0370.2406
Nearest to goodgreat, brings, blessingblessing, brings, lovely
Examples per window1one per context word

Where you train your own Word2Vec

  • Text with its own vocabulary: SMS and chat slang, medical notes, legal contracts, code comments.
  • Languages or domains with no good pretrained vectors.
  • Features for a classifier: Average Word2Vec turns the trained vectors into one row per message.
Watch out. With the default workers=3, two runs with the same seed give different vectors, because the threads finish in a different order. For a run that repeats, pass workers=1 together with seed, as the gensim documentation says.
Try it yourself
  • Train with min_count=2 instead of the default. How large is the vocabulary now?
  • Change epochs=50 to epochs=20 for the CBOW model. Where does the mean cosine land?
  • Print cbow50.wv.similar_by_word("free") and skip20.wv.similar_by_word("free") and compare the two lists.
PreviousSkip-gram

Every expert started right here.