Training Word2Vec with gensim
gensim's Word2Vec class is a library implementation that trains CBOW or skip-gram word vectors on your own tokenized sentences.
Last updated: 07 Oct, 2026 · gensim 4.4 · NLTK 3.10
Google's pretrained vectors in Word2Vec come from news text. They were learned from news articles, where SMS spellings such as ur, lor or txt are rare. Training from scratch gives vectors for the words of your own corpus, and the video's materials train Word2Vec on the SMS spam messages.
Preparing the SMS messages as token lists
gensim's Word2Vec takes a list of sentences, each a list of words. The notebook in the video's materials keeps only the letters of each message, lower-cases it and lemmatizes every word with WordNet; stopwords stay, because they are context for the words around them.
Cleaning and lemmatizing each message
lemmatizer = WordNetLemmatizer()
corpus = []
for message in messages["message"]:
review = re.sub("[^a-zA-Z]", " ", message).lower().split()
corpus.append(" ".join(lemmatizer.lemmatize(word) for word in review))Tokenizing with simple_preprocess
simple_preprocess lower-cases and splits the text, and drops tokens shorter than 2 characters, such as n, e and u. A message that is empty after cleaning gives no token list at all.
words = []
for text in corpus:
for sent in sent_tokenize(text):
words.append(simple_preprocess(sent)) # ['go', 'jurong', 'point', ...]Setting Word2Vec's main parameters
Every setting of a Word2Vec model is a keyword argument. The example reads the defaults from the installed gensim:
import inspect
from gensim.models import Word2Vec
defaults = inspect.signature(Word2Vec).parameters
for name in ["vector_size", "window", "min_count", "sg", "hs", "negative", "sample", "epochs", "workers", "seed"]:
print(f"{name:12} {defaults[name].default}")vector_size 100 window 5 min_count 5 sg 0 hs 0 negative 5 sample 0.001 epochs 5 workers 3 seed 1
What each of them controls:
| Parameter | What it controls |
|---|---|
vector_size | how many numbers each word vector has |
window | the most words taken on each side of the centre word |
min_count | words seen fewer times than this are left out of the vocabulary |
sg | 0 for CBOW, 1 for skip-gram |
hs, negative | hierarchical softmax (hs=1) or negative sampling with that many noise words (negative=5) |
sample | how strongly very frequent words are randomly skipped |
epochs | how many passes over the corpus |
seed, workers | the random seed and the number of threads; workers=1 makes a run repeat exactly |
Replacing the softmax with negative sampling
The networks in CBOW (continuous bag of words) and Skip-gram end in a softmax over the whole vocabulary, which means one score per word for every training pair. With 3 million words that is far too slow. gensim offers two cheaper objectives:
- Negative sampling (the default,
negative=5): for each true (word, context) pair, push their vectors together and push the word away from 5 random noise words, drawn with probability proportional to their count to the power 0.75 (ns_exponent). - Hierarchical softmax (
hs=1, negative=0): arranges the vocabulary as a binary tree, so a probability needs about log₂ V steps instead of V.
Subsampling (sample=0.001) randomly drops very frequent words such as the and to from the windows. They add little information, and dropping them both speeds up training and widens the effective window of the rarer words.
Training the model on the SMS messages
With seed=42 and workers=1 the run repeats exactly. The other settings are the defaults: CBOW, 100 numbers per word, 5 epochs.
import re
import nltk
import numpy as np
import pandas as pd
from gensim.models import Word2Vec
from gensim.utils import simple_preprocess
from nltk import sent_tokenize
from nltk.stem import WordNetLemmatizer
nltk.download("wordnet", quiet=True)
nltk.download("punkt_tab", quiet=True)
url = "https://raw.githubusercontent.com/krishnaik06/Avgword2vec-Implementation/main/smsspamcollection/SMSSpamCollection"
messages = pd.read_csv(url, sep="\t", names=["label", "message"])
lemmatizer = WordNetLemmatizer()
corpus = []
for message in messages["message"]:
review = re.sub("[^a-zA-Z]", " ", message).lower().split()
corpus.append(" ".join(lemmatizer.lemmatize(word) for word in review))
words = []
for text in corpus:
for sent in sent_tokenize(text):
words.append(simple_preprocess(sent))
print(len(messages), "messages ->", len(words), "token lists; first:", words[0])
def mean_cosine(model, n=500):
V = model.wv.vectors[:n] # the n most frequent words
V = V / np.linalg.norm(V, axis=1, keepdims=True)
C = V @ V.T
return C[np.triu_indices(n, 1)].mean()
model = Word2Vec(words, seed=42, workers=1) # gensim defaults: CBOW, 5 epochs
print("corpus_count:", model.corpus_count, " epochs:", model.epochs, " vocabulary:", len(model.wv))
print("vector for 'good':", model.wv["good"].shape)
print("similar to 'good':", [(w, round(s, 3)) for w, s in model.wv.similar_by_word("good", topn=5)])
print("mean cosine of the 500 most frequent words:", round(mean_cosine(model), 4))5572 messages -> 5569 token lists; first: ['go', 'until', 'jurong', 'point', 'crazy', 'available', 'only', 'in', 'bugis', 'great', 'world', 'la', 'buffet', 'cine', 'there', 'got', 'amore', 'wat']
corpus_count: 5569 epochs: 5 vocabulary: 1721
vector for 'good': (100,)
similar to 'good': [('hope', 0.999), ('all', 0.999), ('well', 0.999), ('day', 0.999), ('my', 0.999)]
mean cosine of the 500 most frequent words: 0.9926What the default model shows
- 5,572 messages give 5,569 token lists: three messages contain no letters at all.
- The vocabulary has 1,721 words, those seen at least 5 times (
min_count=5), and each word gets 100 numbers. - Every neighbour of good scores 0.999, and the 500 most frequent words have a mean pairwise cosine of 0.9926: nearly all the vectors point the same way.
- The model is undertrained. Five passes over 5,569 short messages are not enough to pull the words apart, so the similarity lists mean nothing yet.
Training for more epochs
More passes over the same data spread the vectors out. The example retrains CBOW for 50 epochs and skip-gram for 20, with mean_cosine and the token lists from above.
cbow50 = Word2Vec(words, epochs=50, seed=42, workers=1)
skip20 = Word2Vec(words, sg=1, epochs=20, seed=42, workers=1)
for name, m in [("CBOW, 50 epochs", cbow50), ("skip-gram, 20 epochs", skip20)]:
print(name, " mean cosine:", round(mean_cosine(m), 4))
print(" similar to 'good':", [(w, round(s, 3)) for w, s in m.wv.similar_by_word("good", topn=5)])
print(" similar to 'prize':", [w for w, s in m.wv.similar_by_word("prize", topn=5)])CBOW, 50 epochs mean cosine: 0.037
similar to 'good': [('great', 0.519), ('brings', 0.516), ('blessing', 0.497), ('wonderful', 0.476), ('lovely', 0.465)]
similar to 'prize': ['guaranteed', 'final', 'reward', 'bonus', 'claim']
skip-gram, 20 epochs mean cosine: 0.2406
similar to 'good': [('blessing', 0.583), ('brings', 0.555), ('lovely', 0.543), ('great', 0.536), ('ttyl', 0.521)]
similar to 'prize': ['guaranteed', 'reward', 'bonus', 'final', 'caller']- CBOW with 50 epochs drops the mean cosine from 0.9926 to 0.037, and the neighbours of good become great, brings, blessing, wonderful and lovely.
- Skip-gram with 20 epochs lands at a mean cosine of 0.2406, with blessing, brings, lovely and great near good.
- prize finds the vocabulary of spam messages in both models, learned without any labels.
Plotting the trained vectors with PCA
A 100-number vector cannot be drawn, but Principal component analysis (PCA) reduces a set of them to the two directions of largest spread. The plot shows 14 words from the 50-epoch CBOW model.
import matplotlib.pyplot as plt
from sklearn.decomposition import PCA
model = Word2Vec(words, epochs=50, seed=42, workers=1)
show = ["good", "great", "happy", "love", "free", "prize", "claim", "win", "cash",
"call", "txt", "home", "night", "morning"]
pca = PCA(n_components=2)
points = pca.fit_transform(model.wv[show])
print("share of the variance kept by 2 components:", round(float(pca.explained_variance_ratio_.sum()), 3))
plt.figure(figsize=(7, 5))
plt.scatter(points[:, 0], points[:, 1], color="tab:blue")
for w, (x, y) in zip(show, points):
plt.annotate(w, (x, y), textcoords="offset points", xytext=(5, -14) if w == "happy" else (5, 4))
plt.title("Word2Vec trained on the SMS messages, reduced to 2-D with PCA")
plt.xlabel("PC 1")
plt.ylabel("PC 2")
plt.show()share of the variance kept by 2 components: 0.416
The spam words free, prize, claim, win and cash sit on the right with call and txt, away from everyday words such as good, happy, home and night on the left. The two components keep 41.6% of the variance of these 14 vectors, so neighbours in the plot are a rough guide.
Handling words the model has never seen
The vocabulary is fixed when the model is trained. Asking for an unknown word is an error:
model = Word2Vec(words, seed=42, workers=1)
print("ineuron" in model.wv, "good" in model.wv)
model.wv["ineuron"]False True
Traceback (most recent call last):
File "main.py", line 3, in <module>
model.wv["ineuron"]
KeyError: "Key 'ineuron' not present"Check with word in model.wv before the lookup and skip the words that are missing. To keep a model, model.save("sms_word2vec.model") writes it to disk and Word2Vec.load("sms_word2vec.model") reads it back.
CBOW vs skip-gram on the SMS messages
| CBOW, 50 epochs | Skip-gram, 20 epochs | |
|---|---|---|
| Setting | sg=0, epochs=50 | sg=1, epochs=20 |
| Mean cosine of the top 500 words | 0.037 | 0.2406 |
| Nearest to good | great, brings, blessing | blessing, brings, lovely |
| Examples per window | 1 | one per context word |
Where you train your own Word2Vec
- Text with its own vocabulary: SMS and chat slang, medical notes, legal contracts, code comments.
- Languages or domains with no good pretrained vectors.
- Features for a classifier: Average Word2Vec turns the trained vectors into one row per message.
workers=3, two runs with the same seed give different vectors, because the threads finish in a different order. For a run that repeats, pass workers=1 together with seed, as the gensim documentation says.Related
- Previous: Skip-gram
- Next: Average Word2Vec
- Reference: models.word2vec in the gensim documentation
- Train with
min_count=2instead of the default. How large is the vocabulary now? - Change
epochs=50toepochs=20for the CBOW model. Where does the mean cosine land? - Print
cbow50.wv.similar_by_word("free")andskip20.wv.similar_by_word("free")and compare the two lists.
Every expert started right here.