Average Word2Vec
Average Word2Vec is a way to turn a sentence into one fixed-length vector by taking the element-wise mean of the Word2Vec vectors of its words.
Last updated: 07 Oct, 2026 · gensim 4.4 · NLTK 3.10
Training Word2Vec with gensim gave one vector per word. A classifier needs one row per message, with the same number of features for every message, whatever its length.
Averaging the word vectors of a sentence
The video's text data is the one-hot example again: D1 “The food is good” (1), D2 “The food is bad” (0) and D3 “Pizza is Amazing” (1). With Google's pretrained Word2Vec model, every word becomes a vector of 300 dimensions: one for The, one for food, one for is and one for good.
That is four vectors for one sentence, but the model needs a single 300-dimension input for D1 next to its output 1. So the four vectors are averaged, element by element: the first number of the sentence vector is the mean of the four first numbers, and so on down all 300 rows. The result still has 300 dimensions, and the same happens for D2 and D3.
Averaging two documents by hand
The feature table from Word embeddings is small enough to average by hand. Take two two-word documents, “king queen” and “apple mango”.
import numpy as np
words = ["boy", "girl", "king", "queen", "apple", "mango"]
# boy girl king queen apple mango
table = np.array([[-1.00, 1.00, -0.92, 0.93, 0.01, 0.05], # gender
[ 0.01, 0.02, 0.95, 0.96, -0.02, 0.02], # royal
[ 0.03, 0.02, 0.75, 0.68, 0.95, 0.96]]) # age
vec = {w: table[:, j] for j, w in enumerate(words)} # each column is a word's vector
def cosine(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
doc1 = ["king", "queen"]
doc2 = ["apple", "mango"]
avg1 = np.mean([vec[w] for w in doc1], axis=0) # element by element, row by row
avg2 = np.mean([vec[w] for w in doc2], axis=0)
print("average of king and queen:", avg1.round(3))
print("average of apple and mango:", avg2.round(3))
print("cosine of the two documents:", round(cosine(avg1, avg2), 3))average of king and queen: [0.005 0.955 0.715] average of apple and mango: [0.03 0. 0.955] cosine of the two documents: 0.599
- king queen averages to [0.005, 0.955, 0.715]: the opposite gender values cancel, and the shared royal and age values stay.
- apple mango averages to [0.03, 0.0, 0.955]: no gender, no royalty, high age.
- Each document is now 3 numbers, the same length as one word vector, whatever the number of words. The two documents have a cosine of 0.599: they share only the age direction.
Building Average Word2Vec for the SMS messages
Skipping words the model does not know
The function averages only the words in the model's vocabulary. A message with no known word gets a vector of zeros, so every message still has a row.
def avg_word2vec(doc, wv):
vectors = [wv[w] for w in doc if w in wv]
if not vectors:
return np.zeros(wv.vector_size)
return np.mean(vectors, axis=0)Stacking one row per message
np.vstack turns the list of vectors into one matrix, with one token list per message so that row i belongs to message i. The example uses the SMS corpus and words from the gensim lesson and the 50-epoch CBOW model.
def avg_word2vec(doc, wv):
vectors = [wv[w] for w in doc if w in wv] # skip words the model does not know
if not vectors: # no known word at all
return np.zeros(wv.vector_size)
return np.mean(vectors, axis=0)
model = Word2Vec(words, epochs=50, seed=42, workers=1)
docs = [simple_preprocess(text) for text in corpus] # one token list per message
X = np.vstack([avg_word2vec(doc, model.wv) for doc in docs])
empty = sum(not any(w in model.wv for w in doc) for doc in docs)
print("X:", X.shape, " messages with no known word:", empty)
print(docs[0])
print(X[0][:5].round(4))
def cos(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
d1 = avg_word2vec("the food is good".split(), model.wv)
d2 = avg_word2vec("the food is bad".split(), model.wv)
print("good vs bad:", round(cos(model.wv["good"], model.wv["bad"]), 3), " their sentences:", round(cos(d1, d2), 3))X: (5572, 100) messages with no known word: 15 ['go', 'until', 'jurong', 'point', 'crazy', 'available', 'only', 'in', 'bugis', 'great', 'world', 'la', 'buffet', 'cine', 'there', 'got', 'amore', 'wat'] [-0.2619 0.0067 -0.043 -0.1094 -0.1248] good vs bad: 0.308 their sentences: 0.804
What the document vectors show
- X is 5,572 × 100: one row per message, the size of one word vector.
- 15 messages have no known word and get a row of zeros instead of an error.
- The first row is the mean of the vectors of the first message's words that the model knows, such as go, point and great.
- good and bad have a cosine of 0.308, but “the food is good” and “the food is bad” share three of their four words, so their averages score 0.804. Averaging blurs the one word that decides the sentiment.
Using pretrained Google News vectors
The video's practical uses gensim with Google's pretrained model, word2vec-google-news-300: vectors trained on part of the Google News dataset, about 100 billion words, with 300 dimensions for 3 million words and phrases. api.load downloads it once (1662.8 MB), and wv['king'] returns 300 numbers.
import gensim
from gensim.models import Word2Vec, KeyedVectors
import gensim.downloader as api
wv = api.load('word2vec-google-news-300')
vec_king = wv['king']
vec_king.shape # (300,) in the video's notebookavg_word2vec(doc, wv) works unchanged with these vectors: every message becomes 300 numbers. The pretrained model brings meanings learned from 100 billion words, but its words come from news text, where SMS spellings are rare or missing; the w in wv check skips any word it lacks.
Average Word2Vec vs TF-IDF
| TF-IDF | Average Word2Vec | |
|---|---|---|
| Vector length | vocabulary size (thousands) | vector size (100 or 300) |
| Values | mostly 0 | dense |
| Synonyms (good, great) | unrelated columns | close vectors |
| Word order | lost | lost |
| Frequent words | down-weighted by the idf | count as much as any other word |
Where you use Average Word2Vec
- Spam and sentiment classifiers: Spam classifier with Average Word2Vec and Sentiment analysis of Kindle reviews.
- Short-text similarity: comparing questions or tweets that use different words for the same thing.
- A quick baseline before sequence models that keep word order.
Related
- Previous: Training Word2Vec with gensim
- Next: Spam classifier with BoW and TF-IDF
- Reference: models.keyedvectors in the gensim documentation
- In the hand example, average the four words king, queen, apple and mango. Which of the two documents is the result closer to?
- Remove “the” and “is” from the two food sentences and compare their averages again.
- Train the model with
vector_size=50and check the new shape of X.
Little by little, you're building something great.