Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Average Word2Vec

Average Word2Vec is a way to turn a sentence into one fixed-length vector by taking the element-wise mean of the Word2Vec vectors of its words.

Last updated: 07 Oct, 2026 · gensim 4.4 · NLTK 3.10

Training Word2Vec with gensim gave one vector per word. A classifier needs one row per message, with the same number of features for every message, whatever its length.

Averaging word vectors into one sentence vector · from the Complete NLP Machine Learning in One Shot video · 3:39:39 to 3:43:40

Averaging the word vectors of a sentence

The video's text data is the one-hot example again: D1 “The food is good” (1), D2 “The food is bad” (0) and D3 “Pizza is Amazing” (1). With Google's pretrained Word2Vec model, every word becomes a vector of 300 dimensions: one for The, one for food, one for is and one for good.

That is four vectors for one sentence, but the model needs a single 300-dimension input for D1 next to its output 1. So the four vectors are averaged, element by element: the first number of the sentence vector is the mean of the four first numbers, and so on down all 300 rows. The result still has 300 dimensions, and the same happens for D2 and D3.

Average Word2Vec: the mean of the vectors of the words in document d
Average Word2Vec for The food is good: the four 300-number word vectors are averaged row by row into one 300-number document vector, and D1, D2 and D3 each become one such row next to their output.

Averaging two documents by hand

The feature table from Word embeddings is small enough to average by hand. Take two two-word documents, “king queen” and “apple mango”.

ExampleFrom the video, run with NumPy
import numpy as np

words = ["boy", "girl", "king", "queen", "apple", "mango"]
#                     boy   girl  king   queen  apple  mango
table = np.array([[-1.00, 1.00, -0.92, 0.93, 0.01, 0.05],     # gender
                  [ 0.01, 0.02,  0.95, 0.96, -0.02, 0.02],    # royal
                  [ 0.03, 0.02,  0.75, 0.68, 0.95, 0.96]])    # age
vec = {w: table[:, j] for j, w in enumerate(words)}           # each column is a word's vector

def cosine(a, b):
    return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))

doc1 = ["king", "queen"]
doc2 = ["apple", "mango"]
avg1 = np.mean([vec[w] for w in doc1], axis=0)       # element by element, row by row
avg2 = np.mean([vec[w] for w in doc2], axis=0)
print("average of king and queen:", avg1.round(3))
print("average of apple and mango:", avg2.round(3))
print("cosine of the two documents:", round(cosine(avg1, avg2), 3))
  • king queen averages to [0.005, 0.955, 0.715]: the opposite gender values cancel, and the shared royal and age values stay.
  • apple mango averages to [0.03, 0.0, 0.955]: no gender, no royalty, high age.
  • Each document is now 3 numbers, the same length as one word vector, whatever the number of words. The two documents have a cosine of 0.599: they share only the age direction.

Building Average Word2Vec for the SMS messages

Skipping words the model does not know

The function averages only the words in the model's vocabulary. A message with no known word gets a vector of zeros, so every message still has a row.

python
def avg_word2vec(doc, wv):
    vectors = [wv[w] for w in doc if w in wv]
    if not vectors:
        return np.zeros(wv.vector_size)
    return np.mean(vectors, axis=0)

Stacking one row per message

np.vstack turns the list of vectors into one matrix, with one token list per message so that row i belongs to message i. The example uses the SMS corpus and words from the gensim lesson and the 50-epoch CBOW model.

ExampleFrom the video's materials, run on gensim 4.4.0
def avg_word2vec(doc, wv):
    vectors = [wv[w] for w in doc if w in wv]          # skip words the model does not know
    if not vectors:                                     # no known word at all
        return np.zeros(wv.vector_size)
    return np.mean(vectors, axis=0)

model = Word2Vec(words, epochs=50, seed=42, workers=1)
docs = [simple_preprocess(text) for text in corpus]     # one token list per message
X = np.vstack([avg_word2vec(doc, model.wv) for doc in docs])
empty = sum(not any(w in model.wv for w in doc) for doc in docs)
print("X:", X.shape, " messages with no known word:", empty)
print(docs[0])
print(X[0][:5].round(4))

def cos(a, b):
    return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
d1 = avg_word2vec("the food is good".split(), model.wv)
d2 = avg_word2vec("the food is bad".split(), model.wv)
print("good vs bad:", round(cos(model.wv["good"], model.wv["bad"]), 3), " their sentences:", round(cos(d1, d2), 3))

What the document vectors show

  • X is 5,572 × 100: one row per message, the size of one word vector.
  • 15 messages have no known word and get a row of zeros instead of an error.
  • The first row is the mean of the vectors of the first message's words that the model knows, such as go, point and great.
  • good and bad have a cosine of 0.308, but “the food is good” and “the food is bad” share three of their four words, so their averages score 0.804. Averaging blurs the one word that decides the sentiment.
Loading the pretrained Google News Word2Vec model · from the Complete NLP Machine Learning in One Shot video · 3:45:40 to 3:48:47

Using pretrained Google News vectors

The video's practical uses gensim with Google's pretrained model, word2vec-google-news-300: vectors trained on part of the Google News dataset, about 100 billion words, with 300 dimensions for 3 million words and phrases. api.load downloads it once (1662.8 MB), and wv['king'] returns 300 numbers.

python
import gensim
from gensim.models import Word2Vec, KeyedVectors
import gensim.downloader as api

wv = api.load('word2vec-google-news-300')

vec_king = wv['king']
vec_king.shape                    # (300,) in the video's notebook

avg_word2vec(doc, wv) works unchanged with these vectors: every message becomes 300 numbers. The pretrained model brings meanings learned from 100 billion words, but its words come from news text, where SMS spellings are rare or missing; the w in wv check skips any word it lacks.

Average Word2Vec vs TF-IDF

TF-IDFAverage Word2Vec
Vector lengthvocabulary size (thousands)vector size (100 or 300)
Valuesmostly 0dense
Synonyms (good, great)unrelated columnsclose vectors
Word orderlostlost
Frequent wordsdown-weighted by the idfcount as much as any other word

Where you use Average Word2Vec

Watch out. Every word counts the same in the average, so frequent words such as the and is pull every sentence towards the same point, and order is gone: “the food is good” and “the food is bad” end up close. Removing stopwords first, or weighting each word by its TF-IDF, keeps more of what makes a sentence different.
Try it yourself
  • In the hand example, average the four words king, queen, apple and mango. Which of the two documents is the result closer to?
  • Remove “the” and “is” from the two food sentences and compare their averages again.
  • Train the model with vector_size=50 and check the new shape of X.

Little by little, you're building something great.