Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Word2Vec

Word2Vec is a technique, published by researchers at Google in 2013, that trains a shallow neural network on a large corpus to give every word a dense vector, so that words used in similar contexts get similar vectors.

Last updated: 07 Oct, 2026 · gensim 4.4 · NumPy

The feature table in Word embeddings had hand-written values. Word2Vec learns such numbers from raw text: nobody tells it what gender or royal means, and the vectors still end up placing related words together.

Defining Word2Vec

The board's definition reads: Word2Vec is a technique for natural language processing published in 2013. The algorithm uses a neural network model to learn word associations from a large corpus of text. Once trained, such a model can detect synonymous words or suggest additional words for a partial sentence. As the name implies, Word2Vec represents each distinct word with a particular list of numbers called a vector.

The two papers are Mikolov et al., “Efficient Estimation of Word Representations in Vector Space” and “Distributed Representations of Words and Phrases and their Compositionality”, both from 2013.

Learning from the text itself

Word2Vec needs no labelled data. It slides a window over the sentences and makes its own training examples: the words around a centre word, and the centre word itself. That makes it self-supervised: the labels come from the raw text. There are two ways to set up the prediction:

The trained vectors can come from a pretrained model, such as Google's, trained once on a huge corpus and downloaded, or from a model you train from scratch on your own text.

Similar words and king minus man plus woman with Google's vectors · from the Complete NLP Machine Learning in One Shot video · 3:48:47 to 3:52:21

Querying pretrained vectors for similar words and analogies

The video loads Google's pretrained vectors with gensim (the download is 1662.8 MB) and asks them questions. wv holds a 300-number vector for each of 3 million words and phrases.

python
import gensim.downloader as api
wv = api.load('word2vec-google-news-300')

wv['cricket']                          # a vector of 300 numbers
wv.most_similar('cricket')             # the 10 nearest words by cosine similarity
wv.most_similar('happy')
wv.similarity("hockey", "sports")      # the cosine between two words
vec = wv['king'] - wv['man'] + wv['woman']
wv.most_similar([vec])

In the video's notebook (gensim 3.6 on Colab), the nearest words to cricket are cricketing 0.837, cricketers 0.817, Test_cricket 0.809 and Twenty##_cricket 0.807; the Google News vocabulary writes digits as #, so Twenty## is Twenty20. The nearest words to happy start with glad 0.741, pleased 0.663, ecstatic 0.663 and overjoyed 0.660. hockey and sports have a similarity of 0.535.

For king − man + woman, most_similar([vec]) lists king first (0.845) and queen second (0.730). When you pass a raw vector, gensim does not leave out the words that built it, and the result is still close to king. Naming the words lets gensim exclude them, which puts queen first:

python
wv.most_similar(positive=['king', 'woman'], negative=['man'])   # king, man and woman are left out

Doing king − boy + girl on the board's feature table

The same arithmetic works on the feature table. Subtracting boy and adding girl moves king along the gender axis from the male end to the female end while it keeps its royal value, and the nearest word to the result is queen.

ExampleFrom the video, run with NumPy
import numpy as np
import matplotlib.pyplot as plt

words = ["boy", "girl", "king", "queen", "apple", "mango"]
#                     boy   girl  king   queen  apple  mango
table = np.array([[-1.00, 1.00, -0.92, 0.93, 0.01, 0.05],     # gender
                  [ 0.01, 0.02,  0.95, 0.96, -0.02, 0.02],    # royal
                  [ 0.03, 0.02,  0.75, 0.68, 0.95, 0.96]])    # age
vec = {w: table[:, j] for j, w in enumerate(words)}           # each column is a word's vector

def cosine(a, b):
    return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))

result = vec["king"] - vec["boy"] + vec["girl"]
print("king - boy + girl =", result.round(2))
for w in sorted(words, key=lambda w: -cosine(result, vec[w])):
    print(f"  {w:5} {cosine(result, vec[w]):.3f}")

plt.figure(figsize=(6, 4.5))
for w in words:
    plt.scatter(*vec[w][:2], color="tab:blue")
    plt.annotate(w, vec[w][:2], textcoords="offset points", xytext={"apple": (6, -14), "queen": (-44, 6)}.get(w, (6, 6)))
plt.annotate("", xy=vec["queen"][:2], xytext=vec["king"][:2], arrowprops=dict(arrowstyle="->", color="tab:orange"))
plt.annotate("", xy=vec["girl"][:2], xytext=vec["boy"][:2], arrowprops=dict(arrowstyle="->", color="tab:orange"))
plt.scatter(*result[:2], color="tab:orange", marker="x", s=80)
plt.annotate("king - boy + girl", result[:2], textcoords="offset points", xytext=(-110, -18), color="tab:orange")
plt.xlim(-1.2, 1.3)
plt.ylim(-0.1, 1.1)
plt.xlabel("gender")
plt.ylabel("royal")
plt.title("king - boy + girl lands next to queen")
plt.show()
Boy, girl, king, queen, apple and mango plotted by their gender and royal values from the board's table; an arrow from boy to girl runs parallel to an arrow from king to queen, and the point king minus boy plus girl, marked with an orange cross, lands next to queen.

What the analogy shows

  • king − boy + girl = [1.08 0.96 0.74] over gender, royal and age.
  • queen is the nearest word, at 0.998. girl comes next at 0.686, because the result kept king's high royal value.
  • The two arrows in the plot are parallel: boy to girl and king to queen are the same step along gender. An analogy works when a relationship is one direction in the vector space.

Finding words that are close but not synonyms

Word2Vec learns from contexts, not from definitions. In the video's notebook the ninth nearest word to happy is disappointed, at 0.627, one place above excited. Happy and disappointed appear in the same kinds of sentences (“I was ___ with the result”), so their vectors are close even though their meanings are opposite.

Word2Vec vs TF-IDF

TF-IDFWord2Vec
Representsa documenta word
Vector lengthvocabulary sizechosen, for example 300
Learned fromcounts in your corpusword contexts in a large corpus
good vs greatunrelated columnsclose vectors
Word orderlostonly nearness: which words fall inside the window

Where you use Word2Vec

  • Sentence features for a classifier, by averaging word vectors: Average Word2Vec.
  • Query expansion in search: also searching for the nearest words of each query word.
  • Starting weights for a neural network's embedding layer, instead of random ones.
Watch out. A pretrained model knows only its own vocabulary: wv['ineuron'] raises KeyError, and words with digits look different (Twenty##). The Google News model also needs about 3.6 GB of memory once loaded (3 million × 300 numbers × 4 bytes).
Try it yourself
  • Compute vec["queen"] - vec["girl"] + vec["boy"]. Which word is nearest?
  • Try vec["mango"] - vec["apple"] + vec["king"]. Does the result have a sensible nearest word, and why?
  • Change the plot to show the royal and age rows (vec[w][1:]) and look where the fruits move.

You understood something today that you didn't yesterday.