Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Word embeddings

A word embedding is a representation of a word as a dense, real-valued vector, learned so that words with similar meanings sit close together in the vector space.

Last updated: 07 Oct, 2026 · NumPy

One-hot encoding, bag of words and TF-IDF give vectors as long as the vocabulary, mostly zeros, in which two different words have nothing in common. Cosine similarity for documents showed that one-hot vectors are always at right angles. Word embeddings give each word a short vector of real numbers where the distance between two words means something.

Defining word embeddings

The board takes its definition from Wikipedia: in natural language processing, a word embedding is a term used for the representation of words for text analysis, typically in the form of a real-valued vector that encodes the meaning of the word, such that words that are closer in the vector space are expected to be similar in meaning.

The video's example is two words, happy and excited. A word embedding turns each into a vector. Reduced to two dimensions with a technique such as Principal component analysis (PCA) and plotted, the two land close together, which says they are similar words.

Count-based and deep learning word embeddings · from the Complete NLP Machine Learning in One Shot video · 2:47:55 to 2:49:43

Splitting word embeddings into two families

The video sorts every technique for turning words into vectors into two types:

  • Count or frequency: one-hot encoding, bag of words and TF-IDF, the methods of the previous part. They convert words or whole sentences into vectors by counting.
  • Deep learning trained models: a neural network learns the vectors from a large corpus. This family addresses most of the disadvantages of the counting methods, and its best-known member is Word2Vec, which comes in two architectures, CBOW and skip-gram.
The video's taxonomy: word embeddings split into count or frequency methods (one-hot encoding, bag of words, TF-IDF), which give sparse vectors, and deep learning trained models, where Word2Vec learns dense vectors with CBOW or skip-gram.

In most books and libraries, “embedding” means the second family only: dense, learned vectors. The counting methods are usually called sparse vectorisations. Besides Word2Vec, the learned family includes GloVe (Stanford, 2014), trained on global word co-occurrence counts, and fastText (Facebook, 2016), which builds a word's vector from its character n-grams.

Representing a vocabulary by features · from the Complete NLP Machine Learning in One Shot video · 2:54:19 to 2:57:33

Representing words by features

Bag of words and TF-IDF give 1s, 0s or decimals such as 0.25 and 0.6, with zeros everywhere else. Word2Vec works differently. Take a vocabulary, the unique words of a corpus: boy, girl, king, queen, apple and mango.

Each word is converted into a feature representation: one number for each of a set of features such as gender, royal, age and food. Google's pretrained Word2Vec uses 300 features, so every word becomes a vector of 300 numbers, 300 dimensions. In a real trained model the features have no names; gender, royal, age and food are there to build the intuition.

Filling the feature table for boy, girl, king and queen · from the Complete NLP Machine Learning in One Shot video · 2:58:01 to 3:01:00

Filling the feature representation table

Each value expresses how a word relates to a feature. On gender, boy is −1 and girl is +1, the opposite ends. On royal, boy and girl are near 0 (0.01 and 0.02), because there is no relationship. king and queen are −0.92 and +0.93 on gender, and both high on royal: 0.95 and 0.96.

Age relates to king and queen (0.75 and 0.68, an old king) and also to apple and mango (0.95 and 0.96: fruit rots with time). Food is high only for apple and mango: 0.91 and 0.92. Each column of the table is one word's vector.

The board's feature table: gender minus 1 for boy, plus 1 for girl, minus 0.92 for king, plus 0.93 for queen; royal 0.95 and 0.96 only for king and queen; age 0.75 and 0.68 for king and queen and 0.95 and 0.96 for apple and mango; food 0.91 and 0.92 for apple and mango; each column is a word's vector.
FeatureBoyGirlKingQueenAppleMango
Gender−11−0.920.930.010.05
Royal0.010.020.950.96−0.020.02
Age0.030.020.750.680.950.96
Food––––0.910.92

Comparing words by their feature vectors

Three rows of the table, gender, royal and age, have a value for every word. The code turns each column into a vector and measures the cosine between pairs of words.

ExampleFrom the video, run with NumPy
import numpy as np

words = ["boy", "girl", "king", "queen", "apple", "mango"]
#                     boy   girl  king   queen  apple  mango
table = np.array([[-1.00, 1.00, -0.92, 0.93, 0.01, 0.05],     # gender
                  [ 0.01, 0.02,  0.95, 0.96, -0.02, 0.02],    # royal
                  [ 0.03, 0.02,  0.75, 0.68, 0.95, 0.96]])    # age
vec = {w: table[:, j] for j, w in enumerate(words)}           # each column is a word's vector

def cosine(a, b):
    return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))

print("vector for king:", vec["king"])
for a, b in [("boy", "girl"), ("king", "queen"), ("apple", "mango"), ("king", "apple"), ("boy", "king")]:
    print(f"{a:5} - {b:5}  cosine {cosine(vec[a], vec[b]):+.3f}")

What the cosine values show

  • apple and mango score +0.998: their vectors point almost the same way, as two fruits should.
  • boy and girl score −0.998: with only three features, gender dominates, and they sit at its opposite ends.
  • king and queen score +0.248: they agree on royal and age but are opposite on gender. Three features are too few; the real 300 let similar words agree on many more of them.
  • The values spread from −0.998 to +0.998, where one-hot vectors give 0 for every pair of different words.

Sparse vectors vs word embeddings

One-hot, BoW, TF-IDFWord embeddings (Word2Vec)
Length of a word vectorthe vocabulary sizefixed, for example 300
Valuesmostly 0real numbers in every position
Similar wordsunrelated, cosine 0close together
How they are madeby countingby training a neural network on a corpus
Unknown wordsignoredno vector either (fastText can build one)

Where you use word embeddings

  • Features for a classifier: averaging a sentence's word vectors, as in Average Word2Vec.
  • Search and suggestions: finding words or documents that mean the same thing without sharing a word.
  • The first layer of a neural network: an Embedding layer in Keras holds one learned vector per word.
Watch out. The dimensions of a trained embedding are not named features like gender or royal. A single number in a word vector has no meaning of its own; only the distances and directions between vectors do.
Try it yourself
  • Add the pair ("girl", "queen") to the loop. Is it closer than boy and king?
  • Change king's gender value from −0.92 to 0.93 and rerun. What happens to the cosine of king and queen?
  • Delete the age row from table. How do the cosines of apple and mango, and of king and queen, change?

Little by little, you're building something great.