Word embeddings
A word embedding is a representation of a word as a dense, real-valued vector, learned so that words with similar meanings sit close together in the vector space.
Last updated: 07 Oct, 2026 · NumPy
One-hot encoding, bag of words and TF-IDF give vectors as long as the vocabulary, mostly zeros, in which two different words have nothing in common. Cosine similarity for documents showed that one-hot vectors are always at right angles. Word embeddings give each word a short vector of real numbers where the distance between two words means something.
Defining word embeddings
The board takes its definition from Wikipedia: in natural language processing, a word embedding is a term used for the representation of words for text analysis, typically in the form of a real-valued vector that encodes the meaning of the word, such that words that are closer in the vector space are expected to be similar in meaning.
The video's example is two words, happy and excited. A word embedding turns each into a vector. Reduced to two dimensions with a technique such as Principal component analysis (PCA) and plotted, the two land close together, which says they are similar words.
Splitting word embeddings into two families
The video sorts every technique for turning words into vectors into two types:
- Count or frequency: one-hot encoding, bag of words and TF-IDF, the methods of the previous part. They convert words or whole sentences into vectors by counting.
- Deep learning trained models: a neural network learns the vectors from a large corpus. This family addresses most of the disadvantages of the counting methods, and its best-known member is Word2Vec, which comes in two architectures, CBOW and skip-gram.
In most books and libraries, “embedding” means the second family only: dense, learned vectors. The counting methods are usually called sparse vectorisations. Besides Word2Vec, the learned family includes GloVe (Stanford, 2014), trained on global word co-occurrence counts, and fastText (Facebook, 2016), which builds a word's vector from its character n-grams.
Representing words by features
Bag of words and TF-IDF give 1s, 0s or decimals such as 0.25 and 0.6, with zeros everywhere else. Word2Vec works differently. Take a vocabulary, the unique words of a corpus: boy, girl, king, queen, apple and mango.
Each word is converted into a feature representation: one number for each of a set of features such as gender, royal, age and food. Google's pretrained Word2Vec uses 300 features, so every word becomes a vector of 300 numbers, 300 dimensions. In a real trained model the features have no names; gender, royal, age and food are there to build the intuition.
Filling the feature representation table
Each value expresses how a word relates to a feature. On gender, boy is −1 and girl is +1, the opposite ends. On royal, boy and girl are near 0 (0.01 and 0.02), because there is no relationship. king and queen are −0.92 and +0.93 on gender, and both high on royal: 0.95 and 0.96.
Age relates to king and queen (0.75 and 0.68, an old king) and also to apple and mango (0.95 and 0.96: fruit rots with time). Food is high only for apple and mango: 0.91 and 0.92. Each column of the table is one word's vector.
| Feature | Boy | Girl | King | Queen | Apple | Mango |
|---|---|---|---|---|---|---|
| Gender | −1 | 1 | −0.92 | 0.93 | 0.01 | 0.05 |
| Royal | 0.01 | 0.02 | 0.95 | 0.96 | −0.02 | 0.02 |
| Age | 0.03 | 0.02 | 0.75 | 0.68 | 0.95 | 0.96 |
| Food | – | – | – | – | 0.91 | 0.92 |
Comparing words by their feature vectors
Three rows of the table, gender, royal and age, have a value for every word. The code turns each column into a vector and measures the cosine between pairs of words.
import numpy as np
words = ["boy", "girl", "king", "queen", "apple", "mango"]
# boy girl king queen apple mango
table = np.array([[-1.00, 1.00, -0.92, 0.93, 0.01, 0.05], # gender
[ 0.01, 0.02, 0.95, 0.96, -0.02, 0.02], # royal
[ 0.03, 0.02, 0.75, 0.68, 0.95, 0.96]]) # age
vec = {w: table[:, j] for j, w in enumerate(words)} # each column is a word's vector
def cosine(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
print("vector for king:", vec["king"])
for a, b in [("boy", "girl"), ("king", "queen"), ("apple", "mango"), ("king", "apple"), ("boy", "king")]:
print(f"{a:5} - {b:5} cosine {cosine(vec[a], vec[b]):+.3f}")vector for king: [-0.92 0.95 0.75] boy - girl cosine -0.998 king - queen cosine +0.248 apple - mango cosine +0.998 king - apple cosine +0.474 boy - king cosine +0.626
What the cosine values show
- apple and mango score +0.998: their vectors point almost the same way, as two fruits should.
- boy and girl score −0.998: with only three features, gender dominates, and they sit at its opposite ends.
- king and queen score +0.248: they agree on royal and age but are opposite on gender. Three features are too few; the real 300 let similar words agree on many more of them.
- The values spread from −0.998 to +0.998, where one-hot vectors give 0 for every pair of different words.
Sparse vectors vs word embeddings
| One-hot, BoW, TF-IDF | Word embeddings (Word2Vec) | |
|---|---|---|
| Length of a word vector | the vocabulary size | fixed, for example 300 |
| Values | mostly 0 | real numbers in every position |
| Similar words | unrelated, cosine 0 | close together |
| How they are made | by counting | by training a neural network on a corpus |
| Unknown words | ignored | no vector either (fastText can build one) |
Where you use word embeddings
- Features for a classifier: averaging a sentence's word vectors, as in Average Word2Vec.
- Search and suggestions: finding words or documents that mean the same thing without sharing a word.
- The first layer of a neural network: an Embedding layer in Keras holds one learned vector per word.
Related
- Previous: Cosine similarity for documents
- Next: Word2Vec
- Reference: Word embedding on Wikipedia
- Add the pair
("girl", "queen")to the loop. Is it closer than boy and king? - Change king's gender value from −0.92 to 0.93 and rerun. What happens to the cosine of king and queen?
- Delete the age row from
table. How do the cosines of apple and mango, and of king and queen, change?
Little by little, you're building something great.