Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

One-hot encoding for text

One-hot encoding is a text representation that turns each word into a vector as long as the vocabulary, with a 1 in that word's position and 0 everywhere else.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

Named entity recognition (NER) closed the preprocessing steps: the text is now clean tokens. A model needs numbers, so the next step turns words into vectors. One-hot encoding is the simplest way to do it, and its weaknesses explain the techniques that follow.

One-hot encoding of the food sentences · from the Complete NLP Machine Learning in One Shot video · 1:51:34 to 1:56:15

Building the vocabulary of unique words

The video's dataset has three documents, each with an output label: D1 “The food is good” (1), D2 “The food is bad” (0) and D3 “Pizza is Amazing” (1). The words stay as written, without lower-casing.

The vocabulary is the set of unique words across the whole corpus. Reading the documents in order gives seven of them: The, food, is, good, bad, Pizza and Amazing. The vocabulary size is V = 7.

Encoding each word as a vector of size V

Every word becomes a vector of length 7 with a single 1 in its own position: “The” is [1 0 0 0 0 0 0] and “food” is [0 1 0 0 0 0 0]. A document is the stack of its words' vectors, one row per word. D1 has four words, so it becomes a 4 × 7 matrix.

D2, “The food is bad”, shares its first three rows with D1. Its last row puts the 1 in the column of “bad”: [0 0 0 0 1 0 0]. D3 has three words, so its matrix is 3 × 7.

The three documents of the video one-hot encoded over the vocabulary The, food, is, good, bad, Pizza, Amazing: The food is good and The food is bad are 4 by 7 matrices with one 1 per row, and Pizza is Amazing is a 3 by 7 matrix.

Encoding the three documents with NumPy

The code builds the vocabulary in order of first appearance, as on the board, then places one 1 per row at the index of each word.

ExampleFrom the video, run with NumPy
import numpy as np

docs = ["The food is good", "The food is bad", "Pizza is Amazing"]

vocab = []                                   # unique words, in order of first appearance
for d in docs:
    for w in d.split():
        if w not in vocab:
            vocab.append(w)
print("vocabulary:", vocab, "V =", len(vocab))

def one_hot(doc):
    m = np.zeros((len(doc.split()), len(vocab)), dtype=int)
    for row, w in enumerate(doc.split()):
        m[row, vocab.index(w)] = 1           # a single 1 in the word's own column
    return m

for name, d in zip(["D1", "D2", "D3"], docs):
    m = one_hot(d)
    print(name, repr(d), "shape", m.shape)
    print(m)

What the three matrices show

  • D1 and D2 are 4 × 7, D3 is 3 × 7. The number of rows follows the number of words, so the documents do not share one shape.
  • Each row has exactly one 1. D1 holds 4 ones among 28 cells; every other value is 0.
  • D2 differs from D1 only in its last row, where the 1 moves from the column of good to the column of bad.
Why one-hot vectors carry no meaning · from the Complete NLP Machine Learning in One Shot video · 2:04:52 to 2:07:53

Measuring the distance between food, pizza and burger

The third disadvantage on the board is that no semantic meaning is captured: the vectors say nothing about how close two words are in meaning. The video takes a vocabulary of three words, food, pizza and burger, encoded as [1 0 0], [0 1 0] and [0 0 1].

Drawn on three axes, each word sits at the end of its own axis. The distance between any two of them is the same, so the encoding cannot say that pizza is closer to food than to burger. Every pair looks equally unrelated.

The one-hot vectors of food, pizza and burger are the corners of a triangle on three axes, every pair the same distance, square root of 2, apart and with a cosine of 0.

The code measures both the straight-line distance and the cosine of the angle between each pair. The cosine is 1 for vectors pointing the same way and 0 for vectors at right angles; Cosine similarity for documents covers it in full.

ExampleFrom the video, run with NumPy
import numpy as np
from itertools import combinations

vectors = {"food": np.array([1, 0, 0]), "pizza": np.array([0, 1, 0]), "burger": np.array([0, 0, 1])}

for a, b in combinations(vectors, 2):
    u, v = vectors[a], vectors[b]
    distance = np.linalg.norm(u - v)
    cosine = u @ v / (np.linalg.norm(u) * np.linalg.norm(v))
    print(f"{a:6} - {b:6}  distance {distance:.4f}  cosine {cosine:.1f}")

Every pair is √2 = 1.4142 apart with a cosine of 0.0: one-hot vectors are always at right angles to each other, whatever the words mean.

Listing the advantages and disadvantages of one-hot encoding

The board lists one advantage and four disadvantages.

  • Easy to implement. scikit-learn has OneHotEncoder and pandas has pd.get_dummies().
  • Sparse matrix. Almost every value is 0. With a vocabulary of 50,000 words, every word is a vector of 50,000 numbers holding a single 1. That costs memory and computation, and so many dimensions compared with the number of training examples raise the risk of overfitting.
  • No fixed size. D1 is 4 × 7 but D3 is 3 × 7. A machine learning algorithm needs every input to have the same number of features.
  • No semantic meaning. All words are equally far apart, as the food, pizza and burger vectors show.
  • Out of vocabulary (OOV). The test sentence “Burger is bad” contains Burger, which is not in the vocabulary, so it has no vector at all.

Encoding words with scikit-learn's OneHotEncoder

OneHotEncoder is built for categorical columns, so each token goes in its own row. It sorts the categories, and the sort is case-sensitive, so the column order differs from the board's.

Unknown words raise an error by default

Fitting on the three documents and then encoding the test sentence “Burger is bad” shows the out-of-vocabulary problem as a real error.

ExampleFrom the video, run on scikit-learn 1.9.1
import numpy as np
from sklearn.preprocessing import OneHotEncoder

tokens = np.array("The food is good The food is bad Pizza is Amazing".split()).reshape(-1, 1)
encoder = OneHotEncoder().fit(tokens)        # one token per row, like a categorical column
print(encoder.categories_[0])

encoder.transform([["Burger"], ["is"], ["bad"]])   # the test sentence "Burger is bad"

Ignoring unknown words

handle_unknown="ignore" turns an unknown word into a row of zeros instead. pd.get_dummies makes one column per word that appears in the text it is given.

ExampleFrom the video, run on scikit-learn 1.9.1
import pandas as pd

encoder = OneHotEncoder(handle_unknown="ignore").fit(tokens)
print(encoder.transform([["Burger"], ["is"], ["bad"]]).toarray().astype(int))

print(pd.get_dummies(pd.Series("The food is good".split()), dtype=int))

What the encoder printed

  • The categories are sorted: Amazing, Pizza, The, bad, food, good, is. Capital letters sort before lower-case ones, so the columns are not in the board's order.
  • Burger raises ValueError: Found unknown categories ['Burger'] with the default handle_unknown='error'.
  • With handle_unknown="ignore" the Burger row is all zeros, while is and bad get their usual single 1.
  • get_dummies on D1 gives only the four columns The, food, good and is, because it sees only the words of that one document.

One-hot encoding vs bag of words

One-hot encodingBag of words
Unitone vector per wordone vector per sentence
Shape of a documentwords × V, changes with lengthV, the same for every sentence
Valuesa single 1 per rowcounts of each word (or 1/0)
Word meaningnone: every pair at right anglesnone: each word is its own column
Typical usethe index fed to an embedding layerfeatures for a text classifier

Where you use one-hot encoding

  • Input to an embedding layer: a one-hot vector times an embedding matrix picks one row, which is how Embedding layer in Keras looks up a word's vector.
  • Class labels for a softmax output, compared with Categorical cross-entropy.
  • Small categorical features next to text, such as the day of the week a message was sent.
Watch out. OneHotEncoder raises ValueError on any word it did not see during fit, which stops a prediction on new text. Pass handle_unknown="ignore", and remember that the unknown word then adds nothing to the input.
Try it yourself
  • Add a fourth document, “Burger is bad”, to docs in the NumPy example. What is V now, and what shape does it get?
  • Lower-case every document with d.lower() before building the vocabulary. Does V change for these three documents?
  • Change ["Burger"] to ["Pizza"] in the OneHotEncoder example and check that the error goes away.

Little by little, you're building something great.