Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Categorical cross-entropy

Categorical cross-entropy is a multi-class loss function that compares a one-hot label with the softmax probabilities of the output layer and charges −ln of the probability given to the true class.

Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras

Binary cross-entropy handles two classes with one sigmoid output. With three or more classes the output layer has one neuron per class, Softmax turns them into probabilities, and the loss needs the label written as a vector too.

Categorical cross-entropy and one-hot labels · from the Deep Learning In-depth Tutorials in 5 Hours video · 144:30 to 148:47

One-hot encoding the Good, Bad and Neutral labels

The board's dataset has three features and a three-class output. The first step is to one-hot encode the output: one column per class, a 1 in the column of the row's class and 0 everywhere else.

f1f2f3OutputGoodBadNeutral
234Good100
567Bad010
8910Neutral001

Writing the categorical cross-entropy formula

  • i is the row (the record) and j the column (the class); C is the number of categories, 3 here.
  • yi = [yi1, yi2, …, yiC] is the one-hot row, with yij = 1 if the record is in class j and 0 otherwise.
  • ŷij is the predicted probability of class j for record i. It comes from the softmax activation in the output layer, which divides ezj by the sum of ezk over all classes, so each row of ŷ adds up to 1.

Keeping only the true class term

Because the label is one-hot, every term with yij = 0 drops out of the sum. For a Good record, y = [1, 0, 0], and a softmax output ŷ = [0.7, 0.2, 0.1]:

So the loss is −ln of the probability the network gave to the right class. The other probabilities still matter through softmax: raising them lowers the true class's share.

The Good, Bad and Neutral table with its one-hot columns, and for the first row the label 1, 0, 0 lined up with the softmax output 0.7, 0.2, 0.1: the products with 0 vanish, and the loss is minus ln 0.7, which is 0.357.

Computing categorical cross-entropy in NumPy

Labels and softmax outputs

The three board rows, with logits chosen so the first row's softmax is about [0.7, 0.2, 0.1]:

python
import numpy as np

classes = ["Good", "Bad", "Neutral"]
labels = np.array([0, 1, 2])               # rows (2,3,4) Good, (5,6,7) Bad, (8,9,10) Neutral
one_hot = np.eye(3)[labels]                # [1,0,0], [0,1,0], [0,0,1]

def softmax(z):
    e = np.exp(z - z.max(axis=1, keepdims=True))
    return e / e.sum(axis=1, keepdims=True)

logits = np.array([[2.0, 0.75, 0.05], [0.3, 1.6, -0.4], [0.5, 0.9, 1.1]])
y_hat = softmax(logits)                    # one row of probabilities per record

The one-hot loss and the integer-label loss

python
def cce(one_hot, y_hat):
    return -(one_hot * np.log(y_hat)).sum(axis=1)       # one-hot labels

def sparse_cce(labels, y_hat):
    return -np.log(y_hat[np.arange(len(labels)), labels])  # integer labels
ExampleRun on NumPy 2.5.3
print("one-hot:\n", one_hot)
print("y_hat:\n", y_hat.round(3))
print("categorical, one-hot labels:", cce(one_hot, y_hat).round(3))
print("sparse, labels", labels, "       :", sparse_cce(labels, y_hat).round(3))
print("cost:", round(cce(one_hot, y_hat).mean(), 3))

notes = np.array([[0.2, 0.3, 0.5]])         # the notes' softmax output
print("notes' row, true class 2:", cce(np.array([[0, 0, 1]]), notes).round(3), sparse_cce(np.array([2]), notes).round(3))

What the losses printed

  • The first row gets ŷ = [0.7, 0.2, 0.1] and a loss of 0.357, the −ln 0.7 worked above.
  • The other rows cost 0.342 (Bad gets 0.71) and 0.862 (Neutral gets only 0.422, the network is unsure), and the cost is their mean, 0.52.
  • One-hot and integer labels give identical losses, row by row, and the notes' [0.2, 0.3, 0.5] with the true class at index 2 costs −ln 0.5 = 0.693 either way.

Using integer labels with sparse categorical cross-entropy

Sparse categorical cross-entropy is the same loss with the label written as a class index, 2, instead of the one-hot row [0, 0, 1]. It picks −ln ŷ at that index directly. With many classes (10,000 words in a vocabulary) the integer labels save the memory of a huge one-hot matrix.

The notes list a disadvantage, losing the probabilities of the other categories. That describes reading only the top index of a prediction; the sparse loss gives the same number as the one-hot loss, and the model still outputs the full softmax row.

Computing the loss from logits

A softmax written as ez / Σ ez overflows for large logits: e1000 is infinite in floating point, and infinity divided by infinity is nan. Taking the log of softmax in one step, after subtracting the row's largest logit, stays finite:

python
def cce_from_logits(one_hot, z):
    log_p = z - z.max(axis=1, keepdims=True)
    log_p = log_p - np.log(np.exp(log_p).sum(axis=1, keepdims=True))   # log softmax
    return -(one_hot * log_p).sum(axis=1)
ExampleRun on NumPy 2.5.3
big = np.array([[1000.0, 0.0, -1000.0]])
with np.errstate(all="ignore"):
    naive = np.exp(big) / np.exp(big).sum()
print("naive softmax:", naive)
print("from logits, board rows:", cce_from_logits(one_hot, logits).round(3))
print("from logits, big row, true class Bad:", cce_from_logits(np.array([[0, 1, 0]]), big))

The naive softmax prints [nan, 0., 0.]; the log-softmax form gives the board rows the same 0.357, 0.342 and 0.862, and the big row a loss of 1000 for a true class that the network scored 1000 below the top. This is what from_logits=True does in Keras.

Categorical cross-entropy in Keras

python
import keras
from keras import layers

model = keras.Sequential([keras.Input(shape=(3,)),        # f1, f2, f3
                          layers.Dense(8, activation="relu"),
                          layers.Dense(3, activation="softmax")])   # Good, Bad, Neutral
model.compile(optimizer="adam", loss="categorical_crossentropy")    # one-hot labels
# model.compile(optimizer="adam", loss="sparse_categorical_crossentropy")  # labels 0, 1, 2

# logits: no softmax in the last layer, and the loss applies it
# layers.Dense(3) with loss=keras.losses.CategoricalCrossentropy(from_logits=True)

Pick the loss by the label format: one-hot rows go with categorical_crossentropy, integer labels 0, 1, 2 with sparse_categorical_crossentropy. The model is the same.

Categorical vs sparse categorical cross-entropy

Categorical cross-entropySparse categorical cross-entropy
Label for Neutral[0, 0, 1]2
Loss value−ln ŷ of the true classthe same number
Model outputsoftmax over C classessoftmax over C classes
Keras name"categorical_crossentropy""sparse_categorical_crossentropy"
Best forlabels already one-hot, or soft labels such as [0.9, 0.05, 0.05]integer labels, many classes

Where you use categorical cross-entropy

  • Single-label multi-class problems: a review that is Good, Bad or Neutral, a digit from 0 to 9, a type of flower.
  • Image classification: the softmax layer at the end of a convolutional network with one neuron per class.
  • Language models: predicting the next word out of a vocabulary, with sparse labels because the vocabulary is large.
Watch out. Match the loss to the labels. Integer labels with categorical_crossentropy, or one-hot rows with sparse_categorical_crossentropy, stop Keras with a shape error at the first batch; the fix is to change the loss name, not the model.
Try it yourself
  • Change the first row's logits to [4.0, 0.75, 0.05]: its probability for Good rises and its loss falls below 0.1.
  • Give the Neutral record the label Bad (labels = np.array([0, 1, 1])): its loss rises from 0.862 to −ln 0.346 ≈ 1.062.
  • Use the notes' five-class output [0.1, 0.2, 0.3, 0.2, 0.2] with true class 2: the loss is −ln 0.3 ≈ 1.204.

Slow is fine. Stopping is the only problem.