Categorical cross-entropy
Categorical cross-entropy is a multi-class loss function that compares a one-hot label with the softmax probabilities of the output layer and charges −ln of the probability given to the true class.
Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras
Binary cross-entropy handles two classes with one sigmoid output. With three or more classes the output layer has one neuron per class, Softmax turns them into probabilities, and the loss needs the label written as a vector too.
One-hot encoding the Good, Bad and Neutral labels
The board's dataset has three features and a three-class output. The first step is to one-hot encode the output: one column per class, a 1 in the column of the row's class and 0 everywhere else.
| f1 | f2 | f3 | Output | Good | Bad | Neutral |
|---|---|---|---|---|---|---|
| 2 | 3 | 4 | Good | 1 | 0 | 0 |
| 5 | 6 | 7 | Bad | 0 | 1 | 0 |
| 8 | 9 | 10 | Neutral | 0 | 0 | 1 |
Writing the categorical cross-entropy formula
- i is the row (the record) and j the column (the class); C is the number of categories, 3 here.
- yi = [yi1, yi2, …, yiC] is the one-hot row, with yij = 1 if the record is in class j and 0 otherwise.
- ŷij is the predicted probability of class j for record i. It comes from the softmax activation in the output layer, which divides ezj by the sum of ezk over all classes, so each row of ŷ adds up to 1.
Keeping only the true class term
Because the label is one-hot, every term with yij = 0 drops out of the sum. For a Good record, y = [1, 0, 0], and a softmax output ŷ = [0.7, 0.2, 0.1]:
So the loss is −ln of the probability the network gave to the right class. The other probabilities still matter through softmax: raising them lowers the true class's share.

Computing categorical cross-entropy in NumPy
Labels and softmax outputs
The three board rows, with logits chosen so the first row's softmax is about [0.7, 0.2, 0.1]:
import numpy as np
classes = ["Good", "Bad", "Neutral"]
labels = np.array([0, 1, 2]) # rows (2,3,4) Good, (5,6,7) Bad, (8,9,10) Neutral
one_hot = np.eye(3)[labels] # [1,0,0], [0,1,0], [0,0,1]
def softmax(z):
e = np.exp(z - z.max(axis=1, keepdims=True))
return e / e.sum(axis=1, keepdims=True)
logits = np.array([[2.0, 0.75, 0.05], [0.3, 1.6, -0.4], [0.5, 0.9, 1.1]])
y_hat = softmax(logits) # one row of probabilities per recordThe one-hot loss and the integer-label loss
def cce(one_hot, y_hat):
return -(one_hot * np.log(y_hat)).sum(axis=1) # one-hot labels
def sparse_cce(labels, y_hat):
return -np.log(y_hat[np.arange(len(labels)), labels]) # integer labelsprint("one-hot:\n", one_hot)
print("y_hat:\n", y_hat.round(3))
print("categorical, one-hot labels:", cce(one_hot, y_hat).round(3))
print("sparse, labels", labels, " :", sparse_cce(labels, y_hat).round(3))
print("cost:", round(cce(one_hot, y_hat).mean(), 3))
notes = np.array([[0.2, 0.3, 0.5]]) # the notes' softmax output
print("notes' row, true class 2:", cce(np.array([[0, 0, 1]]), notes).round(3), sparse_cce(np.array([2]), notes).round(3))one-hot: [[1. 0. 0.] [0. 1. 0.] [0. 0. 1.]] y_hat: [[0.7 0.201 0.1 ] [0.194 0.71 0.096] [0.232 0.346 0.422]] categorical, one-hot labels: [0.357 0.342 0.862] sparse, labels [0 1 2] : [0.357 0.342 0.862] cost: 0.52 notes' row, true class 2: [0.693] [0.693]
What the losses printed
- The first row gets ŷ = [0.7, 0.2, 0.1] and a loss of 0.357, the −ln 0.7 worked above.
- The other rows cost 0.342 (Bad gets 0.71) and 0.862 (Neutral gets only 0.422, the network is unsure), and the cost is their mean, 0.52.
- One-hot and integer labels give identical losses, row by row, and the notes' [0.2, 0.3, 0.5] with the true class at index 2 costs −ln 0.5 = 0.693 either way.
Using integer labels with sparse categorical cross-entropy
Sparse categorical cross-entropy is the same loss with the label written as a class index, 2, instead of the one-hot row [0, 0, 1]. It picks −ln ŷ at that index directly. With many classes (10,000 words in a vocabulary) the integer labels save the memory of a huge one-hot matrix.
The notes list a disadvantage, losing the probabilities of the other categories. That describes reading only the top index of a prediction; the sparse loss gives the same number as the one-hot loss, and the model still outputs the full softmax row.
Computing the loss from logits
A softmax written as ez / Σ ez overflows for large logits: e1000 is infinite in floating point, and infinity divided by infinity is nan. Taking the log of softmax in one step, after subtracting the row's largest logit, stays finite:
def cce_from_logits(one_hot, z):
log_p = z - z.max(axis=1, keepdims=True)
log_p = log_p - np.log(np.exp(log_p).sum(axis=1, keepdims=True)) # log softmax
return -(one_hot * log_p).sum(axis=1)big = np.array([[1000.0, 0.0, -1000.0]])
with np.errstate(all="ignore"):
naive = np.exp(big) / np.exp(big).sum()
print("naive softmax:", naive)
print("from logits, board rows:", cce_from_logits(one_hot, logits).round(3))
print("from logits, big row, true class Bad:", cce_from_logits(np.array([[0, 1, 0]]), big))naive softmax: [[nan 0. 0.]] from logits, board rows: [0.357 0.342 0.862] from logits, big row, true class Bad: [1000.]
The naive softmax prints [nan, 0., 0.]; the log-softmax form gives the board rows the same 0.357, 0.342 and 0.862, and the big row a loss of 1000 for a true class that the network scored 1000 below the top. This is what from_logits=True does in Keras.
Categorical cross-entropy in Keras
import keras
from keras import layers
model = keras.Sequential([keras.Input(shape=(3,)), # f1, f2, f3
layers.Dense(8, activation="relu"),
layers.Dense(3, activation="softmax")]) # Good, Bad, Neutral
model.compile(optimizer="adam", loss="categorical_crossentropy") # one-hot labels
# model.compile(optimizer="adam", loss="sparse_categorical_crossentropy") # labels 0, 1, 2
# logits: no softmax in the last layer, and the loss applies it
# layers.Dense(3) with loss=keras.losses.CategoricalCrossentropy(from_logits=True)Pick the loss by the label format: one-hot rows go with categorical_crossentropy, integer labels 0, 1, 2 with sparse_categorical_crossentropy. The model is the same.
Categorical vs sparse categorical cross-entropy
| Categorical cross-entropy | Sparse categorical cross-entropy | |
|---|---|---|
| Label for Neutral | [0, 0, 1] | 2 |
| Loss value | −ln ŷ of the true class | the same number |
| Model output | softmax over C classes | softmax over C classes |
| Keras name | "categorical_crossentropy" | "sparse_categorical_crossentropy" |
| Best for | labels already one-hot, or soft labels such as [0.9, 0.05, 0.05] | integer labels, many classes |
Where you use categorical cross-entropy
- Single-label multi-class problems: a review that is Good, Bad or Neutral, a digit from 0 to 9, a type of flower.
- Image classification: the softmax layer at the end of a convolutional network with one neuron per class.
- Language models: predicting the next word out of a vocabulary, with sparse labels because the vocabulary is large.
categorical_crossentropy, or one-hot rows with sparse_categorical_crossentropy, stop Keras with a shape error at the first batch; the fix is to change the loss name, not the model.Related
- Previous: Binary cross-entropy
- Next: Choosing a loss function
- See also: Softmax
- Reference: Probabilistic losses in the Keras documentation
- Change the first row's logits to
[4.0, 0.75, 0.05]: its probability for Good rises and its loss falls below 0.1. - Give the Neutral record the label Bad (
labels = np.array([0, 1, 1])): its loss rises from 0.862 to −ln 0.346 ≈ 1.062. - Use the notes' five-class output
[0.1, 0.2, 0.3, 0.2, 0.2]with true class 2: the loss is −ln 0.3 ≈ 1.204.
Slow is fine. Stopping is the only problem.