Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Choosing a loss function

Choosing a loss function is the step that matches the loss to the problem type and to the output layer's activation: sigmoid with binary cross-entropy, softmax with categorical cross-entropy, and a linear output with MSE, MAE or Huber loss.

Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras

The hidden layers use ReLU or one of its variants whatever the problem, as Choosing an activation function showed. The output layer and the loss are what change, and they always change together. This closes the part with one table to remember.

Activation and loss for each problem · from the Deep Learning In-depth Tutorials in 5 Hours video · 151:00 to 153:35

The board heads the regression row "Linear Regression"; the rule holds for any regression network, not only a linear regression model.

Pairing the output activation with the loss

The video ends the session with these conclusions, and the notes extend them into a "right combination" table that also lists sparse categorical cross-entropy and RMSE:

ProblemHidden layersOutput layerLoss function
Binary classificationReLU or a variantsigmoid, 1 neuronbinary cross-entropy
Multi-class classificationReLU or a variantsoftmax, 1 neuron per classcategorical or sparse categorical cross-entropy
RegressionReLU or a variantlinear, 1 neuronMSE, MAE, Huber loss (RMSE as a metric)
Three small networks with ReLU hidden layers: binary classification ends in one sigmoid neuron and binary cross-entropy, multi-class classification ends in softmax neurons and categorical cross-entropy, and regression ends in one linear neuron and MSE, MAE or Huber loss.

These pairings are the interview question the video closes on: given a problem, name the output activation and the loss.

Adding multi-label problems

Sometimes a record can belong to several classes at once: a photo with a cat and a dog, a film tagged comedy and drama. Softmax forces the classes to compete for one total of 1, so it is the wrong output here. Use one sigmoid per class, each an independent yes or no, and a binary cross-entropy per class, averaged.

Using hinge loss for labels of −1 and +1

The materials' notebook also lists hinge loss, the loss of support vector machines. The label t is −1 or +1 and the output y is a raw score:

A score on the right side by a margin of at least 1 costs 0; inside the margin or on the wrong side the cost grows in a straight line. It is rarely used in neural networks, where binary cross-entropy gives probabilities and smoother slopes, but Keras has it as "hinge".

Checking each pairing in NumPy

The same three raw outputs z of a last layer, passed through each problem's activation and loss:

ExampleRun on NumPy 2.5.3
import numpy as np

z = np.array([1.5, -0.5, 0.2])                  # raw outputs of a last layer
sigmoid = lambda v: 1 / (1 + np.exp(-v))
softmax = lambda v: np.exp(v) / np.exp(v).sum()

# regression: linear output (ŷ = z) with MSE against y = 1.2
print("regression :", round((1.2 - z[0]) ** 2, 4))
# binary: sigmoid on one neuron with binary cross-entropy, y = 1
print("binary     :", round(-np.log(sigmoid(z[0])), 4))
# multi-class: softmax on all three with categorical cross-entropy, true class 0
print("multi-class:", round(-np.log(softmax(z)[0]), 4), " probabilities", softmax(z).round(3))
# multi-label: a sigmoid per neuron, a binary cross-entropy per label, y = [1, 0, 1]
y = np.array([1, 0, 1]); p = sigmoid(z)
print("multi-label:", round(np.mean(-(y * np.log(p) + (1 - y) * np.log(1 - p))), 4), " probabilities", p.round(3))
# two classes: softmax over [z, 0] equals sigmoid(z)
print("softmax([1.5, 0])[0] =", round(softmax(np.array([1.5, 0.0]))[0], 4), " sigmoid(1.5) =", round(sigmoid(1.5), 4))

Hinge loss for a true label of +1

The notebook's hinge cell uses outputs from −3 to 5 and a true label of 1. In NumPy, next to binary cross-entropy from the logit for the same label:

ExampleRun on NumPy 2.5.3
import matplotlib.pyplot as plt

y_pred = np.linspace(-3, 5, 500)            # the notebook's range
t = 1                                       # true label, +1 or -1
hinge = np.maximum(0, 1 - t * y_pred)
bce_curve = np.log1p(np.exp(-y_pred))       # binary cross-entropy from the logit, y = 1
plt.plot(y_pred, hinge, "--", label="hinge")
plt.plot(y_pred, bce_curve, label="binary cross-entropy")
plt.xlabel("raw output"); plt.ylabel("loss"); plt.title("Hinge loss for a true label of +1")
plt.legend(); plt.show()
for v in (-1.0, 0.0, 0.5, 1.0, 2.0):
    print(f"output {v:4}: hinge {max(0.0, 1 - t * v):.2f}")
Hinge loss and binary cross-entropy for a true label of plus 1 against the raw output from minus 3 to 5: the dashed hinge line falls straight to 0 at an output of 1 and stays 0, while binary cross-entropy is a smooth curve that keeps shrinking towards 0.

What the pairings printed

  • Regression: the linear output 1.5 against a target of 1.2 costs 0.3² = 0.09.
  • Binary: σ(1.5) = 0.818 for a true label of 1 costs −ln 0.818 = 0.2014.
  • Multi-class: softmax gives [0.71, 0.096, 0.194], adding up to 1, and the true class 0 costs −ln 0.71 = 0.3421.
  • Multi-label: the sigmoids [0.818, 0.378, 0.55] are independent and do not add up to 1; the mean of the three binary cross-entropies is 0.4245.
  • Two classes: softmax over [1.5, 0] and the sigmoid of 1.5 both give 0.8176, so a softmax with two outputs and categorical cross-entropy is the same model as one sigmoid with binary cross-entropy.
  • Hinge prints 2.00, 1.00, 0.50, 0.00 and 0.00: no cost once the output reaches the margin of 1, while binary cross-entropy is still above 0 there.

Writing the pairings in Keras

python
from keras import layers

# binary:      layers.Dense(1, activation="sigmoid")   loss="binary_crossentropy"
# multi-class: layers.Dense(3, activation="softmax")   loss="categorical_crossentropy"
#              (integer labels: loss="sparse_categorical_crossentropy")
# multi-label: layers.Dense(3, activation="sigmoid")   loss="binary_crossentropy"
# regression:  layers.Dense(1)  (linear)               loss="mse", "mae" or "huber"
# hinge:       layers.Dense(1, activation="tanh")      loss="hinge", labels -1 / +1

Sigmoid with binary cross-entropy vs softmax with categorical cross-entropy

Sigmoid + binary cross-entropySoftmax + categorical cross-entropy
Output neurons1 (or 1 per label)1 per class
Outputs add up to 1no, each is its own probabilityyes
Classes per recordone of two, or several labelsexactly one
Two classesthe usual choicethe same model, written with two outputs

Where you use the combination table

  • Starting a model: read the problem off the target column (a number, two classes, several classes, several labels), and the last layer and the loss follow.
  • Debugging a model that does not learn: a mismatched pair, such as softmax with binary cross-entropy or a linear output with categorical cross-entropy, is a common cause.
  • The churn network in Building an ANN in Keras uses the binary row: ReLU hidden layers, one sigmoid output, binary cross-entropy.
Watch out. Keras does not check that the output activation fits the loss. A softmax over a single neuron always outputs 1.0, so with binary cross-entropy the loss never moves and the model predicts one class for every record, without an error.
Try it yourself
  • Change the regression target from 1.2 to 1.5: the MSE prints 0.0, a perfect prediction.
  • Compute softmax(np.array([0.7])) for a single neuron: it prints [1.] whatever the input, the mistake in the Watch out box.
  • Set t = -1 in the hinge example: the cost is now 0 for outputs of −1 and below and grows for outputs above −1.

This is what real progress feels like.