Choosing a loss function
Choosing a loss function is the step that matches the loss to the problem type and to the output layer's activation: sigmoid with binary cross-entropy, softmax with categorical cross-entropy, and a linear output with MSE, MAE or Huber loss.
Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras
The hidden layers use ReLU or one of its variants whatever the problem, as Choosing an activation function showed. The output layer and the loss are what change, and they always change together. This closes the part with one table to remember.
The board heads the regression row "Linear Regression"; the rule holds for any regression network, not only a linear regression model.
Pairing the output activation with the loss
The video ends the session with these conclusions, and the notes extend them into a "right combination" table that also lists sparse categorical cross-entropy and RMSE:
| Problem | Hidden layers | Output layer | Loss function |
|---|---|---|---|
| Binary classification | ReLU or a variant | sigmoid, 1 neuron | binary cross-entropy |
| Multi-class classification | ReLU or a variant | softmax, 1 neuron per class | categorical or sparse categorical cross-entropy |
| Regression | ReLU or a variant | linear, 1 neuron | MSE, MAE, Huber loss (RMSE as a metric) |

These pairings are the interview question the video closes on: given a problem, name the output activation and the loss.
Adding multi-label problems
Sometimes a record can belong to several classes at once: a photo with a cat and a dog, a film tagged comedy and drama. Softmax forces the classes to compete for one total of 1, so it is the wrong output here. Use one sigmoid per class, each an independent yes or no, and a binary cross-entropy per class, averaged.
Using hinge loss for labels of −1 and +1
The materials' notebook also lists hinge loss, the loss of support vector machines. The label t is −1 or +1 and the output y is a raw score:
A score on the right side by a margin of at least 1 costs 0; inside the margin or on the wrong side the cost grows in a straight line. It is rarely used in neural networks, where binary cross-entropy gives probabilities and smoother slopes, but Keras has it as "hinge".
Checking each pairing in NumPy
The same three raw outputs z of a last layer, passed through each problem's activation and loss:
import numpy as np
z = np.array([1.5, -0.5, 0.2]) # raw outputs of a last layer
sigmoid = lambda v: 1 / (1 + np.exp(-v))
softmax = lambda v: np.exp(v) / np.exp(v).sum()
# regression: linear output (ŷ = z) with MSE against y = 1.2
print("regression :", round((1.2 - z[0]) ** 2, 4))
# binary: sigmoid on one neuron with binary cross-entropy, y = 1
print("binary :", round(-np.log(sigmoid(z[0])), 4))
# multi-class: softmax on all three with categorical cross-entropy, true class 0
print("multi-class:", round(-np.log(softmax(z)[0]), 4), " probabilities", softmax(z).round(3))
# multi-label: a sigmoid per neuron, a binary cross-entropy per label, y = [1, 0, 1]
y = np.array([1, 0, 1]); p = sigmoid(z)
print("multi-label:", round(np.mean(-(y * np.log(p) + (1 - y) * np.log(1 - p))), 4), " probabilities", p.round(3))
# two classes: softmax over [z, 0] equals sigmoid(z)
print("softmax([1.5, 0])[0] =", round(softmax(np.array([1.5, 0.0]))[0], 4), " sigmoid(1.5) =", round(sigmoid(1.5), 4))regression : 0.09 binary : 0.2014 multi-class: 0.3421 probabilities [0.71 0.096 0.194] multi-label: 0.4245 probabilities [0.818 0.378 0.55 ] softmax([1.5, 0])[0] = 0.8176 sigmoid(1.5) = 0.8176
Hinge loss for a true label of +1
The notebook's hinge cell uses outputs from −3 to 5 and a true label of 1. In NumPy, next to binary cross-entropy from the logit for the same label:
import matplotlib.pyplot as plt
y_pred = np.linspace(-3, 5, 500) # the notebook's range
t = 1 # true label, +1 or -1
hinge = np.maximum(0, 1 - t * y_pred)
bce_curve = np.log1p(np.exp(-y_pred)) # binary cross-entropy from the logit, y = 1
plt.plot(y_pred, hinge, "--", label="hinge")
plt.plot(y_pred, bce_curve, label="binary cross-entropy")
plt.xlabel("raw output"); plt.ylabel("loss"); plt.title("Hinge loss for a true label of +1")
plt.legend(); plt.show()
for v in (-1.0, 0.0, 0.5, 1.0, 2.0):
print(f"output {v:4}: hinge {max(0.0, 1 - t * v):.2f}")output -1.0: hinge 2.00 output 0.0: hinge 1.00 output 0.5: hinge 0.50 output 1.0: hinge 0.00 output 2.0: hinge 0.00

What the pairings printed
- Regression: the linear output 1.5 against a target of 1.2 costs 0.3² = 0.09.
- Binary: σ(1.5) = 0.818 for a true label of 1 costs −ln 0.818 = 0.2014.
- Multi-class: softmax gives [0.71, 0.096, 0.194], adding up to 1, and the true class 0 costs −ln 0.71 = 0.3421.
- Multi-label: the sigmoids [0.818, 0.378, 0.55] are independent and do not add up to 1; the mean of the three binary cross-entropies is 0.4245.
- Two classes: softmax over [1.5, 0] and the sigmoid of 1.5 both give 0.8176, so a softmax with two outputs and categorical cross-entropy is the same model as one sigmoid with binary cross-entropy.
- Hinge prints 2.00, 1.00, 0.50, 0.00 and 0.00: no cost once the output reaches the margin of 1, while binary cross-entropy is still above 0 there.
Writing the pairings in Keras
from keras import layers
# binary: layers.Dense(1, activation="sigmoid") loss="binary_crossentropy"
# multi-class: layers.Dense(3, activation="softmax") loss="categorical_crossentropy"
# (integer labels: loss="sparse_categorical_crossentropy")
# multi-label: layers.Dense(3, activation="sigmoid") loss="binary_crossentropy"
# regression: layers.Dense(1) (linear) loss="mse", "mae" or "huber"
# hinge: layers.Dense(1, activation="tanh") loss="hinge", labels -1 / +1Sigmoid with binary cross-entropy vs softmax with categorical cross-entropy
| Sigmoid + binary cross-entropy | Softmax + categorical cross-entropy | |
|---|---|---|
| Output neurons | 1 (or 1 per label) | 1 per class |
| Outputs add up to 1 | no, each is its own probability | yes |
| Classes per record | one of two, or several labels | exactly one |
| Two classes | the usual choice | the same model, written with two outputs |
Where you use the combination table
- Starting a model: read the problem off the target column (a number, two classes, several classes, several labels), and the last layer and the loss follow.
- Debugging a model that does not learn: a mismatched pair, such as softmax with binary cross-entropy or a linear output with categorical cross-entropy, is a common cause.
- The churn network in Building an ANN in Keras uses the binary row: ReLU hidden layers, one sigmoid output, binary cross-entropy.
Related
- Previous: Categorical cross-entropy
- Next: Gradient descent
- Reference: Losses in the Keras documentation
- Change the regression target from 1.2 to 1.5: the MSE prints 0.0, a perfect prediction.
- Compute
softmax(np.array([0.7]))for a single neuron: it prints[1.]whatever the input, the mistake in the Watch out box. - Set
t = -1in the hinge example: the cost is now 0 for outputs of −1 and below and grows for outputs above −1.
This is what real progress feels like.