Choosing an activation function
Choosing an activation function is the decision of which activation each layer of a network uses, made separately for the hidden layers and for the output layer according to the problem type.
Last updated: 05 Oct, 2026 · NumPy
The lessons from Activation functions to Softmax covered each function on its own. The video closes the topic with a rule for putting them together, and it is a common interview question.
Ruling out sigmoid and tanh in hidden layers
The video's first point: sigmoid and tanh both cause vanishing gradients in deep networks, so neither goes in the hidden layers. The hidden layers use ReLU, which the video calls the most efficient activation function. If a network with ReLU does not converge, change it to PReLU or ELU (or Leaky ReLU).
Picking activations for binary classification
Hidden layers: ReLU. Output layer: one neuron with sigmoid, whose value between 0 and 1 is the probability of class 1, with 0.5 as the threshold. Whatever the hidden layers use, the video stresses that the output of a binary classifier is sigmoid.
Picking activations for multi-class classification
Hidden layers: ReLU, or PReLU and the other variants if convergence is slow. Output layer: one neuron per class with softmax, so the outputs add up to 1.
The video says the multi-class output can use softmax or sigmoid. Softmax is the choice when each record has exactly one class. A sigmoid on every output neuron is for multi-label problems, where a record can belong to several classes at once.
Picking activations for regression
Hidden layers: ReLU or a variant. Output layer: one neuron with a linear activation, which leaves z as it is, so the output can be any real number such as a salary or a price. Each of the three problems also has its own loss function, the topic that follows.

Running three output layers on one hidden layer
One record goes through the same ReLU hidden layer into three different output layers, so the only difference between the three results is the output activation. The weights are random, so the numbers have no meaning beyond their shape.
A ReLU hidden layer and an output head
import numpy as np
rng = np.random.default_rng(0)
x = np.array([0.5, -1.2, 2.0]) # one record, 3 features
W1, b1 = rng.normal(0, 0.5, (4, 3)), np.zeros(4)
hidden = np.maximum(0, W1 @ x + b1) # hidden layer: ReLU
def head(n_out):
W, b = rng.normal(0, 0.5, (n_out, 4)), np.zeros(n_out)
return W @ hidden + b # the output layer's z valuesz_bin = head(1)
p = 1 / (1 + np.exp(-z_bin)) # binary: sigmoid
print("binary, sigmoid: ", np.round(p, 4), "-> class", int(p[0] >= 0.5))
z_multi = head(3)
e = np.exp(z_multi - z_multi.max())
probs = e / e.sum() # multi-class: softmax
print("multi-class, softmax:", np.round(probs, 4), "sum", round(probs.sum(), 6))
z_reg = head(1) # regression: linear, z as it is
print("regression, linear: ", np.round(z_reg, 4))
print("hidden layer (ReLU): ", np.round(hidden, 4))binary, sigmoid: [0.2715] -> class 0 multi-class, softmax: [0.2058 0.4222 0.372 ] sum 1.0 regression, linear: [-0.1042] hidden layer (ReLU): [0.7511 0.7092 0. 0.0989]
What the three outputs show
- The sigmoid output is one number between 0 and 1, turned into a class with the 0.5 threshold.
- The softmax output is three numbers that add up to 1, one probability per class.
- The linear output is the raw z, here a negative number, which neither sigmoid nor softmax could ever produce: it is a predicted value, not a probability.
- The hidden layer has a 0: ReLU turned that neuron's negative sum into 0, as in the ReLU lesson.
Building the three networks in Keras
The same choices in Keras, where the activation is an argument of each Dense layer. The Keras version of a full network, its training and its output are in the ANN practical.
import keras
from keras import layers
binary = keras.Sequential([
keras.Input(shape=(3,)),
layers.Dense(8, activation="relu"), # hidden layer: ReLU
layers.Dense(1, activation="sigmoid"), # one probability
])
multiclass = keras.Sequential([keras.Input(shape=(3,)), layers.Dense(8, activation="relu"),
layers.Dense(3, activation="softmax")]) # 3 classes
regression = keras.Sequential([keras.Input(shape=(3,)), layers.Dense(8, activation="relu"),
layers.Dense(1, activation="linear")]) # a numberHidden layer vs output layer activations
| Problem | Hidden layers | Output layer | Output neurons | Loss (next part) |
|---|---|---|---|---|
| Binary classification | ReLU (PReLU, ELU if it does not converge) | Sigmoid | 1 | binary cross-entropy |
| Multi-class classification | ReLU (or a variant) | Softmax | one per class | categorical cross-entropy |
| Multi-label classification | ReLU (or a variant) | Sigmoid on each neuron | one per label | binary cross-entropy per label |
| Regression | ReLU (or a variant) | Linear | 1 per value | MSE, MAE or Huber |
Where you use these choices
- Customer churn (will a customer leave, yes or no): ReLU hidden layers and one sigmoid output.
- Digit recognition (0 to 9): ReLU hidden layers and a 10-neuron softmax output.
- House prices: ReLU hidden layers and one linear output.
Related
- Previous: Softmax
- Next: Loss and cost functions
- Change
head(3)tohead(5)for the multi-class output. Do the five probabilities still add up to 1? - Replace the ReLU in
hiddenwithnp.tanh. Which hidden values change sign? - Put a sigmoid on all three multi-class outputs instead of softmax. What do the three values add up to?
This is what real progress feels like.