Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Choosing an activation function

Choosing an activation function is the decision of which activation each layer of a network uses, made separately for the hidden layers and for the output layer according to the problem type.

Last updated: 05 Oct, 2026 · NumPy

The lessons from Activation functions to Softmax covered each function on its own. The video closes the topic with a rule for putting them together, and it is a common interview question.

Which activation function to use · from the Deep Learning In-depth Tutorials in 5 Hours video · 119:50 to 124:04

Ruling out sigmoid and tanh in hidden layers

The video's first point: sigmoid and tanh both cause vanishing gradients in deep networks, so neither goes in the hidden layers. The hidden layers use ReLU, which the video calls the most efficient activation function. If a network with ReLU does not converge, change it to PReLU or ELU (or Leaky ReLU).

Picking activations for binary classification

Hidden layers: ReLU. Output layer: one neuron with sigmoid, whose value between 0 and 1 is the probability of class 1, with 0.5 as the threshold. Whatever the hidden layers use, the video stresses that the output of a binary classifier is sigmoid.

Picking activations for multi-class classification

Hidden layers: ReLU, or PReLU and the other variants if convergence is slow. Output layer: one neuron per class with softmax, so the outputs add up to 1.

The video says the multi-class output can use softmax or sigmoid. Softmax is the choice when each record has exactly one class. A sigmoid on every output neuron is for multi-label problems, where a record can belong to several classes at once.

Picking activations for regression

Hidden layers: ReLU or a variant. Output layer: one neuron with a linear activation, which leaves z as it is, so the output can be any real number such as a salary or a price. Each of the three problems also has its own loss function, the topic that follows.

Three small networks side by side. Binary classification: ReLU in the hidden layer, or PReLU or ELU, and one sigmoid output neuron. Multi-class classification: ReLU in the hidden layer and three softmax output neurons. Regression: ReLU in the hidden layer and one linear output neuron.

Running three output layers on one hidden layer

One record goes through the same ReLU hidden layer into three different output layers, so the only difference between the three results is the output activation. The weights are random, so the numbers have no meaning beyond their shape.

A ReLU hidden layer and an output head

python
import numpy as np
rng = np.random.default_rng(0)

x = np.array([0.5, -1.2, 2.0])                # one record, 3 features
W1, b1 = rng.normal(0, 0.5, (4, 3)), np.zeros(4)
hidden = np.maximum(0, W1 @ x + b1)           # hidden layer: ReLU

def head(n_out):
    W, b = rng.normal(0, 0.5, (n_out, 4)), np.zeros(n_out)
    return W @ hidden + b                     # the output layer's z values
ExampleRun on NumPy 2.5.3
z_bin = head(1)
p = 1 / (1 + np.exp(-z_bin))                  # binary: sigmoid
print("binary, sigmoid:     ", np.round(p, 4), "-> class", int(p[0] >= 0.5))

z_multi = head(3)
e = np.exp(z_multi - z_multi.max())
probs = e / e.sum()                           # multi-class: softmax
print("multi-class, softmax:", np.round(probs, 4), "sum", round(probs.sum(), 6))

z_reg = head(1)                               # regression: linear, z as it is
print("regression, linear:  ", np.round(z_reg, 4))
print("hidden layer (ReLU): ", np.round(hidden, 4))

What the three outputs show

  • The sigmoid output is one number between 0 and 1, turned into a class with the 0.5 threshold.
  • The softmax output is three numbers that add up to 1, one probability per class.
  • The linear output is the raw z, here a negative number, which neither sigmoid nor softmax could ever produce: it is a predicted value, not a probability.
  • The hidden layer has a 0: ReLU turned that neuron's negative sum into 0, as in the ReLU lesson.

Building the three networks in Keras

The same choices in Keras, where the activation is an argument of each Dense layer. The Keras version of a full network, its training and its output are in the ANN practical.

python
import keras
from keras import layers

binary = keras.Sequential([
    keras.Input(shape=(3,)),
    layers.Dense(8, activation="relu"),       # hidden layer: ReLU
    layers.Dense(1, activation="sigmoid"),    # one probability
])
multiclass = keras.Sequential([keras.Input(shape=(3,)), layers.Dense(8, activation="relu"),
                               layers.Dense(3, activation="softmax")])   # 3 classes
regression = keras.Sequential([keras.Input(shape=(3,)), layers.Dense(8, activation="relu"),
                               layers.Dense(1, activation="linear")])    # a number

Hidden layer vs output layer activations

ProblemHidden layersOutput layerOutput neuronsLoss (next part)
Binary classificationReLU (PReLU, ELU if it does not converge)Sigmoid1binary cross-entropy
Multi-class classificationReLU (or a variant)Softmaxone per classcategorical cross-entropy
Multi-label classificationReLU (or a variant)Sigmoid on each neuronone per labelbinary cross-entropy per label
RegressionReLU (or a variant)Linear1 per valueMSE, MAE or Huber

Where you use these choices

  • Customer churn (will a customer leave, yes or no): ReLU hidden layers and one sigmoid output.
  • Digit recognition (0 to 9): ReLU hidden layers and a 10-neuron softmax output.
  • House prices: ReLU hidden layers and one linear output.
Watch out. The output activation must match the loss. A softmax output with a regression loss, or a linear output with cross-entropy, trains without an error message and gives meaningless numbers. Pick the problem type first, then the output activation, then the loss.
Try it yourself
  • Change head(3) to head(5) for the multi-class output. Do the five probabilities still add up to 1?
  • Replace the ReLU in hidden with np.tanh. Which hidden values change sign?
  • Put a sigmoid on all three multi-class outputs instead of softmax. What do the three values add up to?
PreviousSoftmax

This is what real progress feels like.