Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Activation functions

An activation function is a function applied to a neuron's weighted sum z that sets the neuron's output and adds the non-linearity a network needs to learn curved patterns.

Last updated: 05 Oct, 2026 · NumPy

The Vanishing gradient problem lesson ended with the video's fix: use another activation function. This lesson covers why a neuron needs one at all, then the first two the video compares, sigmoid and tanh, each with its derivative, the part backpropagation uses.

Seeing why a network needs an activation

Without an activation, every layer computes z = Wx + b, a linear function. A linear function of a linear function is still linear, so a stack of such layers equals one layer, however deep it is. The run builds two layers with no activation between them and the single layer they collapse into:

ExampleRun on NumPy 2.5.3
import numpy as np
rng = np.random.default_rng(0)

x = rng.normal(0, 1, 3)                       # one record, 3 inputs
W1, b1 = rng.normal(0, 1, (4, 3)), rng.normal(0, 1, 4)
W2, b2 = rng.normal(0, 1, (1, 4)), rng.normal(0, 1, 1)

two_layers = W2 @ (W1 @ x + b1) + b2          # no activation in between
W, b = W2 @ W1, W2 @ b1 + b2                  # one layer with merged weights
print("two linear layers:", np.round(two_layers, 6))
print("one linear layer: ", np.round(W @ x + b, 6))

Both lines print the same numbers. With a non-linear activation between the layers that merge is impossible, which is what lets a network bend its decision boundary; this is the "solve non-linear problems" point the video makes for the sigmoid in each neuron.

Sigmoid and zero-centred outputs · from the Deep Learning In-depth Tutorials in 5 Hours video · 105:13 to 109:06

Plotting the sigmoid and its derivative

The video's notebook shows each activation as two plots: on the left the function, used in the forward pass, and on the right its derivative, used in backpropagation. For sigmoid:

ExampleRun on NumPy 2.5.3 and matplotlib 3.11.2
import numpy as np
import matplotlib.pyplot as plt

z = np.linspace(-10, 10, 401)
s = 1 / (1 + np.exp(-z))                      # sigmoid
ds = s * (1 - s)                              # its derivative

fig, (a, b) = plt.subplots(1, 2, figsize=(9, 3.4))
a.plot(z, s)
a.axhline(0.5, ls="--", c="grey")
a.set_title("sigmoid σ(z)")
a.set_xlabel("z")
b.plot(z, ds, c="tab:orange")
b.set_title("derivative σ′(z)")
b.set_xlabel("z")
plt.tight_layout()
plt.show()
print("σ(0) =", s[200], "  largest σ′ =", ds.max(), "at z =", z[ds.argmax()])
print("σ′ at z = 5:", round(ds[300], 4), "  at z = 10:", round(ds[400], 6))
Left: the sigmoid S-curve rising from 0 to 1 and crossing 0.5 at z = 0. Right: its derivative, a bell-shaped bump that peaks at 0.25 at z = 0 and is nearly 0 beyond z = ±5.
  • σ(0) = 0.5 and the largest σ′ is 0.25 at z = 0, the ceiling behind the vanishing gradient.
  • σ′ is 0.0066 at z = 5 and 0.000045 at z = 10: once the input is a little away from the origin, the gradient is almost 0. A neuron there barely learns.

The notebook's list for sigmoid: its advantages are a smooth gradient, an output between 0 and 1 for every neuron, and clear predictions close to 1 or 0. Its three disadvantages are that it is prone to vanishing gradients, its output is not zero centred, and the exponential is slow to compute (the 1/(1 + e^(−z)) formula takes more time than a comparison or a multiplication).

Understanding zero-centred outputs

The video draws a curve that passes through the origin, with negative values on one side and positive on the other: a zero-centred curve. Sigmoid never goes below 0, so it is not zero centred. The video's point is that a zero-centred curve makes the weight update more efficient. The notes draw the same idea with data: a cloud of points far from the origin, moved onto it by standardization.

Left: the sigmoid curve stays between 0 and 1 while the tanh curve runs from −1 to 1 through the origin. Right: a cloud of points far from the origin and the same cloud after subtracting its mean, centred on the origin.

Here is why. A neuron's weight gradient is ∂L/∂wᵢ = δ · xᵢ, where δ = ∂L/∂z is one number for the whole neuron and xᵢ are its inputs, the previous layer's outputs. If every input is positive, as sigmoid outputs are, every weight gradient of that neuron has the sign of δ:

ExampleRun on NumPy 2.5.3
import numpy as np

delta = -0.3                                  # dL/dz of one neuron in the next layer
inputs_sigmoid = np.array([0.2, 0.7, 0.9])    # sigmoid outputs: all positive
inputs_tanh = np.array([-0.6, 0.4, 0.7])      # tanh outputs: both signs
print("weight gradients, sigmoid inputs:", delta * inputs_sigmoid)
print("weight gradients, tanh inputs:   ", delta * inputs_tanh)

With sigmoid inputs all three gradients are negative, so all three weights must rise together on this step and fall together on another. Reaching a point where one weight should go up and another down then takes a zig-zag of steps. With zero-centred inputs the signs can differ, so each weight moves its own way.

tanh and its derivative · from the Deep Learning In-depth Tutorials in 5 Hours video · 109:53 to 112:10

Comparing tanh with sigmoid

The second activation is tanh, the hyperbolic tangent. Its output runs from −1 to 1 and its derivative from 0 to 1, against 0 to 0.25 for sigmoid:

The video's question: does tanh prevent the vanishing gradient problem? Not in a very deep network. Its derivative is 1 only at z = 0 and falls toward 0 in both tails, so a long chain of tanh derivatives can still be tiny. What tanh does fix is the centring: it passes through the origin, so it is zero centred. In a binary classification problem the notebook's advice is tanh in the hidden layers and sigmoid in the output layer.

Reading the notebook aloud, the video says the two curves are "relatively smaller" and that a small gradient "is conducive to weight update". The notebook says the curves are relatively similar, and that a small gradient for large or small inputs is not conducive to weight updates.

ExampleRun on NumPy 2.5.3 and matplotlib 3.11.2
import numpy as np
import matplotlib.pyplot as plt

z = np.linspace(-10, 10, 401)
t = np.tanh(z)
dt = 1 - t ** 2                               # tanh derivative
s = 1 / (1 + np.exp(-z))
ds = s * (1 - s)                              # sigmoid derivative

fig, (a, b) = plt.subplots(1, 2, figsize=(9, 3.4))
a.plot(z, t)
a.set_title("tanh(z)")
a.set_xlabel("z")
b.plot(z, dt, label="tanh′(z)")
b.plot(z, ds, label="σ′(z)", c="tab:orange")
b.set_title("derivatives")
b.set_xlabel("z")
b.legend()
plt.tight_layout()
plt.show()
for zv in [0, 1, 2, 4]:
    i = 200 + zv * 20
    print(f"z = {zv}: tanh′ = {dt[i]:.4f}   σ′ = {ds[i]:.4f}")
Left: the tanh curve from −1 to 1 through the origin. Right: the tanh derivative peaking at 1 at z = 0 next to the much lower sigmoid derivative peaking at 0.25; both fall to near 0 by z = ±5.

What the two derivatives show

  • At z = 0, tanh′ is 1 and σ′ is 0.25: four times more gradient passes through a tanh neuron near the origin.
  • At z = 2, tanh′ is 0.0707 and σ′ is 0.105: tanh's derivative falls faster, and away from the origin it is not the larger one.
  • At z = 4 both are close to 0 (0.0013 and 0.0177): saturated neurons pass almost no gradient with either function, so deep tanh networks can still vanish.

Sigmoid vs tanh

Sigmoidtanh
Formula1 / (1 + e^(−z))(e^z − e^(−z)) / (e^z + e^(−z))
Output range0 to 1−1 to 1
Zero centrednoyes
Derivative range0 to 0.250 to 1
Vanishing gradientyesyes, in very deep networks
Costone exponentialexponentials (slower)
Typical placeoutput layer, binary classificationhidden layers of small networks

Where you use sigmoid and tanh

  • Binary classification output: one sigmoid neuron gives a value between 0 and 1, and 0.5 is the usual threshold; in Keras, Dense(1, activation="sigmoid").
  • Small networks: tanh in the hidden layers of a shallow network trains better than sigmoid because its outputs are zero centred.
  • Gates inside recurrent cells: LSTM and GRU cells use sigmoid to squash a gate into 0 to 1 and tanh for the values that pass through it.
Watch out. Neither sigmoid nor tanh belongs in the hidden layers of a deep network. Both saturate, both derivatives fall to near 0 a few units from the origin, and the chain rule multiplies those small numbers layer after layer.
Try it yourself
  • In the sigmoid run, print ds[220], the derivative at z = 1. Is it closer to 0.25 or to 0?
  • Check that tanh is a stretched sigmoid: compare np.tanh(1.0) with 2 / (1 + np.exp(-2.0)) - 1.
  • In the two-layer run, put np.tanh around W1 @ x + b1. Do the two lines still match?

Little by little, you're building something great.