Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

ReLU and its variants

ReLU (rectified linear unit) is an activation function that outputs max(0, z): the input itself when it is positive and 0 when it is negative.

Last updated: 05 Oct, 2026 · NumPy

The Activation functions lesson showed that sigmoid and tanh both saturate, so a deep network's gradient still vanishes. ReLU is the next answer, and today it is the most popular activation function in hidden layers.

ReLU and the dead neuron · from the Deep Learning In-depth Tutorials in 5 Hours video · 112:43 to 114:28

Computing ReLU and its derivative

A negative input becomes 0 and a positive one passes through unchanged: the video's examples are 1 → 1, 2 → 2 and 3 → 3. The derivative is either 0 or 1. At exactly z = 0 the derivative is not defined; code takes 0 or 1 there (the notebook calls this a sub-gradient), and it rarely matters because z is almost never exactly 0.

Deriving the dead neuron

A derivative of 1 is good news: the gradient passes through without shrinking, so ReLU does not cause vanishing gradients for positive inputs. A derivative of 0 is the problem. The notes write the chain for the network x1 → O1 → O2 → O3 → loss, where O3 = ReLU(z3) with z3 = O2·w3 + b3:

If z3 is negative, ReLU′(z3) = 0, so the whole chain is 0 and w2_new = w2_old − η · 0 = w2_old. The weights behind that neuron get no update. If z3 is negative for every record, the neuron outputs 0 for all of them and never recovers: a dead neuron, also called the dead ReLU problem. The video says the neuron dies when "one of the weights" becomes zero; it is the derivative that is 0, because the neuron's z is negative.

A chain x1 to O1 to O2 to O3 to the loss, with O3 greyed out as a dead neuron: z3 is negative, ReLU outputs 0 and its derivative is 0, so the gradient stops at O3 and w2_new equals w2_old.
ReLU advantages and Leaky ReLU · from the Deep Learning In-depth Tutorials in 5 Hours video · 114:28 to 117:34

Weighing ReLU's advantages

  • No exponential: max(0, z) is a comparison, much quicker than sigmoid or tanh.
  • No saturation for positive inputs: the derivative is 1, so gradients do not vanish there.
  • Not zero centred: it passes through 0 but never outputs a negative value. The video says ReLU also solves tanh's problem and then, a moment later, that it is not zero centric; the second is right.
  • Dead neurons: a negative input gives a derivative of 0, the disadvantage above.

Fixing dead neurons with Leaky ReLU

Leaky ReLU keeps a small slope α for negative inputs instead of 0, so the output for z < 0 is a small negative number and the derivative is α, never 0:

The video uses α = 0.01, f(z) = max(0.01z, z). The notes treat α as a hyperparameter you pick, with 0.01, 0.02 and 0.03 as values to try.

The run below makes the dead neuron happen. One neuron starts with bias −3, so z < 0 for every record. It trains for 500 steps on a target it could fit, once with ReLU (α = 0) and once with Leaky ReLU (α = 0.01):

ExampleRun on NumPy 2.5.3
import numpy as np
rng = np.random.default_rng(1)
X = rng.uniform(0, 1, (50, 2))                # 50 records, 2 inputs between 0 and 1
y = X @ np.array([2.0, -1.0]) + 0.5           # the target the neuron should learn

def train(alpha, steps=500, eta=0.5):
    w, b = np.array([0.5, 0.5]), -3.0         # bias -3: z < 0 for every record
    for _ in range(steps):
        z = X @ w + b
        out = np.where(z > 0, z, alpha * z)   # ReLU when alpha = 0
        g = (out - y) * np.where(z > 0, 1.0, alpha)   # dL/dz, through the derivative
        w = w - eta * (X.T @ g) / len(y)
        b = b - eta * g.mean()
    return w, b

for alpha in [0.0, 0.01]:
    w, b = train(alpha)
    print(f"alpha = {alpha}: w = {np.round(w, 3)}, b = {b:.3f}")
  • With ReLU, w and b never move: w stays [0.5, 0.5] and b stays −3.0 after 500 steps. Every gradient is 0, so this neuron is dead.
  • With Leaky ReLU they learn: the 0.01 slope lets a small gradient through, b climbs past 0, and the weights reach about [2.003, −1.0] with b = 0.498, close to the target's 2, −1 and 0.5.
A learned slope, ELU and PReLU · from the Deep Learning In-depth Tutorials in 5 Hours video · 117:34 to 119:21

Weighing Leaky ReLU against ReLU

The notebook then makes two points about Leaky ReLU. The slope does not have to be 0.01: a parameter-based method lets it take other values, or learn it, which is where PReLU below comes from. And a caution: in theory Leaky ReLU has all of ReLU's advantages and no dead neurons, but in practice it has not been fully proved that Leaky ReLU is always better than ReLU.

Smoothing the corner with ELU

ELU (exponential linear unit) replaces the negative side with an exponential curve that flattens out at −α:

It has no dead ReLU issue, its negative outputs pull the mean output close to 0 (close to zero centred), and its cost is the exponential, so it is slightly more expensive to compute. With α = 1, the plotted case, the derivative runs from 0 to 1.

Learning the slope with PReLU

Parametric ReLU (PReLU) is Leaky ReLU's shape with the slope for negative inputs turned into a parameter the network learns in backpropagation, one aᵢ per neuron:

  • aᵢ = 0 gives ReLU.
  • aᵢ a small fixed number (such as 0.01) gives Leaky ReLU.
  • aᵢ learnable gives PReLU.

The video reads the condition as "w of i greater than zero"; the variable in the condition is the input zᵢ, and aᵢ is the slope.

Using Swish, the self-gated function

Swish multiplies the input by its own sigmoid, so the sigmoid acts as a gate on the value:

For large positive z it behaves like ReLU, and for negative z it dips a little below 0 before returning to 0. The video says Swish has one problem, that its derivative at zero cannot be found. That is the case for ReLU, not Swish: Swish is smooth everywhere, and its derivative at 0 is σ(0) = 0.5.

Computing the ReLU family in NumPy

ReLU and Leaky ReLU

python
import numpy as np

def sigmoid(z):
    return 1 / (1 + np.exp(-z))

def relu(z):
    return np.maximum(0, z)

def leaky_relu(z, alpha=0.01):
    return np.where(z > 0, z, alpha * z)

ELU and Swish

python
def elu(z, alpha=1.0):
    return np.where(z > 0, z, alpha * (np.exp(z) - 1))

def swish(z):
    return z * sigmoid(z)                     # x times sigmoid(x), smooth everywhere

Their derivatives

python
def d_relu(z):
    return np.where(z > 0, 1.0, 0.0)          # 0 for z < 0, 1 for z > 0

def d_leaky(z, alpha=0.01):
    return np.where(z > 0, 1.0, alpha)

def d_elu(z, alpha=1.0):
    return np.where(z > 0, 1.0, alpha * np.exp(z))

def d_swish(z):
    s = sigmoid(z)
    return s + z * s * (1 - s)
ExampleRun on NumPy 2.5.3
z = np.array([-2.0, -0.5, 0.0, 0.5, 2.0])
print("z          ", z)
print("relu       ", relu(z), "  derivative", d_relu(z))
print("leaky 0.01 ", leaky_relu(z), "  derivative", d_leaky(z))
print("elu        ", np.round(elu(z), 4), "  derivative", np.round(d_elu(z), 4))
print("swish      ", np.round(swish(z), 4), "  derivative", np.round(d_swish(z), 4))
print("relu on 1, 2, 3:", relu(np.array([1, 2, 3])))

What the values show

  • ReLU turns −2 and −0.5 into 0 and keeps 0.5 and 2; its derivative is 0 or 1 (0 at z = 0 in this code).
  • Leaky ReLU keeps −0.02 and −0.005, and its derivative on the negative side is 0.01, never 0.
  • ELU gives −0.8647 at z = −2, heading for −1, with derivative 0.1353 = e^(−2).
  • Swish gives −0.2384 at z = −2, and its derivative at 0 is 0.5: defined and smooth.

Plotting the ReLU family and its derivatives

The same four functions on one figure, the function on top and its derivative below. Leaky ReLU is drawn with α = 0.1 because a 0.01 slope looks flat at this scale.

ExampleRun on NumPy 2.5.3 and matplotlib 3.11.2
import matplotlib.pyplot as plt

z = np.linspace(-5, 5, 501)
rows = [("ReLU", relu, d_relu), ("Leaky ReLU, α = 0.1", lambda v: leaky_relu(v, 0.1), lambda v: d_leaky(v, 0.1)),
        ("ELU, α = 1", elu, d_elu), ("Swish", swish, d_swish)]
fig, axes = plt.subplots(2, 4, figsize=(12, 5))
for col, (name, f, df) in enumerate(rows):
    axes[0, col].plot(z, f(z))
    axes[0, col].set_title(name)
    axes[1, col].plot(z, df(z), c="tab:orange")
    axes[1, col].set_title("derivative")
    axes[1, col].set_ylim(-0.2, 1.2)
    for ax in axes[:, col]:
        ax.axhline(0, c="grey", lw=0.6)
        ax.axvline(0, c="grey", lw=0.6)
plt.tight_layout()
plt.show()
i = np.argmin(swish(z))
print(f"swish minimum: {swish(z)[i]:.4f} at z = {z[i]:.2f}")
Top row: ReLU, Leaky ReLU with alpha 0.1, ELU with alpha 1 and Swish between z = −5 and 5. Bottom row: their derivatives; ReLU jumps from 0 to 1, Leaky ReLU from 0.1 to 1, ELU rises smoothly to 1, and Swish rises smoothly through 0.5 at z = 0 and slightly overshoots 1.

Swish's lowest value is −0.2785 at z = −1.28, the dip below 0. ReLU and Leaky ReLU have a corner at 0, where the derivative jumps; ELU and Swish have none.

In Keras each of these is one argument or one layer. Set Leaky ReLU's slope yourself, because Keras's default is not the video's 0.01:

python
from keras import layers

layers.Dense(8, activation="relu")            # ReLU
layers.Dense(8, activation="elu")             # ELU, alpha = 1
layers.Dense(8, activation="swish")           # Swish (Keras also calls it silu)
layers.LeakyReLU(negative_slope=0.01)         # Leaky ReLU with the video's 0.01
layers.PReLU()                                # PReLU: the slope is learned

ReLU vs Leaky ReLU vs PReLU vs ELU vs Swish

For z < 0Derivative for z < 0Dead neuronsCost
ReLU00yescheapest
Leaky ReLUαz, α = 0.01αnocheap
PReLUaᵢz, aᵢ learnedaᵢnocheap, one extra parameter per neuron
ELUα(e^z − 1)αe^znoan exponential
Swishzσ(z), dips to −0.28smooth, smallnoan exponential

Where you use ReLU and its variants

  • Hidden layers by default: ReLU is the first choice for the hidden layers of a deep network.
  • When training stalls: if many neurons output 0 for every record, switch to Leaky ReLU, PReLU or ELU.
  • Deeper modern networks: Swish (Keras's "swish" or "silu") appears in many large image and language models.
Watch out. A large learning rate or a large negative bias can push a ReLU neuron's z below 0 for every record, and once there it never comes back, as the α = 0 run showed. Count the neurons that output 0 for a whole batch; if it is a large share, lower the learning rate or use Leaky ReLU.
Try it yourself
  • In the dead-neuron run, change the starting bias to -0.2. Does plain ReLU learn now?
  • Train with alpha = 0.1 and steps=100. Does a larger slope revive the neuron faster?
  • Print d_swish(np.array([-1.28])). Why is it close to 0 at Swish's lowest point?

You understood something today that you didn't yesterday.