Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Weight initialization

Weight initialization is the step that sets every weight of a neural network to a starting value before training, chosen small, different from neuron to neuron and with a spread that suits the layer's size.

Last updated: 05 Oct, 2026 · NumPy

The Adam optimizer decides how the weights move during training. Where they start decides whether training gets going at all: weights that start too large make the gradients explode, weights that start identical make every neuron learn the same thing. The notes in the video's materials repo cover it in three pages: the problem large weights cause, three key points for good starting weights and three families of techniques.

Starting all the weights at zero

The Weights and bias lesson asks what happens if every weight starts at 0: x1 × 0, x2 × 0 and x3 × 0 are all 0, so the neuron passes on nothing. The video adds a bias to get round that. A bias does not fix the deeper problem, though. If all the weights into a layer are equal, zero or not, every neuron in the layer computes the same output, gets the same gradient and receives the same update. They stay copies of each other for the whole of training, so a layer of 100 neurons learns no more than one neuron. This is the symmetry problem, and random starting weights are what break it.

Backpropagating once from equal weights

The code runs one record with three inputs through a 3-3-1 network of sigmoid neurons and computes the gradient of every first-layer weight, once with all weights at 0 and once with all at 0.5.

python
h = 1 / (1 + np.exp(-(x @ W1)))         # hidden layer of 3 sigmoid neurons
d_h = (W2[:, 0] * d_out) * h * (1 - h)  # error reaching each hidden neuron
gW1 = np.outer(x, d_h)                  # one gradient per weight, column = neuron
ExampleRun with NumPy
import numpy as np

x = np.array([1.0, 2.0, 3.0])            # one record with 3 inputs
y = 1.0                                  # its label

for name, W1, W2 in [("all zeros", np.zeros((3, 3)), np.zeros((3, 1))),
                     ("all 0.5", np.full((3, 3), 0.5), np.full((3, 1), 0.5))]:
    h = 1 / (1 + np.exp(-(x @ W1)))      # hidden layer of 3 sigmoid neurons
    yhat = 1 / (1 + np.exp(-(h @ W2)))   # output neuron
    d_out = yhat - y                     # dL/dz at the output (log loss + sigmoid)
    d_h = (W2[:, 0] * d_out) * h * (1 - h)
    gW1 = np.outer(x, d_h)               # gradient for every weight into the hidden layer
    print(name, "| hidden outputs:", np.round(h, 4))
    print("gradient of W1, one column per hidden neuron:")
    print(np.round(gW1, 4))
  • All zeros: the three hidden neurons all output 0.5 (the sigmoid of 0) and every first-layer gradient is 0, because the zero weights into the output send no error back. The output weights do get a gradient, the same −0.25 for each, so from the next step the first-layer weights start to move, but in lockstep, like the 0.5 case below.
  • All 0.5: the gradients are not zero, but the three columns are identical (−0.0044, −0.0087, −0.0131). Every hidden neuron gets the same update and stays a copy of the others.

Exploding gradients from large weights

The notes open with the opposite problem. Take one path x1 → O1 → O2 → O3 → ŷ, each neuron a sigmoid with its own bias. Backpropagation updates a weight with wnew = wold − η · ∂L/∂wold, and the Chain rule of derivatives writes that slope as one factor per hop:

Each middle factor is the sigmoid's derivative times the next weight. With z = O2·w3 + b3:

In the Vanishing gradient problem lesson the weights were small, the factors were below 0.25 and the product shrank towards 0. Here the notes start w3 at 500 to 1000. Then one factor is up to 0.25 × 500 = 125, and big × big × big gives a huge slope. The update wnew = wold − η · (huge) jumps far past the minimum, the next slope is bigger still, and the loss zig-zags upward instead of settling. This is the exploding gradient problem, and its cause is weights that start with a very high value.

A chain x1 to O1 to O2 to O3 to y-hat with weights w1, w2, w3; the chain rule multiplies one factor per hop, each factor is sigma prime (at most 0.25) times a weight, so a weight of 500 gives 125 per hop and the gradient explodes; the key points say weights should be small, not all the same, and have a good variance.

Following the three key points

The notes sum up both problems in three rules for the starting weights:

  1. Weights should be small, so the chain-rule factors do not blow up.
  2. Weights should not be the same, so neurons in a layer can learn different things.
  3. Weights should have a good variance: not so tiny that signals fade, not so wide that they explode. The right spread depends on how many inputs a neuron has.

Every technique below draws random weights (rule 2) centred on 0 (rule 1) with a spread computed from the layer's size (rule 3). Two numbers describe the size: input (fan-in), the number of neurons feeding the layer, and output (fan-out), the number of neurons it feeds. The notes draw two cases: a 1 → 3 → 1 network has input = 1 and output = 1 for its hidden layer, and a 3-3-3 network has input = 3 and output = 3.

Drawing weights from a uniform distribution

The simplest rule draws every weight Wij from a uniform distribution whose limit shrinks with the number of inputs. For three inputs that is the range [−1/√3, 1/√3], about ±0.577.

Using Xavier (Glorot) initialization

Xavier Glorot, the researcher the method is named after, worked out a spread that keeps the variance of the signal about the same going forward and the gradient going backward, which needs both the input and the output size. It comes in a normal and a uniform version. In N(0, σ) the σ is the standard deviation, not the variance.

Using He (Kaiming) initialization

Kaiming He's version is built for ReLU. ReLU sets about half of its inputs to 0, which halves the variance at every layer, so He doubles the variance to make up for it and uses only the input size.

A table of five weight initialization techniques: uniform with limit 1 over root input (0.577 for 3 inputs), Xavier normal with sigma root of 2 over input plus output (0.577), Xavier uniform with limit root 6 over root of input plus output (1.0), He normal with sigma root of 2 over input (0.816) and He uniform with limit root of 6 over input (1.414), with their Keras names.

Computing the limits for three inputs

The five formulas in NumPy

python
fan_in, fan_out = 3, 3
xavier_sigma = np.sqrt(2 / (fan_in + fan_out))        # normal: standard deviation
xavier_limit = np.sqrt(6) / np.sqrt(fan_in + fan_out)  # uniform: range [-a, a]
he_sigma, he_limit = np.sqrt(2 / fan_in), np.sqrt(6 / fan_in)

Checking the spread of uniform draws

rng.uniform(-a, a, size=...) draws weights from U[−a, a]. A uniform range has a standard deviation of a/√3, which the run compares with the draws.

ExampleRun with NumPy
import numpy as np

fan_in, fan_out = 3, 3                   # the 3-3-3 network of the notes
limits = {
    "uniform        a": 1 / np.sqrt(fan_in),
    "xavier normal  σ": np.sqrt(2 / (fan_in + fan_out)),
    "xavier uniform a": np.sqrt(6) / np.sqrt(fan_in + fan_out),
    "he normal      σ": np.sqrt(2 / fan_in),
    "he uniform     a": np.sqrt(6 / fan_in),
}
for name, v in limits.items():
    print(f"{name} = {v:.3f}")

rng = np.random.default_rng(0)
a = np.sqrt(6) / np.sqrt(fan_in + fan_out)
w = rng.uniform(-a, a, size=100_000)     # many Xavier-uniform draws
print("xavier uniform draws: min", w.min().round(3), "max", w.max().round(3), "std", w.std().round(3))
print("a / sqrt(3) =", round(a / np.sqrt(3), 3))
  • Uniform a = 0.577 is 1/√3, the notes' [−1/√3, 1/√3].
  • Xavier normal σ = 0.577 and Xavier uniform a = 1.000: for input = output = 3, √(2/6) and √6/√6.
  • He normal σ = 0.816 and He uniform a = 1.414 are √(2/3) and √2, larger than Xavier because He divides by the input size only.
  • The uniform draws have a standard deviation of 0.577, the same as a/√3 and the same as the Xavier normal σ. The √6 in the uniform formula is there to make the two versions spread equally; the same holds for He (1.414/√3 = 0.816).

Measuring activations through ten layers

The point of a good variance shows up in a deep network. The code pushes 1000 random records through ten Dense layers of 256 ReLU neurons, starting the weights four ways, and prints the standard deviation of the activations after layers 1, 5 and 10.

python
a = x
for layer in range(10):
    a = np.maximum(0, a @ init(256, 256))   # one Dense layer + ReLU
    stds.append(a.std())                    # how spread out the signal is
ExampleRun with NumPy
import numpy as np
import matplotlib.pyplot as plt

rng = np.random.default_rng(0)
x = rng.standard_normal((1000, 256))     # 1000 records, 256 features
inits = {
    "N(0, 0.01)":    lambda fi, fo: rng.normal(0, 0.01, (fi, fo)),
    "N(0, 1)":       lambda fi, fo: rng.normal(0, 1.0, (fi, fo)),
    "Xavier normal": lambda fi, fo: rng.normal(0, np.sqrt(2 / (fi + fo)), (fi, fo)),
    "He normal":     lambda fi, fo: rng.normal(0, np.sqrt(2 / fi), (fi, fo)),
}
for name, init in inits.items():
    a, stds = x, []
    for layer in range(10):              # ten Dense layers of 256 ReLU neurons
        a = np.maximum(0, a @ init(256, 256))
        stds.append(a.std())
    print(f"{name:14s} layer 1: {stds[0]:.3g}  layer 5: {stds[4]:.3g}  layer 10: {stds[9]:.3g}")
    plt.semilogy(range(1, 11), stds, marker="o", label=name)

plt.title("Spread of ReLU activations through 10 layers")
plt.xlabel("layer")
plt.ylabel("standard deviation of the activations (log scale)")
plt.legend()
plt.show()
A log-scale plot of the spread of ReLU activations through 10 layers: N(0, 0.01) falls towards 1e-10, N(0, 1) climbs past 1e10, Xavier normal shrinks steadily to about 0.03 and He normal stays flat near 0.8.

What the ten layers show

  • N(0, 0.01), small weights chosen without regard to the layer's size, shrinks the signal about 9 times per layer, to 3.04e-10 by layer 10. The later layers receive almost nothing, and the gradients flowing back shrink the same way.
  • N(0, 1) grows it about 11 times per layer, to 2.65e+10: the exploding case from the notes, with weights that are too large for 256 inputs.
  • Xavier normal keeps the first layer healthy (0.584) but loses about 30% of the spread at each ReLU layer, 0.175 by layer 5 and 0.0308 by layer 10. Xavier was designed for tanh and sigmoid, which do not zero half their inputs.
  • He normal holds the spread near 0.8 at every depth (0.826, 0.771, 0.865). Doubling the variance makes up for the half that ReLU switches off.

Setting the initializer in Keras

In Keras each layer takes a kernel_initializer for its weights and a bias_initializer for its biases. The Dense page that the video opens in the ANN practical shows the defaults: kernel_initializer="glorot_uniform" and bias_initializer="zeros". Biases can start at 0 because the random weights already break the symmetry. The normal versions in Keras, glorot_normal and he_normal, draw from a truncated normal: values more than two standard deviations from 0 are drawn again, and the σ is scaled up so the spread still matches the formula.

python
from tensorflow.keras.layers import Dense

Dense(64, activation="relu", kernel_initializer="he_normal")      # ReLU layer: He
Dense(64, activation="tanh", kernel_initializer="glorot_normal")  # tanh layer: Xavier
Dense(1, activation="sigmoid")                 # default glorot_uniform, zero bias

Xavier vs He initialization

Xavier (Glorot)He (Kaiming)
Usesinput and output sizeinput size only
Normal σ√(2/(input+output))√(2/input)
Uniform limit√6/√(input+output)√(6/input)
Built forsigmoid and tanhReLU and its variants
Keras namesglorot_normal, glorot_uniform (the default)he_normal, he_uniform
In the ten-layer ReLU runspread falls to 0.0308spread stays near 0.8

Where you use weight initialization

  • Deep ReLU networks, including CNNs: switch the layers to he_normal or he_uniform when a deep model's loss does not move in the first epochs.
  • Debugging a loss that turns into nan early in training: oversized starting weights (or a too-large learning rate) are the first suspects.
  • Shallow networks, like the three-hidden-layer churn ANN later in this course: the default Glorot uniform trains them well, and the choice starts to matter as the network gets deeper.
Watch out. Setting all weights to the same value, zero or not, leaves every neuron in a layer identical forever, and training still runs without an error, so the bug hides behind a loss that falls slowly and then stalls. Always start weights randomly; only biases may start at 0.
Try it yourself
  • In the ten-layer run, change np.maximum(0, ...) to np.tanh(...) and compare Xavier normal with He normal again.
  • Set fan_in, fan_out = 1, 1 for the notes' 1 → 3 → 1 network and read the five limits.
  • In the symmetry run, start W1 with np.random.default_rng(0).normal(0, 0.5, (3, 3)) and check that the three gradient columns differ.

Little by little, you're building something great.