Dropout
Dropout is a regularization layer that switches off a random fraction of a layer's neurons on every training step, so the network cannot lean on any single neuron and overfits less.
Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras
Good Weight initialization lets a deep network train. A deep network that trains well can also learn its training data too well, and dropout is one of the standard ways to stop that.
Reducing overfitting with a dropout layer
The board in the video draws a network with three inputs, two hidden layers and an output. Many interconnected neurons, trained with forward and backward propagation and an optimizer, can solve complex problems, but sometimes the network overfits: accuracy on the training data is very high while accuracy on the test data goes down. The notes give numbers for it: training accuracy 90%, test accuracy 60%.
Dropout works like the L1 and L2 regularization of Lasso regression and Ridge regression, by making the model less able to memorise. A dropout ratio of 0.3 on a layer means 30% of that layer's neurons are deactivated while training; 0.5 deactivates half of them. A deactivated neuron is like a dead neuron: its connections are cut for that step, and only the remaining neurons do forward and backward propagation.
The neurons are picked at random and the pick changes on every training step, so each neuron is off in some steps and on in others. A ratio of 0.3 on a layer of 4 neurons drops 1.2 neurons on average: one in some steps, none or two in others.

Sampling neurons like a random forest
The notes explain why this helps with a machine learning analogy. A single Decision tree classifier grown to full depth overfits. A Random forest trains many trees, each on a sample of the rows and a sample of the features, and combines them into a generalized model. Dropout does the same inside one network: every training step trains a different thinned network made of the neurons that survived, and the full network at test time behaves roughly like the average of all those thinned networks. No neuron can rely on one particular partner being present, so each one has to learn something useful on its own.
Choosing the dropout ratio p
The ratio p lies between 0 and 1 (0 ≤ p ≤ 1) and is a hyperparameter: you choose it and compare validation scores. The notes use p = 0.5 on the input layer, p = 2/5 on the first hidden layer (2 of its 5 neurons crossed out) and p = 0.5 on the second. Common choices are 0.2 to 0.5 on hidden layers and a smaller value, around 0.2, on inputs, since dropping half of the raw features throws away a lot of information. The output layer gets no dropout.
Scaling the weights at test time
On test data every neuron is on. That changes the size of the signal: during training a neuron in the next layer received input from only the kept fraction of the layer, about 3 of the 5 neurons when p = 2/5, and at test time it receives all 5. To keep the two the same, each weight leaving a dropout layer is multiplied by the keep fraction 1 − p at test time.
The notes write W × 2/5 for the first hidden layer; with 2 of 5 neurons dropped, 3/5 are kept, so the weights are scaled by 3/5. For the two layers with p = 0.5 the keep fraction is also 0.5, so W × 0.5 is right there.
Running a dropout mask in NumPy
A random mask per step
rng.random(5) >= p gives five True/False values, each True with probability 1 − p. Multiplying the outputs by it zeroes the dropped neurons.
mask = rng.random(5) >= p # True = neuron stays on this step
dropped_h = h * mask # dropped neurons output 0
z_next = np.sum(w * dropped_h) # what one HL2 neuron receivesComparing training with the two test-time scalings
train_avg = np.mean((masks * h) @ w) # average over many training steps
test_keep = np.sum(w * (1 - p) * h) # all on, W x 3/5
test_drop = np.sum(w * p * h) # all on, W x 2/5 (the notes' slip)import numpy as np
rng = np.random.default_rng(42)
h = np.array([0.8, 1.2, 0.5, 2.0, 1.5]) # outputs of the 5 neurons of HL1
w = np.array([0.4, -0.3, 0.9, 0.2, 0.5]) # their weights into one HL2 neuron
p = 2 / 5 # drop fraction, as in the notes
keep = 1 - p
for step in range(1, 5): # a fresh random mask every step
mask = rng.random(5) >= p # True = the neuron stays on
print(f"step {step}: mask {mask.astype(int)} input to HL2 neuron {np.sum(w * h * mask):.3f}")
masks = rng.random((100_000, 5)) >= p # 100,000 training steps
print("average fraction kept:", masks.mean().round(3))
print("training, average input:", round(np.mean((masks * h) @ w), 3))
print("test, all on, W x 2/5: ", round(np.sum(w * p * h), 3))
print("test, all on, W x 3/5: ", round(np.sum(w * keep * h), 3))
inv = np.mean((masks * h / keep) @ w) # inverted dropout: divide kept outputs by keep
print("inverted dropout, training average:", round(inv, 3), " test, no scaling:", round(np.sum(w * h), 3))step 1: mask [1 1 1 1 0] input to HL2 neuron 0.810 step 2: mask [1 1 1 0 1] input to HL2 neuron 1.160 step 3: mask [0 1 1 1 1] input to HL2 neuron 1.240 step 4: mask [0 1 0 1 1] input to HL2 neuron 0.790 average fraction kept: 0.599 training, average input: 0.934 test, all on, W x 2/5: 0.624 test, all on, W x 3/5: 0.936 inverted dropout, training average: 1.556 test, no scaling: 1.56
What the masks and the averages show
- Each step drops a different set: one neuron in steps 1 to 3, two in step 4, and a different one each time, so the HL2 neuron's input jumps between 0.79 and 1.24.
- About 60% of the neurons are kept over 100,000 steps (0.599), the keep fraction 3/5.
- Scaling by 3/5 matches training: the test input with W × 3/5 is 0.936 against a training average of 0.934. W × 2/5 gives 0.624, a third too small.
- Inverted dropout divides the kept outputs by 0.6 during training instead. The training average becomes 1.556, the same as the unscaled test input 1.56, so nothing has to change at test time.
Adding dropout in Keras
In Keras, dropout is a layer of its own placed after the layer whose outputs it drops; it is not an argument of Dense. Dropout(rate) takes p, the fraction to drop. Keras uses inverted dropout: during fit() it zeroes a random fraction rate of the outputs and divides the rest by 1 − rate, and in predict() and evaluate() the layer passes everything through unchanged.
from tensorflow.keras.layers import Dense, Dropout
classifier.add(Dense(units=7, activation="relu"))
classifier.add(Dropout(0.2)) # drop 20% of these 7 outputs per step
classifier.add(Dense(units=6, activation="relu"))
classifier.add(Dropout(0.3)) # drop 30% of these 6 outputs per stepThese are the layers the video adds to its churn network in Training and evaluating an ANN.
Dropout vs L2 regularization
| Dropout | L2 regularization | |
|---|---|---|
| How it limits the model | switches random neurons off each step | adds λ·Σw² to the loss, keeping weights small |
| Setting | rate p per layer, 0 ≤ p ≤ 1 | one λ per layer or model |
| At test time | all neurons on (Keras: no change) | no change |
| Effect on training | noisier loss, often more epochs | smooth loss |
| In Keras | Dropout(0.2) layer | Dense(..., kernel_regularizer="l2") |
Where you use dropout
- Large Dense layers that reach a high training accuracy while the validation accuracy stalls, like the 90% vs 60% case in the notes.
- The fully connected head of a CNN, where most of the parameters sit.
- Small tabular datasets such as the churn data, where a few thousand rows are easy to memorise.
fit() is measured with dropout switched on, so with a high rate it can sit below the validation accuracy. That is not a bug; compare models by their validation and test scores, which are computed with every neuron on.Related
- Previous: Weight initialization
- Next: Preparing the churn data
- See also: Overfitting and underfitting, Random forest
- Reference: Dropout layer in the Keras API
- Set
p = 0.5in the mask run and check that W × p and W × (1 − p) now give the same test input. - Change
rng.random(5) >= ptorng.random(5) >= 0.0and confirm that nothing is ever dropped. - Print
masks.sum(axis=1)[:20]to see how many of the 5 neurons survive in each of the first 20 steps.
You understood something today that you didn't yesterday.