Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Forward propagation

Forward propagation is the pass that carries one input row through a neural network, layer by layer, with a weighted sum plus bias and an activation in every neuron, until the output layer produces the prediction ŷ.

Last updated: 05 Oct, 2026 · NumPy

Weights and bias explained the two steps inside one neuron. Forward propagation chains them: the output of one neuron becomes the input of the next, and the last one gives the network's answer.

Passing the signal to the next layer

Forward propagation through the layers · from the Deep Learning In-depth Tutorials in 5 Hours video · 42:16 to 44:54

The sigmoid's job is to say whether a neuron is activated or not, and the same two steps happen in every neuron: the sum of the weights times the inputs plus the bias, then the activation. In the next layer another weight is initialised, w4, and the hidden neuron's output O1 is multiplied by it. The output neuron does its own two steps and gives 0 or 1 after the threshold.

The sigmoid fits binary classification. For regression, the output neuron uses a linear activation function, which passes the sum through unchanged. So far that makes weights, neurons, hidden layers and activation functions. The whole operation, take the input, multiply by the weights, add a bias and activate, all the way to the output, is called forward propagation.

Working the pass by hand

The video names forward propagation without numbers; the notes work it out. The network is the 3-1-1 network from the board: three inputs, one hidden neuron with bias b1, and an output neuron with bias b2. The dataset says whether a student passed from the IQ, the study hours and the play hours:

x1 IQx2 Study hoursx3 Play hoursPass/Fail
95441
100521
95270

The weights are w1, w2, w3 = 0.01, 0.02, 0.03 with b1 = 0.001, and w4 = 0.02 with b2 = 0.03. The first row goes in.

Hidden layer, step 1 and step 2

Output layer, step 1 and step 2

The student passed, so y = 1 and the error is 1 − 0.51130 ≈ 0.49. A prediction of 0.511 is barely on the pass side: the made-up weights know nothing yet.

The worked forward pass: inputs 95, 4 and 4 with weights 0.01, 0.02 and 0.03 and bias 0.001 give z = 1.151 and a sigmoid output of 0.7597; that times w4 = 0.02 plus b2 = 0.03 gives z = 0.0452 and an output of about 0.511, against a true value of 1, an error of about 0.49.

Two slips in the notes: the bias term is written 1 × 0.01, but the bias is 0.001, and only 0.001 gives 1.151 (0.01 would give 1.16). And the second neuron is labelled hidden layer 2, but it is the output layer.

Adding non-linearity in every layer

Forward propagation in every layer · from the Deep Learning In-depth Tutorials in 5 Hours video · 61:38 to 65:12

The second day of the video starts from the same network: three inputs, one neuron in the hidden layer, one in the output layer, a bias added in each layer, and a loss, the difference between y and ŷ, to reduce. It adds two points to the forward pass.

First, the weights are assigned randomly at the start, and there are various ways to do it. Second, the sum wᵀx + b on its own is a linear regression; the activation on top of it, written z = σ(y) on the board, is what gives the network non-linear properties, so it can solve non-linear problems. The output ŷ of the last layer then goes into the loss function.

Coding the forward pass

The weights from the notes

python
import numpy as np

x = np.array([95, 4, 4])            # IQ, study hours, play hours (passed: y = 1)
w = np.array([0.01, 0.02, 0.03])    # w1, w2, w3 into the hidden neuron
b1 = 0.001                          # hidden neuron's bias
w4, b2 = 0.02, 0.03                 # weight and bias of the output neuron

The sigmoid

python
def sigmoid(z):
    return 1 / (1 + np.exp(-z))

Running the notes' forward pass

ExampleThe notes' worked example, run on NumPy 2.5.3
z1 = x @ w + b1                  # hidden layer, step 1
o1 = sigmoid(z1)                 # hidden layer, step 2
z2 = o1 * w4 + b2                # output layer, step 1
y_hat = sigmoid(z2)              # output layer, step 2

print("z1    =", round(z1, 3))
print("O1    =", round(o1, 5))
print("z2    =", round(z2, 5))
print("y_hat =", round(y_hat, 5))
print("error =", round(1 - y_hat, 5), " squared:", round((1 - y_hat) ** 2, 4))

Reading the forward pass

  • z1 = 1.151 and O1 = 0.75969: the hidden neuron's sum and its sigmoid, as in the notes.
  • z2 = 0.04519, ŷ = 0.5113: the output neuron barely moves away from 0.5, because w4 = 0.02 makes its sum small whatever O1 is.
  • The error 0.4887 is the notes' 0.49. Squared, as a real loss function does, it is 0.2388; Loss and cost functions explains why the square is used.

Passing all three students at once

A real network passes many rows together. Stack the rows into a matrix and the same lines compute every prediction in one go:

ExampleThe notes' three students, run on NumPy 2.5.3
X = np.array([[95, 4, 4], [100, 5, 2], [95, 2, 7]])
y = np.array([1, 1, 0])

y_hat = sigmoid(sigmoid(X @ w + b1) * w4 + b2)   # both layers, every row
print("predictions:", np.round(y_hat, 4))
print("classes:    ", (y_hat >= 0.5).astype(int))
print("true:       ", y)
  • All three students come out at about 0.511, so all are called a pass, and the third student, who failed, is wrong.
  • The weights decide everything. Nothing in the forward pass looks at y; the pass only computes. Making the predictions right means changing the weights, the job of the backward pass in How a neural network learns.

Forward vs backward propagation

Forward propagationBackward propagation
DirectionInput layer to output layerOutput layer back to the input layer
What it computesWeighted sums, activations, then ŷHow much each weight caused the error
Changes the weightsNoYes, through the optimizer
Needs the true label yNo (only to score the result)Yes, through the loss
Used whenTraining and predictionTraining only

Where you use forward propagation

  • Every prediction. model.predict in Keras is a forward pass, and a trained network only ever runs this pass.
  • Every training step starts with one, since the loss needs ŷ.
  • Debugging a network by hand: push one row through, layer by layer, and check each z and activation against what you expect, as done here.
Watch out. Rounding in the middle of a forward pass changes the end. The notes round O1 to 0.759 and get 0.51129; carrying 0.75969 gives 0.51130. In code, keep full precision and round only what you print.
Try it yourself
  • Use the notes' slip, b1 = 0.01, and check that z1 becomes 1.16 and O1 becomes 0.7613.
  • Set w4 = 2.0, then w4 = -2.0, and run the three students again: all three move together, to about 0.825 and then 0.183, because the hidden outputs 0.760, 0.762 and 0.769 hardly differ. No w4 alone can separate them.
  • Replace the output sigmoid with a linear activation, y_hat = o1 * w4 + b2, the regression version, and read the raw number it gives.

Little by little, you're building something great.