Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Perceptron

A perceptron is the simplest neural network: one neuron that multiplies its inputs by weights, adds a bias, and passes the sum through an activation function to output a class, 0 or 1.

Last updated: 05 Oct, 2026 · NumPy

Every network later in the course, from the churn ANN to a CNN, is built from this unit repeated many times. The perceptron is the first neural network, and the video starts deep learning with it: first the single-layer network, then the multi-layer one.

Drawing a single-layer network

The perceptron and the student dataset · from the Deep Learning In-depth Tutorials in 5 Hours video · 22:52 to 26:52

Draw three circles on the left: they are the input layer, one circle per input. Since this is a single-layer network, the next circle is hidden layer 1, and the last circle is the output layer. Every dot is connected to the next layer.

The example is a dataset that says whether a student passed. Each record has study, play and sleep hours, and the output is pass or fail. Two outcomes make it a binary classification problem, so pass is written as 1 and fail as 0:

  • Studies 7 hours, plays 3, sleeps 7: pass (1).
  • Studies 2 hours, plays 5, sleeps 8: fail (0).
  • Studies 4 hours, plays 3, sleeps 7: pass (1).

The inputs go in record by record. Study is x1, play is x2 and sleep is x3, so the first pass feeds in 7, 3 and 7.

Three inputs, study, play and sleep hours, connect to one hidden neuron and then an output neuron that gives y hat; the dataset beside it has the rows 7, 3, 7 pass, 2, 5, 8 fail and 4, 3, 7 pass, and the first row is fed into the inputs.

The board names the middle neuron "hidden layer 1" and adds a separate output neuron. A perceptron in its strict sense has no hidden layer: the inputs feed one neuron, and that neuron's output is the prediction. The two drawings compute the same kind of thing, a weighted sum and an activation; the code below uses the strict form.

Seeing, processing and training like a brain

The eye, the brain and training a baby · from the Deep Learning In-depth Tutorials in 5 Hours video · 26:52 to 30:39

The video maps the network onto how you recognise a camera. Your eyes are the input layer: the picture of the camera is the input signal. The signal goes to neurons in the hidden layer, and a biological neuron does some kind of signal processing. The output layer is the part at the back of the brain that decides what you are looking at: a camera, a television or a mobile phone. That final neuron is the output neuron.

A newborn baby who sees a camera for the first time cannot say what it is. Someone has to tell it, again and again: this is the camera, this is food, this is your bed, this is milk. After a few months it recognises the milk bottle and the phone. A neural network built from scratch is the same: it must be trained on inputs with their right outputs before it predicts anything useful. Like the brain, a network can have any number of hidden layers, and the signal passes from neuron to neuron through them.

Splitting the neuron into two steps

The notes for this topic list what a perceptron is made of, an input layer, a hidden layer, weights and an activation function, and split the neuron into two steps:

Step 2 passes z through an activation function, which turns it into the output. The notes draw two of them:

The notes write the sigmoid's rule as "1 if z > 0.5". The 0.5 threshold applies to the sigmoid's output σ(z), not to z: σ(z) ≥ 0.5 happens exactly when z ≥ 0, so both functions switch class at z = 0. The difference is that the step jumps from 0 to 1, while the sigmoid rises smoothly and gives a value in between that can be read as a probability.

Separating classes with a straight line

The rule z = w1x1 + w2x2 + b = 0 is a straight line (a plane with three inputs). Points on one side get class 1 and points on the other get class 0, so a perceptron is a linear classifier. If one straight line can split the two classes, the data is linearly separable and a perceptron can learn it. If one class surrounds the other, no line works, and the notes' answer is a multi-layer neural network.

Two scatter plots: on the left two groups of points split by one straight line, which a perceptron can learn; on the right one group surrounds the other, so no straight line separates them and a multi-layer network is needed.

Training a perceptron on the student table

The classic way to train one neuron with a step function is the perceptron learning rule: feed each record in, and when the prediction is wrong, move every weight by the learning rate × the error × that input. When the prediction is right, the error is 0 and nothing changes.

The student table

python
import numpy as np

# the video's student table: study, play and sleep hours, pass (1) or fail (0)
X = np.array([[7, 3, 7], [2, 5, 8], [4, 3, 7]])
y = np.array([1, 0, 1])

The step function and a prediction

python
def step(z):
    return (z > 0).astype(int)     # 1 when z > 0, else 0

def predict(X, w, b):
    return step(X @ w + b)          # step 1: sum, step 2: activation

The perceptron learning rule

python
def train(lr, epochs=3):
    w, b = np.zeros(3), 0.0                  # start from zero weights
    for epoch in range(1, epochs + 1):
        for xi, yi in zip(X, y):             # one record at a time
            error = yi - predict(xi, w, b)   # 0 when right, +1 or -1 when wrong
            w = w + lr * error * xi          # move the weights towards the answer
            b = b + lr * error
        print("epoch", epoch, "w", np.round(w, 2), "b", round(b, 2), "predictions", predict(X, w, b))
    return w, b

Learning the pass and fail rows

ExampleThe video's student table, run on NumPy 2.5.3 and scikit-learn 1.9.1
w, b = train(lr=0.1)

from sklearn.linear_model import Perceptron
clf = Perceptron(shuffle=False, random_state=0).fit(X, y)   # the same rule, learning rate 1
print("scikit-learn w", clf.coef_[0], "b", clf.intercept_[0], "predictions", clf.predict(X))

Reading the training run

  • Epoch 1 already fixes it. The first record (7, 3, 7) is predicted 0 by the zero weights, so w moves to 0.1 × 7, 3, 7. The second record (2, 5, 8) is then wrongly a pass, and w moves back by 0.1 × 2, 5, 8, giving 0.5, −0.2, −0.1.
  • Epochs 2 and 3 change nothing: every prediction is right, every error is 0, so the rule stops moving the weights.
  • scikit-learn's Perceptron lands on 5, −2, −1: the same direction, ten times larger, because it uses a learning rate of 1. From zero weights, the learning rate only scales the line; it does not change which side each student falls on.
  • Study gets a positive weight, play and sleep negative ones: more study pushes towards a pass.

Comparing the step function and the sigmoid

ExampleRun on NumPy 2.5.3 and matplotlib 3.11.2
import numpy as np
import matplotlib.pyplot as plt

z = np.linspace(-6, 6, 241)
fig, axes = plt.subplots(1, 2, figsize=(9, 3.4))
axes[0].plot(z, (z > 0).astype(int), color="tab:red")
axes[0].set_title("Step function: threshold at z = 0")
axes[1].plot(z, 1 / (1 + np.exp(-z)), color="tab:blue")
axes[1].axhline(0.5, color="gray", linestyle="--")
axes[1].set_title("Sigmoid: threshold at 0.5")
for ax in axes:
    ax.set_xlabel("z")
    ax.set_ylabel("output")
plt.tight_layout()

for v in [-0.1, 0.0, 0.1]:
    print(f"z = {v:5}: step {int(v > 0)}, sigmoid {1 / (1 + np.exp(-v)):.3f}")
plt.show()
Two plots side by side: the step function jumps from 0 to 1 at z = 0, and the sigmoid rises smoothly from 0 to 1 through 0.5 at z = 0, with a dashed line at 0.5.

Near z = 0 the step function flips from 0 to 1 at once, while the sigmoid moves from 0.475 to 0.525. That smooth change is what lets a network learn by small adjustments, the subject of How a neural network learns.

Single-layer vs multi-layer perceptron

Single-layer perceptronMulti-layer perceptron (ANN)
LayersInputs straight into one neuronOne or more hidden layers between inputs and outputs
Decision boundaryOne straight line or planeCurved, any shape with enough neurons
Data it can learnLinearly separable onlyNon-linearly separable too
TrainingThe perceptron learning ruleForward propagation, loss, backpropagation, optimizer
ActivationStep function (or sigmoid)Sigmoid, tanh, ReLU and others per layer

Where you use a perceptron

  • As the unit of every network. A Keras Dense layer is a row of these neurons with a smooth activation in place of the step.
  • Simple linear yes/no rules on a few numeric inputs, such as pass or fail from study hours, where a straight line is enough.
  • Interviews: "what is a perceptron and what can it not learn?" The answer is a linear classifier that fails on data that is not linearly separable.
Watch out. A single perceptron never stops making mistakes on data that is not linearly separable. The classic case is XOR: (0, 0) and (1, 1) are class 0, (0, 1) and (1, 0) are class 1, and no straight line splits them. The learning rule then keeps moving the weights in a circle, epoch after epoch; the fix is a hidden layer.
Try it yourself
  • Train on XOR: set X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]]), y = np.array([0, 1, 1, 0]) and np.zeros(2) inside train, then run train(lr=0.1, epochs=10). Every epoch ends on the same weights and the predictions 1, 1, 0, 0: the rule goes round in a circle and never reaches 0, 1, 1, 0.
  • Call train(lr=1) and check that the hand-written rule now prints the same weights as scikit-learn.
  • Add a student who studies 3 hours, plays 6 and sleeps 6, predict with the trained w and b, and say whether the answer matches your own guess.

This is what real progress feels like.