Weights and bias
A weight is a number on each connection that scales how strongly an input drives a neuron, and a bias is a constant added to the neuron's weighted sum that shifts the point where the neuron activates, like the intercept c in y = mx + c.
Last updated: 05 Oct, 2026 · NumPy
The perceptron lesson drew the lines between neurons. Each line carries a weight, each neuron adds a bias, and training means finding good values for these numbers.
Adding up inputs times weights
Each connection gets its own weight: w1 on the line from x1, w2 from x2, w3 from x3. Inside the neuron two operations happen. The first operation is the summation of xi wi:
This is the equation of linear regression again: β0 + β1x, written for many inputs as βᵀx, or the straight line y = mx + c. The second operation passes the result to an activation function.
Why weights? Put a hot object in your right hand and you pull the right hand away, not the left. The neurons on the path from the right hand are activated and pass the signal to the brain. Weights play that part in a network: they decide how much a neuron should be activated or deactivated by each input. Training sets them so the right neurons fire to the right level.

Adding the bias and the sigmoid
Suppose every weight starts at 0. Then x1 × 0, x2 × 0 and x3 × 0 are all 0, and the neuron passes 0 whatever the input. To handle that, one more parameter is added to the sum: the bias. It is a constant term, the intercept, the c in y = mx + c, and it is updated in training like the weights. The video adds that a bias is needed in every hidden layer; to be exact, every neuron has its own bias.
The second step applies an activation function on top of z. The example is the sigmoid, used for binary classification:
The sigmoid gives a value between 0 and 1, not only 0 or 1, as the video corrects itself in the clip. A condition turns it into a class: ≥ 0.5 → 1, < 0.5 → 0. Read this way, the sigmoid says whether the neuron is activated or not.
Two slips in the clip: the spoken formula "1 e to the power of minus y" drops the "1 +" that the board has, 1 / (1 + e^(−y)). And the video calls the sum y; this page calls it z, because y is the true label later on. The bias also does more than rescue zero weights: it shifts where the neuron switches on, so the neuron can fire even when the weighted sum is 0. Starting every weight at 0 has a deeper problem, every neuron in a layer learning the same thing, covered in Weight initialization.
Changing the weights and the bias of one neuron
One neuron as a function
import numpy as np
X = np.array([[7, 3, 7], [2, 5, 8], [4, 3, 7]]) # study, play, sleep hours
def neuron(x, w, b):
z = x @ w + b # step 1: weighted sum + bias
return 1 / (1 + np.exp(-z)) # step 2: sigmoid activationRunning the student rows through four neurons
settings = {
"zero weights, b = 0": (np.zeros(3), 0.0),
"zero weights, b = -1": (np.zeros(3), -1.0),
"w = 0.6, -0.5, -0.25": (np.array([0.6, -0.5, -0.25]), 2.1),
"w1 doubled to 1.2": (np.array([1.2, -0.5, -0.25]), 2.1),
}
for name, (w, b) in settings.items():
out = neuron(X, w, b)
print(f"{name:22} outputs {np.round(out, 3)} classes {(out >= 0.5).astype(int)}")zero weights, b = 0 outputs [0.5 0.5 0.5] classes [1 1 1] zero weights, b = -1 outputs [0.269 0.269 0.269] classes [0 0 0] w = 0.6, -0.5, -0.25 outputs [0.955 0.231 0.777] classes [1 0 1] w1 doubled to 1.2 outputs [0.999 0.5 0.975] classes [1 1 1]
Reading the four neurons
- Zero weights give the same output for every student: z = b, so all three rows get σ(0) = 0.5, which the ≥ 0.5 rule calls a pass, and then σ(−1) = 0.269, all fails. The bias moves the output, but with no weights the neuron cannot tell the students apart.
- The weights 0.6, −0.5, −0.25 with bias 2.1, close to what the logistic regression learned in AI vs ML vs DL vs data science, give 0.955, 0.231 and 0.777: the classes 1, 0, 1, the right answers for the table.
- Doubling the study weight pushes the two passing students to 0.999 and 0.975, but it lifts the failing student, who studies 2 hours, to 0.5 as well, and the threshold now calls a pass. A larger weight makes that input drive the neuron harder, like the hot object in the right hand; too large, and it drowns out the other inputs.
Shifting the sigmoid with the bias
import numpy as np
import matplotlib.pyplot as plt
hours = np.linspace(0, 10, 201) # one input: study hours, weight 1
for b in [-2, -4, -6]:
plt.plot(hours, 1 / (1 + np.exp(-(hours + b))), label=f"b = {b}")
print(f"b = {b}: the neuron reaches 0.5 at {-b} study hours")
plt.axhline(0.5, color="gray", linestyle="--")
plt.xlabel("study hours")
plt.ylabel("sigmoid output")
plt.title("Sigmoid of (hours + b): the bias moves the threshold")
plt.legend()
plt.show()b = -2: the neuron reaches 0.5 at 2 study hours b = -4: the neuron reaches 0.5 at 4 study hours b = -6: the neuron reaches 0.5 at 6 study hours

The weight sets how steep the curve is, and the bias sets where it crosses 0.5. With b = −4, the neuron says pass from 4 study hours. That is the intercept at work: the same line, slid along.
Weights vs bias
| Weight | Bias | |
|---|---|---|
| How many | One per connection (per input of each neuron) | One per neuron |
| Multiplies | Its input | Nothing; it is added |
| Effect | How strongly that input drives the neuron, and the curve's steepness | Where the neuron switches on |
| In y = mx + c | m, the slope | c, the intercept |
| Learned in training | Yes | Yes |
Where you use weights and bias
- Reading a model. Keras keeps each
Denselayer's weights as a matrix of shape (inputs, neurons) and its biases as a vector of length neurons;layer.get_weights()returns both. - Counting parameters for a layer: inputs × neurons weights plus neurons biases, the number
model.summary()prints. - Interviews: "what do weights do?" (how much a neuron activates) and "why add a bias?" (to shift the activation, like an intercept).
Related
- Previous: Perceptron
- Next: Forward propagation
- See also: Simple linear regression, where y = mx + c comes from
- Set every weight to 0 and the bias to 3 in a new setting: which class does each student get, and why can no bias alone fix the second row?
- In the plot, change the weight from 1 to 3 (
3 * hours + b) and see the curves get steeper while the 0.5 crossings move to −b / 3: 0.67, 1.33 and 2 hours. Change the print to{-b / 3:.2f}to match. - In the setting
w = 0.6, -0.5, -0.25, flip the sign of the play weight to +0.5 and see the failing student, who plays 5 hours, become a pass.
Every expert started right here.