Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Gradient descent

Gradient descent is an optimizer that updates every weight with wnew = wold − η·∂L/∂wold, using the whole training set for each update.

Last updated: 05 Oct, 2026 · NumPy

Choosing a loss function picked the number a network should make small. An optimizer is the rule that changes the weights to make it small. Gradient descent is the first one, and every optimizer in this part is a fix for one of its problems. For a refresher on the update rule itself on a single line, see Gradient descent in the Machine Learning course.

Updating the weights with gradient descent

The board starts from the weight update formula that backpropagation uses:

Plot the loss against one weight w and the curve is a U. The bottom is the global minima. A point on either arm moves towards it: on the right arm the slope is positive and w decreases, on the left arm the slope is negative and w increases. The learning rate η sets how big each move is.

The network on the board has three inputs, one hidden layer (HL1) with two neurons and one output. Forward propagation multiplies the inputs by the weights, applies the activation and gives ŷ. The loss is the mean squared error, and backward propagation updates w1, w2, w3, w4 to make it smaller.

The video corrects its earlier sessions here: the MSE needs the n, so it is 1/(2n), not 1/2. The ½ is there so that the 2 from the square cancels when you take the derivative.

A network with three inputs, two hidden neurons and one output, weights w1 to w4, a red forward arrow and a grey backward arrow; all 1,000,000 records go forward and back once, which is one epoch, and the cost is 1 over 2n times the sum of squared errors; beside it the U-shaped loss curve over w with the global minima at the bottom.
Epoch and the gradient descent resource problem · from the Deep Learning In-depth Tutorials in 5 Hours video · 160:28 to 163:39

Defining an epoch

An epoch is one forward propagation plus one backward propagation over the training data. Say the dataset has 1,000,000 records. Gradient descent sends all one million through the forward pass at once, gets one ŷ for each, and the cost function's n is one million. Then backward propagation updates every weight once, based on that cost.

From here the video calls it the cost function rather than the loss, because it covers many records: the loss is the error of one record, the cost is the average over the n records in the pass.

So in gradient descent one epoch is one weight update. Two more words come with the next optimizers. The batch size is how many records go through the network before one weight update; in gradient descent it is the whole dataset. An iteration is one such update, one batch forward and back, so iterations per epoch = records ÷ batch size. A network trains for many epochs; the video uses 100 as its example.

Three rows for 1,000,000 records. Gradient descent: one batch of 1,000,000, so 1 iteration per epoch and 100 in 100 epochs. Mini-batch SGD: batch size 1,000, so 1,000 iterations per epoch and 100,000 in 100 epochs. SGD: batch size 1, so 1,000,000 iterations per epoch and 100,000,000 in 100 epochs.

Why gradient descent is resource intensive

The video names one major disadvantage. To pass a million records in every epoch the machine must hold them all at once: it needs a huge RAM, and to compute that many forward passes in parallel it may need GPUs too. The upside is that each update uses the exact average over all records, so the path to the global minima is smooth and direct.

The question that leads to the next lesson: can an optimizer work well with fewer resources?

Counting iterations and running gradient descent

First the board's numbers. The same arithmetic gives the iterations for any batch size, and a rough memory figure shows why a million records is heavy: a record with 100 features stored as 4-byte floats is 400 bytes.

ExampleThe board's 1,000,000 records, counted in Python
import math

records, epochs, features = 1_000_000, 100, 100
for name, batch_size in [("gradient descent", records), ("mini-batch SGD", 1000), ("SGD", 1)]:
    per_epoch = math.ceil(records / batch_size)          # iterations in one epoch
    print(f"{name:16} batch {batch_size:>9,}  per epoch {per_epoch:>9,}  in {epochs} epochs {per_epoch * epochs:>11,}")

batch_bytes = records * features * 4                     # float32 inputs of one gradient descent batch
print("inputs held for one gradient descent update:", batch_bytes / 1e6, "MB")

The records and the gradients

Now gradient descent for real, on a single neuron ŷ = w·x + b with no activation, which is linear regression. 1,000 records stand in for the million. The gradients are the derivatives of the 1/(2n) cost.

python
import numpy as np

rng = np.random.default_rng(42)
n = 1000                                     # 1,000 records stand in for the board's 1,000,000
x = rng.uniform(0, 2, n)
y = 3 * x + 2 + rng.normal(0, 0.5, n)        # the true line has w = 3 and b = 2
python
def gradients(w, b, xb, yb):
    error = w * xb + b - yb                  # ŷ − y for every record in the batch
    return np.mean(error * xb), np.mean(error)   # ∂C/∂w and ∂C/∂b

def cost(w, b):
    return np.mean((w * x + b - y) ** 2) / 2     # C = 1/(2n) Σ (y − ŷ)²

One update per epoch

Every epoch passes all 1,000 records, so every epoch is one iteration:

ExampleGradient descent on 1,000 records, run with NumPy
w, b, lr = 0.0, 0.0, 0.5
for epoch in range(1, 101):
    gw, gb = gradients(w, b, x, y)           # all 1,000 records, one forward and backward pass
    w, b = w - lr * gw, b - lr * gb          # one weight update per epoch
    if epoch in (1, 2, 5, 10, 50, 100):
        print(f"epoch {epoch:3}: w = {w:.3f}  b = {b:.3f}  cost = {cost(w, b):.4f}")
print("records read:", 100 * n, " weight updates:", 100)

Reading the counts and the run

  • Gradient descent does 1 iteration per epoch, 100 in 100 epochs. Mini-batch SGD with a batch of 1,000 does 1,000 per epoch, and SGD does 1,000,000 per epoch, 100,000,000 in 100 epochs: the board's numbers.
  • 400 MB of inputs for one gradient descent update, before any hidden layer's outputs and gradients, which also scale with the batch. That is the huge RAM the board warns about.
  • The cost falls every epoch and the weights settle near w = 3.05, b = 1.91, the best line for these noisy records (the true line is w = 3, b = 2; the noise moves the best fit a little).
  • 100 updates cost 100,000 record reads. Each update looks at every record. The next optimizers trade this exactness for cheaper updates.

Epoch vs iteration vs batch size

Batch sizeIterationEpoch
What it countsrecords in one forward and backward passone weight update, one batchone pass over every record
Gradient descent, 1,000,000 records1,000,0001 per epoch1 weight update
Mini-batch SGD, batch 1,0001,0001,000 per epoch1,000 weight updates
Set in Keras bybatch_size in fitrecords ÷ batch size, rounded upepochs in fit

Where you use gradient descent

  • Small datasets that fit in memory, where a full pass is cheap and the smooth path is a bonus.
  • Checking a new model: full-batch updates have no noise, so a cost that does not fall points to a bug, not to bad luck in a batch.
  • As the base of every other optimizer: SGD, momentum, Adagrad, RMSprop and Adam all keep this update and change what goes into it.
Watch out. The batch size, not the optimizer's name, decides whether training is gradient descent. In Keras, model.fit uses mini-batches unless you pass batch_size=len(X), so a model trained with the optimizer called SGD is usually running mini-batch SGD.
Try it yourself
  • Change lr to 1.5 in the gradient descent run: the cost grows each epoch and the weights blow up, because the step overshoots the bottom of the U.
  • Set features = 1000 in the counting example: one gradient descent batch now needs 4,000 MB of inputs.
  • Change lr to 0.05: after 100 epochs the cost is still 0.1324, above the 0.1281 that η = 0.5 reaches, because each of the 100 updates is a smaller step.

Every expert started right here.