Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Loss and cost functions

A loss function is a formula that scores how wrong a neural network's prediction ŷ is for one record, and a cost function is the same score averaged over a batch of records.

Last updated: 05 Oct, 2026 · NumPy

After Forward propagation the network has a prediction ŷ, and the error y − ŷ is what backpropagation pushes down. The formula that turns that error into one number depends on the problem, and Choosing an activation function already tied the output layer to the problem type. This part picks the loss to go with it.

Regression and classification losses, loss vs cost · from the Deep Learning In-depth Tutorials in 5 Hours video · 125:32 to 130:12

Splitting problems into regression and classification

An artificial neural network (ANN) solves two kinds of problems, and the board starts with one dataset for each.

  • Regression: years of experience and degree give a salary. Someone with 10 years and a PhD earns a high salary, and salary is a continuous output.
  • Classification: hours played and hours studied give the result. Playing 10 and studying 2 gives Fail, 4 and 3 Fail, 5 and 5 Maybe, 2 and 7 Pass. With Maybe in the column there are three classes, so this is a multi-class problem.
An ANN splits into regression, with an experience, degree and salary table whose salary column is continuous, and classification, with a play hours, study hours and result table whose results Fail, Fail, Maybe and Pass are three classes.

The rows are made up to show the shape of each dataset. Each kind of problem gets its own losses:

ProblemOutputLoss functions
Regressiona continuous numbermean squared error (MSE), mean absolute error (MAE), Huber loss
Classificationa classcross entropy: binary for two classes, categorical for more

Computing the loss for one record

Take 100 records. The simplest training loop passes one record forward, computes ŷ, computes the loss and runs backpropagation, then moves to the next record. For a regression network the loss of that one record is the squared error:

Updating the weights after every single record is slow and noisy, which the video calls not an efficient scenario.

Averaging the loss over a batch into a cost

Instead, set a batch size, say 10. Each forward pass now carries 10 records, the loss is computed for each of them, and the 10 losses are combined into one number, the cost. In the video's words: in the loss function you provide one data point, in the cost function a batch of data points.

The board writes the cost as ½ Σ (y − ŷ)² with no 1/n; a cost is a mean over the batch, so it needs the 1/n (some books write 1/(2n) to tidy the derivative). The later notes write it with 1/n.

The same network twice: on the left one record goes forward and gives one loss, the squared error; on the right a batch of 10 of the 100 records goes forward and the 10 losses are averaged into one cost.

Counting batches, iterations and epochs

Right after the clip the video says a batch is passed in every epoch. One batch is one iteration, one weight update. An epoch is one full pass over all 100 records, so a batch of 10 means 10 iterations per epoch. Gradient descent builds on these three words.

Computing loss and cost in NumPy

The 100 records

A salary that grows by about 1.5 lakhs per year of experience, and a network whose predictions grow by 1.4:

python
import numpy as np

rng = np.random.default_rng(0)
experience = rng.uniform(0, 10, size=100).round(1)                     # 100 records
salary = (3 + 1.5 * experience + rng.normal(0, 1, size=100)).round(2)  # in lakhs
predicted = 3 + 1.4 * experience                                       # one network's ŷ

Loss of one record and cost of a batch

python
def loss(y, y_hat):
    return (y - y_hat) ** 2              # one record

def cost(y, y_hat):
    return np.mean((y - y_hat) ** 2)     # (1/n) Σ over a batch of n records

Comparing a mean with a half sum

ExampleRun on NumPy 2.5.3
print("loss, record 0:", round(loss(salary[0], predicted[0]), 3))
print("cost, records 0-9:", round(cost(salary[:10], predicted[:10]), 3))

for n in (10, 50, 100):
    half_sum = 0.5 * np.sum((salary[:n] - predicted[:n]) ** 2)
    print(f"batch of {n:3d}: half sum = {half_sum:7.3f}   mean = {cost(salary[:n], predicted[:n]):.3f}")

batch_size = 10
print("iterations per epoch:", len(salary) // batch_size)

What the loss and cost printed

  • Record 0 alone has a loss of 0.49: its salary is off by 0.7 lakhs, and 0.7² = 0.49.
  • The first batch of 10 has a cost of 0.848, the average of 10 such losses. One weight update now uses all 10 records.
  • The half sum grows with the batch: 4.241, 29.682, 61.464 for 10, 50 and 100 records, while the mean stays near 1 (0.848, 1.187, 1.229). With the mean, the size of the number does not depend on the batch size.
  • 100 records with a batch of 10 give 10 iterations per epoch.

Loss function vs cost function

Loss functionCost function
Inputone recorda batch of n records
Regression formula(y − ŷ)²(1/n) Σ (yᵢ − ŷᵢ)²
One update usesone predictionthe batch's average error
Typical usethe formula you choose (MSE, cross entropy)what the optimizer minimises during training

Where you use loss and cost functions

  • Compiling a model: the loss you pass to Keras is the per-record formula; Keras averages it over each batch, which is the cost.
  • Reading a training log: the loss printed per epoch is a cost, averaged over batches, so it can be compared across batch sizes.
  • Interviews: "what is the difference between a loss function and a cost function" is a standard question, and the answer is one record vs a batch.
Watch out. A cost written as a plain or half sum grows with the batch size, so changing the batch from 10 to 100 also makes every gradient about 10 times larger, as if the learning rate had changed. Averaging with 1/n keeps the scale the same.
Try it yourself
  • Change predicted to 3 + 1.5 * experience: the cost of all 100 records drops to about 0.95, the noise left in the salaries, because the slope is now right.
  • Set batch_size = 32: len(salary) // batch_size prints 3, and 4 records are left for a smaller last batch.
  • Print cost(salary, predicted) for all 100 records and compare it with the mean of ten batch costs of 10 records each: they are equal.

Every expert started right here.