Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Binary cross-entropy

Binary cross-entropy is a classification loss function that scores a predicted probability ŷ against a label of 0 or 1: it costs −log ŷ when the label is 1 and −log(1 − ŷ) when the label is 0.

Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras

The regression losses in MSE, MAE and Huber loss measure a distance between numbers. A classifier outputs a probability, and its loss comes from the cross entropy family. Binary cross-entropy is the same log loss that logistic regression uses, covered in Log loss.

Cross entropy and binary cross-entropy · from the Deep Learning In-depth Tutorials in 5 Hours video · 140:35 to 144:30

The board's box writes ŷ = 1/(1 + e^−x); the sigmoid is applied to the last neuron's weighted sum z, so ŷ = 1/(1 + e^−z).

Splitting cross entropy into binary and categorical

  • Binary cross-entropy for binary classification: two classes, such as pass or fail, spam or not spam.
  • Categorical cross-entropy for multi-class classification: three or more classes, covered in Categorical cross-entropy.

Writing the binary cross-entropy formula

y is the true label, 0 or 1, and ŷ is the predicted probability that the label is 1. For the cost of a batch the video adds a sum over the records i = 1 to n; dividing that sum by n makes it the average.

Reading the two branches

Put each label into the formula. With y = 0 the first term is 0 · log ŷ = 0 and 1 − y = 1, so only −log(1 − ŷ) is left. With y = 1 the second term vanishes and −log ŷ is left:

The log is the natural log. A confident right answer costs almost nothing, and a confident wrong answer costs a lot:

ŷLoss if y = 1: −ln ŷLoss if y = 0: −ln(1 − ŷ)
0.90.1052.303
0.50.6930.693
0.12.3030.105
Left: the last neuron's weighted sum z goes through the sigmoid to give y hat, which goes into binary cross-entropy. Right: the two branches against y hat from 0 to 1, minus log y hat for a label of 1 falling from high values to 0 at 1, and minus log of 1 minus y hat for a label of 0 rising from 0, with y hat 0.9 costing 0.105 or 2.303.

Getting ŷ from the sigmoid in the last layer

The video's question is how ŷ is calculated. The output layer of a binary classifier has one neuron with the sigmoid activation, so ŷ = σ(z) lies between 0 and 1. Binary cross-entropy is applied on top of it, and backpropagation starts from that loss.

Computing binary cross-entropy in NumPy

The sigmoid and the loss

python
import numpy as np

def sigmoid(z):
    return 1 / (1 + np.exp(-z))

def bce(y, y_hat):
    return -(y * np.log(y_hat) + (1 - y) * np.log(1 - y_hat))

Four records through the last neuron

ExampleRun on NumPy 2.5.3
z = np.array([2.2, -1.4, 0.0, -3.0])    # the last neuron's weighted sums
y = np.array([1, 0, 1, 1])              # true labels
y_hat = sigmoid(z)
print("y_hat:", y_hat.round(3))
print("loss :", bce(y, y_hat).round(3))
print("cost :", round(bce(y, y_hat).mean(), 3))
for p in (0.9, 0.5, 0.1):
    print(f"y_hat = {p}: loss if y = 1 is {bce(1, p):.3f}, if y = 0 is {bce(0, p):.3f}")

A very confident wrong answer

With z = 40 the sigmoid rounds to exactly 1.0 in floating point, so 1 − ŷ is 0 and log 0 is minus infinity. NumPy prints a divide warning and the loss comes out as inf, which would break training. Writing the loss directly in terms of z avoids forming 1 − ŷ:

python
def bce_from_logits(y, z):
    # the same loss written with z, never forming 1 - ŷ
    return np.maximum(z, 0) - z * y + np.log1p(np.exp(-np.abs(z)))
ExampleRun on NumPy 2.5.3
z, y = 40.0, 0                          # very confident, and wrong
print("from y_hat:", bce(y, sigmoid(z)))
print("from z    :", bce_from_logits(y, z))
print("the four records from z:", bce_from_logits(np.array([1, 0, 1, 1]), np.array([2.2, -1.4, 0.0, -3.0])).round(3))

The slope a wrong record sends back

Why not train a classifier with MSE on ŷ? Compare the slope of each loss with respect to z for a record that is confidently wrong:

ExampleRun on NumPy 2.5.3
z, y = -6.0, 1                          # confidently wrong: y_hat is about 0.0025
s = sigmoid(z)
print("dL/dz, binary cross-entropy:", round(s - y, 4))
print("dL/dz, squared error       :", round(2 * (s - y) * s * (1 - s), 4))

What the losses printed

  • The four records get ŷ = 0.9, 0.198, 0.5 and 0.047. The last one is labelled 1 but predicted 0.047, so it costs 3.049, far more than the others (0.105, 0.22, 0.693). The cost is their mean, 1.017.
  • The table values come out of the same function: 0.105, 0.693 and 2.303, mirrored between the two labels.
  • From ŷ, z = 40 gives inf; from z it gives 40.0, the correct loss (about z for a large wrong z). The four records give the same losses from z as from ŷ.
  • The slope for a wrong record is −0.9975 with binary cross-entropy and −0.0049 with squared error. The sigmoid's slope σ(1 − σ) is tiny at z = −6 and multiplies the squared error's gradient, so the network barely learns from its worst mistakes. With binary cross-entropy the slope is ŷ − y, large exactly when the prediction is wrong.

Binary cross-entropy in Keras

The output layer gets one neuron with a sigmoid, and the loss is named in compile. With from_logits=True the last layer has no activation and the loss works from z, the stable form above:

python
import keras
from keras import layers

model = keras.Sequential([keras.Input(shape=(3,)),
                          layers.Dense(4, activation="relu"),
                          layers.Dense(1, activation="sigmoid")])   # ŷ between 0 and 1
model.compile(optimizer="adam", loss="binary_crossentropy", metrics=["accuracy"])

# or: no sigmoid in the last layer, and the loss applies it to z
# layers.Dense(1) with loss=keras.losses.BinaryCrossentropy(from_logits=True)

Binary cross-entropy vs MSE for classification

Binary cross-entropyMSE on a sigmoid output
Loss for y = 1, ŷ = 0.12.3030.81
Slope with respect to zŷ − y2(ŷ − y) · ŷ(1 − ŷ)
Confidently wrong recordlarge slope, fast correctionslope near 0, slow learning
Shape over a logistic model's weightsconvexnot convex

Where you use binary cross-entropy

  • Yes or no outputs: customer churn, spam detection, pass or fail, with one sigmoid neuron at the end.
  • Multi-label outputs: one sigmoid per label (a photo can contain a cat and a dog), with a binary cross-entropy for each label; see Choosing a loss function.
  • Imbalanced data: the same loss with class weights, so the rare class counts more.
Watch out. Do not apply the sigmoid twice. With from_logits=True the last layer must have no activation; a sigmoid layer plus from_logits=True squeezes the output again, and the loss and gradients are wrong. Training does not stop: at most a warning is printed, on the TensorFlow backend.
Try it yourself
  • Change the last label in y from 1 to 0: that record's loss drops from 3.049 to about 0.049.
  • Call bce(1, 1e-12) and bce_from_logits(1, -27.6): both give about 27.6, the loss for a near-zero probability.
  • Set z = 6.0 with y = 1 in the slope check: both slopes are close to 0, because the prediction is right.

You understood something today that you didn't yesterday.