Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Softmax

Softmax is an activation function for the output layer that turns K raw scores, one per class, into K probabilities that are positive and add up to 1.

Last updated: 05 Oct, 2026 · NumPy

The ReLU and its variants lesson covered the hidden layers. For a multi-class problem the output layer needs something else: one probability per class. The video sets softmax aside until the loss functions; this lesson works it out on the notes' four-class example.

Writing the softmax formula

The output layer has one neuron per class. Each gives a raw score z, called a logit: its weighted sum, with no activation yet. Softmax raises e to each logit and divides by the sum of all of them:

  • e^z is always positive, so every probability is above 0.
  • Dividing by the sum makes the K values add up to exactly 1.
  • A larger logit gets a larger share, and the exponential stretches the gaps: a logit 2 higher gets e² ≈ 7.4 times the share.

Working the cat, dog, monkey and horse example

The notes' network has four output neurons for four animals, with logits −1 for cat, 0 for dog, 3 for monkey and 5 for horse. The exponentials are e^(−1) = 0.3679, e^0 = 1, e^3 = 20.0855 and e^5 = 148.4132, and their sum is 169.87. Dividing each by the sum:

A last hidden layer feeds four output neurons with logits −1, 0, 3 and 5 for cat, dog, monkey and horse. A table lists e to the z, summing to 169.87, and the probabilities 0.00217, 0.00589, 0.11824 and 0.87370, which add up to 1.

The notes, like the video's board, write the denominator as e^(−1+0+3+5), e raised to the sum of the logits. The denominator is the sum of the exponentials, e^(−1) + e^0 + e^3 + e^5. With the wrong denominator the notes get 0.00033, 0.0024, 0.0183 and 0.1353, which do not add up to 1 (and 0.0024 for dog should be 0.00091 even on their own formula). Their last step divides 0.1353 by the sum of the four and gets about 86%; the correct probability for horse is 87.4%.

The video's own board example has the same denominator slip, e^10 / e^(10+20+30+40+50), and gives probabilities 0.4, 0.5 and 0.6, which add up to 1.5. It also feeds five numbers into two output neurons; softmax takes one logit per output neuron, and its outputs always add up to 1.

Computing softmax in NumPy

The softmax function

python
import numpy as np

def softmax(z):
    e = np.exp(z)                             # one exponential per class
    return e / e.sum()                        # divided by the SUM of the exponentials
ExampleFrom the video's notes, run on NumPy 2.5.3
z = np.array([-1.0, 0.0, 3.0, 5.0])           # cat, dog, monkey, horse
print("e^z:  ", np.round(np.exp(z), 4), " sum:", round(np.exp(z).sum(), 4))
p = softmax(z)
print("probs:", np.round(p, 5), " sum:", p.sum())
names = ["cat", "dog", "monkey", "horse"]
print("prediction:", names[p.argmax()])
print("the notes' denominator e^(sum of z):", np.round(np.exp(z) / np.exp(z.sum()), 5))
print("same probs after adding 10 to every z:", np.round(softmax(z + 10), 5))

What the probabilities show

  • Cat 0.00217, dog 0.00589, monkey 0.11824, horse 0.87370, and they add up to 1. The prediction is the class with the largest probability: horse.
  • The notes' denominator gives 0.00034, 0.00091, 0.01832 and 0.13534: these are e^(−8), e^(−7), e^(−4) and e^(−2), and they add up to about 0.155, not 1.
  • Adding 10 to every logit changes nothing: e^(z+10) = e^10 · e^z, and the e^10 cancels between top and bottom. Softmax only cares about the differences between logits.

Avoiding overflow with the max trick

That last property fixes a real bug. e^1000 is too large for a float, so the plain function fails on large logits:

ExampleRun on NumPy 2.5.3
print(softmax(np.array([1000.0, 1001.0])))

NumPy also prints an overflow warning: e^1000 becomes inf, and inf / inf is nan. Subtracting the largest logit first leaves the probabilities unchanged and keeps every exponent at 0 or below:

ExampleRun on NumPy 2.5.3
def softmax(z):
    e = np.exp(z - z.max())                   # subtract the largest score first
    return e / e.sum()

print(softmax(np.array([1000.0, 1001.0])))
print(np.round(softmax(np.array([-1.0, 0.0, 3.0, 5.0])), 5))

Checking that two-class softmax equals sigmoid

With K = 2, softmax gives P(class 1) = e^(z1) / (e^(z0) + e^(z1)). Divide top and bottom by e^(z1):

So a two-neuron softmax output and a one-neuron sigmoid output describe the same model; the sigmoid neuron's z plays the role of z1 − z0. The video's notebook shows the same derivation. The check below uses the stable softmax from the run above:

ExampleRun on NumPy 2.5.3
def sigmoid(t):
    return 1 / (1 + np.exp(-t))

for z0, z1 in [(0.0, 2.0), (1.5, -0.5), (3.0, 3.0)]:
    p = softmax(np.array([z0, z1]))
    print(f"z = ({z0}, {z1}): softmax P(class 1) = {p[1]:.5f}   sigmoid(z1 - z0) = {sigmoid(z1 - z0):.5f}")

Softmax vs sigmoid outputs

Softmax outputSigmoid output
NeuronsK, one per class1 (binary) or K (multi-label)
Each valuebetween 0 and 1between 0 and 1
Values add up to 1yesno, each is separate
Problemmulti-class: exactly one class is truebinary, or multi-label: several classes can be true
Pairs with the losscategorical cross-entropybinary cross-entropy

Where you use softmax

  • Multi-class classification: the output layer of a digit classifier (10 classes) or an animal classifier, in Keras Dense(K, activation="softmax").
  • Reading a model's confidence: the largest probability is the prediction, and the gap to the second tells how sure the model is.
  • Language models: the next-word probabilities over a whole vocabulary come from a softmax over one logit per word.
Watch out. Use softmax only when exactly one class is true. When an image can be both a "dog" and "outdoors", softmax forces the two to compete for one share of 1; that multi-label case needs one sigmoid per class.
Try it yourself
  • Change horse's logit from 5 to 3, the same as monkey's. What are the two probabilities now?
  • Multiply every logit by 2 with softmax(2 * z). Does horse's probability go up or down, and why is this different from adding 2?
  • Call the plain softmax on np.array([-1000.0, -1001.0]). What goes wrong this time, and does the max trick fix it?

Slow is fine. Stopping is the only problem.