Softmax
Softmax is an activation function for the output layer that turns K raw scores, one per class, into K probabilities that are positive and add up to 1.
Last updated: 05 Oct, 2026 · NumPy
The ReLU and its variants lesson covered the hidden layers. For a multi-class problem the output layer needs something else: one probability per class. The video sets softmax aside until the loss functions; this lesson works it out on the notes' four-class example.
Writing the softmax formula
The output layer has one neuron per class. Each gives a raw score z, called a logit: its weighted sum, with no activation yet. Softmax raises e to each logit and divides by the sum of all of them:
- e^z is always positive, so every probability is above 0.
- Dividing by the sum makes the K values add up to exactly 1.
- A larger logit gets a larger share, and the exponential stretches the gaps: a logit 2 higher gets e² ≈ 7.4 times the share.
Working the cat, dog, monkey and horse example
The notes' network has four output neurons for four animals, with logits −1 for cat, 0 for dog, 3 for monkey and 5 for horse. The exponentials are e^(−1) = 0.3679, e^0 = 1, e^3 = 20.0855 and e^5 = 148.4132, and their sum is 169.87. Dividing each by the sum:

The notes, like the video's board, write the denominator as e^(−1+0+3+5), e raised to the sum of the logits. The denominator is the sum of the exponentials, e^(−1) + e^0 + e^3 + e^5. With the wrong denominator the notes get 0.00033, 0.0024, 0.0183 and 0.1353, which do not add up to 1 (and 0.0024 for dog should be 0.00091 even on their own formula). Their last step divides 0.1353 by the sum of the four and gets about 86%; the correct probability for horse is 87.4%.
The video's own board example has the same denominator slip, e^10 / e^(10+20+30+40+50), and gives probabilities 0.4, 0.5 and 0.6, which add up to 1.5. It also feeds five numbers into two output neurons; softmax takes one logit per output neuron, and its outputs always add up to 1.
Computing softmax in NumPy
The softmax function
import numpy as np
def softmax(z):
e = np.exp(z) # one exponential per class
return e / e.sum() # divided by the SUM of the exponentialsz = np.array([-1.0, 0.0, 3.0, 5.0]) # cat, dog, monkey, horse
print("e^z: ", np.round(np.exp(z), 4), " sum:", round(np.exp(z).sum(), 4))
p = softmax(z)
print("probs:", np.round(p, 5), " sum:", p.sum())
names = ["cat", "dog", "monkey", "horse"]
print("prediction:", names[p.argmax()])
print("the notes' denominator e^(sum of z):", np.round(np.exp(z) / np.exp(z.sum()), 5))
print("same probs after adding 10 to every z:", np.round(softmax(z + 10), 5))e^z: [ 0.3679 1. 20.0855 148.4132] sum: 169.8666 probs: [0.00217 0.00589 0.11824 0.8737 ] sum: 1.0 prediction: horse the notes' denominator e^(sum of z): [0.00034 0.00091 0.01832 0.13534] same probs after adding 10 to every z: [0.00217 0.00589 0.11824 0.8737 ]
What the probabilities show
- Cat 0.00217, dog 0.00589, monkey 0.11824, horse 0.87370, and they add up to 1. The prediction is the class with the largest probability: horse.
- The notes' denominator gives 0.00034, 0.00091, 0.01832 and 0.13534: these are e^(−8), e^(−7), e^(−4) and e^(−2), and they add up to about 0.155, not 1.
- Adding 10 to every logit changes nothing: e^(z+10) = e^10 · e^z, and the e^10 cancels between top and bottom. Softmax only cares about the differences between logits.
Avoiding overflow with the max trick
That last property fixes a real bug. e^1000 is too large for a float, so the plain function fails on large logits:
print(softmax(np.array([1000.0, 1001.0])))[nan nan]
NumPy also prints an overflow warning: e^1000 becomes inf, and inf / inf is nan. Subtracting the largest logit first leaves the probabilities unchanged and keeps every exponent at 0 or below:
def softmax(z):
e = np.exp(z - z.max()) # subtract the largest score first
return e / e.sum()
print(softmax(np.array([1000.0, 1001.0])))
print(np.round(softmax(np.array([-1.0, 0.0, 3.0, 5.0])), 5))[0.26894142 0.73105858] [0.00217 0.00589 0.11824 0.8737 ]
Checking that two-class softmax equals sigmoid
With K = 2, softmax gives P(class 1) = e^(z1) / (e^(z0) + e^(z1)). Divide top and bottom by e^(z1):
So a two-neuron softmax output and a one-neuron sigmoid output describe the same model; the sigmoid neuron's z plays the role of z1 − z0. The video's notebook shows the same derivation. The check below uses the stable softmax from the run above:
def sigmoid(t):
return 1 / (1 + np.exp(-t))
for z0, z1 in [(0.0, 2.0), (1.5, -0.5), (3.0, 3.0)]:
p = softmax(np.array([z0, z1]))
print(f"z = ({z0}, {z1}): softmax P(class 1) = {p[1]:.5f} sigmoid(z1 - z0) = {sigmoid(z1 - z0):.5f}")z = (0.0, 2.0): softmax P(class 1) = 0.88080 sigmoid(z1 - z0) = 0.88080 z = (1.5, -0.5): softmax P(class 1) = 0.11920 sigmoid(z1 - z0) = 0.11920 z = (3.0, 3.0): softmax P(class 1) = 0.50000 sigmoid(z1 - z0) = 0.50000
Softmax vs sigmoid outputs
| Softmax output | Sigmoid output | |
|---|---|---|
| Neurons | K, one per class | 1 (binary) or K (multi-label) |
| Each value | between 0 and 1 | between 0 and 1 |
| Values add up to 1 | yes | no, each is separate |
| Problem | multi-class: exactly one class is true | binary, or multi-label: several classes can be true |
| Pairs with the loss | categorical cross-entropy | binary cross-entropy |
Where you use softmax
- Multi-class classification: the output layer of a digit classifier (10 classes) or an animal classifier, in Keras
Dense(K, activation="softmax"). - Reading a model's confidence: the largest probability is the prediction, and the gap to the second tells how sure the model is.
- Language models: the next-word probabilities over a whole vocabulary come from a softmax over one logit per word.
Related
- Previous: ReLU and its variants
- Next: Choosing an activation function
- Change horse's logit from 5 to 3, the same as monkey's. What are the two probabilities now?
- Multiply every logit by 2 with
softmax(2 * z). Does horse's probability go up or down, and why is this different from adding 2? - Call the plain softmax on
np.array([-1000.0, -1001.0]). What goes wrong this time, and does the max trick fix it?
Slow is fine. Stopping is the only problem.