Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Entropy and Gini impurity

Entropy and Gini impurity are purity measures that score how mixed the classes in a decision tree node are: 0 for a pure node, and higher the more evenly the classes are mixed.

Last updated: 04 Oct, 2026 · scikit-learn 1.9.1

The Decision tree classifier split play tennis on Outlook and got one node with only Yes days and two mixed ones. To decide whether to split a node again, the tree needs a number for that mix.

Pure and impure splits · from the Complete Machine Learning in 6 Hours video · 209:02 to 212:55

Telling a pure split from an impure split

Look again at the three Outlook nodes. Overcast is 4 Yes and 0 No. Every Overcast day is a Yes, so if day 15 is Overcast, the answer is Yes without asking anything else. A node whose records all have one class is pure; it becomes a pure leaf node and is not split again.

Sunny is 2 Yes and 3 No, and Rain is 3 Yes and 2 No. Mixed nodes are impure. The tree takes another feature (Temperature, say, under Sunny) and keeps splitting until it reaches pure leaves.

That leaves two questions. How do we measure purity? With entropy or Gini impurity, both covered here. Which feature do we split on? With information gain, covered in Information gain. The board writes "Gini coefficient"; the right name, as the video says, is Gini impurity.

Entropy of a pure node · from the Complete Machine Learning in 6 Hours video · 212:55 to 216:44

Computing entropy for a pure node

For two classes, write p₊ for the probability of Yes in the node and p₋ for the probability of No. The entropy of the sample S is:

The video's example: a feature f1 at the root holds 6 Yes and 3 No. It has two categories, so it splits into C1 with 3 Yes and 3 No and C2 with 3 Yes and 0 No. The children add back up: 3 + 3 = 6 Yes and 3 + 0 = 3 No.

For C2, p₊ = 3/3 = 1 and p₋ = 0/3 = 0:

log₂ 0 itself is undefined; the term 0 · log₂ 0 is taken as 0, because p · log₂ p shrinks to 0 as p does. So a pure split always has entropy 0.

A root f1 with 6 Yes and 3 No splits into C1 with 3 Yes and 3 No, an impure node with entropy 1 and Gini 0.5, and C2 with 3 Yes and 0 No, a pure node with entropy 0 and Gini 0.
The entropy curve · from the Complete Machine Learning in 6 Hours video · 216:44 to 220:34

Reading the entropy curve

Now C1, with 3 Yes and 3 No. Here p₊ = 3/6 = 0.5 and p₋ = 0.5 (since p₋ = 1 − p₊):

Plot H(S) against p₊ and you get the board's curve. It is 0 at p₊ = 0 and at p₊ = 1 (pure nodes), and it peaks at 1 when p₊ = 0.5, the most mixed a two-class node can be. For two classes, entropy always lies between 0 and 1. For any other mix you read the value off the curve: at p₊ = 0.3 it is about 0.88.

Gini impurity and when to use it · from the Complete Machine Learning in 6 Hours video · 229:10 to 233:32

Computing Gini impurity

Gini impurity adds up the squared class probabilities and subtracts them from 1. With n output classes:

The video's example is a node with 2 Yes and 2 No, so p₊ = p₋ = 1/2:

The same node has entropy 1. So for two classes, Gini impurity runs from 0 (pure) to 0.5 (a 50/50 mix), and its curve sits under the entropy curve, peaking at 0.5 at p₊ = 0.5.

Both measures work for more than two classes. Gini impurity already sums over all n classes. Entropy adds one term per class, so for an output with three categories c₁, c₂ and c₃ it reads:

Choosing between Gini and entropy

The difference is speed. Entropy needs a logarithm; Gini needs only squares. A tree computes impurity again and again, for every candidate split of every feature, so with many features (100 or 200) the cheaper formula adds up. The video's advice: with many features use Gini impurity, with a small set entropy is fine. scikit-learn's DecisionTreeClassifier uses Gini by default.

The number of records points the same way. Every record adds candidate splits to score, so a small dataset can afford entropy, and a large one is a reason to prefer Gini impurity.

The board writes "Gini >> Entropy" for speed; Gini is faster, but by a modest margin, and the two measures often choose the same splits.

Computing entropy and Gini in Python

An entropy function

python
import numpy as np

def entropy(counts):
    p = np.array(counts) / sum(counts)
    p = p[p > 0]                      # 0 * log2(0) counts as 0
    return float((p * np.log2(1 / p)).sum())   # same as -sum(p * log2(p))

A Gini impurity function

python
def gini(counts):
    p = np.array(counts) / sum(counts)
    return float(1 - (p ** 2).sum())
ExampleFrom the video, run on scikit-learn 1.9.1
nodes = {"f1 root 6Y/3N": [6, 3], "C1 3Y/3N": [3, 3], "C2 3Y/0N": [3, 0],
         "2Y/2N": [2, 2], "p+ = 0.3": [3, 7], "play tennis 9Y/5N": [9, 5],
         "three equal classes": [1, 1, 1]}
for name, counts in nodes.items():
    print(f"{name:20} entropy {entropy(counts):.3f}   gini {gini(counts):.3f}")

Plotting the entropy and Gini curves

ExampleFrom the video's board, run on matplotlib 3.11.2
import numpy as np
import matplotlib.pyplot as plt

p = np.linspace(0.001, 0.999, 300)                   # P(+), the share of Yes
H = -p * np.log2(p) - (1 - p) * np.log2(1 - p)       # entropy
G = 1 - p ** 2 - (1 - p) ** 2                        # Gini impurity

plt.figure(figsize=(7, 4))
plt.plot(p, H, label="Entropy H(S)")
plt.plot(p, G, label="Gini impurity", color="green")
plt.axvline(0.5, linestyle="--", color="grey")
plt.title("Entropy and Gini impurity for two classes")
plt.xlabel("P(+)")
plt.ylabel("Impurity")
plt.legend()
plt.show()
print("peak entropy:", round(H.max(), 3), " peak gini:", round(G.max(), 3))
The entropy curve rises from 0 to a peak of 1 at P(+) = 0.5 and falls back to 0; the Gini curve has the same shape with a peak of 0.5.

What the impurity values show

  • C2 3Y/0N gives 0 and 0: a pure node has no impurity on either measure.
  • C1 3Y/3N gives entropy 1 and Gini 0.5: the two-class maximums, as on the board.
  • The root 6Y/3N gives 0.918 and 0.444: a mix between pure and 50/50. Play tennis's 9Y/5N gives 0.940, the 0.94 used in the information gain lesson.
  • p₊ = 0.3 gives entropy 0.881: the value read off the curve.
  • Three equal classes give entropy 1.585 and Gini 0.667: with more classes both maximums grow. The 0 to 1 and 0 to 0.5 ranges hold for two classes only.

Entropy vs Gini impurity

EntropyGini impurity
Formula−Σ pᵢ log₂ pᵢ1 − Σ pᵢ²
Pure node00
50/50, two classes10.5
Three equal classeslog₂ 3 = 1.5851 − 3 × (1/3)² = 0.667
CostNeeds a logarithmOnly squares, a little faster
In DecisionTreeClassifiercriterion="entropy" or "log_loss"criterion="gini" (the default)

Where you use entropy and Gini impurity

  • Choosing the criterion: DecisionTreeClassifier(criterion="gini") or criterion="entropy"; in scikit-learn, "entropy" and "log_loss" both mean the Shannon entropy.
  • Reading a fitted tree: plot_tree prints the Gini value of every node, so you can see where the tree is still unsure.
  • Scoring a split: information gain subtracts the children's entropy from the parent's.
Watch out. The ceilings 1 for entropy and 0.5 for Gini are for two classes. With three classes a 50/50/50 node has Gini 0.667 and entropy 1.585, so a Gini value above 0.5 on a three-class tree is correct, not a bug.
Try it yourself
  • Add "Sunny 2Y/3N": [2, 3] to the nodes and compare it with Rain's [3, 2]. Why are they equal?
  • Try four equal classes, [1, 1, 1, 1]. Is the Gini value 1 − 4 × (1/4)² = 0.75?
  • Mark p₊ = 0.3 on the plot with plt.scatter([0.3], [0.881], color="red") before plt.show().

You understood something today that you didn't yesterday.