Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Decision tree classifier

A decision tree is a supervised learning algorithm that predicts by asking a chain of yes or no questions about the features, the same way nested if-else statements do.

Last updated: 04 Oct, 2026 · scikit-learn 1.9.1

K nearest neighbours (KNN) compares every new point with all the training points. A decision tree learns a short set of questions once; answering them for a new point leads straight to the prediction. The same idea solves both classification and regression.

Nested if-else as a decision tree · from the Complete Machine Learning in 6 Hours video · 200:16 to 203:39

Turning nested if-else into a tree

The video starts from plain code. If age is 18 or less, print "College". Else, if age is above 18 and at most 35, print "Work". Otherwise print "Retire".

Each condition becomes a node. The first condition, age ≤ 18, is the root node, where every prediction starts. Its Yes branch ends in College: a node with no more questions is a leaf node, and it holds the answer. Its No branch goes to the next condition, 18 < age ≤ 35, whose Yes branch is the Work leaf and whose No branch is the Retire leaf.

The second node on the board reads "<18 & <=35"; the condition is age > 18 and age ≤ 35, as the video corrects aloud. The board also leaves the Retire leaf undrawn; it is the last No branch.

Nested if-else code for age drawn as a decision tree: the root asks age <= 18, Yes leads to the College leaf, and No leads to a node asking 18 < age <= 35 with Work and Retire leaves.
Splitting the play-tennis data on Outlook · from the Complete Machine Learning in 6 Hours video · 204:10 to 208:54

Splitting play tennis on Outlook

The video brings back the play-tennis table from the Naive Bayes worked example: 14 days, the inputs Outlook, Temperature, Humidity and Wind, and the output PlayTennis with 9 Yes and 5 No. The root node holds all 14 days, so it is written 9Y / 5N.

Pick Outlook as the first feature. It has three categories, so the root gets three children. Count the days in each: Sunny is 2 Yes and 3 No, Overcast is 4 Yes and 0 No, and Rain is 3 Yes and 2 No. The children add back up to 9 Yes and 5 No.

Why Outlook and not another feature? For now the video takes it at random and promises the answer later: Information gain is the rule that picks it. Overcast already holds only Yes days, so it needs no more questions. Sunny and Rain are mixed, so each is split again by another feature. How to measure that mix is the subject of Entropy and Gini impurity.

The play-tennis root Outlook with 9 Yes and 5 No splits into Sunny 2Y/3N and Rain 3Y/2N, which are split again, and Overcast 4Y/0N, which always answers Yes.

Learning the age rules with DecisionTreeClassifier

scikit-learn's DecisionTreeClassifier learns the questions from labelled examples. Give it a few ages labelled by the if-else rules and it finds the same two boundaries on its own.

The age rules as a function

python
def life_stage(age):
    if age <= 18:
        return "College"
    elif age > 18 and age <= 35:
        return "Work"
    else:
        return "Retire"

Learning the same rules from examples

python
from sklearn.tree import DecisionTreeClassifier, export_text

# Ten example people, labelled by the if-else rules
ages = [[10], [15], [18], [19], [25], [35], [36], [50], [60], [70]]
labels = [life_stage(a[0]) for a in ages]
age_tree = DecisionTreeClassifier(random_state=0).fit(ages, labels)
ExampleThe video's example, run on scikit-learn 1.9.1
print(export_text(age_tree, feature_names=["age"]))
for age in [16, 30, 40]:
    print(age, "if-else:", life_stage(age), "| tree:", age_tree.predict([[age]])[0])

What the learned age tree shows

  • The thresholds are 18.5 and 35.5: the tree puts each cut halfway between two training ages it has seen (18 and 19, 35 and 36), so it agrees with the if-else on every whole age.
  • The root asks age ≤ 35.5 first, not age ≤ 18. A tree is free to ask the questions in a different order, as long as the leaves give the same answers.
  • 16, 30 and 40 get College, Work and Retire from both the function and the tree.

Fitting a tree on the play-tennis table

scikit-learn trees need numbers, so each category becomes its own 0/1 column first (one-hot encoding).

Typing the play-tennis table

python
import pandas as pd

# The 14 days of the play-tennis table, one string per column
outlook = "Sunny Sunny Overcast Rain Rain Rain Overcast Sunny Sunny Rain Sunny Overcast Overcast Rain".split()
temperature = "Hot Hot Hot Mild Cool Cool Cool Mild Cool Mild Mild Mild Hot Mild".split()
humidity = "High High High High Normal Normal Normal High Normal Normal Normal High Normal High".split()
wind = "Weak Strong Weak Weak Weak Strong Strong Weak Weak Weak Strong Strong Weak Strong".split()
play = "No No Yes Yes Yes No Yes No Yes Yes Yes Yes Yes No".split()

df = pd.DataFrame({"Outlook": outlook, "Temperature": temperature,
                   "Humidity": humidity, "Wind": wind, "PlayTennis": play})

One-hot encoding the categories

python
from sklearn.tree import DecisionTreeClassifier, export_text

# One 0/1 column per category: Outlook_Sunny, Outlook_Overcast, ...
X = pd.get_dummies(df.drop(columns="PlayTennis"), dtype=int)
y = df["PlayTennis"]
ExampleFrom the video, run on scikit-learn 1.9.1
print(pd.crosstab(df["Outlook"], df["PlayTennis"]), "\n")

tennis_tree = DecisionTreeClassifier(criterion="entropy", random_state=0).fit(X, y)
print(export_text(tennis_tree, feature_names=list(X.columns)))
print("depth:", tennis_tree.get_depth(), "leaves:", tennis_tree.get_n_leaves(),
      "training accuracy:", tennis_tree.score(X, y))

What the play-tennis tree shows

  • The Outlook counts match the board: Overcast 0 No and 4 Yes, Rain 2 and 3, Sunny 3 and 2.
  • The first question is Outlook_Overcast ≤ 0.5: the tree starts by separating the 4 Overcast days, which all say Yes, from the rest. It picks Outlook first, as the video's tree does.
  • Every question has two answers: scikit-learn asks "is it Overcast or not" instead of splitting Outlook three ways, so the other branches need extra questions.
  • Depth 4, 7 leaves, training accuracy 1.0: the tree keeps splitting until every leaf is pure, so it fits all 14 days.

Multiway split vs binary split

Tree classifiers come in two families. ID3 gives a node one branch per category, as the video's Outlook tree does. CART (classification and regression trees) always makes two branches, and scikit-learn uses an optimised version of CART.

The video's tree (one branch per category)scikit-learn's tree (two branches)
Outlook at the rootSunny, Overcast and Rain at onceOvercast or not, then Rain or not later
Children per nodeAs many as the categoriesAlways 2
Numeric featuresThresholds, as in the information gain lessonThresholds such as age <= 18.5
InputCategory namesNumbers, so categories are one-hot encoded
Algorithm familyID3 styleCART (classification and regression trees)

Where you use decision trees

  • Rules people must read: loan approval or triage rules, where the reason for each prediction has to be shown as a list of questions.
  • Mixed tabular data: trees need no feature scaling, so ages, prices and 0/1 flags can sit side by side.
  • The building block of ensembles: random forest and boosting combine many trees.
Watch out. A tree with no limits keeps splitting until every leaf is pure, as the play-tennis tree did with training accuracy 1.0. That fits the noise in the training data and does worse on new data. Decision tree regression and pruning shows how to stop it.
Try it yourself
  • Add the ages 30 and 33 to the age examples. Do the thresholds 18.5 and 35.5 move?
  • Fit the play-tennis tree with max_depth=1. What is the single question, and what does it predict on each side?
  • Predict a new day with the play-tennis tree: build a one-row DataFrame for (Rain, Mild, High, Weak), run it through pd.get_dummies, and line its columns up with .reindex(columns=X.columns, fill_value=0) before predict.

Little by little, you're building something great.