Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Random forest

Random forest is a bagging algorithm that trains many decision trees, each on its own sample of rows and features, and combines their outputs by majority vote for classification or by the average for regression.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Bagging and boosting combined any models with a vote. A random forest uses one kind of model, the decision tree, and is the bagging algorithm used most often.

The problem with one decision tree · from the Complete Machine Learning in 6 Hours video · 261:28 to 264:23

Fixing the high variance of a decision tree

A Decision tree classifier grown with no hyperparameters keeps splitting until every leaf is pure. That leads to overfitting: the accuracy is high on the training data and lower on the test data. In the terms of Bias and variance, one full tree has low bias and high variance.

Pruning can help, but it is hard work with 100 features or a million rows, and pre-pruning gives no guarantee. The goal is a generalised model with low bias and low variance. A random forest gets there by training many trees and taking the majority vote: each tree has high variance, and the vote over many of them turns that into low variance.

Row sampling and feature sampling in a random forest · from the Complete Machine Learning in 6 Hours video · 264:23 to 267:34

Sampling rows and features for every tree

Every model M1, M2, M3, M4 in a random forest is a decision tree. Each one gets row sampling plus feature sampling: a sample of the rows and a sample of the features. Rows can repeat between trees, and so can features. A tree trained on its own slice of the data becomes an expert on that slice, the way the physics and chemistry teachers each know their own subject.

For a new test point the four trees on the board predict 0, 1, 0 and 0. The majority vote gives 0. With four features, one tree might get two of them, another three, another all four. The regressor is the same forest with one change: the output is the average of the trees' outputs.

Left: one fully grown decision tree has low bias and high variance. Right: a random forest gives each of four decision trees its own sample of rows and features; their outputs 0, 1, 0, 0 give a majority vote of 0, and a regressor averages the outputs.

scikit-learn draws the feature sample at every split of a tree, not once per tree. The share is set by max_features; for the classifier the default is the square root of the number of features.

Comparing one deep tree with a forest

The breast cancer dataset has 569 tumours, 30 measurements each, labelled malignant or benign. One fully grown tree and a forest of 100 trees train on the same 70% and are tested on the same 30%.

Building the forest with RandomForestClassifier

python
from sklearn.ensemble import RandomForestClassifier

# 100 trees; each sees a bootstrap sample of rows, and each split tries sqrt(30) of the features
forest = RandomForestClassifier(n_estimators=100, max_features="sqrt", random_state=0)
forest.fit(X_train, y_train)

Measuring variance across ten splits

Variance is how much a model changes when the data changes. Training each model on ten different train and test splits and looking at the spread of its test accuracy measures it.

python
accs = []
for seed in range(10):
    a, b, c, d = train_test_split(X, y, test_size=0.3, random_state=seed)
    accs.append(Model(random_state=0).fit(a, c).score(b, d))
print(np.mean(accs), np.std(accs))   # the average and the spread
ExampleRun on scikit-learn 1.9.1
import numpy as np
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0)

tree = DecisionTreeClassifier(random_state=0).fit(X_train, y_train)
forest = RandomForestClassifier(n_estimators=100, max_features="sqrt", random_state=0)
forest.fit(X_train, y_train)
for name, m in [("one deep tree", tree), ("random forest", forest)]:
    print(f"{name:14} train {m.score(X_train, y_train):.3f}  test {m.score(X_test, y_test):.3f}")

for name, Model in [("one deep tree", DecisionTreeClassifier), ("random forest", RandomForestClassifier)]:
    accs = []
    for seed in range(10):
        a, b, c, d = train_test_split(X, y, test_size=0.3, random_state=seed)
        accs.append(Model(random_state=0).fit(a, c).score(b, d))
    print(f"{name:14} 10 splits: mean {np.mean(accs):.3f}, spread {np.std(accs):.3f}")

What the tree and the forest scores mean

  • Both score 1.000 on the training data. Every tree in the forest is grown fully too, so low bias holds for both.
  • The test accuracy differs: 0.912 for the tree, 0.959 for the forest. The gap between train and test is the overfitting the video describes.
  • The spread over ten splits is 0.017 for the tree and 0.012 for the forest, and the forest's average is higher (0.960 against 0.926). That is the high variance turned into lower variance.
Random forest interview questions · from the Complete Machine Learning in 6 Hours video · 267:34 to 269:36

Answering the random forest interview questions

The video closes with the questions interviewers ask most often about random forests.

  1. Is normalisation needed for a random forest or a decision tree? No. A split compares one feature with a threshold, and scaling the feature moves the threshold with it, so the same rows land on each side.
  2. Is standardisation needed for KNN? Yes. K nearest neighbours (KNN) measures Euclidean or Manhattan distance, and a feature with large numbers takes over the distance.
  3. Is a random forest affected by outliers? Mostly no, for outliers in the features: a split only cares about the order of the values, so one huge value ends up on one side like any other large value. Outliers in the target of a RandomForestRegressor do pull the average in their leaf.
  4. Is KNN affected by outliers? Yes. An outlier sits among the neighbours of nearby points and shifts the distances.

Bagging is not limited to forests. A custom bagging model uses any combination of algorithms and combines their outputs, as the M1 to M4 example in Bagging and boosting did.

Scaling the features for a forest and for KNN

python
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler().fit(X_train)         # learn the mean and std on the training rows
Xs_train, Xs_test = scaler.transform(X_train), scaler.transform(X_test)
ExampleThe video's interview questions, run on scikit-learn 1.9.1
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier

scaler = StandardScaler().fit(X_train)
Xs_train, Xs_test = scaler.transform(X_train), scaler.transform(X_test)
for name, m in [("random forest", RandomForestClassifier(random_state=0)), ("KNN", KNeighborsClassifier())]:
    raw = m.fit(X_train, y_train).score(X_test, y_test)
    scaled = m.fit(Xs_train, y_train).score(Xs_test, y_test)
    print(f"{name:13} raw {raw:.3f}   scaled {scaled:.3f}")

What scaling does to the forest and to KNN

  • The forest scores 0.959 raw and scaled. Scaling changed every number in the data and none of the tree splits.
  • KNN moves from 0.947 to 0.959 once the features are standardised, because its distances change.
Vote or average
The video describes a majority vote over the trees' classes. scikit-learn's RandomForestClassifier averages the trees' class probabilities instead and picks the class with the highest average (a soft vote). The answer is almost always the same; the difference shows only when the vote is close.

Predicting used car prices with RandomForestRegressor

A random forest regressor is the same forest with one change: every tree predicts a number, and the forest returns the average of the trees' numbers. The used car dataset has 15,411 cars sold on cardekho.com, with each car's model, age, kilometres driven, seller type, fuel, transmission, mileage, engine, power, seats and its selling price in rupees. The task is to predict the selling price.

Loading the car data and encoding the text columns

The file loads from its URL. car_name and brand repeat what model already says, so they are dropped. model has 120 values and becomes one number per model (label encoding). Seller type, fuel type and transmission have 2 to 5 values each and become 0 or 1 columns (one-hot encoding). The forest needs no scaling.

python
url = ("https://raw.githubusercontent.com/krishnaik06/The-Grand-Complete-Data-Science-Materials/"
       "main/Machine%20Learning/8-Random%20Forest/Projects/data/cardekho_imputated.csv")
df = pd.read_csv(url, index_col=[0]).drop(columns=["car_name", "brand"])
df["model"] = df["model"].astype("category").cat.codes        # one number per car model
X = pd.get_dummies(df.drop(columns="selling_price"), drop_first=True,
                   columns=["seller_type", "fuel_type", "transmission_type"])
y = df["selling_price"]

Averaging the trees' prices for one car

A fitted forest keeps its trees in estimators_, so the average can be checked by hand.

python
forest = RandomForestRegressor(n_estimators=100, random_state=0).fit(X_train, y_train)
car = X_test.iloc[[0]].to_numpy()
each = [tree.predict(car)[0] for tree in forest.estimators_]   # 100 prices for one car
ExampleThe used car dataset, run on scikit-learn 1.9.1
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import RandomForestRegressor

url = ("https://raw.githubusercontent.com/krishnaik06/The-Grand-Complete-Data-Science-Materials/"
       "main/Machine%20Learning/8-Random%20Forest/Projects/data/cardekho_imputated.csv")
df = pd.read_csv(url, index_col=[0]).drop(columns=["car_name", "brand"])
df["model"] = df["model"].astype("category").cat.codes        # one number per car model
X = pd.get_dummies(df.drop(columns="selling_price"), drop_first=True,
                   columns=["seller_type", "fuel_type", "transmission_type"])
y = df["selling_price"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print("cars:", len(df), "  columns after encoding:", X.shape[1])

models = {"linear regression": LinearRegression(),
          "one deep tree": DecisionTreeRegressor(random_state=0),
          "random forest": RandomForestRegressor(n_estimators=100, random_state=0)}
for name, m in models.items():
    m.fit(X_train, y_train)
    print(f"{name:17} train R2 {m.score(X_train, y_train):.3f}   test R2 {m.score(X_test, y_test):.3f}")

forest = models["random forest"]
car = X_test.iloc[[0]].to_numpy()
each = [tree.predict(car)[0] for tree in forest.estimators_]
print("first four trees for the first test car:", [int(v) for v in each[:4]])
print("mean of the 100 trees:", int(np.mean(each)), "  forest prediction:", int(forest.predict(car)[0]))
print("true price:", int(y_test.iloc[0]))

What the car price models show

  • Linear regression scores a test R² of 0.665: one straight line cannot follow how the price depends on the model, age and power together.
  • One deep tree fits the training cars almost exactly (0.999) and drops to 0.880 on the test cars, the low bias and high variance of a single tree.
  • The forest scores 0.931 on the test cars, the best of the three, with a training R² of 0.981.
  • The trees disagree about the first test car (211,000, 211,000, 235,000 and 335,000 rupees from the first four), and the mean of all 100 trees, 246,180, is exactly the forest's prediction. The car sold for 190,000.

Scoring a forest with out-of-bag rows

Each tree's bootstrap sample makes as many draws as there are rows, d, but with replacement: some rows are drawn twice and others never. The rows a tree never draws are its out-of-bag (OOB) rows. They work like a validation set for that tree. Scoring every training row with only the trees that did not train on it gives the out-of-bag score, a test score that needs no rows set aside.

A dataset of d rows sends a bootstrap sample of rows and features to each of four trees DT1 to DT4; the rows a tree never draws, about one third of d, are its out-of-bag rows, and scoring every row with only the trees that did not see it gives the out-of-bag score, like a validation set of about 368 out of 1,000 rows.

How many rows stay out? One draw misses a given row with probability 1 − 1/d, so all d draws miss it with probability (1 − 1/d)d. For a large d this tends to e−1 ≈ 0.368: about 63.2% of the rows go into the bag and 36.8% stay out. The board rounds it to one third, as in a split of 1,000 training rows into 2/3 for training and 1/3 for validation.

Counting the rows a bootstrap sample leaves out

ExampleRun with NumPy
import numpy as np

rng = np.random.default_rng(0)
for d in (10, 100, 1000, 10000):
    rows = rng.choice(d, size=d, replace=True)          # one bootstrap sample of d draws
    oob = d - len(np.unique(rows))                      # rows that were never drawn
    print(f"d = {d:>5}: out of bag {oob:>4} rows = {oob / d:.3f}   formula {(1 - 1 / d) ** d:.3f}")
print("e^-1 =", round(float(np.exp(-1)), 3))

Reading oob_score_ on the car forest

oob_score=True makes the forest remember which rows each tree left out. After fitting, oob_score_ holds the R² of the out-of-bag predictions and oob_prediction_ the predictions themselves, all computed from the training rows alone.

ExampleThe used car forest, run on scikit-learn 1.9.1
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_absolute_error

forest = RandomForestRegressor(n_estimators=100, oob_score=True, random_state=0)
forest.fit(X_train, y_train)
oob_mae = mean_absolute_error(y_train, forest.oob_prediction_)
test_mae = mean_absolute_error(y_test, forest.predict(X_test))
print(f"out-of-bag: R2 {forest.oob_score_:.3f}   MAE {oob_mae:,.0f}")
print(f"test:       R2 {forest.score(X_test, y_test):.3f}   MAE {test_mae:,.0f}")
print(f"most expensive car: training rows {y_train.max():,}   test rows {y_test.max():,}")

What the out-of-bag numbers show

  • The share left out settles at 0.368: 370 of 1,000 rows and 3,658 of 10,000, the e−1 of the formula. With only 10 rows the share is rough (3 rows here, against 0.349 from the formula).
  • The out-of-bag MAE, 102,415 rupees, is close to the test MAE of 101,812, so the training rows alone gave a fair estimate of the typical error on new cars.
  • The out-of-bag R² is lower, 0.843 against 0.931. The training rows hold the most expensive car, 39,500,000 rupees, and the test rows stop at 13,000,000. R² squares the errors, so a few very expensive training cars weigh heavily in it, while the typical error in rupees is about the same.

Random forest vs one decision tree

One decision treeRandom forest
Bias and variancelow bias, high variancelow bias, lower variance
Data per treeall rows, all featuresbootstrap rows, a sample of features at each split
Breast cancer test accuracy0.9120.959
Used car price test R²0.8800.931
Can you read it?yes, a white box modelno, 100 trees make a black box model
Training timeone treeabout n_estimators times one tree

Where you use a random forest

  • A first strong model on tabular data, since it needs no scaling and little tuning.
  • Feature importance: forest.feature_importances_ ranks which columns the splits used most.
  • Data with outliers in the features, where distance-based models such as KNN struggle.
Watch out. A random forest still overfits noisy data if its trees are very deep and the dataset is small. A training accuracy of 1.000 says nothing about test accuracy; check the test score or oob_score_, and tune max_depth or min_samples_leaf if the gap is large.
Try it yourself
  • Change max_features="sqrt" to max_features=None so every split sees all 30 features, and compare the forest's test accuracy.
  • Set n_estimators=300 in the out-of-bag forest and check whether oob_score_ moves away from 0.843.
  • Print forest.feature_importances_.argmax() to find the feature the trees used most.

You understood something today that you didn't yesterday.