Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Bagging and boosting

Bagging and boosting are ensemble techniques that combine several models to solve one problem: bagging trains the models side by side on samples of the rows, and boosting trains them one after another so each one covers what the last one missed.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Every algorithm so far solved a classification or a regression problem with one model at a time. An ensemble asks a different question: can several models work on the same problem and give a better answer together than any one of them alone?

Ensemble techniques and bagging · from the Complete Machine Learning in 6 Hours video · 249:46 to 254:16

Combining models with bagging

The video starts from a dataset D with features f1, f2, f3 and an output column. Bagging builds several models, M1, M2, M3 and M4. They do not have to be the same algorithm: on the board M1 is logistic regression, M2 a decision tree, M3 KNN and M4 a decision tree or naive Bayes.

Each model trains on its own sample of the rows. If D has 10,000 rows, M1 might get d′ = 1,000 of them, M2 another 1,000 called d″, then d‴ and d⁗. This is row sampling, and a sample is always much smaller than D (d′ ≪ D). Some rows show up in more than one sample. The sampling is done with replacement: a row that was picked goes back into the pool and can be picked again. A sample drawn that way is called a bootstrap sample.

Majority vote and mean in bagging · from the Complete Machine Learning in 6 Hours video · 254:16 to 257:30

Aggregating the outputs with a majority vote or a mean

Once every model is trained, a new test point goes to all four of them. On the board M1 predicts 0, and M2, M3 and M4 predict 1. The outputs are aggregated by majority voting: three of the four models say 1, so the answer is 1. Sampling with replacement and then aggregating gives the method its full name, bootstrap aggregating, shortened to bagging.

What happens when two models say 0 and two say 1? In practice a bagging model has 100 to 200 models, so a tie is very unlikely. A random forest uses only decision trees, but bagging as an idea can mix different algorithms, as the board does.

For a regression problem each model outputs a number. On the board they give 120, 140, 122 and 148, and the bagging output is their mean.

A dataset D of 10,000 rows sends four row samples of 1,000 rows to models M1 to M4, whose classes 0, 1, 1, 1 give a majority vote of 1 and whose values 120, 140, 122, 148 give a mean of 132.5.
Boosting with weak learners in sequence · from the Complete Machine Learning in 6 Hours video · 257:30 to 261:20

Chaining weak learners with boosting

In bagging the models are parallel and independent of each other. Boosting is a sequential combination: the training data goes to M1, then to M2, then M3, then M4, and each model hands on to the next.

Every model in the chain is a weak learner: on its own its predictions are poor. Combined in sequence, the weak learners become a strong learner. The video's example is a group of teachers. If the physics teacher cannot solve a problem, the chemistry, maths or geography teacher may, and together their expertise gives a good answer.

Boosting chains four weak learners M1 to M4 one after another into a strong learner, like a physics, chemistry, maths and geography teacher each helping with what the last one missed.

The board ends with the algorithms of each family. Bagging gives the Random forest classifier and regressor. Boosting gives AdaBoost, Gradient boosting and extreme gradient boosting, XGBoost classifier.

Computing the board's vote and mean

The four outputs from the board, for classification and for regression, combined the way the video combines them.

ExampleFrom the video, run on scikit-learn 1.9.1
from collections import Counter
import numpy as np

# The four models' outputs for one test point, from the board
class_votes = [0, 1, 1, 1]        # M1 logistic, M2 tree, M3 KNN, M4 naive Bayes
values = [120, 140, 122, 148]     # the same four models on a regression problem

# Classification: the class that most models predict
winner, count = Counter(class_votes).most_common(1)[0]
print("majority vote:", winner, f"({count} of {len(class_votes)} models)")

# Regression: the mean of the outputs
print("mean:", np.mean(values))

Bagging four different models on 10,000 rows

The board's setup can run as it is drawn: a dataset of 10,000 rows, four different algorithms, and a bootstrap sample of 1,000 rows for each. The data comes from make_classification, which generates a labelled dataset of any size.

Drawing a bootstrap sample of 1,000 rows

rng.choice with replace=True picks 1,000 row numbers and lets a number repeat, which is sampling with replacement.

python
rng = np.random.default_rng(0)
# 1,000 row numbers out of the training rows; a number can repeat
rows = rng.choice(len(X_train), size=1000, replace=True)
model.fit(X_train[rows], y_train[rows])

Taking the majority vote of four predictions

Stack the four models' predictions into one array, one row per model. A test point is class 1 when at least three of the four models say 1. Two against two is a tie, which this code sends to 0.

python
votes = np.array(preds)                   # shape: (4 models, number of test rows)
ones = votes.sum(axis=0)                  # how many models said 1, per test row
majority = (ones >= 3).astype(int)        # 3 or 4 votes for 1 -> class 1
ExampleThe board's setup (10,000 rows, samples of 1,000) on generated data, run on scikit-learn 1.9.1
import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.neighbors import KNeighborsClassifier
from sklearn.naive_bayes import GaussianNB

X, y = make_classification(n_samples=10000, n_features=10, n_informative=5, random_state=0)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)

models = {"M1 logistic": LogisticRegression(), "M2 tree": DecisionTreeClassifier(random_state=0),
          "M3 KNN": KNeighborsClassifier(), "M4 naive Bayes": GaussianNB()}
rng = np.random.default_rng(0)
preds = []
for name, model in models.items():
    rows = rng.choice(len(X_train), size=1000, replace=True)   # d' = 1,000 rows
    model.fit(X_train[rows], y_train[rows])
    preds.append(model.predict(X_test))
    print(f"{name:15} test accuracy {model.score(X_test, y_test):.3f}")

votes = np.array(preds)
ones = votes.sum(axis=0)
majority = (ones >= 3).astype(int)
print("votes for the first test point:", votes[:, 0], "-> majority", majority[0])
print("unique rows in the last sample:", len(np.unique(rows)), "of 1000")
print("ties (2 against 2):", (ones == 2).sum(), "of", len(y_test))
print(f"majority vote test accuracy {(majority == y_test).mean():.3f}")

What the four models and the vote produced

  • Each model alone scores between 0.832 and 0.869 on the 2,000 test rows, trained on only 1,000 sampled rows.
  • The first test point gets the votes 1, 0, 1, 1, so the majority says 1, the same pattern as the board with the 0 in a different place.
  • The last sample holds 930 different rows out of 1,000 draws. The other 70 draws are repeats, which is what sampling with replacement does.
  • 133 test points are ties with four models, the problem the video answers with 100 to 200 models. Even with ties sent to 0, the vote scores 0.874, above every single model.

Bagging and boosting in scikit-learn

scikit-learn has both families ready made. BaggingClassifier trains many copies of one algorithm, each on its own bootstrap sample, and combines them. AdaBoostClassifier is the boosting chain of weak learners that AdaBoost works through by hand.

BaggingClassifier with 100 trees

python
from sklearn.ensemble import BaggingClassifier

# 100 decision trees, each trained on a bootstrap sample of 1,000 rows
bag = BaggingClassifier(estimator=DecisionTreeClassifier(), n_estimators=100,
                        max_samples=1000, random_state=0)

AdaBoostClassifier with a chain of stumps

A stump is a decision tree of depth 1, the usual weak learner. staged_score returns the test accuracy after the first model, the first two, the first three and so on, so it shows the chain getting stronger.

python
from sklearn.ensemble import AdaBoostClassifier

stump = DecisionTreeClassifier(max_depth=1)
boost = AdaBoostClassifier(estimator=stump, n_estimators=200, random_state=0)
boost.fit(X_train, y_train)
scores = list(boost.staged_score(X_test, y_test))   # accuracy after 1, 2, 3, ... stumps
ExampleRun on scikit-learn 1.9.1, same data as above
from sklearn.ensemble import BaggingClassifier, AdaBoostClassifier

bag = BaggingClassifier(estimator=DecisionTreeClassifier(), n_estimators=100,
                        max_samples=1000, random_state=0)
bag.fit(X_train, y_train)
print(f"bagging, 100 trees: {bag.score(X_test, y_test):.3f}")

stump = DecisionTreeClassifier(max_depth=1)
print(f"one stump alone:    {stump.fit(X_train, y_train).score(X_test, y_test):.3f}")
boost = AdaBoostClassifier(estimator=stump, n_estimators=200, random_state=0)
boost.fit(X_train, y_train)
scores = list(boost.staged_score(X_test, y_test))
for n in (1, 5, 20, 50, 200):
    print(f"boosting, {n:3} stumps in sequence: {scores[n - 1]:.3f}")

What the bagged trees and the chain of stumps show

  • 100 bagged trees reach 0.900, higher than the four-model vote above. More models also make ties rare.
  • One stump scores 0.777. It splits on a single feature once, so it is a weak learner.
  • The chain improves as stumps are added: 0.825 after 5, 0.837 after 20, 0.850 after 200. Each stump is weak, the sequence is stronger.

Bagging vs boosting

BaggingBoosting
Order of trainingparallel, independent modelssequential, one after another
Data for each modela bootstrap sample of the rowsall rows, reweighted toward past mistakes
Typical modela full decision tree (low bias, high variance)a weak learner such as a stump
Combining outputsmajority vote or meana weighted sum of the learners
Main effectlowers variancelowers bias
Algorithmsrandom forest, BaggingClassifierAdaBoost, gradient boosting, XGBoost

Where you use bagging and boosting

  • Tabular data competitions, where the video notes that most winning Kaggle solutions use a bagging or boosting model.
  • A model that overfits: a deep decision tree that scores far better on training data than on test data is the classic case for bagging.
  • Several different models you already have: a custom ensemble that votes over logistic regression, KNN and a tree, like the board's M1 to M4.
Watch out. Bagging only helps when the models make different mistakes. Four copies of the same model trained on the same rows always vote together, so the vote adds nothing. The bootstrap samples (and, in a random forest, the feature samples) are what make the models differ.
Try it yourself
  • Change the tie rule to ones >= 2 so a tie goes to 1, and compare the majority vote accuracy.
  • Change size=1000 to size=8000 in the bootstrap sample and see how many unique rows the last sample has.
  • Set n_estimators=10 in the BaggingClassifier and compare its accuracy with 100 trees.

Little by little, you're building something great.