Bagging and boosting
Bagging and boosting are ensemble techniques that combine several models to solve one problem: bagging trains the models side by side on samples of the rows, and boosting trains them one after another so each one covers what the last one missed.
Last updated: 05 Oct, 2026 · scikit-learn 1.9.1
Every algorithm so far solved a classification or a regression problem with one model at a time. An ensemble asks a different question: can several models work on the same problem and give a better answer together than any one of them alone?
Combining models with bagging
The video starts from a dataset D with features f1, f2, f3 and an output column. Bagging builds several models, M1, M2, M3 and M4. They do not have to be the same algorithm: on the board M1 is logistic regression, M2 a decision tree, M3 KNN and M4 a decision tree or naive Bayes.
Each model trains on its own sample of the rows. If D has 10,000 rows, M1 might get d′ = 1,000 of them, M2 another 1,000 called d″, then d‴ and d⁗. This is row sampling, and a sample is always much smaller than D (d′ ≪ D). Some rows show up in more than one sample. The sampling is done with replacement: a row that was picked goes back into the pool and can be picked again. A sample drawn that way is called a bootstrap sample.
Aggregating the outputs with a majority vote or a mean
Once every model is trained, a new test point goes to all four of them. On the board M1 predicts 0, and M2, M3 and M4 predict 1. The outputs are aggregated by majority voting: three of the four models say 1, so the answer is 1. Sampling with replacement and then aggregating gives the method its full name, bootstrap aggregating, shortened to bagging.
What happens when two models say 0 and two say 1? In practice a bagging model has 100 to 200 models, so a tie is very unlikely. A random forest uses only decision trees, but bagging as an idea can mix different algorithms, as the board does.
For a regression problem each model outputs a number. On the board they give 120, 140, 122 and 148, and the bagging output is their mean.

Chaining weak learners with boosting
In bagging the models are parallel and independent of each other. Boosting is a sequential combination: the training data goes to M1, then to M2, then M3, then M4, and each model hands on to the next.
Every model in the chain is a weak learner: on its own its predictions are poor. Combined in sequence, the weak learners become a strong learner. The video's example is a group of teachers. If the physics teacher cannot solve a problem, the chemistry, maths or geography teacher may, and together their expertise gives a good answer.

The board ends with the algorithms of each family. Bagging gives the Random forest classifier and regressor. Boosting gives AdaBoost, Gradient boosting and extreme gradient boosting, XGBoost classifier.
Computing the board's vote and mean
The four outputs from the board, for classification and for regression, combined the way the video combines them.
from collections import Counter
import numpy as np
# The four models' outputs for one test point, from the board
class_votes = [0, 1, 1, 1] # M1 logistic, M2 tree, M3 KNN, M4 naive Bayes
values = [120, 140, 122, 148] # the same four models on a regression problem
# Classification: the class that most models predict
winner, count = Counter(class_votes).most_common(1)[0]
print("majority vote:", winner, f"({count} of {len(class_votes)} models)")
# Regression: the mean of the outputs
print("mean:", np.mean(values))majority vote: 1 (3 of 4 models) mean: 132.5
Bagging four different models on 10,000 rows
The board's setup can run as it is drawn: a dataset of 10,000 rows, four different algorithms, and a bootstrap sample of 1,000 rows for each. The data comes from make_classification, which generates a labelled dataset of any size.
Drawing a bootstrap sample of 1,000 rows
rng.choice with replace=True picks 1,000 row numbers and lets a number repeat, which is sampling with replacement.
rng = np.random.default_rng(0)
# 1,000 row numbers out of the training rows; a number can repeat
rows = rng.choice(len(X_train), size=1000, replace=True)
model.fit(X_train[rows], y_train[rows])Taking the majority vote of four predictions
Stack the four models' predictions into one array, one row per model. A test point is class 1 when at least three of the four models say 1. Two against two is a tie, which this code sends to 0.
votes = np.array(preds) # shape: (4 models, number of test rows)
ones = votes.sum(axis=0) # how many models said 1, per test row
majority = (ones >= 3).astype(int) # 3 or 4 votes for 1 -> class 1import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.neighbors import KNeighborsClassifier
from sklearn.naive_bayes import GaussianNB
X, y = make_classification(n_samples=10000, n_features=10, n_informative=5, random_state=0)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)
models = {"M1 logistic": LogisticRegression(), "M2 tree": DecisionTreeClassifier(random_state=0),
"M3 KNN": KNeighborsClassifier(), "M4 naive Bayes": GaussianNB()}
rng = np.random.default_rng(0)
preds = []
for name, model in models.items():
rows = rng.choice(len(X_train), size=1000, replace=True) # d' = 1,000 rows
model.fit(X_train[rows], y_train[rows])
preds.append(model.predict(X_test))
print(f"{name:15} test accuracy {model.score(X_test, y_test):.3f}")
votes = np.array(preds)
ones = votes.sum(axis=0)
majority = (ones >= 3).astype(int)
print("votes for the first test point:", votes[:, 0], "-> majority", majority[0])
print("unique rows in the last sample:", len(np.unique(rows)), "of 1000")
print("ties (2 against 2):", (ones == 2).sum(), "of", len(y_test))
print(f"majority vote test accuracy {(majority == y_test).mean():.3f}")M1 logistic test accuracy 0.860 M2 tree test accuracy 0.832 M3 KNN test accuracy 0.869 M4 naive Bayes test accuracy 0.848 votes for the first test point: [1 0 1 1] -> majority 1 unique rows in the last sample: 930 of 1000 ties (2 against 2): 133 of 2000 majority vote test accuracy 0.874
What the four models and the vote produced
- Each model alone scores between 0.832 and 0.869 on the 2,000 test rows, trained on only 1,000 sampled rows.
- The first test point gets the votes 1, 0, 1, 1, so the majority says 1, the same pattern as the board with the 0 in a different place.
- The last sample holds 930 different rows out of 1,000 draws. The other 70 draws are repeats, which is what sampling with replacement does.
- 133 test points are ties with four models, the problem the video answers with 100 to 200 models. Even with ties sent to 0, the vote scores 0.874, above every single model.
Bagging and boosting in scikit-learn
scikit-learn has both families ready made. BaggingClassifier trains many copies of one algorithm, each on its own bootstrap sample, and combines them. AdaBoostClassifier is the boosting chain of weak learners that AdaBoost works through by hand.
BaggingClassifier with 100 trees
from sklearn.ensemble import BaggingClassifier
# 100 decision trees, each trained on a bootstrap sample of 1,000 rows
bag = BaggingClassifier(estimator=DecisionTreeClassifier(), n_estimators=100,
max_samples=1000, random_state=0)AdaBoostClassifier with a chain of stumps
A stump is a decision tree of depth 1, the usual weak learner. staged_score returns the test accuracy after the first model, the first two, the first three and so on, so it shows the chain getting stronger.
from sklearn.ensemble import AdaBoostClassifier
stump = DecisionTreeClassifier(max_depth=1)
boost = AdaBoostClassifier(estimator=stump, n_estimators=200, random_state=0)
boost.fit(X_train, y_train)
scores = list(boost.staged_score(X_test, y_test)) # accuracy after 1, 2, 3, ... stumpsfrom sklearn.ensemble import BaggingClassifier, AdaBoostClassifier
bag = BaggingClassifier(estimator=DecisionTreeClassifier(), n_estimators=100,
max_samples=1000, random_state=0)
bag.fit(X_train, y_train)
print(f"bagging, 100 trees: {bag.score(X_test, y_test):.3f}")
stump = DecisionTreeClassifier(max_depth=1)
print(f"one stump alone: {stump.fit(X_train, y_train).score(X_test, y_test):.3f}")
boost = AdaBoostClassifier(estimator=stump, n_estimators=200, random_state=0)
boost.fit(X_train, y_train)
scores = list(boost.staged_score(X_test, y_test))
for n in (1, 5, 20, 50, 200):
print(f"boosting, {n:3} stumps in sequence: {scores[n - 1]:.3f}")bagging, 100 trees: 0.900 one stump alone: 0.777 boosting, 1 stumps in sequence: 0.777 boosting, 5 stumps in sequence: 0.825 boosting, 20 stumps in sequence: 0.837 boosting, 50 stumps in sequence: 0.849 boosting, 200 stumps in sequence: 0.850
What the bagged trees and the chain of stumps show
- 100 bagged trees reach 0.900, higher than the four-model vote above. More models also make ties rare.
- One stump scores 0.777. It splits on a single feature once, so it is a weak learner.
- The chain improves as stumps are added: 0.825 after 5, 0.837 after 20, 0.850 after 200. Each stump is weak, the sequence is stronger.
Bagging vs boosting
| Bagging | Boosting | |
|---|---|---|
| Order of training | parallel, independent models | sequential, one after another |
| Data for each model | a bootstrap sample of the rows | all rows, reweighted toward past mistakes |
| Typical model | a full decision tree (low bias, high variance) | a weak learner such as a stump |
| Combining outputs | majority vote or mean | a weighted sum of the learners |
| Main effect | lowers variance | lowers bias |
| Algorithms | random forest, BaggingClassifier | AdaBoost, gradient boosting, XGBoost |
Where you use bagging and boosting
- Tabular data competitions, where the video notes that most winning Kaggle solutions use a bagging or boosting model.
- A model that overfits: a deep decision tree that scores far better on training data than on test data is the classic case for bagging.
- Several different models you already have: a custom ensemble that votes over logistic regression, KNN and a tree, like the board's M1 to M4.
Related
- Previous: Decision tree in scikit-learn
- Next: Random forest
- See also: Bias and variance
- Reference: scikit-learn user guide, Ensembles
- Change the tie rule to
ones >= 2so a tie goes to 1, and compare the majority vote accuracy. - Change
size=1000tosize=8000in the bootstrap sample and see how many unique rows the last sample has. - Set
n_estimators=10in the BaggingClassifier and compare its accuracy with 100 trees.
Little by little, you're building something great.