Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

AdaBoost

AdaBoost (adaptive boosting) is a boosting algorithm that trains a chain of decision stumps, raising the weight of the records each stump gets wrong so that the next stump focuses on them.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Bagging and boosting described boosting as weak learners in sequence. AdaBoost is the first boosting algorithm the video works through with numbers.

AdaBoost weights and stumps · from the Complete Machine Learning in 6 Hours video · 269:57 to 273:32

Giving every record an equal weight

The board's dataset has four features f1 to f4 and an output of yes or no, with seven records. The first step gives every record the same sample weight, a number between 0 and 1, with all the weights adding up to 1. With seven records each weight is 1/7.

Training a stump as the weak learner

Next AdaBoost picks one feature, the one with the best split by Information gain and entropy or Gini, and builds a decision tree on it with a depth of one. A tree of depth one is a stump. A stump is a weak learner: one split on one feature cannot predict well.

A seven-row AdaBoost table with features f1 to f4, output and a weight of 1/7 on every row; a stump on f1 gets row 4 wrong, so the total error is 1/7 and the performance of the stump is one half of ln 6, which is 0.896.
Total error and performance of stump · from the Complete Machine Learning in 6 Hours video · 273:32 to 276:22

Calculating the total error and the performance of stump

All seven records go through the trained stump. On the board it gets one record wrong, record 4.

Step 1 is the total error (TE): the sum of the weights of the wrong records. One record is wrong and its weight is 1/7, so TE = 1/7.

Step 2 is the performance of stump, which says how much this stump's vote will count:

The board rounds this to 0.895; ½·ln 6 is 0.8959, which rounds to 0.896. A stump with a small error gets a large performance, and a stump that is right only half the time (TE = 0.5) gets 0.

Updating the sample weights · from the Complete Machine Learning in 6 Hours video · 276:22 to 280:06

Updating the weights of correct and wrong records

Step 3 updates the weights. The correct records should count less from now on and the wrong record more, so the next stump pays attention to it.

For the six correct records that is 1/7 × e−0.896 = 0.0583. For record 4 it is 1/7 × e0.896 = 0.3499. The board writes 0.05 for the correct records; the product is 0.058. Its 0.349 for the wrong record is right.

Normalised weights and buckets · from the Complete Machine Learning in 6 Hours video · 280:06 to 284:44

Normalising the weights and filling the buckets

The new weights no longer add up to 1. Their sum is 6 × 0.0583 + 0.3499 = 0.6999, so each weight is divided by 0.6999. That gives 0.0833 (1/12) for each correct record and 0.5 for record 4, and the seven add up to exactly 1. The board carries its rounded 0.05 forward, so its sum (0.649) and its normalised weights (0.07 and 0.537) come out slightly off.

The normalised weights become buckets on the line from 0 to 1: record 1 owns 0 to 0.083, record 2 owns 0.083 to 0.167, and so on. Record 4's bucket runs from 0.25 to 0.75, half the line. To build the training data for the second stump, AdaBoost draws random numbers between 0 and 1 and takes the record whose bucket each number lands in. Record 4 is drawn about half the time, so the second stump trains mostly on the record the first one got wrong. The board writes the fifth bucket as 0.747 to 0.751; it runs from 0.75 to 0.833.

The seven new weights 0.0583 and 0.3499 sum to 0.6999; after normalising, the correct rows hold 0.0833 each and the wrong row 4 holds 0.5, so its bucket runs from 0.25 to 0.75 on the line from 0 to 1 and row 4 is drawn often for the second stump.

Combining the stumps with a weighted vote

The steps repeat: stump 2 makes its own mistakes, the weights update again, and stump 3 follows, for 100 or more stumps. For a new record each stump gives an output, for example 0, 1, 1, 1. The video combines them by majority vote for classification and by the average for regression. AdaBoost weights each stump's vote by its performance of stump instead, so a stump with a small error counts more.

For example, four stumps judge a new applicant with a salary of 50K or less and good credit. They say Yes, No, Yes and No, and their performances are 0.896, 0.65, 0.24 and −0.30. The Yes votes add up to 0.896 + 0.24 = 1.136 and the No votes to 0.65 + (−0.30) = 0.35, so the answer is Yes.

A negative performance belongs to a stump whose total error is above 0.5, worse than a coin toss (here about 0.65), so its No counts as a small Yes. With Yes as +1 and No as −1 the score is the same decision in one sum:

Running the seven records through one round

The board keeps the feature values as dashes. The code needs only which records the stump got right, so it runs the three steps, the normalising and the buckets on the board's seven records.

The total error and the performance of stump

python
w = np.full(7, 1 / 7)                                   # equal weights, sum 1
te = w[~correct].sum()                                  # 1. total error
alpha = 0.5 * np.log((1 - te) / te)                     # 2. performance of stump

The new weights and the buckets

np.where applies the minus exponent to correct records and the plus exponent to the wrong one. np.cumsum turns the normalised weights into the upper edge of each bucket, and np.searchsorted finds the bucket a random number falls in.

python
new_w = np.where(correct, w * np.exp(-alpha), w * np.exp(alpha))   # 3. new weights
norm_w = new_w / new_w.sum()                                       # sum back to 1
upper = np.cumsum(norm_w)                                          # bucket edges
picked = np.searchsorted(upper, draws) + 1                         # record per draw
ExampleFrom the video, run with NumPy
import numpy as np

correct = np.array([True, True, True, False, True, True, True])   # the stump gets record 4 wrong
w = np.full(7, 1 / 7)

te = w[~correct].sum()
alpha = 0.5 * np.log((1 - te) / te)
new_w = np.where(correct, w * np.exp(-alpha), w * np.exp(alpha))
norm_w = new_w / new_w.sum()
upper = np.cumsum(norm_w)

print("total error:", round(te, 4))
print("performance of stump:", round(alpha, 4))
print("new weights:", np.round(new_w, 4), "sum", round(new_w.sum(), 4))
print("normalised: ", np.round(norm_w, 4))
print("bucket upper edges:", np.round(upper, 3))

draws = np.random.default_rng(0).random(7)          # 7 random numbers in [0, 1)
picked = np.searchsorted(upper, draws) + 1
print("random numbers:", np.round(draws, 2))
print("records drawn for stump 2:", picked)

What the weights and the draws show

  • Total error 0.1429 is 1/7, and the performance of stump 0.8959 is ½·ln 6.
  • The new weights are 0.0583 for the correct records and 0.3499 for record 4; their sum is 0.6999.
  • After normalising, record 4 holds 0.5 and every other record 0.0833. The wrong records always end up with half the total weight after one AdaBoost update.
  • Three of the seven draws land in bucket 4 (0.64, 0.27 and 0.61, all between 0.25 and 0.75), so record 4 appears three times in the second stump's data. Record 1 is drawn twice and records 5 and 6 once; records 2, 3 and 7 are not drawn this time.

Running the round on the loan table

The board keeps the feature values as dashes. The same round on a table with values: seven loan applications with salary (50K or less, or more than 50K), credit (bad B, good G or normal N) and approval. The XGBoost classifier works on the same table.

#SalaryCreditApproval
1≤50KBNo
2≤50KGYes
3≤50KGYes
4>50KBNo
5>50KGYes
6>50KNYes
7≤50KNNo

Two candidate stumps compete. Salary ≤50K puts 2 Yes and 2 No on one side and 2 Yes and 1 No on the other. Credit = G puts 3 Yes and 0 No on one side and 1 Yes and 3 No on the other, so its entropy is lower and it becomes stump 1. It predicts No for record 6 (>50K, N, Yes), its only mistake.

From there the numbers are the board's: TE = 1/7, performance 0.896, new weights 0.0583 and 0.3499, normalised weights 0.0833 and 0.5. Record 6 owns the bucket from 0.417 to 0.917. Seven random numbers, 0.50, 0.10, 0.60, 0.75, 0.24, 0.32 and 0.87, pick records 6, 2, 6, 6, 3, 4 and 6, so record 6 fills four of the seven rows the second stump trains on.

The seven loan records with salary, credit and approval each start at weight 1/7; the stump credit equals G gets record 6 (over 50K, normal credit, approved) wrong, so its total error is 1/7 and its performance 0.896; record 6's weight rises to 0.3499, and after normalising it holds 0.5 with the bucket 0.417 to 0.917, so the draws 0.50, 0.10, 0.60, 0.75, 0.24, 0.32 and 0.87 pick records 6, 2, 6, 6, 3, 4 and 6.

Fitting AdaBoostClassifier on the loan table

The text columns become 0 or 1 columns so that a stump can split on them: salary_over_50k, credit_good and credit_bad (0 in both credit columns means normal).

Reading the errors and the stump weights

After fitting, estimator_errors_ holds each stump's total error and estimator_weights_ holds its vote weight. scikit-learn's weight is ln((1 − TE)/TE), twice the board's performance of stump. Doubling every stump's weight does not change which way the vote goes, and it makes the update simpler: scikit-learn multiplies only the wrong records by eweight and then normalises, which gives the same 0.0833 and 0.5. It also passes the weights to the next stump directly, instead of drawing records from buckets.

python
from sklearn.ensemble import AdaBoostClassifier
from sklearn.tree import DecisionTreeClassifier

ada = AdaBoostClassifier(estimator=DecisionTreeClassifier(max_depth=1), n_estimators=4,
                         random_state=0).fit(X, y)
ada.estimator_errors_, ada.estimator_weights_   # total error and vote weight per stump
ExampleThe loan table, run on scikit-learn 1.9.1
import numpy as np
from sklearn.ensemble import AdaBoostClassifier
from sklearn.tree import DecisionTreeClassifier

salary = np.array(["<=50K", "<=50K", "<=50K", ">50K", ">50K", ">50K", "<=50K"])
credit = np.array(["B", "G", "G", "B", "G", "N", "N"])
y = np.array([0, 1, 1, 0, 1, 1, 0])                        # approval: 1 = Yes, 0 = No
X = np.column_stack([salary == ">50K", credit == "G", credit == "B"]).astype(int)
names = ["salary_over_50k", "credit_good", "credit_bad"]

ada = AdaBoostClassifier(estimator=DecisionTreeClassifier(max_depth=1), n_estimators=4,
                         random_state=0).fit(X, y)
print("stump 1 splits on:", names[ada.estimators_[0].tree_.feature[0]])
print("stump 1 predicts:", ada.estimators_[0].predict(X), " true:", y)
print("total error per stump:", np.round(ada.estimator_errors_, 4))
print("weight per stump:", np.round(ada.estimator_weights_, 4))
print("half of stump 1's weight:", round(ada.estimator_weights_[0] / 2, 4))

# The weighted vote: each stump says +1 or -1, times its weight
votes = np.array([np.where(s.predict(X) == 1, 1, -1) for s in ada.estimators_])
score = ada.estimator_weights_ @ votes
print("weighted vote:", np.round(score, 2))
print("final class:  ", (score > 0).astype(int), " predict():", ada.predict(X))
print("new applicant, <=50K with good credit:", ada.predict([[0, 1, 0]]))

What the four stumps learned

  • Stump 1 splits on credit_good, the Credit = G stump, and gets only record 6 wrong; its total error is 0.1429 = 1/7.
  • Its weight 1.7918 is ln 6; half of it is 0.8959, the performance of stump.
  • Stump 2 has a total error of 0.0833, one correct record's normalised weight: it gets record 6 right and one of the others wrong.
  • The weighted vote is positive for records 2, 3, 5 and 6 and negative for the rest, so the four stumps together classify all seven records. Record 6 is the closest call, at 0.78.
  • The new applicant with a salary of 50K or less and good credit is approved (class 1), the same answer as the four-stump vote worked by hand above.

AdaBoost vs random forest

AdaBoostRandom forest
Familyboostingbagging
Treesstumps (depth 1), weak learnersfull trees, strong but high variance
Orderone after another; each depends on the lastindependent, can train in parallel
How a tree sees the dataall records, weighted toward past mistakesa bootstrap sample of records and features
Final answervote weighted by each stump's performanceplain vote or average
Sensitive tonoisy labels and outliers, which keep getting heavierlittle; averaging softens noise

Where you use AdaBoost

  • Small, clean tabular datasets where a few simple rules separate the classes well.
  • Face detection: the Viola-Jones detector is a chain of boosted simple features.
  • A baseline before gradient boosting, since it has few settings: the number of stumps and the learning rate.
Watch out. AdaBoost keeps raising the weight of records it cannot fit. A mislabelled record or an extreme outlier ends up with a huge weight and pulls later stumps toward it. Clean the labels, or use fewer stumps and a learning_rate below 1.
Try it yourself
  • Mark two records as wrong in correct and check that the wrong records still end up with 0.5 of the weight between them.
  • Change n_estimators=4 to n_estimators=1 and see which record predict() gets wrong.
  • Set learning_rate=0.5 in AdaBoostClassifier and compare estimator_weights_ with the run above.
PreviousRandom forest

Slow is fine. Stopping is the only problem.