AdaBoost
AdaBoost (adaptive boosting) is a boosting algorithm that trains a chain of decision stumps, raising the weight of the records each stump gets wrong so that the next stump focuses on them.
Last updated: 05 Oct, 2026 · scikit-learn 1.9.1
Bagging and boosting described boosting as weak learners in sequence. AdaBoost is the first boosting algorithm the video works through with numbers.
Giving every record an equal weight
The board's dataset has four features f1 to f4 and an output of yes or no, with seven records. The first step gives every record the same sample weight, a number between 0 and 1, with all the weights adding up to 1. With seven records each weight is 1/7.
Training a stump as the weak learner
Next AdaBoost picks one feature, the one with the best split by Information gain and entropy or Gini, and builds a decision tree on it with a depth of one. A tree of depth one is a stump. A stump is a weak learner: one split on one feature cannot predict well.

Calculating the total error and the performance of stump
All seven records go through the trained stump. On the board it gets one record wrong, record 4.
Step 1 is the total error (TE): the sum of the weights of the wrong records. One record is wrong and its weight is 1/7, so TE = 1/7.
Step 2 is the performance of stump, which says how much this stump's vote will count:
The board rounds this to 0.895; ½·ln 6 is 0.8959, which rounds to 0.896. A stump with a small error gets a large performance, and a stump that is right only half the time (TE = 0.5) gets 0.
Updating the weights of correct and wrong records
Step 3 updates the weights. The correct records should count less from now on and the wrong record more, so the next stump pays attention to it.
For the six correct records that is 1/7 × e−0.896 = 0.0583. For record 4 it is 1/7 × e0.896 = 0.3499. The board writes 0.05 for the correct records; the product is 0.058. Its 0.349 for the wrong record is right.
Normalising the weights and filling the buckets
The new weights no longer add up to 1. Their sum is 6 × 0.0583 + 0.3499 = 0.6999, so each weight is divided by 0.6999. That gives 0.0833 (1/12) for each correct record and 0.5 for record 4, and the seven add up to exactly 1. The board carries its rounded 0.05 forward, so its sum (0.649) and its normalised weights (0.07 and 0.537) come out slightly off.
The normalised weights become buckets on the line from 0 to 1: record 1 owns 0 to 0.083, record 2 owns 0.083 to 0.167, and so on. Record 4's bucket runs from 0.25 to 0.75, half the line. To build the training data for the second stump, AdaBoost draws random numbers between 0 and 1 and takes the record whose bucket each number lands in. Record 4 is drawn about half the time, so the second stump trains mostly on the record the first one got wrong. The board writes the fifth bucket as 0.747 to 0.751; it runs from 0.75 to 0.833.

Combining the stumps with a weighted vote
The steps repeat: stump 2 makes its own mistakes, the weights update again, and stump 3 follows, for 100 or more stumps. For a new record each stump gives an output, for example 0, 1, 1, 1. The video combines them by majority vote for classification and by the average for regression. AdaBoost weights each stump's vote by its performance of stump instead, so a stump with a small error counts more.
For example, four stumps judge a new applicant with a salary of 50K or less and good credit. They say Yes, No, Yes and No, and their performances are 0.896, 0.65, 0.24 and −0.30. The Yes votes add up to 0.896 + 0.24 = 1.136 and the No votes to 0.65 + (−0.30) = 0.35, so the answer is Yes.
A negative performance belongs to a stump whose total error is above 0.5, worse than a coin toss (here about 0.65), so its No counts as a small Yes. With Yes as +1 and No as −1 the score is the same decision in one sum:
Running the seven records through one round
The board keeps the feature values as dashes. The code needs only which records the stump got right, so it runs the three steps, the normalising and the buckets on the board's seven records.
The total error and the performance of stump
w = np.full(7, 1 / 7) # equal weights, sum 1
te = w[~correct].sum() # 1. total error
alpha = 0.5 * np.log((1 - te) / te) # 2. performance of stumpThe new weights and the buckets
np.where applies the minus exponent to correct records and the plus exponent to the wrong one. np.cumsum turns the normalised weights into the upper edge of each bucket, and np.searchsorted finds the bucket a random number falls in.
new_w = np.where(correct, w * np.exp(-alpha), w * np.exp(alpha)) # 3. new weights
norm_w = new_w / new_w.sum() # sum back to 1
upper = np.cumsum(norm_w) # bucket edges
picked = np.searchsorted(upper, draws) + 1 # record per drawimport numpy as np
correct = np.array([True, True, True, False, True, True, True]) # the stump gets record 4 wrong
w = np.full(7, 1 / 7)
te = w[~correct].sum()
alpha = 0.5 * np.log((1 - te) / te)
new_w = np.where(correct, w * np.exp(-alpha), w * np.exp(alpha))
norm_w = new_w / new_w.sum()
upper = np.cumsum(norm_w)
print("total error:", round(te, 4))
print("performance of stump:", round(alpha, 4))
print("new weights:", np.round(new_w, 4), "sum", round(new_w.sum(), 4))
print("normalised: ", np.round(norm_w, 4))
print("bucket upper edges:", np.round(upper, 3))
draws = np.random.default_rng(0).random(7) # 7 random numbers in [0, 1)
picked = np.searchsorted(upper, draws) + 1
print("random numbers:", np.round(draws, 2))
print("records drawn for stump 2:", picked)total error: 0.1429 performance of stump: 0.8959 new weights: [0.0583 0.0583 0.0583 0.3499 0.0583 0.0583 0.0583] sum 0.6999 normalised: [0.0833 0.0833 0.0833 0.5 0.0833 0.0833 0.0833] bucket upper edges: [0.083 0.167 0.25 0.75 0.833 0.917 1. ] random numbers: [0.64 0.27 0.04 0.02 0.81 0.91 0.61] records drawn for stump 2: [4 4 1 1 5 6 4]
What the weights and the draws show
- Total error 0.1429 is 1/7, and the performance of stump 0.8959 is ½·ln 6.
- The new weights are 0.0583 for the correct records and 0.3499 for record 4; their sum is 0.6999.
- After normalising, record 4 holds 0.5 and every other record 0.0833. The wrong records always end up with half the total weight after one AdaBoost update.
- Three of the seven draws land in bucket 4 (0.64, 0.27 and 0.61, all between 0.25 and 0.75), so record 4 appears three times in the second stump's data. Record 1 is drawn twice and records 5 and 6 once; records 2, 3 and 7 are not drawn this time.
Running the round on the loan table
The board keeps the feature values as dashes. The same round on a table with values: seven loan applications with salary (50K or less, or more than 50K), credit (bad B, good G or normal N) and approval. The XGBoost classifier works on the same table.
| # | Salary | Credit | Approval |
|---|---|---|---|
| 1 | ≤50K | B | No |
| 2 | ≤50K | G | Yes |
| 3 | ≤50K | G | Yes |
| 4 | >50K | B | No |
| 5 | >50K | G | Yes |
| 6 | >50K | N | Yes |
| 7 | ≤50K | N | No |
Two candidate stumps compete. Salary ≤50K puts 2 Yes and 2 No on one side and 2 Yes and 1 No on the other. Credit = G puts 3 Yes and 0 No on one side and 1 Yes and 3 No on the other, so its entropy is lower and it becomes stump 1. It predicts No for record 6 (>50K, N, Yes), its only mistake.
From there the numbers are the board's: TE = 1/7, performance 0.896, new weights 0.0583 and 0.3499, normalised weights 0.0833 and 0.5. Record 6 owns the bucket from 0.417 to 0.917. Seven random numbers, 0.50, 0.10, 0.60, 0.75, 0.24, 0.32 and 0.87, pick records 6, 2, 6, 6, 3, 4 and 6, so record 6 fills four of the seven rows the second stump trains on.

Fitting AdaBoostClassifier on the loan table
The text columns become 0 or 1 columns so that a stump can split on them: salary_over_50k, credit_good and credit_bad (0 in both credit columns means normal).
Reading the errors and the stump weights
After fitting, estimator_errors_ holds each stump's total error and estimator_weights_ holds its vote weight. scikit-learn's weight is ln((1 − TE)/TE), twice the board's performance of stump. Doubling every stump's weight does not change which way the vote goes, and it makes the update simpler: scikit-learn multiplies only the wrong records by eweight and then normalises, which gives the same 0.0833 and 0.5. It also passes the weights to the next stump directly, instead of drawing records from buckets.
from sklearn.ensemble import AdaBoostClassifier
from sklearn.tree import DecisionTreeClassifier
ada = AdaBoostClassifier(estimator=DecisionTreeClassifier(max_depth=1), n_estimators=4,
random_state=0).fit(X, y)
ada.estimator_errors_, ada.estimator_weights_ # total error and vote weight per stumpimport numpy as np
from sklearn.ensemble import AdaBoostClassifier
from sklearn.tree import DecisionTreeClassifier
salary = np.array(["<=50K", "<=50K", "<=50K", ">50K", ">50K", ">50K", "<=50K"])
credit = np.array(["B", "G", "G", "B", "G", "N", "N"])
y = np.array([0, 1, 1, 0, 1, 1, 0]) # approval: 1 = Yes, 0 = No
X = np.column_stack([salary == ">50K", credit == "G", credit == "B"]).astype(int)
names = ["salary_over_50k", "credit_good", "credit_bad"]
ada = AdaBoostClassifier(estimator=DecisionTreeClassifier(max_depth=1), n_estimators=4,
random_state=0).fit(X, y)
print("stump 1 splits on:", names[ada.estimators_[0].tree_.feature[0]])
print("stump 1 predicts:", ada.estimators_[0].predict(X), " true:", y)
print("total error per stump:", np.round(ada.estimator_errors_, 4))
print("weight per stump:", np.round(ada.estimator_weights_, 4))
print("half of stump 1's weight:", round(ada.estimator_weights_[0] / 2, 4))
# The weighted vote: each stump says +1 or -1, times its weight
votes = np.array([np.where(s.predict(X) == 1, 1, -1) for s in ada.estimators_])
score = ada.estimator_weights_ @ votes
print("weighted vote:", np.round(score, 2))
print("final class: ", (score > 0).astype(int), " predict():", ada.predict(X))
print("new applicant, <=50K with good credit:", ada.predict([[0, 1, 0]]))stump 1 splits on: credit_good stump 1 predicts: [0 1 1 0 1 0 0] true: [0 1 1 0 1 1 0] total error per stump: [0.1429 0.0833 0.1364 0.1579] weight per stump: [1.7918 2.3979 1.8458 1.674 ] half of stump 1's weight: 0.8959 weighted vote: [-7.71 4.02 4.02 -4.02 7.71 0.78 -2.91] final class: [0 1 1 0 1 1 0] predict(): [0 1 1 0 1 1 0] new applicant, <=50K with good credit: [1]
What the four stumps learned
- Stump 1 splits on credit_good, the Credit = G stump, and gets only record 6 wrong; its total error is 0.1429 = 1/7.
- Its weight 1.7918 is ln 6; half of it is 0.8959, the performance of stump.
- Stump 2 has a total error of 0.0833, one correct record's normalised weight: it gets record 6 right and one of the others wrong.
- The weighted vote is positive for records 2, 3, 5 and 6 and negative for the rest, so the four stumps together classify all seven records. Record 6 is the closest call, at 0.78.
- The new applicant with a salary of 50K or less and good credit is approved (class 1), the same answer as the four-stump vote worked by hand above.
AdaBoost vs random forest
| AdaBoost | Random forest | |
|---|---|---|
| Family | boosting | bagging |
| Trees | stumps (depth 1), weak learners | full trees, strong but high variance |
| Order | one after another; each depends on the last | independent, can train in parallel |
| How a tree sees the data | all records, weighted toward past mistakes | a bootstrap sample of records and features |
| Final answer | vote weighted by each stump's performance | plain vote or average |
| Sensitive to | noisy labels and outliers, which keep getting heavier | little; averaging softens noise |
Where you use AdaBoost
- Small, clean tabular datasets where a few simple rules separate the classes well.
- Face detection: the Viola-Jones detector is a chain of boosted simple features.
- A baseline before gradient boosting, since it has few settings: the number of stumps and the learning rate.
learning_rate below 1.Related
- Previous: Random forest
- Next: Gradient boosting
- See also: Entropy and Gini impurity
- Reference: AdaBoostClassifier in the scikit-learn API
- Mark two records as wrong in
correctand check that the wrong records still end up with 0.5 of the weight between them. - Change
n_estimators=4ton_estimators=1and see which recordpredict()gets wrong. - Set
learning_rate=0.5in AdaBoostClassifier and compareestimator_weights_with the run above.
Slow is fine. Stopping is the only problem.