Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Gradient boosting

Gradient boosting is a boosting algorithm that starts from the average of the target and adds decision trees one after another, each trained on the residuals of the model so far and scaled down by a learning rate.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

AdaBoost made the next stump focus on the records the last one got wrong by raising their weights. Gradient boosting hands the next tree the errors themselves. The XGBoost classifier and XGBoost regressor are an optimised form of it, so the steps here are the ones those lessons build on.

Starting from the average salary

Base model and residuals · from the Gradient Boosting In Depth Intuition video · 1:39 to 4:29

The worked example predicts salary from experience and degree. Gradient boosting works for regression and for classification; the regression form is the easiest to follow by hand.

ExpDegreeSalaryBase model ŷResidual R1
2B.E50K75K−25K
3Masters70K75K−5K
5Masters80K75K5K
6PhD100K75K25K

Step 1 creates a base model that predicts the average for every record: (50 + 70 + 80 + 100) / 4 = 75K. Step 2 computes the residuals, the error of that prediction: R1 = y − ŷ, so −25, −5, 5 and 25.

Fitting a tree to the residuals

Learning rate and the next tree · from the Gradient Boosting In Depth Intuition video · 4:29 to 8:30

Step 3 trains a decision tree DT1 with experience and degree as the inputs and R1 as the target. Say the tree predicts the residuals as −23, −3, 3 and 20: close to −25, −5, 5 and 25, but not equal, as a small tree's leaves rarely are.

Adding the tree's output in full gives 75 + (−23) = 52 for the first record, almost its 50K. The video calls this overfitting: a fit this close to the training data has low bias but high variance, so new data gets worse predictions. So each tree's output is multiplied by a learning rate α between 0 and 1. With α = 0.1 the first record becomes 75 + 0.1 × (−23) = 72.7 and the second 75 + 0.1 × (−3) = 74.7. The new residuals, R2, are the targets of the next tree DT2. The video works the first record out as 73.7; 75 − 2.3 is 72.7.

A table of experience, degree and salary 50K, 70K, 80K and 100K: the base model predicts the average 75K, the residuals are −25, −5, 5 and 25, the first tree predicts −23, −3, 3 and 20, and with a learning rate of 0.1 the new predictions are 72.7, 74.7, 75.3 and 77.0 with residuals −22.7, −4.7, 4.7 and 23.0; the final model is the base plus the learning rate times every tree.
Salaryŷ after DT1R2 = salary − ŷ
50K75 + 0.1 × (−23) = 72.7−22.7
70K75 + 0.1 × (−3) = 74.7−4.7
80K75 + 0.1 × 3 = 75.34.7
100K75 + 0.1 × 20 = 77.023.0

Every residual moved a little toward 0. Each new tree trains on the latest residuals and adds a small step, so the model reaches the targets slowly, over many trees.

Writing the final function

Adding the trees into one function · from the Gradient Boosting In Depth Intuition video · 8:30 to 10:40

The model is the base model h₀ plus every tree hᵢ, each scaled by the learning rate:

The video writes h₀(x) + α₁h₁(x) + … + αₙhₙ(x), one α per tree. In scikit-learn the base model is added once, unscaled, and one learning_rate is shared by every tree.

Seeing the residual as a gradient

Pseudo-residuals as the negative gradient · from the Gradient Boosting Complete Maths video · 8:17 to 11:26

The name comes from the loss. For the squared error L = ½(y − F)², the derivative with respect to the prediction F is −(y − F). The video works it on three records with salaries 50, 70 and 60K: the base model predicts their mean, 60, so the residuals are −10, 10 and 0. The residual is the negative gradient, so each tree takes one step downhill, like Gradient descent, but on the predictions instead of on a weight. The video calls it the pseudo-residual, and a different loss gives a different one: for log loss in classification it is label − probability, the residual of the XGBoost classifier.

Running three rounds by hand and with GradientBoostingRegressor

A depth-1 tree fitted to the residuals, three times, with α = 0.1, on the four records. The trees here are real fits, so their outputs are not the board's assumed −23, −3, 3 and 20.

One round of gradient boosting

python
F = np.full(len(salary), salary.mean())          # base model: the average, 75
r = salary - F                                   # residuals
tree = DecisionTreeRegressor(max_depth=1).fit(X, r)
F = F + 0.1 * tree.predict(X)                    # add the tree's output times the learning rate

GradientBoostingRegressor with the same settings

staged_predict yields the predictions after one tree, two trees, three trees, and so on.

python
from sklearn.ensemble import GradientBoostingRegressor

gbr = GradientBoostingRegressor(n_estimators=3, learning_rate=0.1, max_depth=1).fit(X, salary)
stages = list(gbr.staged_predict(X))             # predictions after 1, 2 and 3 trees
ExampleThe salary table, run on scikit-learn 1.9.1
import numpy as np
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import GradientBoostingRegressor

X = np.array([[2, 0], [3, 1], [5, 1], [6, 2]])   # experience, degree (0 = B.E, 1 = Masters, 2 = PhD)
salary = np.array([50.0, 70, 80, 100])

F = np.full(len(salary), salary.mean())
print("base model:", F)
for round_ in (1, 2, 3):
    r = salary - F
    tree = DecisionTreeRegressor(max_depth=1).fit(X, r)
    F = F + 0.1 * tree.predict(X)
    print(f"round {round_}: residuals {r.round(2)}  tree says {tree.predict(X).round(2)}  new prediction {F.round(2)}")

gbr = GradientBoostingRegressor(n_estimators=3, learning_rate=0.1, max_depth=1).fit(X, salary)
stages = list(gbr.staged_predict(X))
print("GradientBoostingRegressor after 3 trees:", stages[-1].round(2))
print("same as by hand:", np.allclose(stages[-1], F))

What the three rounds show

  • The base model predicts 75 for every record, and the first residuals are the board's −25, −5, 5 and 25.
  • The first tree splits experience between 3 and 5 and predicts −15 and 15, the average residual on each side. With α = 0.1 the predictions move from 75 to 73.5 and 76.5.
  • Each round shrinks the residuals a little: −25 becomes −23.5 and then −21.15, and 25 becomes 23.5 and then 22.72. Three trees with α = 0.1 are still far from the salaries, which is why gradient boosting uses many trees.
  • GradientBoostingRegressor gives the same predictions after three trees, 70.39, 73.53, 76.53 and 79.55, so the hand loop is the algorithm.

Predicting used car prices with gradient boosting

The used car data from Random forest, loaded and split the same way. The forest scored a test R² of 0.931 there. Gradient boosting with depth-3 trees and α = 0.1 runs for 300 trees, and staged_predict scores it after 1, 10, 100 and 300 of them.

ExampleThe used car dataset, run on scikit-learn 1.9.1
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.metrics import r2_score

gbr = GradientBoostingRegressor(n_estimators=300, learning_rate=0.1, max_depth=3, random_state=0)
gbr.fit(X_train, y_train)
stages = list(gbr.staged_predict(X_test))
for n in (1, 10, 100, 300):
    print(f"{n:3} trees: test R2 {r2_score(y_test, stages[n - 1]):.3f}")
print(f"train R2 after 300 trees: {gbr.score(X_train, y_train):.3f}")

Reading the car price scores

  • One tree scores 0.118: with α = 0.1 it moves every price only a tenth of the way from the average.
  • 10 trees reach 0.668 and 100 trees 0.912; each tree fixes part of what the earlier ones left.
  • 300 trees score 0.933, level with the random forest's 0.931, with a training R² of 0.971. The gap between train and test is small, so these 300 trees have not started to overfit.

Gradient boosting vs AdaBoost

Gradient boostingAdaBoost
What the next tree learnsthe residuals of the model so farthe same records, reweighted toward mistakes
Base modelthe average of the target (regression)none; the first stump starts the chain
Treessmall trees, often depth 3stumps of depth 1
Combiningbase + learning rate × each tree's outputvote weighted by each stump's performance
scikit-learnGradientBoostingRegressor, GradientBoostingClassifierAdaBoostRegressor, AdaBoostClassifier

Where you use gradient boosting

  • Price and demand prediction on tabular data, such as the used car prices above.
  • Ranking and click prediction, where many weak signals across columns add up.
  • A step before XGBoost or LightGBM: the same idea with fewer settings, in scikit-learn itself.
Watch out. The learning rate and the number of trees work against each other. A learning rate of 1 lets each tree jump straight to the residuals, the 75 + (−23) = 52 overfit; a small rate needs many more trees. Tune the two together with Cross-validation, or set n_iter_no_change so training stops when a validation score stops improving.
Try it yourself
  • Change 0.1 to 1.0 in the hand loop and see how close round 1 gets to the salaries.
  • Set max_depth=2 in the hand loop and in GradientBoostingRegressor, and check they still match.
  • Set learning_rate=0.5 on the car data and compare the test R² after 10 trees.
PreviousAdaBoost

This is what real progress feels like.