Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

XGBoost regressor

The XGBoost regressor is the regression form of XGBoost: it starts from the average of the target, then adds binary trees one after another, each fitted to the residuals and split by the gain in similarity weight.

Last updated: 05 Oct, 2026 · XGBoost 3.4.1

The XGBoost classifier started from a probability of 0.5. The regressor starts, like Gradient boosting, from the average of the target, and follows the same steps as the classifier with numbers: residuals, similarity weights and gain, with a simpler similarity formula.

The salary table and the base model · from the Complete Machine Learning in 6 Hours video · 367:29 to 371:32

Starting from the average salary

The video predicts salary (in thousands) from years of experience and whether the person had a gap in their career.

ExperienceGapSalaryResidual
2Yes40K−11
2.5Yes42K−9
3No52K1
4No60K9
4.5Yes62K11

The base model outputs the average salary for every input: 51K. Each residual is the salary minus 51. The root of the first tree holds all five: [−11, −9, 1, 9, 11].

The video first writes 41K in the second row, which makes the average exactly 51, and then changes it to 42K to make the arithmetic easier while keeping 51 as the base. With 42K the average is 51.2, so the residuals add up to 1 instead of 0. The numbers here follow the board: 42K with a base of 51.

The first split is on experience, a continuous feature, so it is a threshold: ≤2 and >2. The left branch gets [−11], the right branch [−9, 1, 9, 11].

Similarity weights and gain of the first split · from the Complete Machine Learning in 6 Hours video · 371:32 to 374:50

Calculating the similarity weight with λ = 1

For regression the bottom of the similarity weight is the number of residuals plus λ. A larger λ penalises the similarity more. The video first works the left leaf with λ = 0, which gives 121 / 1 = 121, then sets λ = 1 for the rest.

  • ≤2: (−11)² / (1 + 1) = 121 / 2 = 60.5.
  • >2: (−9 + 1 + 9 + 11)² / (4 + 1) = 144 / 5 = 28.8.
  • Root: (−11 − 9 + 1 + 9 + 11)² / (5 + 1) = 1 / 6 ≈ 0.17.

The board writes 121 / 2 as 65.5 and the gain as 98.34. It corrects 65.5 to 60.5 but not the gain; with 60.5 the gain is 89.13, the number the video says aloud.

A better split and the regression output · from the Complete Machine Learning in 6 Hours video · 374:50 to 378:49

Choosing the better split at 2.5

The next candidate threshold is ≤2.5 and >2.5. The left branch gets [−11, −9]: (−20)² / (2 + 1) = 400 / 3 = 133.33. The right branch gets [1, 9, 11]: 21² / (3 + 1) = 441 / 4 = 110.25. The gain is 133.33 + 110.25 − 0.17 = 243.42, far above 89.13, so the tree splits at 2.5.

Two candidate splits of the XGBoost regressor on experience: splitting at 2 gives similarity weights 60.5 and 28.8 and a gain of 89.13; splitting at 2.5 gives 133.33 and 110.25 and a gain of 243.42, so the split at 2.5 wins.

Predicting a salary from the base value and the trees

A new record starts at the base value 51, then follows the tree and adds the learning rate times the leaf's value. More trees add more terms: 51 + α₁·tree₁ + α₂·tree₂ + … + αₙ·treeₙ.

The video uses the average of each leaf as its value: −10 for [−11, −9] and 7 for [1, 9, 11]. XGBoost's leaf output divides by the count plus λ, so with λ = 1 the values are −20 / 3 = −6.67 and 21 / 4 = 5.25. The video's averages are the leaf outputs for λ = 0. The library also uses one learning rate for all trees.

Computing every split of the salary table

The table typed in and a loop over the four thresholds between the experience values.

The similarity weight for regression

python
base = 51                       # the board's base model output
res = salary - base
lam = 1

def sim(r):
    return r.sum() ** 2 / (len(r) + lam)    # (sum of residuals)^2 / (count + lambda)
ExampleFrom the video, run with NumPy
import numpy as np

exp = np.array([2, 2.5, 3, 4, 4.5])
salary = np.array([40, 42, 52, 60, 62])

base = 51
res = salary - base
lam = 1

def sim(r):
    return r.sum() ** 2 / (len(r) + lam)

print("residuals:", res, "  root similarity:", round(sim(res), 3))
for t in (2, 2.5, 3, 4):
    left, right = res[exp <= t], res[exp > t]
    print(f"Exp <= {t:<3}  left {sim(left):6.2f}  right {sim(right):6.2f}  gain {sim(left) + sim(right) - sim(res):6.2f}")

left, right = res[exp <= 2.5], res[exp > 2.5]
print("leaf outputs (lambda = 1):", round(left.sum() / (len(left) + lam), 2), round(right.sum() / (len(right) + lam), 2))
print("leaf averages (the video):", left.mean(), right.mean())

What the four candidate splits show

  • The ≤2 split gives 60.5, 28.8 and a gain of 89.13; the ≤2.5 split gives 133.33, 110.25 and 243.42, the board's corrected numbers.
  • The ≤3 split scores 223.42 and ≤4 only 80.33, so 2.5 is the best threshold of the four.
  • The leaf outputs are −6.67 and 5.25 with λ = 1, smaller than the averages −10 and 7.

Splitting on gap and computing the second-round residuals

The right branch [1, 9, 11] can be split again, on gap: Yes gets [11] (record 5) and No gets [1, 9] (records 3 and 4). With the board's leaf values, the average residual in each leaf, the first tree outputs −10, 11 and 5. With a learning rate of 0.1 the predictions are 51 + 0.1 × (−10) = 50.0 for records 1 and 2, 51 + 0.1 × 5 = 51.5 for records 3 and 4, and 51 + 0.1 × 11 = 52.1 for record 5.

The first XGBoost regressor tree splits experience at 2.5 and the right branch on gap: the leaves hold the residuals [−11, −9], [11] and [1, 9] with values −10, 11 and 5; with a learning rate of 0.1 the predictions are 50.0, 50.0, 51.5, 51.5 and 52.1, leaving second-round residuals of −10, −8, 0.5, 8.5 and 9.9.

The new residuals, salary minus prediction, are −10, −8, 0.5, 8.5 and 9.9. Each one is a little closer to 0 than the first round's −11, −9, 1, 9 and 11, and the second tree is trained on them.

λ also decides whether a split is made at all. With λ = 1 the gap split scores 121/2 + 100/3 − 441/4 = 60.5 + 33.33 − 110.25 = −16.42. A negative gain means the split makes the tree worse, so XGBoost keeps [1, 9, 11] as one leaf or splits it another way. With λ = 0 the same split gains 121 + 50 − 147 = 24.

The second round and the gap split as code

python
gap = np.array([1, 1, 0, 0, 1])                                       # 1 = Yes
leaf = np.where(exp <= 2.5, -10.0, np.where(gap == 1, 11.0, 5.0))      # the board's leaf values
pred = base + 0.1 * leaf                                              # learning rate 0.1
res2 = salary - pred                                                  # the targets of tree 2
ExampleFrom the video's salary table, run with NumPy
gap = np.array([1, 1, 0, 0, 1])
leaf = np.where(exp <= 2.5, -10.0, np.where(gap == 1, 11.0, 5.0))
pred = base + 0.1 * leaf
res2 = salary - pred
print("predictions after tree 1:", pred)
print("second-round residuals:  ", res2)

right, g = res[exp > 2.5], gap[exp > 2.5]
for lam_value in (1, 0):
    s = lambda r: r.sum() ** 2 / (len(r) + lam_value)
    print(f"lambda = {lam_value}: gain of the gap split {s(right[g == 1]) + s(right[g == 0]) - s(right):.2f}")

What the second round shows

  • The predictions are 50, 50, 51.5, 51.5 and 52.1, and the residuals −10, −8, 0.5, 8.5 and 9.9, the numbers in the diagram.
  • With λ = 1 the gap split has a gain of −16.42, so XGBoost would not make it; with λ = 0 it gains 24. The board's tree uses the averages, which are the λ = 0 leaf values.

Fitting XGBRegressor on the salary table

XGBRegressor comes from the same xgboost package as the classifier; Installing scikit-learn has the install line and the extra brew install libomp step on macOS.

Matching the board with one tree of depth 1

base_score=51 is the base model, reg_lambda=1 is λ, and learning_rate=1.0 keeps the leaf values unscaled so they match the formula. max_depth=1 allows one split. Left unset, base_score would be estimated from the data as the mean, 51.2.

python
from xgboost import XGBRegressor

reg = XGBRegressor(n_estimators=1, max_depth=1, learning_rate=1.0, reg_lambda=1,
                   base_score=51, tree_method="exact")
reg.fit(X, salary)
ExampleFrom the video, run on XGBoost 3.4.1
import pandas as pd
from xgboost import XGBRegressor

X = pd.DataFrame({"exp": exp, "gap": [1, 1, 0, 0, 1]})     # gap: 1 = Yes, 0 = No
reg = XGBRegressor(n_estimators=1, max_depth=1, learning_rate=1.0, reg_lambda=1,
                   base_score=51, tree_method="exact")
reg.fit(X, salary)
print(reg.get_booster().get_dump(with_stats=True)[0])
print("predictions:", reg.predict(X).round(2))

for lam_value in (0, 1, 10):
    r = XGBRegressor(n_estimators=1, max_depth=1, learning_rate=1.0, reg_lambda=lam_value,
                     base_score=51, tree_method="exact").fit(X, salary)
    print(f"reg_lambda={lam_value:<3} leaf outputs {np.round(r.predict(X)[[0, 4]] - 51, 2)}")

What the one-tree regressor printed

  • The split is exp < 2.75 with gain 243.4. XGBoost places the threshold halfway between 2.5 and 3, the same partition as ≤2.5, and its gain is the board's 243.42.
  • The leaves are −6.67 and 5.25, so the predictions are 51 − 6.67 = 44.33 and 51 + 5.25 = 56.25.
  • λ shrinks the leaves: with λ = 0 they are −10 and 7, the video's averages; with λ = 10 they fall to −1.67 and 1.62.

Predicting diabetes progression with more and more trees

On the diabetes dataset (442 patients, 10 measurements, a disease progression score) the number of trees decides how far the boosting goes. r2_score takes the true values first.

ExampleRun on XGBoost 3.4.1
from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split
from sklearn.metrics import r2_score
from xgboost import XGBRegressor

X, y = load_diabetes(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0)
for n in (1, 10, 100, 300):
    reg = XGBRegressor(n_estimators=n, max_depth=2, learning_rate=0.1, reg_lambda=1,
                       random_state=0).fit(X_train, y_train)
    print(f"{n:3} trees: test R2 {r2_score(y_test, reg.predict(X_test)):.3f}")

Reading the R2 scores

  • One tree scores 0.079: with a learning rate of 0.1 it moves the predictions only a little away from the base value.
  • 10 to 100 trees bring the test R2 to 0.337 and 0.340.
  • 300 trees drop to 0.260. The trees keep fitting the training residuals, and the test score falls: boosting overfits when it runs too long, which is why the number of trees is tuned.

XGBoost regressor vs XGBoost classifier

RegressorClassifier
Base modelaverage of the target (51)probability 0.5, log-odds 0
Residualvalue − predictionlabel − probability
Bottom of the similarity weightnumber of residuals + λΣ p(1 − p) + λ
Leaf outputΣ residuals / (count + λ)Σ residuals / (Σ p(1 − p) + λ)
Final outputbase + η · Σ tree outputssigmoid of 0 + η · Σ tree outputs

Where you use the XGBoost regressor

  • Salary and price prediction from mixed columns, like the experience and gap table.
  • Demand forecasting with features such as day, store and promotion.
  • Any tabular regression where a random forest is good and you want to try a boosted model with tuning.
Watch out. More trees are not always better. On the diabetes data 300 trees scored lower than 100. Use early_stopping_rounds with a validation set, or tune n_estimators together with learning_rate by Cross-validation.
Try it yourself
  • Change 42 back to 41 in salary and set base = salary.mean(); check that the residuals now add up to 0.
  • Set max_depth=2 in the one-tree XGBRegressor and print the dump: with λ = 1 the right branch splits on exp<3.5, not on gap.
  • Set max_depth=3 in the diabetes loop and see whether 300 trees overfit sooner.

Little by little, you're building something great.