Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

XGBoost classifier

XGBoost (extreme gradient boosting) is a boosting algorithm that builds binary decision trees one after another on the residuals of the model so far, choosing each split by the gain in similarity weight.

Last updated: 05 Oct, 2026 · XGBoost 3.4.1

Gradient boosting fits each new tree to the residuals of the model so far. XGBoost follows the same plan with its own way of building each tree: a similarity weight, a gain and a penalty λ. It solves both classification and regression; this lesson is the classifier, on the same loan table as AdaBoost.

The loan approval table and the base model · from the Complete Machine Learning in 6 Hours video · 343:29 to 347:20

Starting from a base model and residuals

The video's dataset decides loan approval from salary and credit score. Salary is either ≤50K or >50K; credit is bad (B), good (G) or normal (N); approval is 0 or 1.

SalaryCreditApprovalResidual
≤50KB0−0.5
≤50KG10.5
≤50KG10.5
>50KB0−0.5
>50KG10.5
>50KN10.5
≤50KN0−0.5

XGBoost first creates a base model, a weak learner that outputs a probability of 0.5 for every record whatever the input. The residual of a record is its approval minus that probability: 0 − 0.5 = −0.5 and 1 − 0.5 = 0.5. The next trees are trained on these residuals, in sequence.

Building the binary tree on salary · from the Complete Machine Learning in 6 Hours video · 347:20 to 351:27

Building a binary tree on the residuals

The steps are: (1) build a binary decision tree using the features, (2) calculate the similarity weight of each node, (3) calculate the gain of the split. The tree is always binary, so a feature with three categories, like credit, is split into two groups.

The video starts with salary. The root holds all seven residuals, [−0.5, 0.5, 0.5, −0.5, 0.5, 0.5, −0.5]. The ≤50K branch gets the residuals of records 1, 2, 3 and 7: [−0.5, 0.5, 0.5, −0.5]. The >50K branch gets records 4, 5 and 6: [−0.5, 0.5, 0.5].

The video once says “probability minus approval”; the board and the numbers use approval minus probability, the correct order.

Similarity weight of each branch · from the Complete Machine Learning in 6 Hours video · 351:27 to 355:48

Calculating the similarity weight

The similarity weight scores how alike the residuals in a node are. The residuals are added first and the sum is squared. The bottom adds p(1 − p) for every record in the node, where p is the probability from the base model, plus λ (lambda), a hyperparameter that stops the tree from overfitting. The video sets λ = 0 for the worked numbers.

The board writes the bottom as Σ(p(1 − p) + λ), which would add λ once per record; λ is added once. For the ≤50K branch the residuals cancel: (−0.5 + 0.5 + 0.5 − 0.5)² = 0, so its similarity weight is 0. For the >50K branch, (−0.5 + 0.5 + 0.5)² = 0.25 and 3 × 0.5 × 0.5 = 0.75, so it is 0.25 / 0.75 = 0.33.

Similarity of the root and the gain · from the Complete Machine Learning in 6 Hours video · 355:48 to 358:11

Calculating the gain of the salary split

The root's residuals add up to 0.5, so its similarity weight is 0.25 / (7 × 0.25) = 0.25 / 1.75 = 0.14. The gain of a split is the similarity of the two branches minus the similarity of the node they came from.

The board calls this information gain, the same name as in Information gain, but here it is computed from similarity weights. XGBoost compares the gain of every candidate split and keeps the highest. The video assumes salary gives the highest gain so that it can show the steps; the code below checks that assumption.

Splitting the left branch on credit · from the Complete Machine Learning in 6 Hours video · 358:11 to 361:45

Splitting the ≤50K branch on credit

The ≤50K branch is split again, on credit. Credit has three categories, so the binary split puts B on one side and G, N on the other. B gets record 1, residual [−0.5]: similarity 0.25 / 0.25 = 1. G, N gets records 2, 3 and 7, [0.5, 0.5, −0.5]: similarity 0.25 / 0.75 = 0.33. The node above has similarity 0, so the gain is 1 + 0.33 − 0 = 1.33.

The >50K branch can be split on credit too, with B and G on one side and N on the other: [−0.5, 0.5] has similarity 0 / 0.5 = 0 and [0.5] has 0.25 / 0.25 = 1, so the gain is 0 + 1 − 0.33 = 0.67. The tree now has four leaves.

The XGBoost classifier tree: the Salary root holds all seven residuals with a similarity weight of 0.14 and the split has a gain of 0.19; the 50K-or-less branch splits on credit B against G, N with weights 1 and 0.33 for a gain of 1.33, and the over-50K branch splits on credit B, G against N with weights 0 and 1 for a gain of 0.67; the four leaves output −2, 0.67, 0 and 2.
Predicting with the base model, trees and sigmoid · from the Complete Machine Learning in 6 Hours video · 361:56 to 365:27

Predicting a probability with the sigmoid

To predict record 1, it first goes to the base model. The probability 0.5 is turned into log-odds, log(p / (1 − p)) = log(0.5 / 0.5) = 0. The record then follows the tree: ≤50K, then credit B. The leaf's output is multiplied by the learning rate α and added. A sigmoid turns the total into a probability between 0 and 1. With more trees the output is σ(0 + α·tree₁ + α·tree₂ + … + α·treeₙ): each tree's output is added slowly, which is what makes it boosting.

Two points differ from the board. First, the number a leaf adds is its output value, Σ residuals / (Σ p(1 − p) + λ), not its similarity weight. For the B leaf that is −0.5 / 0.25 = −2, so record 1 gets σ(0 + α·(−2)), not σ(0 + α·1). The similarity weight only scores splits. Second, XGBoost uses one learning rate for every tree, and it builds the trees one after another, each on the residuals left by the trees before it. The video says the trees are built in parallel; XGBoost runs in parallel only inside one tree, when it searches for the best split.

A record starts at the base model with probability 0.5, which is log-odds 0; each tree adds the learning rate times its leaf output, and a sigmoid turns the sum into a probability between 0 and 1.

XGBoost is a black box model: with hundreds of trees nobody follows the calculation by hand. It still overfits, so its trees are limited in advance (pre-pruning) with settings such as max_depth, and λ is tuned with cross-validation.

Computing the board's similarity weights and gains

The table typed in, a function for the similarity weight, and the two splits from the board. The last line scores credit at the root, the split the video does not try.

The similarity weight as a function

python
p = 0.5                                    # the base model's probability for every record
res = approval - p                         # residual = approval - probability

def sim(r, lam=0.0):
    # (sum of residuals) squared, over the sum of p(1 - p) plus lambda
    return r.sum() ** 2 / (len(r) * p * (1 - p) + lam)
ExampleFrom the video, run with NumPy
import numpy as np

approval = np.array([0, 1, 1, 0, 1, 1, 0])
salary = np.array(["<=50K", "<=50K", "<=50K", ">50K", ">50K", ">50K", "<=50K"])
credit = np.array(["B", "G", "G", "B", "G", "N", "N"])

p = 0.5
res = approval - p

def sim(r, lam=0.0):
    return r.sum() ** 2 / (len(r) * p * (1 - p) + lam)

root = sim(res)
left, right = res[salary == "<=50K"], res[salary == ">50K"]
print("residuals:", res)
print(f"root {root:.2f}   <=50K {sim(left):.2f}   >50K {sim(right):.2f}")
print(f"gain of the salary split: {sim(left) + sim(right) - root:.2f}")

cl = credit[salary == "<=50K"]
b, gn = left[cl == "B"], left[cl != "B"]
print(f"under <=50K: B {sim(b):.2f}   G,N {sim(gn):.2f}   gain {sim(b) + sim(gn) - sim(left):.2f}")
print(f"leaf output of B: {b.sum() / (len(b) * p * (1 - p)):.1f}")

cr = credit[salary == ">50K"]
bg, n = right[cr != "N"], right[cr == "N"]
print(f"under >50K: B,G {sim(bg):.2f}   N {sim(n):.2f}   gain {sim(bg) + sim(n) - sim(right):.2f}")

b_all, gn_all = res[credit == "B"], res[credit != "B"]
print(f"credit B vs G,N at the root: gain {sim(b_all) + sim(gn_all) - root:.2f}")

What the similarity weights and gains show

  • The board's numbers come out exactly: root 0.14, ≤50K 0, >50K 0.33, gain 0.19; under ≤50K, B 1, G,N 0.33, gain 1.33.
  • The split of the >50K branch gives B,G 0 and N 1, for a gain of 0.67.
  • The B leaf's output is −2.0, the number a prediction uses, while its similarity weight is 1.
  • Credit at the root has a gain of 3.66, far above salary's 0.19. A real XGBoost run would start with credit, which is what the library does below.

Computing the second-round residuals

The first tree changes every record's probability, and the second tree trains on what is left. With a learning rate of 0.1, record 1 (≤50K, B) lands in the leaf with output −2. Its log-odds become 0 + 0.1 × (−2) = −0.2, its probability σ(−0.2) = 0.450, and its new residual 0 − 0.450 = −0.450, a little closer to 0 than −0.5. Record 2 (≤50K, G) lands in the leaf with output 0.67: 0 + 0.1 × 0.67 = 0.067, probability 0.517, residual 1 − 0.517 = 0.483.

#SalaryCreditApprovalR1Leaf outputp after tree 1R2
1≤50KB0−0.5−20.450−0.450
2≤50KG10.50.670.5170.483
3≤50KG10.50.670.5170.483
4>50KB0−0.500.500−0.500
5>50KG10.500.5000.500
6>50KN10.520.5500.450
7≤50KN0−0.50.670.517−0.517

Records 4 and 5 keep their residuals, because their leaf outputs 0; record 7 moves away from its label, because it shares a leaf with records 2 and 3. The second tree uses the same formulas, but p(1 − p) is no longer 0.25 for everyone: for record 1 it is 0.450 × 0.550 = 0.248.

The second-round residuals as code

python
leaves = [(salary == "<=50K") & (credit == "B"), (salary == "<=50K") & (credit != "B"),
          (salary == ">50K") & (credit != "N"), (salary == ">50K") & (credit == "N")]
log_odds = np.zeros(7)                           # the base model: log(0.5 / 0.5) = 0
for leaf in leaves:
    output = res[leaf].sum() / (leaf.sum() * p * (1 - p))     # the leaf output, lambda = 0
    log_odds[leaf] += 0.1 * output                             # learning rate 0.1
p1 = 1 / (1 + np.exp(-log_odds))                 # the sigmoid
res2 = approval - p1                             # the residuals tree 2 trains on
ExampleThe loan table, run with NumPy
leaves = [(salary == "<=50K") & (credit == "B"), (salary == "<=50K") & (credit != "B"),
          (salary == ">50K") & (credit != "N"), (salary == ">50K") & (credit == "N")]
log_odds = np.zeros(7)
for leaf in leaves:
    output = res[leaf].sum() / (leaf.sum() * p * (1 - p))
    log_odds[leaf] += 0.1 * output
p1 = 1 / (1 + np.exp(-log_odds))
res2 = approval - p1
print("log-odds after tree 1:  ", log_odds.round(3))
print("probabilities:          ", p1.round(3))
print("second-round residuals: ", res2.round(3))
print("p(1 - p) for tree 2:    ", (p1 * (1 - p1)).round(3))

What the second round starts from

  • The log-odds after tree 1 are 0.1 times the leaf output: −0.2, 0.067, 0 and 0.2.
  • The probabilities are 0.45, 0.517, 0.5 and 0.55, and the second-round residuals −0.45, 0.483, −0.5, 0.5, 0.45 and −0.517, as in the table above.
  • p(1 − p) is 0.248 for records 1 and 6 and 0.25 for the rest; it goes into the bottom of the second tree's similarity weights.

Fitting XGBClassifier on the loan table

The XGBoost library's scikit-learn style classes are XGBClassifier and XGBRegressor. It installs with pip install xgboost==3.4.1; on macOS it also needs brew install libomp, as Installing scikit-learn explains.

Matching the board's settings

Each setting maps to a part of the board. base_score=0.5 is the base model. reg_lambda=0 is λ = 0. learning_rate is the α the video multiplies by. min_child_weight=0 lets a leaf hold one record, because one record's p(1 − p) is 0.25, below the default minimum of 1. tree_method="exact" tries every split value.

python
from xgboost import XGBClassifier

model = XGBClassifier(n_estimators=1, max_depth=2, learning_rate=0.3, reg_lambda=0,
                      min_child_weight=0, base_score=0.5, tree_method="exact")

Printing the tree with its gains

get_booster().get_dump(with_stats=True) prints each tree as text: the split, its gain, and each leaf's value. cover is the node's Σ p(1 − p). The printed leaf value already includes the learning rate.

ExampleFrom the video, run on XGBoost 3.4.1
import pandas as pd
from xgboost import XGBClassifier

X = pd.DataFrame({"salary_over_50k": (salary == ">50K").astype(int),
                  "credit_bad": (credit == "B").astype(int)})
model = XGBClassifier(n_estimators=1, max_depth=2, learning_rate=0.3, reg_lambda=0,
                      min_child_weight=0, base_score=0.5, tree_method="exact")
model.fit(X, approval)
print(model.get_booster().get_dump(with_stats=True)[0])

p1 = model.predict_proba(X.iloc[[0]])[0, 1]
print("record 1, probability of approval:", round(float(p1), 4))
print("by hand, sigmoid(0 + 0.3 * -2):    ", round(1 / (1 + np.exp(0.6)), 4))

What the fitted tree says

  • The root splits on credit_bad with gain 3.657, the 3.66 computed by hand, and cover 1.75 = 7 × 0.25.
  • The credit-bad leaf is −0.6: the output −2 times the learning rate 0.3.
  • Under good and normal credit the tree splits on salary (gain 0.533), the other order of the video's two splits.
  • Record 1's probability is 0.3543 both from the library and from σ(0 + 0.3 × (−2)) by hand, so the leaf output, not the similarity weight, is what goes into the prediction.

Classifying tumours with 200 trees

On a real dataset XGBoost runs with many trees and the defaults for the base score. The breast cancer split is the same as in Random forest, so the two scores compare.

python
clf = XGBClassifier(n_estimators=200, max_depth=3, learning_rate=0.1, reg_lambda=1,
                    random_state=0)
clf.fit(X_train, y_train)
ExampleRun on XGBoost 3.4.1
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from xgboost import XGBClassifier

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0)

clf = XGBClassifier(n_estimators=200, max_depth=3, learning_rate=0.1, reg_lambda=1,
                    random_state=0)
clf.fit(X_train, y_train)
print(f"test accuracy: {clf.score(X_test, y_test):.3f}")
print("probabilities of benign, first 3 test tumours:", clf.predict_proba(X_test[:3])[:, 1].round(3))

Reading the tumour results

  • Test accuracy 0.965, a little above the random forest's 0.959 on the same split.
  • The probabilities are the sigmoid outputs of 200 trees added together; a value above 0.5 is predicted benign.

XGBoost vs gradient boosting

XGBoostGradient boosting in scikit-learn
Each new tree learnsthe residuals of the model so farthe residuals of the model so far
How a split is chosengain in similarity weight, which uses Σ p(1 − p) and λhow much the split lowers the squared error of the residuals (criterion "friedman_mse")
Penalties inside the treereg_lambda (λ), gamma (a minimum gain), min_child_weightmax_depth, min_samples_leaf
Base modelbase_score: 0.5 in the video, estimated from the labels when unsetthe log-odds of the class share (classifier), the mean (regressor)
Where it livesthe xgboost package: XGBClassifier, XGBRegressorsklearn.ensemble: GradientBoostingClassifier, GradientBoostingRegressor

Where you use the XGBoost classifier

  • Credit and loan approval on tabular data, like the video's salary and credit example.
  • Fraud and churn prediction, where many weak signals across columns add up.
  • Kaggle-style tabular problems, where gradient boosted trees are the usual strong baseline.
Watch out. When base_score is not set, XGBoost 3.4.1 estimates it from the labels instead of starting at 0.5: on the loan table it starts from 0.571, the share of approvals. Pass base_score=0.5 when you want the video's base model and the board's residuals.
Try it yourself
  • Call sim(left, lam=1) and sim(right, lam=1) to see how λ = 1 shrinks the similarity weights.
  • Set learning_rate=1.0 in the one-tree model and check that the credit-bad leaf prints −2.
  • Remove min_child_weight=0 and print the dump again: every split would leave a leaf with a cover below 1, so the tree is a single leaf.

Every expert started right here.