Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Cross-validation

Cross-validation is a model evaluation method that splits the training data into several folds, trains on all but one fold, tests on the one left out, and repeats until every fold has been the test fold once.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

One train and test split gives one score, and a lucky or unlucky split moves it. Cross-validation gives several scores from different splits, so their mean is steadier.

Scoring with cross_val_score

cross_val_score with five folds · from the Complete Machine Learning in 6 Hours video · 144:26 to 148:07

The video imports LinearRegression from sklearn.linear_model and cross_val_score from sklearn.model_selection, and explains the idea on 100 records with five folds: in the first round the first block is the test data and the rest is training data, in the second round the next block is the test data, and so on, five times.

Five-fold cross-validation on 100 records: in each of five rounds a different block of 20 records is the test fold and the other 80 train the model, giving five scores that are averaged.

The call takes the model, X and y, a scoring name and cv, the number of folds. The video scores with neg_mean_squared_error and cv=5, gets five values, and averages them with np.mean. The minus sign comes from the "neg" scorer: Boston's mean squared error is 37.13. The video passes the whole of X and y, and says the better practice is to do the train and test split first and cross-validate on the training data only.

The video uses load_boston(), which scikit-learn removed in 1.2; the code here uses fetch_california_housing(). The steps are the same, the numbers are not.

Why the scorer is negative

scikit-learn's scorers follow one rule: greater is better. An error is better when smaller, so its scorer returns the error with a minus sign. Read -0.52 as a mean squared error of 0.52, and closer to zero is better.

The folds behind cv=5

For a regression model, cv=5 means KFold(n_splits=5) without shuffling: five blocks in row order, as in the video's 100-record picture.

ExampleThe video's 100 records and five folds
import numpy as np
from sklearn.model_selection import KFold

records = np.arange(100)
for fold, (train_idx, test_idx) in enumerate(KFold(n_splits=5).split(records), start=1):
    print(f"fold {fold}: test rows {test_idx.min()}-{test_idx.max()}, {len(train_idx)} train rows")

Cross-validating the whole table

python
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_val_score

lin_reg = LinearRegression()
mse = cross_val_score(lin_reg, X, y, scoring="neg_mean_squared_error", cv=5)

Cross-validating California housing as in the video

ExampleThe video's call on California housing, run on scikit-learn 1.9.1
print("five scores:", mse.round(4))
print("mean:", round(np.mean(mse), 4))

Predicting after cross_val_score

The video then calls lin_reg.predict and gets an error. cross_val_score fits copies of the model, one per fold, and leaves lin_reg itself untouched. This run first splits the table with the same call as Multiple linear regression, train_test_split(X, y, test_size=0.33, random_state=42):

ExampleThe not-fitted error from the video's practical
mse_train = cross_val_score(lin_reg, X_train, y_train, scoring="neg_mean_squared_error", cv=5)
lin_reg.predict(X_test)        # the next step in the video

Splitting first, then cross-validating the training rows

ExampleThe video's best practice: split first, cross-validate on the training data
from sklearn.metrics import mean_squared_error

mse_train = cross_val_score(lin_reg, X_train, y_train, scoring="neg_mean_squared_error", cv=5)
print("five scores on the training rows:", mse_train.round(4))
print("mean:", round(np.mean(mse_train), 4))

lin_reg.fit(X_train, y_train)  # fit the model itself before predicting
print("test MSE, scored once at the end:", round(mean_squared_error(y_test, lin_reg.predict(X_test)), 4))

What the fold scores show

  • The folds take turns: rows 0-19 test first, then 20-39, and so on; each round trains on 80 rows.
  • On the whole table the five scores spread from about −0.48 to −0.65. California's rows are stored region by region, and unshuffled folds test one region at a time.
  • On the shuffled training rows the scores sit closer, from about −0.50 to −0.55, with a mean MSE of 0.523.
  • The NotFittedError is the video's moment: cross_val_score never fits lin_reg itself. One fit call fixes it, and the test MSE comes out close to the cross-validated mean.

Holding out a validation set

The notes start from 1000 records. 70% go to training and 30% to the test set, which is kept back to check the final model. The training part is split again into train and validation data, used to train the model and to tune its hyperparameters. Which rows land in validation depends on the random_state: one split may score 85%, another 92%, another 78%. Cross-validation stops that luck from deciding, by giving every training row a turn in validation.

Choosing among the types of cross-validation

The notes list five ways to give rows their turn, on a training set of 500 records:

Leave one out (LOOCV)

Leave one out and leave p out · from the All Type Of Cross Validation With Python video · 11:51 to 14:06

Each experiment validates on one record and trains on the other 499: experiment 1 leaves out the first record, experiment 2 the second, up to experiment 500. The 500 scores are averaged. The cost is 500 trainings, and the notes add a risk: every model trains on nearly all the data, so validation scores can look better than new test data will. scikit-learn's guide puts it this way: leave one out often gives a high-variance estimate of the test error, and 5-fold or 10-fold is usually preferred.

Leave p out

Leave p records out instead of one, with p = 3, 4 or 5 in the video and 10, 20 or 30 in the notes, and try every possible group of p. The number of experiments grows very fast with p, so the video calls it hardly ever done on real data with millions of records.

K-fold

With k = 5 and 500 records, each fold holds 500 / 5 = 100 records. Experiment 1 validates on the first 100, experiment 2 on the next 100, and so on; the five accuracies are averaged. This is cv=5 above.

Stratified k-fold

Stratified k-fold on imbalanced data · from the All Type Of Cross Validation With Python video · 9:17 to 11:51

For classification, plain k-fold can fill a fold with mostly one class. The video's example has 900 records of one class and 100 of the other, a 9:1 ratio, and stratified k-fold keeps each class's share inside every fold. If the training data has 60 ones for every 40 zeros, stratified k-fold keeps that 60:40 ratio inside every fold. For a classifier with a binary or multiclass output, scikit-learn's cv=5 already means StratifiedKFold(5); for a regressor it means KFold(5).

Time series cross-validation

The notes' example is product sentiment analysis on reviews from January to December. Time order matters: a model must not train on December and validate on March. So training takes day 1 to day 4 and validation the days after, then the window moves forward. Shuffling is never used here.

Four splitting schemes drawn as rows of experiments: leave one out, where a single record validates in each of 500 experiments; k-fold with k = 5, where each block of 100 rows validates once; stratified k-fold, where every fold keeps the 60 to 40 ratio of ones and zeros; and time series, where the model trains on the earlier days and validates on the next ones.

The splitters

python
import numpy as np
from sklearn.model_selection import LeaveOneOut, LeavePOut, KFold, StratifiedKFold, TimeSeriesSplit

rows = np.zeros((500, 1))   # 500 training records
y = np.array([1] * 300 + [0] * 200)   # 60% ones, 40% zeros, sorted by class
days = np.arange(12).reshape(-1, 1)   # 12 days in time order

Generating each split in scikit-learn

ExampleThe notes' 500 records and 60:40 labels, run on scikit-learn 1.9.1
print("leave one out:", LeaveOneOut().get_n_splits(rows), "experiments")
print("leave 2 out:  ", LeavePOut(p=2).get_n_splits(rows), "experiments")

for name, cv in [("KFold          ", KFold(n_splits=5)), ("StratifiedKFold", StratifiedKFold(n_splits=5))]:
    ones = [int(y[val].sum()) for _, val in cv.split(rows, y)]
    print(name, "ones in each validation fold of 100:", ones)

for train, val in TimeSeriesSplit(n_splits=4).split(days):
    print("train days", train.min(), "to", train.max(), "-> validate days", val.tolist())

Reading the splitters

  • Leave one out makes 500 experiments on 500 records, and leave 2 out makes 124,750: every pair of records. Leave p out is rarely practical beyond tiny tables.
  • Plain KFold on sorted labels gives folds of all ones or all zeros: 100, 100, 100, then 0, 0. A model validated on a fold with no zeros learns nothing about its errors on zeros.
  • StratifiedKFold puts 60 ones in every fold of 100, the 60:40 ratio of the whole set.
  • TimeSeriesSplit always validates on later days, and the training window grows: days 0 to 3 validate on 4 and 5, then days 0 to 5 on 6 and 7, and so on.

Cross-validation vs a single train and test split

Single splitCross-validation (cv=5)
ScoresOneFive, then their mean
Models trainedOneFive copies, plus a final fit
Each row is testedOnly if it lands in the test setExactly once
Sensitive to one lucky splitYesMuch less
CostFastAbout five times the training time

Comparing the types of cross-validation

TypeValidation rows per experimentExperiments on 500 rowsUse it when
Leave one out1500The table is very small
Leave p outpEvery group of p: 124,750 for p = 2Almost never; tiny tables
K-foldn / kkRegression, the default choice
Stratified k-foldn / k, same class ratiokClassification
Time seriesThe next block in timeThe number of splitsData in time order

Where you use cross-validation

  • Choosing between models on the training data, keeping the test set for one final check.
  • Tuning hyperparameters: Hyperparameter tuning with GridSearchCV runs this same scoring and cv for every candidate value.
  • Small datasets, where a single test set is too small to trust.
  • Imbalanced classes or data in time order: pass cv=StratifiedKFold(5) or cv=TimeSeriesSplit(5) to the same cross_val_score call.
Watch out. Cross-validating on all of X and y and then reporting a test score on rows from the same X lets the test rows into the folds, so the score looks better than the model is. Split first, cross-validate on X_train, and keep X_test for one final score.
Try it yourself
  • Change cv=5 to cv=10 on the training rows and compare the mean.
  • Use scoring="r2": the five scores become R² values, all positive and near 0.6.
  • On the whole table, pass cv=KFold(n_splits=5, shuffle=True, random_state=42): the five scores move much closer together.
  • Change TimeSeriesSplit(n_splits=4) to n_splits=3 and read which days validate now.

This is what real progress feels like.