Cross-validation
Cross-validation is a model evaluation method that splits the training data into several folds, trains on all but one fold, tests on the one left out, and repeats until every fold has been the test fold once.
Last updated: 05 Oct, 2026 · scikit-learn 1.9.1
One train and test split gives one score, and a lucky or unlucky split moves it. Cross-validation gives several scores from different splits, so their mean is steadier.
Scoring with cross_val_score
The video imports LinearRegression from sklearn.linear_model and cross_val_score from sklearn.model_selection, and explains the idea on 100 records with five folds: in the first round the first block is the test data and the rest is training data, in the second round the next block is the test data, and so on, five times.

The call takes the model, X and y, a scoring name and cv, the number of folds. The video scores with neg_mean_squared_error and cv=5, gets five values, and averages them with np.mean. The minus sign comes from the "neg" scorer: Boston's mean squared error is 37.13. The video passes the whole of X and y, and says the better practice is to do the train and test split first and cross-validate on the training data only.
The video uses load_boston(), which scikit-learn removed in 1.2; the code here uses fetch_california_housing(). The steps are the same, the numbers are not.
Why the scorer is negative
scikit-learn's scorers follow one rule: greater is better. An error is better when smaller, so its scorer returns the error with a minus sign. Read -0.52 as a mean squared error of 0.52, and closer to zero is better.
The folds behind cv=5
For a regression model, cv=5 means KFold(n_splits=5) without shuffling: five blocks in row order, as in the video's 100-record picture.
import numpy as np
from sklearn.model_selection import KFold
records = np.arange(100)
for fold, (train_idx, test_idx) in enumerate(KFold(n_splits=5).split(records), start=1):
print(f"fold {fold}: test rows {test_idx.min()}-{test_idx.max()}, {len(train_idx)} train rows")fold 1: test rows 0-19, 80 train rows fold 2: test rows 20-39, 80 train rows fold 3: test rows 40-59, 80 train rows fold 4: test rows 60-79, 80 train rows fold 5: test rows 80-99, 80 train rows
Cross-validating the whole table
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_val_score
lin_reg = LinearRegression()
mse = cross_val_score(lin_reg, X, y, scoring="neg_mean_squared_error", cv=5)Cross-validating California housing as in the video
print("five scores:", mse.round(4))
print("mean:", round(np.mean(mse), 4))five scores: [-0.4849 -0.6225 -0.6462 -0.5432 -0.4947] mean: -0.5583
Predicting after cross_val_score
The video then calls lin_reg.predict and gets an error. cross_val_score fits copies of the model, one per fold, and leaves lin_reg itself untouched. This run first splits the table with the same call as Multiple linear regression, train_test_split(X, y, test_size=0.33, random_state=42):
mse_train = cross_val_score(lin_reg, X_train, y_train, scoring="neg_mean_squared_error", cv=5)
lin_reg.predict(X_test) # the next step in the videoTraceback (most recent call last):
File "main.py", line 2, in <module>
lin_reg.predict(X_test) # the next step in the video
sklearn.exceptions.NotFittedError: This LinearRegression instance is not fitted yet. Call 'fit' with appropriate arguments before using this estimator.Splitting first, then cross-validating the training rows
from sklearn.metrics import mean_squared_error
mse_train = cross_val_score(lin_reg, X_train, y_train, scoring="neg_mean_squared_error", cv=5)
print("five scores on the training rows:", mse_train.round(4))
print("mean:", round(np.mean(mse_train), 4))
lin_reg.fit(X_train, y_train) # fit the model itself before predicting
print("test MSE, scored once at the end:", round(mean_squared_error(y_test, lin_reg.predict(X_test)), 4))five scores on the training rows: [-0.541 -0.4987 -0.5048 -0.52 -0.5507] mean: -0.523 test MSE, scored once at the end: 0.537
What the fold scores show
- The folds take turns: rows 0-19 test first, then 20-39, and so on; each round trains on 80 rows.
- On the whole table the five scores spread from about −0.48 to −0.65. California's rows are stored region by region, and unshuffled folds test one region at a time.
- On the shuffled training rows the scores sit closer, from about −0.50 to −0.55, with a mean MSE of 0.523.
- The NotFittedError is the video's moment:
cross_val_scorenever fitslin_regitself. Onefitcall fixes it, and the test MSE comes out close to the cross-validated mean.
Holding out a validation set
The notes start from 1000 records. 70% go to training and 30% to the test set, which is kept back to check the final model. The training part is split again into train and validation data, used to train the model and to tune its hyperparameters. Which rows land in validation depends on the random_state: one split may score 85%, another 92%, another 78%. Cross-validation stops that luck from deciding, by giving every training row a turn in validation.
Choosing among the types of cross-validation
The notes list five ways to give rows their turn, on a training set of 500 records:
Leave one out (LOOCV)
Each experiment validates on one record and trains on the other 499: experiment 1 leaves out the first record, experiment 2 the second, up to experiment 500. The 500 scores are averaged. The cost is 500 trainings, and the notes add a risk: every model trains on nearly all the data, so validation scores can look better than new test data will. scikit-learn's guide puts it this way: leave one out often gives a high-variance estimate of the test error, and 5-fold or 10-fold is usually preferred.
Leave p out
Leave p records out instead of one, with p = 3, 4 or 5 in the video and 10, 20 or 30 in the notes, and try every possible group of p. The number of experiments grows very fast with p, so the video calls it hardly ever done on real data with millions of records.
K-fold
With k = 5 and 500 records, each fold holds 500 / 5 = 100 records. Experiment 1 validates on the first 100, experiment 2 on the next 100, and so on; the five accuracies are averaged. This is cv=5 above.
Stratified k-fold
For classification, plain k-fold can fill a fold with mostly one class. The video's example has 900 records of one class and 100 of the other, a 9:1 ratio, and stratified k-fold keeps each class's share inside every fold. If the training data has 60 ones for every 40 zeros, stratified k-fold keeps that 60:40 ratio inside every fold. For a classifier with a binary or multiclass output, scikit-learn's cv=5 already means StratifiedKFold(5); for a regressor it means KFold(5).
Time series cross-validation
The notes' example is product sentiment analysis on reviews from January to December. Time order matters: a model must not train on December and validate on March. So training takes day 1 to day 4 and validation the days after, then the window moves forward. Shuffling is never used here.

The splitters
import numpy as np
from sklearn.model_selection import LeaveOneOut, LeavePOut, KFold, StratifiedKFold, TimeSeriesSplit
rows = np.zeros((500, 1)) # 500 training records
y = np.array([1] * 300 + [0] * 200) # 60% ones, 40% zeros, sorted by class
days = np.arange(12).reshape(-1, 1) # 12 days in time orderGenerating each split in scikit-learn
print("leave one out:", LeaveOneOut().get_n_splits(rows), "experiments")
print("leave 2 out: ", LeavePOut(p=2).get_n_splits(rows), "experiments")
for name, cv in [("KFold ", KFold(n_splits=5)), ("StratifiedKFold", StratifiedKFold(n_splits=5))]:
ones = [int(y[val].sum()) for _, val in cv.split(rows, y)]
print(name, "ones in each validation fold of 100:", ones)
for train, val in TimeSeriesSplit(n_splits=4).split(days):
print("train days", train.min(), "to", train.max(), "-> validate days", val.tolist())leave one out: 500 experiments leave 2 out: 124750 experiments KFold ones in each validation fold of 100: [100, 100, 100, 0, 0] StratifiedKFold ones in each validation fold of 100: [60, 60, 60, 60, 60] train days 0 to 3 -> validate days [4, 5] train days 0 to 5 -> validate days [6, 7] train days 0 to 7 -> validate days [8, 9] train days 0 to 9 -> validate days [10, 11]
Reading the splitters
- Leave one out makes 500 experiments on 500 records, and leave 2 out makes 124,750: every pair of records. Leave p out is rarely practical beyond tiny tables.
- Plain KFold on sorted labels gives folds of all ones or all zeros: 100, 100, 100, then 0, 0. A model validated on a fold with no zeros learns nothing about its errors on zeros.
- StratifiedKFold puts 60 ones in every fold of 100, the 60:40 ratio of the whole set.
- TimeSeriesSplit always validates on later days, and the training window grows: days 0 to 3 validate on 4 and 5, then days 0 to 5 on 6 and 7, and so on.
Cross-validation vs a single train and test split
| Single split | Cross-validation (cv=5) | |
|---|---|---|
| Scores | One | Five, then their mean |
| Models trained | One | Five copies, plus a final fit |
| Each row is tested | Only if it lands in the test set | Exactly once |
| Sensitive to one lucky split | Yes | Much less |
| Cost | Fast | About five times the training time |
Comparing the types of cross-validation
| Type | Validation rows per experiment | Experiments on 500 rows | Use it when |
|---|---|---|---|
| Leave one out | 1 | 500 | The table is very small |
| Leave p out | p | Every group of p: 124,750 for p = 2 | Almost never; tiny tables |
| K-fold | n / k | k | Regression, the default choice |
| Stratified k-fold | n / k, same class ratio | k | Classification |
| Time series | The next block in time | The number of splits | Data in time order |
Where you use cross-validation
- Choosing between models on the training data, keeping the test set for one final check.
- Tuning hyperparameters: Hyperparameter tuning with GridSearchCV runs this same
scoringandcvfor every candidate value. - Small datasets, where a single test set is too small to trust.
- Imbalanced classes or data in time order: pass
cv=StratifiedKFold(5)orcv=TimeSeriesSplit(5)to the samecross_val_scorecall.
Related
- Previous: Polynomial regression
- Next: Overfitting and underfitting
- See also: Hyperparameter tuning with GridSearchCV
- Reference: Cross-validation in the scikit-learn user guide
- Change
cv=5tocv=10on the training rows and compare the mean. - Use
scoring="r2": the five scores become R² values, all positive and near 0.6. - On the whole table, pass
cv=KFold(n_splits=5, shuffle=True, random_state=42): the five scores move much closer together. - Change
TimeSeriesSplit(n_splits=4)ton_splits=3and read which days validate now.
This is what real progress feels like.