Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Hyperparameter tuning with GridSearchCV

Hyperparameter tuning is the search for the settings a model does not learn by itself, such as the alpha of Ridge, by scoring every candidate with cross-validation and keeping the best one.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Ridge regression and Lasso regression both depend on λ. GridSearchCV tries every value you list, scores each with cross-validation, and reports the winner.

Ridge, GridSearchCV and alpha · from the Complete Machine Learning in 6 Hours video · 148:24 to 151:26

The video's notebook uses load_boston(), which scikit-learn removed in 1.2; the code here uses fetch_california_housing().

Choosing what to tune

A fitted model predicts with .predict on the test values; the video notes this, then sets linear regression aside for Ridge. Two imports do the work: Ridge from sklearn.linear_model and GridSearchCV from sklearn.model_selection. The parameter to tune is alpha, the λ of "alpha multiplied by slope square". max_iter, how many times the slopes get updated, can be tuned too.

About the score: the video uses scoring="neg_mean_squared_error". scikit-learn scorers always treat a higher number as better, so the mean squared error comes back with a minus sign. A score of −37 means an MSE of 37; closer to 0 is better.

Building the alpha grid · from the Complete Machine Learning in 6 Hours video · 151:26 to 155:21

Writing the parameter grid

The grid is a dictionary: the parameter name as the key, the list of values to try as the value. The video's list runs from tiny to large: 1e-15, 1e-10, 1e-8, 1e-3, 1e-2, 1, 5, 10, 20. GridSearchCV takes the model, the grid, the scoring and cv, the number of cross-validation folds. It fits every value on every fold and keeps the best.

While writing the fit, the video adds: "you can first of all do train test split on X and Y and then probably only do this on X train and Y train". The code below does that.

The first run in the video fails: the grid is written with square brackets.

ExampleFrom the video: the grid written with list brackets
params = ['alpha': [1e-15, 1e-10, 1e-8, 1e-3, 1e-2, 1, 5, 10, 20]]

A key and a value separated by a colon only make sense inside curly braces. "It has become a list; I'm going to make this as dictionary": params = {'alpha': [...]} fixes it.

GridSearchCV takes the alpha grid and the training data, scores every alpha with 5-fold cross-validation, reports the best alpha and its score as best_params_ and best_score_, and refits that model on the whole training set.

GridSearchCV

python
from sklearn.linear_model import Ridge
from sklearn.model_selection import GridSearchCV

params = {"alpha": [1e-15, 1e-10, 1e-8, 1e-3, 1e-2, 1, 5, 10, 20]}
ridge_regressor = GridSearchCV(Ridge(), params, scoring="neg_mean_squared_error", cv=5)
ridge_regressor.fit(X_train, y_train)      # 9 alphas x 5 folds = 45 fits, then one refit
print(ridge_regressor.best_params_)        # the winning alpha
print(ridge_regressor.best_score_)         # its mean cross-validated score

Keeping a validation set for tuning

Choosing alpha needs data the model was not fitted on, and the test set cannot be it: once alpha is picked by its test score, that score is no longer a fair check. So the data is split three ways. Out of 1,000 data points, 70% (700) become the training dataset that trains the model and 30% (300) the test dataset that tests it. The training dataset is split once more: one part trains the model, the other, the validation dataset, scores each hyperparameter value.

A dataset of 1000 data points splits 70 to 30 into a training dataset of 700 that trains the model and a test dataset of 300 that tests it; the training dataset splits again into a train part that trains the model and a validation part used for hyperparameter tuning.

GridSearchCV makes the second split for you. With cv=5 each fold takes its turn as the validation part, so every row of X_train trains some fits and validates others, and the test set stays untouched until the final score.

best_params_, best_score_ and Lasso · from the Complete Machine Learning in 6 Hours video · 155:21 to 160:01

Tuning Ridge and Lasso on California housing

In the video, on the Boston data: linear regression scored −37.13, Ridge picked alpha 20 at −32.38, and Lasso picked alpha 1 at −35.53. With a longer grid ending at 100, Ridge picked alpha 100 at −29.91. Those searches ran on all of X and y. Here the data is split first (as in Train and test split), the search runs on the training part only, and the test part is kept for one final score.

ExampleFrom the video, run on scikit-learn 1.9.1
import numpy as np
import pandas as pd
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split, GridSearchCV, cross_val_score
from sklearn.linear_model import LinearRegression, Ridge, Lasso
from sklearn.metrics import r2_score

X, y = fetch_california_housing(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.33, random_state=42)

mse = cross_val_score(LinearRegression(), X_train, y_train, scoring="neg_mean_squared_error", cv=5)
print("LinearRegression:", round(np.mean(mse), 5))

params = {"alpha": [1e-15, 1e-10, 1e-8, 1e-3, 1e-2, 1, 5, 10, 20]}
ridge_regressor = GridSearchCV(Ridge(), params, scoring="neg_mean_squared_error", cv=5)
ridge_regressor.fit(X_train, y_train)
print("Ridge:", ridge_regressor.best_params_, round(ridge_regressor.best_score_, 5))

lasso_regressor = GridSearchCV(Lasso(), params, scoring="neg_mean_squared_error", cv=5)
lasso_regressor.fit(X_train, y_train)
print("Lasso:", lasso_regressor.best_params_, round(lasso_regressor.best_score_, 5))

results = pd.DataFrame(ridge_regressor.cv_results_)[["param_alpha", "mean_test_score", "rank_test_score"]]
print(results.round({"mean_test_score": 5}).to_string(index=False))

y_pred = ridge_regressor.predict(X_test)     # the best alpha, refit on all of X_train
print("Ridge test R2:", round(r2_score(y_test, y_pred), 4))

Lasso prints ConvergenceWarning messages for the tiny alphas, as it did in the video: with alpha near 0 its solver needs more iterations. The search still finishes and the results stand.

Why the smallest alpha won

  • best_params_ is the smallest alpha. Ridge picks 1e-15 and Lasso 1e-08, and both best scores equal linear regression's −0.52305 to five decimals.
  • The table shows the reason. The mean score barely moves from 1e-15 to 1 and gets slightly worse at 5, 10 and 20. With 13,828 training rows and 8 slopes, a straight line does not overfit, so a penalty only adds error.
  • The video saw the same once it split the data. Searching on X_train, its Ridge picked alpha 0.01 and Lasso 1e-08, both at −25.47, against −25.19 for linear regression: no better than the plain line.
  • The test R² is the honest final number. 0.597 on rows the search never saw. GridSearchCV refits the best alpha on the whole training set, so predict uses that refit model.

GridSearchCV vs cross_val_score

cross_val_scoreGridSearchCV
Scoresone model with fixed settingsone model per value in the grid
Returnsan array of fold scoresa fitted search: best_params_, best_score_, cv_results_
Refitsnoyes, the best setting on all the data it was given
Can predictnoyes, with the refit best model
Use it forchecking one modelchoosing alpha, C, max_iter, depth

Ridge, Lasso and ElasticNet also have a search built in, RidgeCV, LassoCV and ElasticNetCV, which ElasticNet regression uses on the forest fires data.

Where you use GridSearchCV

  • Choosing λ. The alpha of Ridge and Lasso, as here.
  • Tuning a classifier. C and max_iter of logistic regression in Logistic regression in scikit-learn, with scoring="f1".
  • Any model with settings. Tree depth, number of neighbours, number of trees: the same three lines, a different grid.
Watch out. If the best value sits at the edge of the grid (the smallest here, 20 in the video's first run), the true best may lie beyond it. The video extended its grid to 100 for that reason. And search only on the training data: a score from a search that saw the test rows is too optimistic.
Try it yourself
  • Add the video's longer grid, 30, 35, 40, 45, 50, 55, 100, to params and check whether best_params_ changes.
  • Set cv=10 in both searches, as the video does, and compare best_score_.
  • Change scoring to "r2". The best score becomes an R², where higher is better without a minus sign.

Every expert started right here.