Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Ridge regression

Ridge regression is a linear regression that adds λ times the sum of the squared slopes to the cost function, so the fitted line cannot become steep enough to overfit.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

The two-point line in Overfitting and underfitting had a cost of 0 and still failed on new points. Ridge changes the cost so that a zero error alone no longer wins.

Ridge adds lambda times slope squared · from the Complete Machine Learning in 6 Hours video · 75:25 to 79:32

The board first writes 0 + 1(2) = 3, then corrects it to 0 + 1 × 2² = 4.

Adding a penalty to the cost function

Back to the two points. The cost is 1/2m Σ (hθ(x(i)) − y(i))²; writing hθ(x) as ŷ, each term is (ŷ(i) − y(i))², the difference between the predicted and the real value. Here every difference is 0, so the cost is 0, "and this is still overfitting".

Ridge, also called L2 regularization, adds one more term to the cost: λ multiplied by the slope squared. With the line through the origin θ₀ = 0, so hθ(x) = θ₁x and θ₁ is the slope.

The video's numbers: λ = 1 and a slope of 2, the slope of the line through both points. The error is 0 and the penalty is 1 × 2² = 4, so the cost is 4, not 0. Gradient descent keeps going: it changes θ₁ to bring the 4 down.

A flatter line and lambda as a hyperparameter · from the Complete Machine Learning in 6 Hours video · 79:32 to 83:52

Trading a small error for a flatter line

Changing θ₁ gives the next best fit line, a less steep one. The video takes a slope of 1.5. This line no longer passes through the points, so the error becomes a small value, and the penalty is 1 × 1.5² = 2.25. The cost is that small value plus 2.25, which the video puts at about 3. It is less than 4, so Ridge prefers the flatter line.

The lesson of the picture: "the slope should not be steep". A steep slope most of the time leads to overfitting. New points land closer to the flatter line, so it generalizes better, low bias and low variance instead of the overfit condition.

A steep red line of slope 2 passes through two training points with zero error, so its Ridge cost is 0 plus 1 times 2 squared, which is 4; a flatter green line of slope 1.5 has a small error plus 2.25, which is less than 4, so Ridge prefers it.

λ is a hyperparameter: a value you choose, not one the model learns. It sets how hard the steepness is pushed down. The number of iterations, how many times θ₁ is updated, is a hyperparameter too. The Hyperparameter tuning with GridSearchCV lesson chooses λ by search.

Checking the video's numbers in code

The board does not give the two points, only the slope of 2. Any two points on the line y = 2x work; the code takes (1, 2) and (2, 4) and adds up the error and the penalty the way the board does, with λ = 1.

The Ridge class

scikit-learn calls λ alpha. Ridge minimises ‖y − Xw‖² + alpha ‖w‖²: the sum of the squared errors plus alpha times the sum of the squared slopes, with no 1/2m in front, the form Σ(ŷ − y)² + λ slope² the board uses for its numbers. fit_intercept=False keeps θ₀ = 0, as in the video.

python
from sklearn.linear_model import Ridge

# alpha is the λ of the board; fit_intercept=False keeps θ0 = 0
ridge = Ridge(alpha=1, fit_intercept=False)
ridge.fit(X, y)          # X has one column per feature
print(ridge.coef_)       # the slopes after the penalty

Scoring slope 2, slope 1.5 and the Ridge slope

ExampleFrom the video, run on scikit-learn 1.9.1
import numpy as np
from sklearn.linear_model import Ridge

x = np.array([1.0, 2.0])
y = np.array([2.0, 4.0])      # two points on the line y = 2x
lam = 1

def ridge_cost(slope):
    error = np.sum((slope * x - y) ** 2)   # sum of (y_hat - y)^2
    penalty = lam * slope ** 2             # lambda * slope^2
    return error, penalty, error + penalty

for slope in [2, 1.5]:
    error, penalty, cost = ridge_cost(slope)
    print(f"slope {slope}: error {error:.2f} + penalty {penalty:.2f} = cost {cost:.2f}")

ridge = Ridge(alpha=lam, fit_intercept=False).fit(x.reshape(-1, 1), y)
best = ridge.coef_[0]
error, penalty, cost = ridge_cost(best)
print(f"Ridge slope {best:.3f}: error {error:.3f} + penalty {penalty:.3f} = cost {cost:.3f}")

Reading the Ridge slope

  • Slope 2 costs 4.00. No error, all penalty: the board's 0 + 1 × 2² = 4.
  • Slope 1.5 costs 3.50. For these two points the "small value" is 1.25, so the board's "about 3" is 3.5 here. It is below 4, as the video says.
  • Ridge finds slope 1.667 with cost 3.333. That is the lowest cost of all slopes: Ridge gives up a small error (0.556) to cut the penalty. Neither line from the board is the exact minimum; the minimum sits between them.
  • The slope that fits the points exactly no longer wins. Once λ is above 0, a perfect fit on the training data is not the cheapest answer.

Moving the lowest cost with λ

The cost can be drawn as a curve over every slope θ₁, the bowl that gradient descent walks down. With λ = 0 it is the plain squared error, and its lowest point, the global minimum, sits at slope 2: the overfit line. Adding λθ₁² lifts the curve more the farther θ₁ is from 0, so as λ grows the bowl rises and its lowest point slides toward 0: λ up, slope down. For the two points the lowest point has a short formula, θ₁ = 10 / (5 + λ).

ExampleRun on scikit-learn 1.9.1
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import Ridge

x = np.array([1.0, 2.0])
y = np.array([2.0, 4.0])
theta1 = np.linspace(-0.2, 2.4, 261)              # the slopes to try

for lam, color in [(0, "black"), (10, "orangered"), (30, "goldenrod")]:
    cost = np.array([np.sum((t * x - y) ** 2) + lam * t ** 2 for t in theta1])
    ridge = Ridge(alpha=lam, fit_intercept=False).fit(x.reshape(-1, 1), y)
    print(f"lambda {lam:2}: lowest cost at slope {theta1[cost.argmin()]:.2f}, Ridge slope {ridge.coef_[0]:.3f}")
    plt.plot(theta1, cost, color=color, label=f"lambda = {lam}")

plt.title("Ridge cost for three values of lambda")
plt.xlabel("slope theta1")
plt.ylabel("J(theta1)")
plt.legend()
plt.show()
Three Ridge cost curves over the slope: the black curve for lambda 0 has its lowest point at slope 2, the orange curve for lambda 10 bottoms out near 0.67 and the yellow curve for lambda 30 near 0.29, each one higher and closer to 0.

What the three curves show

  • λ = 0 keeps the overfit line. The lowest cost is at slope 2, where the line passes through both points.
  • λ = 10 moves it to 0.667 and λ = 30 to 0.286. The search over the curve and scikit-learn's Ridge agree, and both match 10 / (5 + λ).
  • The slope never reaches 0. 10 / (5 + λ) stays above 0 for every λ, so Ridge flattens the line without removing the feature.

Shrinking coefficients on California housing

With many features, every slope gets the same treatment. A fitted line ŷ = 0.34 + 0.52x₁ + 0.48x₂ + 0.24x₃ might become ŷ = 0.34 + 0.40x₁ + 0.38x₂ + 0.14x₃ under Ridge: every slope smaller, none of them 0, and the intercept left alone. The California housing data shows it for real. It has 20,640 districts, 8 features (median income MedInc, house age, average rooms, average bedrooms, population, average occupancy, latitude, longitude) and the median house value as the target. StandardScaler first rescales each feature to mean 0 and standard deviation 1, so that one alpha pushes on every slope equally.

ExampleRun on scikit-learn 1.9.1
import pandas as pd
from sklearn.datasets import fetch_california_housing
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge

data = fetch_california_housing(as_frame=True)
X = StandardScaler().fit_transform(data.data)
y = data.target

coefs = {}
for alpha in [1, 100, 1000, 10000]:
    coefs[f"alpha={alpha}"] = Ridge(alpha=alpha).fit(X, y).coef_
table = pd.DataFrame(coefs, index=data.feature_names).round(3)
print(table)
print("sum of |coef|:", table.abs().sum().round(2).tolist())

What a larger alpha does to the slopes

  • The slopes shrink toward 0. The sum of the absolute coefficients falls from 3.33 at alpha 1 to 1.01 at alpha 10000.
  • None of them becomes exactly 0. Squaring makes the penalty on a small slope tiny, so Ridge never removes a feature. Population ends at 0.001, small but not zero.
  • alpha 1 and alpha 100 are almost the same. scikit-learn's Ridge adds the penalty to a sum of squared errors over 20,640 rows, so alpha has to be large before it matters on a big data set.
  • Strongly linked features move together. Latitude and Longitude shrink from −0.899 and −0.870 to −0.162 and −0.122 at the same pace, and AveRooms changes sign: with a heavy penalty, Ridge spreads the weight across features that carry the same information.

Ridge vs linear regression

Linear regressionRidge regression
Costsquared errorsquared error + λ Σ θⱼ²
Two-point lineslope 2, cost 0slope 1.667 with λ = 1
Coefficientswhatever fits the training data bestpulled toward 0, never exactly 0
Extra settingnonealpha (λ), chosen by cross-validation
alpha = 0the same as linear regression

Where you use Ridge regression

  • An overfit linear model. When the training score is much higher than the test score, a Ridge penalty is the first thing to try.
  • Many polynomial or one-hot features. Like the degree 15 model in Overfitting and underfitting, where the slopes grow huge without a penalty.
  • Correlated features. When two features carry the same information (multicollinearity, in Linear regression assumptions), Ridge keeps their slopes small and steady.
Watch out. Scale the features before Ridge. The penalty treats every slope the same, so a feature measured in large units (population in people) gets a tiny slope and almost no penalty, while one in small units is punished hard. StandardScaler first, then Ridge.
Try it yourself
  • Set lam = 0 in the first example. The penalty disappears and slope 2 becomes the cheapest line again.
  • Change lam to 0.25. The Ridge slope works out to 10 / (5 + 0.25) = 1.905, closer to 2.
  • Add 100000 to the alpha list in the California run and see how close to 0 every slope gets.

You understood something today that you didn't yesterday.