Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Linear regression assumptions

Linear regression assumptions are the conditions on the data under which a linear regression trains well and its slopes can be trusted: roughly normal features, scaled features, a linear relationship and no multicollinearity.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

A linear regression always returns a line, even when a line is the wrong shape for the data. These checks, from the end of the video's regression session, say when to trust it.

Assumptions of linear regression · from the Complete Machine Learning in 6 Hours video · 89:45 to 93:08

Checking the four assumptions

The four assumptions from the video, normal features, standardization, linearity and no multicollinearity, with a sketch of features X1 and X2 that are 95 percent correlated next to X3 and the target Y.

Normal features and feature transformation

If the features follow a normal (Gaussian) distribution, the model trains well. If a feature does not, apply a mathematical function to it to bring it closer to normal, which the video calls feature transformation; taking the log of a long-tailed feature is the usual one. Strictly, the classical assumption is about the errors of the fit being normal, not the features; normal-looking features help, but they are not required.

Standardization with the z-score

Standardization (the StandardScaler) rescales a feature with the z-score so its mean μ becomes 0 and its standard deviation σ becomes 1. Wherever gradient descent is involved, it helps: on scaled data the cost surface is closer to round, and the steps reach the global minimum quickly instead of zig-zagging across a long valley. It is not compulsory, but it makes training faster.

Linearity

Linear regression works well when the output changes in a straight-line way with each feature. If the relationship bends, a straight line underfits; polynomial features (as in Overfitting and underfitting) or another model are the fix.

Multicollinearity and VIF

Multicollinearity means two features move together. In the video's sketch X1 and X2 are 95% correlated. Using both adds little: "we can drop this particular feature" and keep one. The variance inflation factor (VIF) measures it per feature: fit the feature from all the other features, take that R², and compute 1 / (1 − R²). A VIF of 1 means no overlap; values above about 5 to 10 are usually treated as a problem.

The video also names homoscedasticity: the errors should have about the same spread for small and large predictions. A plot of the residuals against the predicted values shows it; a funnel shape, narrow on one side and wide on the other, means the assumption fails.

Checking the assumptions on California housing

StandardScaler

fit learns μ and σ of each column, transform applies the z-score; fit_transform does both.

python
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)    # each column: (x - mean) / std
print(scaler.mean_, scaler.scale_)    # the μ and σ it learned

Skew and the z-score

Skew measures how lopsided a distribution is: 0 is symmetric, a large positive value is a long right tail. The first lines check one feature against the normal assumption; the rest compute the z-score by hand and with StandardScaler.

ExampleRun on scikit-learn 1.9.1
import numpy as np
from sklearn.datasets import fetch_california_housing
from sklearn.preprocessing import StandardScaler

X, y = fetch_california_housing(return_X_y=True, as_frame=True)

# assumption 1: is the feature close to normal? skew 0 = symmetric
print("Population skew:", round(X["Population"].skew(), 2))
print("log(1 + Population) skew:", round(np.log1p(X["Population"]).skew(), 2))

# assumption 2: z = (x - mu) / sigma, by hand and with StandardScaler
income = X["MedInc"]
mu, sigma = income.mean(), income.std(ddof=0)
by_hand = (income - mu) / sigma
scaled = StandardScaler().fit_transform(X[["MedInc"]])[:, 0]
print("mu =", round(mu, 3), " sigma =", round(sigma, 3))
print("first three z by hand:", by_hand[:3].round(3).tolist())
print("first three z scaler: ", scaled[:3].round(3).tolist())
print("after scaling: mean", abs(round(scaled.mean(), 3)), " std", round(scaled.std(), 3))

Correlation and VIF with numpy

For standardized features the VIF of every column is on the diagonal of the inverse of the correlation matrix, so numpy computes all of them in one line. The last lines check one value the long way, with a LinearRegression of AveRooms on the other seven features.

ExampleRun on scikit-learn 1.9.1
import numpy as np
import pandas as pd
from sklearn.datasets import fetch_california_housing
from sklearn.linear_model import LinearRegression

X, y = fetch_california_housing(return_X_y=True, as_frame=True)

corr = X.corr()
print("AveRooms vs AveBedrms:", round(corr.loc["AveRooms", "AveBedrms"], 2))
print("Latitude vs Longitude:", round(corr.loc["Latitude", "Longitude"], 2))

# VIF of every feature = the diagonal of the inverse correlation matrix
vif = pd.Series(np.diag(np.linalg.inv(corr.to_numpy())), index=X.columns)
print(vif.round(2))

# the same number for AveRooms, from its R2 against the other features
others = X.drop(columns=["AveRooms"])
r2 = LinearRegression().fit(others, X["AveRooms"]).score(others, X["AveRooms"])
print("1 / (1 - R2) for AveRooms:", round(1 / (1 - r2), 2))

# drop one feature of the correlated pair and recompute
X_small = X.drop(columns=["AveBedrms"])
vif_small = pd.Series(np.diag(np.linalg.inv(X_small.corr().to_numpy())), index=X_small.columns)
print("AveRooms VIF after dropping AveBedrms:", round(vif_small["AveRooms"], 2))

What the checks found

  • Population is far from normal. A skew of 4.94, a long right tail of crowded districts. log(1 + x) brings it to −1.04, much closer to symmetric: the feature transformation the video describes.
  • StandardScaler is the z-score. With μ = 3.871 and σ = 1.9, the first incomes become 2.345, 2.332 and 1.783 both by hand and with the scaler, and the scaled column has mean 0 and standard deviation 1.
  • AveRooms and AveBedrms are the video's X1 and X2. They are 0.85 correlated, and their VIFs are 8.34 and 6.99. The long way gives the same 8.34 for AveRooms.
  • Dropping one fixes the other. Without AveBedrms, the VIF of AveRooms falls from 8.34 to 1.26.
  • Latitude and Longitude also score high, 9.30 and 8.96, with a correlation of −0.92. They are a different case: together they locate a district, so dropping one would lose information.

Assumption checks vs fixes

AssumptionHow to checkIf it fails
Normal features (normal errors)skew, a histogram, a residual histogramtransform the feature, for example log(1 + x)
Scaled featuresmean and standard deviation of each columnStandardScaler
Linearityscatter plot of each feature against ypolynomial features or a non-linear model
No multicollinearitycorrelation matrix, VIF above about 5 to 10drop one of the pair, or use Ridge
Homoscedasticityresiduals against predictions: no funneltransform the target, for example log(y)

Where you use these checks

  • Before trusting the slopes. With high VIF the individual coefficients swing from sample to sample, even when the predictions stay fine.
  • Before gradient descent or regularization. Scaling makes training faster and makes one alpha mean the same for every slope.
  • When the fit is poor. A low R² on a curved relationship points to the linearity assumption, not to the algorithm.
Watch out. Fit the scaler on the training data only and reuse it on the test data. Computing μ and σ on all the data lets the test rows leak into training; a Pipeline with StandardScaler handles this for you.
Try it yourself
  • Print the skew of X["MedInc"] and of np.log1p(X["MedInc"]).
  • Drop Latitude instead of AveBedrms and recompute the VIF table: Longitude's VIF falls close to 1.
  • Replace StandardScaler with MinMaxScaler from sklearn.preprocessing and print the minimum and maximum of the scaled column.

Little by little, you're building something great.