Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Ordinary least squares

Ordinary least squares (OLS) is a method that finds the intercept and slope of a linear regression line directly from a formula, by setting the derivatives of the squared error to zero instead of stepping towards the minimum.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Gradient descent walks down the cost curve one small step at a time. The bottom of that curve is the one point where its slope is zero, so OLS solves for that point with algebra and lands there in one go.

Setting the slope of the error to zero

The notes write the line as hθ(x) = β₀ + β₁x, so ŷᵢ = β₀ + β₁xᵢ. The error to minimise is the mean squared error, called S here:

S is a bowl with one global minimum. At the minimum the slope is zero in both directions, so both partial derivatives are set to 0. Each one uses the rule that the derivative of a square, u², is 2u times the derivative of u:

Left: the four age and weight points with the red least squares line, its residuals and the mean point (24.25, 64.75); the slope is -31.75 / 18.75, about -1.69, and the intercept about 105.81. Right: a bowl-shaped error curve where gradient descent takes many small steps while ordinary least squares solves for the point where the slope is zero.

Solving equation 1 for the intercept

Multiply equation 1 by −n/2, which leaves the sum equal to 0. Split the sum: Σyᵢ − nβ₀ − β₁Σxᵢ = 0. Move nβ₀ to the other side and divide by n. Σyᵢ/n is the mean ȳ and Σxᵢ/n is the mean x̄:

Solving equation 2 for the slope

Multiply equation 2 by −n/2 as well and put in β₀ = ȳ − β₁x̄. Each bracket becomes (yᵢ − ȳ) − β₁(xᵢ − x̄), still multiplied by xᵢ:

The deviations from a mean always add up to zero: Σ(yᵢ − ȳ) = 0 and Σ(xᵢ − x̄) = 0. So subtracting x̄ from the multiplier xᵢ changes nothing, and the xᵢ can be written as (xᵢ − x̄). Solving for β₁ gives the slope:

The notes drop the multiplier xᵢ midway and end with β₁ = Σ(yᵢ − ȳ) / Σ(xᵢ − x̄). Both of those sums are always 0, so that version cannot be computed; the slope needs the products (xᵢ − x̄)(yᵢ − ȳ) on top and the squares (xᵢ − x̄)² below, as above. The code further down prints both sums.

Working the formulas on the age and weight table

The notes end with a table to fill in: x, y, the deviations, then β₁ and β₀. On the video's table, x̄ = (24 + 25 + 21 + 27) / 4 = 24.25 and ȳ = (62 + 63 + 72 + 62) / 4 = 64.75:

Age xWeight yx − x̄y − ȳ(x − x̄)(y − ȳ)(x − x̄)²
2462−0.25−2.750.68750.0625
25630.75−1.75−1.31250.5625
2172−3.257.25−23.562510.5625
27622.75−2.75−7.56257.5625
Sum−31.7518.75

These are the θ₁ ≈ −1.69 and θ₀ ≈ 105.81 that Simple linear regression printed and that Gradient descent reached after 575,671 steps.

Computing OLS in NumPy

The two formulas

python
import numpy as np

x = np.array([24, 25, 21, 27], dtype=float)   # age
y = np.array([62, 63, 72, 62], dtype=float)   # weight
x_bar, y_bar = x.mean(), y.mean()

beta1 = np.sum((x - x_bar) * (y - y_bar)) / np.sum((x - x_bar) ** 2)   # slope
beta0 = y_bar - beta1 * x_bar                                          # intercept

Matching LinearRegression

ExampleThe video's age and weight table, run on scikit-learn 1.9.1
from sklearn.linear_model import LinearRegression

print("x̄ =", x_bar, " ȳ =", y_bar)
print("OLS by hand:      β1 =", round(beta1, 4), " β0 =", round(beta0, 4))
model = LinearRegression().fit(x.reshape(-1, 1), y)
print("LinearRegression: β1 =", round(model.coef_[0], 4), " β0 =", round(model.intercept_, 4))
print("the notes' sums:  Σ(y − ȳ) =", np.sum(y - y_bar), " Σ(x − x̄) =", np.sum(x - x_bar))

Reading the OLS run

  • The hand formula and LinearRegression agree to four decimals: β₁ = −1.6933 and β₀ = 105.8133. scikit-learn's LinearRegression is ordinary least squares: its docs describe it as plain OLS (scipy.linalg.lstsq) wrapped as a predictor object.
  • x̄ = 24.25 and ȳ = 64.75 match the table above, and the line passes through that point.
  • Both of the notes' sums print 0.0, which is why the products and squares are needed.

Ordinary least squares vs gradient descent

Ordinary least squaresGradient descent
How it finds θSolves derivative = 0 with a formulaRepeats small steps down the slope
Learning rateNoneNeeds α; too big diverges
Work on the age tableTwo sums575,671 steps
With many featuresMatrix algebra, slow when there are very many featuresScales to huge tables and neural networks
In scikit-learnLinearRegressionSGDRegressor

Where you use ordinary least squares

  • Every LinearRegression fit in this course: the coefficients come from least squares, not from gradient descent.
  • Reading a regression report: the statsmodels library's OLS prints the same coefficients with standard errors and p-values, as the practical notebooks do.
  • Interviews: derive β₀ = ȳ − β₁x̄ and the slope formula on paper, the way the notes do.
Watch out. statsmodels' OLS adds no intercept unless you add a column of ones with sm.add_constant(X) (statsmodels OLS docs). The practical notebooks call sm.OLS(y_train, X_train) without it: the slope it prints, 17.2982, still matches LinearRegression, because the standardised weights average 0, but every prediction misses the 156.47 cm intercept and comes out near 0 (5.79 for a 75 kg person).
Try it yourself
  • Add a fifth person, age 30 and weight 75, to x and y: work out the new x̄ and ȳ, then check β₁ against LinearRegression.
  • Run np.polyfit(x, y, 1): it returns the same slope and intercept, in that order.
  • Replace the slope line with the notes' version, np.sum(y - y_bar) / np.sum(x - x_bar), and read the result: 0 divided by 0 gives nan.

This is what real progress feels like.