Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Overfitting and underfitting

Overfitting is a modelling failure in which a model fits its training data almost perfectly but predicts new data badly; underfitting is the opposite failure, in which the model is poor on the training data and on new data alike.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

A model is only useful on data it has not seen. Comparing its score on the training data with its score on the test data tells you whether it learned the pattern, memorised the points, or learned too little.

Overfitting on two training points · from the Complete Machine Learning in 6 Hours video · 68:07 to 71:45

Fitting a line through two points

The video starts from the cost function of the Cost function lesson, J(θ₀, θ₁) = 1/2m Σ (hθ(x(i)) − y(i))², and a data set with only two points. The best fit line passes through both points and through the origin, so θ₀ = 0 and every error is zero. The cost is J = 0, the smallest value it can take.

These two points are the training data, the data the model learns from. Now new points arrive, the test data. For a new point the line predicts a value on the line, while the real value sits far below it. The gap between the predicted point and the real point is large. That condition is called overfitting.

A red best fit line passes exactly through two training points so the cost is zero, while three new test points sit away from it; the gap between a green predicted point and a blue test point is the error on new data.

In the words of the video, an overfit model "performs well with training data but it fails to perform well with test data". Doing well on the training data is called low bias. Failing on the test data is called high variance. The Bias and variance lesson defines both terms properly.

Underfitting and three models compared · from the Complete Machine Learning in 6 Hours video · 71:45 to 75:25

The board labels the underfit model "high bias, high variance"; the usual reading is high bias with low variance, because its training and test scores (70% and 65%) are close together.

Comparing train and test accuracy

Underfitting is the second failure: the model accuracy is bad on the training data and also bad on the test data. The model has not learned the pattern at all. The video then puts three models side by side.

Three models from the video: 90 percent train and 80 percent test is overfitting with low bias and high variance, 92 and 91 is a generalized model with low bias and low variance, and 70 and 65 is underfitting with high bias and low variance.
  • Model 1, 90% train and 80% test. Good on the training data, a clear drop on the test data: overfitting, low bias and high variance.
  • Model 2, 92% train and 91% test. Good on both, and the two scores are almost equal: a generalized model, low bias and low variance.
  • Model 3, 70% train and 65% test. Poor on both: underfitting, high bias. The small gap between the scores means low variance.

The generalized model is the one you want, because when new data comes it still gives good output. Back on the two-point example, the red line has low bias and high variance: perfect on its own two points, far off on every new one.

Overfitting a polynomial on noisy data

The two-point line is the smallest case. In code the same thing appears when a model has far more freedom than the data needs. Below, 60 points follow the curve y = 0.5x² + x + 2 plus random noise, and half of them are held back as test data with train_test_split (see Train and test split). Three models are fitted on the training half: a straight line (degree 1), a curve with an x² term (degree 2) and a degree 15 polynomial.

The scores are R², from the R squared and adjusted R squared lesson: 1 is a perfect fit, 0 is no better than always predicting the mean, and a value below 0 is worse than predicting the mean.

Making polynomial features

PolynomialFeatures turns the one column x into the columns 1, x, x², up to the chosen degree. LinearRegression then fits one coefficient per column, so a straight-line model can bend. StandardScaler puts the columns on one scale, which keeps the degree 15 fit numerically stable. make_pipeline chains the three steps: fit runs them in order and passes each step's output to the next.

python
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures, StandardScaler
from sklearn.linear_model import LinearRegression

# degree 2: the columns 1, x and x² go into one linear regression
model = make_pipeline(PolynomialFeatures(degree=2), StandardScaler(), LinearRegression())
model.fit(X_train, y_train)
print(model.score(X_test, y_test))     # R² on the test data

Scoring degree 1, 2 and 15

The loop fits the three models on the same training points, prints both scores and draws each fitted curve over the data.

ExampleRun on scikit-learn 1.9.1
import numpy as np
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures, StandardScaler
from sklearn.linear_model import LinearRegression

rng = np.random.RandomState(0)
X = rng.uniform(-3, 3, 60).reshape(-1, 1)
y = 0.5 * X[:, 0] ** 2 + X[:, 0] + 2 + rng.normal(0, 1, 60)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.5, random_state=42)

grid = np.linspace(-3, 3, 300).reshape(-1, 1)
fig, axes = plt.subplots(1, 3, figsize=(12, 3.6), sharey=True)
for ax, degree in zip(axes, [1, 2, 15]):
    model = make_pipeline(PolynomialFeatures(degree=degree), StandardScaler(), LinearRegression())
    model.fit(X_train, y_train)
    train_r2 = model.score(X_train, y_train)
    test_r2 = model.score(X_test, y_test)
    print(f"degree {degree:2}: train R2 = {train_r2:.2f}, test R2 = {test_r2:.2f}")
    ax.scatter(X_train, y_train, color="green", label="train")
    ax.scatter(X_test, y_test, color="blue", marker="x", label="test")
    ax.plot(grid, model.predict(grid), color="red")
    ax.set_title(f"degree {degree}")
    ax.set_xlabel("x")
    ax.set_ylim(-2, 12)
axes[0].set_ylabel("y")
axes[0].legend()
plt.show()
Three panels of the same noisy points: a straight red line misses the curve, a degree 2 curve follows it, and a degree 15 curve wiggles through the green training points and swings away from the blue test points.

What the three scores show

  • Degree 1 underfits. Train R² 0.42 and test R² 0.45: both low and close together. A straight line cannot follow the bend, which is high bias with low variance, the video's model 3.
  • Degree 2 generalizes. Train 0.77 and test 0.82: good on both, with no drop on new points. The model has the same shape as the curve that made the data, the video's model 2.
  • Degree 15 overfits. It has the best training score, 0.88, and a test R² of −2.35, worse than predicting the mean. It bends through the training points and swings far away between them, the two-point line again.
  • The training score alone misleads. It rises with every extra degree. Only the test score shows that degree 15 is the worst model of the three.

Overfitting vs underfitting

OverfittingGeneralized modelUnderfitting
Training scorehighhighlow
Test scoremuch lower than trainingclose to traininglow, close to training
Bias / variancelow bias, high variancelow bias, low variancehigh bias, low variance
Typical causemodel too flexible for the dataflexibility matches the patternmodel too simple
What helpsmore data, a simpler model, Ridge or Lassokeep itmore features, a more flexible model

Where you use the train and test comparison

  • After every fit. Print the training and the test score together; the gap tells you which failure you have before you tune anything.
  • Choosing model complexity. The polynomial degree here, the depth of a decision tree later: raise it while the test score improves, stop when it drops.
  • Choosing regularization. Ridge regression and Lasso regression exist to pull an overfit model back toward the generalized one.
Watch out. A training score near 100% is not good news on its own. The two-point line had a cost of exactly 0 and was the worst model in the example. Always look at the test score next to it.
Try it yourself
  • Change the degrees to [1, 2, 5]. Degree 5 has a little more freedom than the curve needs; compare its test R² with degree 2.
  • Change 60 to 600 in both rng.uniform and rng.normal. With ten times more training points, degree 15 has much less room to overfit.
  • Change the noise rng.normal(0, 1, 60) to rng.normal(0, 0.2, 60) and see how the three train scores move.

Every expert started right here.