Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Principal component analysis (PCA)

Principal component analysis (PCA) is a dimensionality reduction technique that turns many correlated features into a few new features, the principal components, that keep as much of the variance as possible.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Curse of dimensionality showed how too many features make a model worse and slower, and how feature selection drops the unimportant ones. When every feature matters, nothing can be dropped. PCA extracts two or three new features from all of the old ones instead; the breast cancer data, with 30 features, goes down to two.

Projecting onto the axis with the most spread

Projecting 2D data onto one axis · from the PCA in depth video · 31:37 to 36:41

The video's housing example has two features, the size of the house and the number of rooms, and both drive the price. Bigger houses have more rooms, so the points rise together. The task is to turn these two dimensions into one.

The plain way is to project every point onto the size axis. The spread of the points along size is kept, and spread is what variance measures: the wider the spread, the higher the variance. The spread in the number of rooms is thrown away, so information about rooms is lost.

Finding new axes: PC1 and PC2

New axes and maximum variance · from the PCA in depth video · 37:11 to 42:04

PCA instead transforms the axes. A new axis, size′, runs along the cloud of points, and a second one, rooms′, stands at a right angle to it. Projected onto size′, the points keep most of their spread; little is left along rooms′, so dropping it loses little.

The new axes are the principal components. PC1 captures the largest variance, PC2 the next largest, and with three features PC3 the rest. To go from 3D to 2D, keep PC1 and PC2; to go to 1D, keep PC1.

Two panels of ten standardised houses, size against number of rooms: projecting onto the size axis keeps 50% of the variance, projecting onto the first principal component keeps 93.4%.

The diagram uses ten made-up houses, sizes 600 to 1800 square feet with 1 to 5 rooms, standardised. Projected onto the size axis they keep 50% of the total variance; projected onto PC1 they keep 93.4%.

The board notes call this variance the cost function. Every unit vector u gives one set of projected numbers p′₁ to p′ₙ and one variance; the goal is the unit vector with the largest one.

Eigenvectors of the covariance matrix

The PCA steps: standardise, covariance, eigenvectors · from the PCA in depth video · 70:51 to 73:54

Trying every possible line would never end, so PCA solves for the best one. The video's steps:

  1. Standardise the data, so every feature has mean 0 and standard deviation 1.
  2. Compute the covariance matrix of the features: variances on the diagonal, covariances off it.
  3. Find its eigenvectors and eigenvalues, the solutions of A v = λ v.
  4. The eigenvector with the largest eigenvalue is PC1, the next is PC2, and so on.

The video describes the eigenvalue as the magnitude of the eigenvector. The eigenvectors here have length 1; the eigenvalue λ is the variance of the data along that eigenvector, which is why the largest λ marks PC1. The share of the variance a component keeps is its eigenvalue divided by the sum of all of them.

Computing PC1 for the house data with NumPy

ExampleThe video's steps on ten houses, run on NumPy 2.5.3 and scikit-learn 1.9.1
import numpy as np
from sklearn.decomposition import PCA

# Ten houses: size in square feet and number of rooms
size = np.array([600, 750, 800, 950, 1100, 1200, 1350, 1500, 1650, 1800])
rooms = np.array([1, 2, 1, 3, 2, 3, 4, 3, 5, 4])
X = np.column_stack([size, rooms])

Xs = (X - X.mean(axis=0)) / X.std(axis=0)      # step 1: standardise
cov = np.cov(Xs, rowvar=False)                  # step 2: covariance matrix
values, vectors = np.linalg.eigh(cov)           # step 3: eigenvalues and eigenvectors
order = np.argsort(values)[::-1]                # largest eigenvalue first
values, vectors = values[order], vectors[:, order]
print("covariance matrix:\n", cov.round(3))
print("eigenvalues:", values.round(3))
print("PC1 direction:", vectors[:, 0].round(3))

print("variance kept on the size axis:", round(Xs[:, 0].var(ddof=1) / values.sum(), 3))
print("variance kept on PC1          :", round((Xs @ vectors[:, 0]).var(ddof=1) / values.sum(), 3))

pca = PCA(n_components=2).fit(Xs)
print("PCA explained_variance_:", pca.explained_variance_.round(3))
  • The covariance matrix has 1.111 on the diagonal (standardised features, divided by n − 1) and 0.964 off it: size and rooms rise together.
  • The eigenvalues are 2.075 and 0.148, and PCA's explained_variance_ prints the same two numbers.
  • PC1 points along [0.707, 0.707], the diagonal of the cloud, and keeps 93.4% of the variance against 50% for the size axis.

Reducing the breast cancer data to two components

PCA on the breast cancer data · from the PCA in depth video · 80:04 to 83:49

The practical loads the breast cancer data (569 rows, 30 features) into a DataFrame, standardises it with StandardScaler, then extracts two components with PCA(n_components=2). The video stresses that PCA needs the scaling step.

Standardising the 30 features

python
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
scaler.fit(df)                          # mean and standard deviation per feature
scaled_data = scaler.transform(df)      # every feature now has mean 0, std 1

Fitting PCA with n_components=2

python
from sklearn.decomposition import PCA

pca = PCA(n_components=2)                  # keep PC1 and PC2
data_pca = pca.fit_transform(scaled_data)  # 569 rows, 2 columns
print(pca.explained_variance_)             # variance along PC1 and PC2
ExampleFrom the PCA in depth video, run on scikit-learn 1.9.1
import numpy as np
import pandas as pd
from sklearn.datasets import load_breast_cancer
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA

cancer_dataset = load_breast_cancer()
df = pd.DataFrame(cancer_dataset["data"], columns=cancer_dataset["feature_names"])

scaler = StandardScaler()
scaler.fit(df)
scaled_data = scaler.transform(df)

pca = PCA(n_components=2)
data_pca = pca.fit_transform(scaled_data)
print(df.shape, "->", data_pca.shape)
print(data_pca[:3].round(3))
print("explained_variance_:", pca.explained_variance_.round(3))
print("explained_variance_ratio_:", pca.explained_variance_ratio_.round(3))

# The same two numbers from the eigenvalues of the covariance matrix
eigenvalues = np.linalg.eigvalsh(np.cov(scaled_data, rowvar=False))[::-1]
print("top two eigenvalues:", eigenvalues[:2].round(3))
  • The first rows match the video's notebook: 9.193 and 1.949, then 2.388 and −3.768, then 5.734 and −1.075.
  • explained_variance_ is [13.305, 5.701], the same as the video, and equals the top two eigenvalues of the covariance matrix.
  • explained_variance_ratio_ is [0.443, 0.19]: two components keep 63.2% of the variance of 30 features.

How much variance each component keeps

Explained variance and the 2D plot · from the PCA in depth video · 83:49 to 87:16

The video says the variances of all the components add up to near 200 and calls the data 11 to 12 features. The breast cancer data has 30 features, and after standardising, the 30 variances add up to 30.05: each standardised feature contributes a variance of about 1. The video prints explained_variance_; explained_variance_ratio_ turns the same numbers into shares.

ExampleRun on scikit-learn 1.9.1
full = PCA()                         # no n_components: keep all 30
full.fit(scaled_data)
print("components:", full.n_components_)
print("sum of explained_variance_:", round(full.explained_variance_.sum(), 2))
print("first five ratios:", full.explained_variance_ratio_[:5].round(3))
kept = np.cumsum(full.explained_variance_ratio_)
for k in [2, 3, 6, 10]:
    print(f"variance kept by {k:>2} components: {kept[k - 1]:.3f}")

Plotting the two components

ExampleFrom the PCA in depth video, run on scikit-learn 1.9.1
import matplotlib.pyplot as plt

plt.figure(figsize=(8, 6))
plt.scatter(data_pca[:, 0], data_pca[:, 1], c=cancer_dataset["target"], cmap="plasma")
plt.xlabel("First principal component")
plt.ylabel("Second Principal Component")
plt.title("Breast cancer data after PCA (0 = malignant, 1 = benign)")
plt.show()
print("malignant:", (cancer_dataset["target"] == 0).sum(), " benign:", (cancer_dataset["target"] == 1).sum())
A scatter plot of the breast cancer data on the first two principal components: benign tumours (yellow) cluster on the left and malignant tumours (dark blue) spread to the right, with a little overlap.

What the two components show

  • Two components keep 63.2% of the variance; three keep 72.6%, six keep 88.8% and ten keep 95.2%.
  • The 212 malignant and 357 benign tumours fall mostly on opposite sides of PC1, so two extracted features already separate the classes well.
  • PC1 is not one of the 30 features: it is a weighted mix of all of them.

Feature extraction vs feature selection

Feature extraction (PCA)Feature selection
What it keepsnew features built from all the old onesa subset of the original features
Example from the videosize and rooms become one new featurefountain size dropped, house size kept
Information lostonly the smallest varianceseverything in the dropped features
Readable featuresno: each component mixes all featuresyes: the names stay

Where you use PCA

  • Plotting data with many features in 2D or 3D, as with the breast cancer plot.
  • Faster training: train on 10 components that keep 95% of the variance instead of 30 features.
  • Before clustering: K-means and DBSCAN measure distances, which work better on a few uncorrelated components than on many correlated features.
Watch out. Scale before PCA. Without StandardScaler, PC1 of the breast cancer data keeps 98.2% of the variance and points along worst area, whose values run into the thousands: PCA then measures units, not structure.
Try it yourself
  • Change PCA(n_components=2) to PCA(n_components=3) in the breast cancer example and read the three explained_variance_ratio_ values.
  • Pass PCA(n_components=0.95) instead and print pca.n_components_ to see how many components keep 95%.
  • Remove the StandardScaler step and fit PCA on df directly: compare explained_variance_ratio_.

Every expert started right here.