StatisticsSciPy 1.18 · pandas 3.0 · statsmodels 0.15 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
57 small wins to finish your pathNext lesson →

Standardization and normalization

Standardization and normalization are feature scaling methods that put columns measured in different units on one common scale: standardization turns each value into its z-score, so a column gets mean 0 and standard deviation 1, and min-max normalization maps each column onto the range 0 to 1.

Last updated: 07 Oct, 2026 · SciPy 1.18

Z-score and the standard normal distribution measured one value in standard deviations. Machine learning does the same to whole columns, because columns in years, rupees and kilograms cannot be compared as raw numbers.

Seeing why columns need scaling

The video's practical application is a machine learning dataset with three features: age in years (24, 25, 26, 27), salary in rupees (40K, 80K, 60K, 70K) and weight in kilograms (70, 80, 55, 45). The units differ, and so do the sizes: a difference of 1 in salary is one rupee, a difference of 1 in age is a whole year. A model that measures distances between rows, such as K nearest neighbours (KNN), would be driven almost entirely by salary.

Standardizing each column with the z-score

Standardization applies the z-score to every value of a column, using that column's own mean and standard deviation:

Standardization of one column

Every standardized column has mean 0 and standard deviation 1, whatever its units. Age 24 becomes −1.3416: 1.34 standard deviations below the mean age of 25.5. The σ here is the population standard deviation (divide by n), the convention scikit-learn's StandardScaler uses.

Normalization with min-max scaling · from the Complete Statistics for Data Science in 6 Hours video · 1:45:37 to 1:47:58

Normalizing to 0 to 1 with min-max scaling

Normalization, as the video uses the word, moves the values into a range you choose, usually 0 to 1 and sometimes −1 to 1. The tool for it is the min-max scaler:

Min-max scaling to the range 0 to 1

The smallest value becomes 0 and the largest becomes 1. For a range from a to b the formula is x' = a + (x − min)(b − a)/(max − min). In the notes that go with the video, the feature f1 = 2, 5, 6, 8, 1 has min 1 and max 8, so 2 becomes (2 − 1)/(8 − 1) = 1/7 = 0.143, 5 becomes 4/7 = 0.571, 6 becomes 5/7 = 0.714, 8 becomes 1 and 1 becomes 0.

Three number lines for the feature f1 with values 2, 5, 6, 8 and 1. Min-max scaling maps them to 0.143, 0.571, 0.714, 1 and 0; standardization with mean 4.4 and population standard deviation 2.577 maps them to minus 0.931, plus 0.233, plus 0.621, plus 1.397 and minus 1.319. The order of the points is the same on all three lines.

The video's deep learning case is images. A CNN for image classification takes pixels from 0 to 255, and before training each pixel is divided by 255, so 0 stays 0 and 255 becomes 1. That is min-max scaling with the bounds fixed at the pixel range instead of read from the image; when an image holds both a 0 and a 255 pixel, the two give the same numbers.

The word normalization has other meanings too. scikit-learn's Normalizer rescales each row to length 1, which is a different operation, so "min-max scaling" is the clearer name for this one.

Scaling the age, salary and weight table

Standardizing with pandas

pandas .std() divides by n − 1 by default (ddof=1). ddof=0 divides by n, the same as numpy and StandardScaler.

python
standardized = (df - df.mean()) / df.std(ddof=0)    # z-score per column, divide by n

Min-max scaling with pandas

python
minmax = (df - df.min()) / (df.max() - df.min())     # each column onto 0 to 1
ExampleThe video's table, run on pandas 3.0.6
import pandas as pd

df = pd.DataFrame({"age": [24, 25, 26, 27],
                   "salary": [40_000, 80_000, 60_000, 70_000],
                   "weight": [70, 80, 55, 45]})
standardized = (df - df.mean()) / df.std(ddof=0)
minmax = (df - df.min()) / (df.max() - df.min())
print(standardized.round(4))
print(minmax.round(4))
print("means after standardizing:", standardized.mean().round(4).tolist())
print("SDs after standardizing:  ", standardized.std(ddof=0).round(4).tolist())

What the scaled table shows

  • Standardized age is −1.3416, −0.4472, 0.4472, 1.3416: evenly spaced ages stay evenly spaced.
  • Standardized salary is −1.5213, 1.1832, −0.169, 0.5071: 80K is 1.18 SDs above the mean of 62,500, and the rupee figures are now on the same footing as age and weight.
  • Every standardized column has mean 0.0 and SD 1.0.
  • Min-max salary is 0.0, 1.0, 0.5, 0.75: 40K, the smallest, becomes 0 and 80K, the largest, becomes 1. Weight becomes 0.7143, 1.0, 0.2857, 0.0.

Standardization keeps the shape of the data

Standardizing changes the numbers on the axis, never the shape of the distribution. The tips bills from the video's Python session show it:

ExampleThe tips data from the video, run on seaborn 0.13.2
import matplotlib.pyplot as plt
import seaborn as sns
from scipy.stats import skew

bill = sns.load_dataset("tips")["total_bill"]
z = (bill - bill.mean()) / bill.std(ddof=0)

fig, axes = plt.subplots(1, 2, figsize=(10, 4))
axes[0].hist(bill, bins=20, color="steelblue", edgecolor="white")
axes[0].set_title("total_bill in dollars")
axes[1].hist(z, bins=20, color="seagreen", edgecolor="white")
axes[1].set_title("total_bill standardized")
plt.show()
print("skewness before:", round(skew(bill), 3), " after:", round(skew(z), 3))
print("range before:", bill.min(), "to", bill.max(), " after:", round(z.min(), 3), "to", round(z.max(), 3))
Two histograms of the tips total_bill with the same right-skewed shape: on the left in dollars from about 3 to 51, on the right standardized from about minus 1.88 to plus 3.49.

The two histograms are the same picture with a new axis, and the skewness is 1.126 before and after. Standardized data has mean 0 and SD 1, but it is not standard normal unless the original column was normal.

Min-max scaling and outliers

The outlier dataset from Outlier detection with IQR and z-score has 31 values between 10 and 19 and three far above them:

ExampleThe outlier data from the video, run on NumPy 2.5.3
import numpy as np

dataset = np.array([11, 10, 12, 14, 12, 15, 14, 13, 15, 102, 12, 14, 17, 19, 107,
                    10, 13, 12, 14, 12, 108, 12, 11, 14, 13, 15, 10, 15, 12, 10, 14, 13, 15, 10])
scaled = (dataset - dataset.min()) / (dataset.max() - dataset.min())
print("min-max of 102, 107, 108:", scaled[dataset > 100].round(3))
print("largest of the other 31: ", scaled[dataset < 100].max().round(3))
print("share of values below 0.1:", np.mean(scaled < 0.1).round(3))

The three outliers take the top of the range and squeeze the other 31 values into 0 to 0.092, so 91.2% of the data ends up below 0.1. Outliers also inflate the standard deviation that standardization divides by, so handle them before scaling.

Fitting the scaler on the training data only

In a machine learning pipeline, scikit-learn's StandardScaler and MinMaxScaler apply these two formulas. The mean, SD, min and max must come from the training rows only, and the test rows are then transformed with those same numbers. Fitting on all the rows lets the test set leak into training and makes the test score look better than it will be on new data. Train and test split covers the split itself.

python
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import MinMaxScaler, StandardScaler

X_train, X_test = train_test_split(df, test_size=0.25, random_state=42)
scaler = StandardScaler()                          # or MinMaxScaler()
X_train_scaled = scaler.fit_transform(X_train)     # learns mean and SD from train
X_test_scaled = scaler.transform(X_test)           # reuses them on test

StandardScaler divides by the population standard deviation, the same as ddof=0 above, so on the training rows it gives the numbers the pandas formula gives.

Standardization vs normalization

StandardizationMin-max normalization
Formula(x − μ)/σ(x − min)/(max − min)
ResultMean 0, SD 1, no fixed rangeRange 0 to 1, or a range you choose
Shape of the dataUnchangedUnchanged
OutliersInflate σ and pull the other values togetherSqueeze the other values into a small part of the range
scikit-learnStandardScalerMinMaxScaler
Typical useLinear and logistic regression, SVM, PCA, KNN, k-meansImage pixels for CNNs, inputs that must stay bounded

Where you use standardization and normalization

  • Distance-based models. KNN and k-means compare rows by distance, so every feature must be on one scale; see K nearest neighbours (KNN).
  • Gradient descent. Features on similar scales let Gradient descent move evenly in every direction instead of zigzagging.
  • Images. Pixels are divided by 255 before a CNN is trained, as in the video.
Watch out. Fit the scaler on the training set only and reuse it for the test set and for new data. A scaler fitted on all the rows, or refitted on the test rows, leaks information and gives test scores you will not see in use.
Try it yourself
  • Replace df.std(ddof=0) with the pandas default df.std() in the table example: age 24 becomes −1.1619 instead of −1.3416.
  • Scale salary to the range −1 to 1 with -1 + 2 * minmax["salary"]: 40K becomes −1, 60K becomes 0 and 80K becomes 1.
  • Drop 102, 107 and 108 from dataset (dataset[dataset < 100]) and min-max scale again: the 31 values now spread over the whole range, and the four 13s sit at 0.333.

Slow is fine. Stopping is the only problem.