Max pooling
Max pooling is a layer that slides a window, usually 2×2 with a stride of 2, over a feature map and keeps only the largest value in each window, so the map shrinks and its strongest features stay.
Last updated: 05 Oct, 2026 · NumPy
After the convolution and ReLU of Convolution and filters, a CNN usually adds a pooling layer. It has no weights; it only picks values.
Keeping the strongest feature in each window
The board's image holds three cats. After convolution and ReLU, one filter responds to the cats' round eyes and gives a 3×3 map: 1 2 3 / 4 5 7 / 0 3 5. Other filters pick up vertical edges, horizontal edges and more.
The video explains why pooling is needed with location invariance: the objects can be anywhere in the image, and as the data passes through more layers the network should extract clearer and clearer information. The usual name is translation invariance: a feature that moves a little still gives the same output.
There are three types: average pooling, min pooling and max pooling. Max pooling places a 2×2 window on the map and keeps the highest number, 5, “one of the cat eyes is clearly visible”. With a stride of 2 the window jumps two cells and picks 7, then moves down and picks 3 and 5. Average pooling takes the mean of each window and min pooling the minimum; which one fits depends on the problem.
A 2×2 window with a stride of 2 does not fit a 3×3 map: ⌊(3 − 2) / 2⌋ + 1 = 1, so without padding (Keras' default padding='valid') the output is the single value 5. The board's 5 7 / 3 5 needs padding='same', where the windows run past the right and bottom edges, as the red boxes on the board do.

Comparing max, min and average pooling on a 4×4 map
A 4×4 map fits four 2×2 windows exactly. The map below keeps the board's answer for max pooling, 5 7 / 3 5, so the three types can be compared on the same windows.

The output is ⌊(4 − 2) / 2⌋ + 1 = 2 per side: a quarter of the values. This is how VGG16 goes from 224×224 to 112×112 after its first block.
Pooling the feature map in NumPy
The pooling function
The function takes each k×k window with step s and applies fn: np.max, np.min or np.mean.
def pool(a, fn, k=2, s=2):
n = (a.shape[0] - k) // s + 1 # floor((n - k) / s) + 1
return np.array([[fn(a[i*s:i*s+k, j*s:j*s+k]) for j in range(n)] for i in range(n)])Same padding for the board's map
Padding for max pooling uses −∞, so a padded cell can never be the maximum.
same = np.pad(board.astype(float), ((0, 1), (0, 1)), constant_values=-np.inf) # right and bottom
pool(same, np.max)Max, min and average pooling in one run
import numpy as np
def pool(a, fn, k=2, s=2):
n = (a.shape[0] - k) // s + 1 # floor((n - k) / s) + 1
return np.array([[fn(a[i*s:i*s+k, j*s:j*s+k]) for j in range(n)] for i in range(n)])
fmap = np.array([[1, 2, 3, 1], [4, 5, 7, 2], [0, 3, 5, 1], [2, 1, 0, 4]])
print("max:")
print(pool(fmap, np.max))
print("min:")
print(pool(fmap, np.min))
print("average:")
print(pool(fmap, np.mean))
board = np.array([[1, 2, 3], [4, 5, 7], [0, 3, 5]])
print("board, valid:", pool(board, np.max))
same = np.pad(board.astype(float), ((0, 1), (0, 1)), constant_values=-np.inf)
print("board, same:")
print(pool(same, np.max).astype(int))
print("board, stride 1:")
print(pool(board, np.max, s=1))
shifted = fmap.copy()
shifted[0, 2], shifted[1, 2] = fmap[1, 2], fmap[0, 2] # move the 7 up one cell
print("7 moved, max pool unchanged:", np.array_equal(pool(shifted, np.max), pool(fmap, np.max)))
print("13x13 map pooled:", pool(np.zeros((13, 13)), np.max).shape)max: [[5 7] [3 5]] min: [[1 1] [0 0]] average: [[3. 3.25] [1.5 2.5 ]] board, valid: [[5]] board, same: [[5 7] [3 5]] board, stride 1: [[5 7] [5 7]] 7 moved, max pool unchanged: True 13x13 map pooled: (6, 6)
What the pooled maps show
- Max pooling gives 5 7 / 3 5, min pooling 1 1 / 0 0 and average pooling 3 3.25 / 1.5 2.5 on the same four windows.
- The board's 3×3 map gives [[5]] with valid padding and 5 7 / 3 5 with same padding; a stride of 1 gives 5 7 / 5 7.
- Moving the 7 one cell inside its window leaves the max-pooled map unchanged: the translation invariance the video describes.
- A 13×13 map becomes 6×6: the last row and column are dropped, as in the CIFAR-10 model's summary (13 → 6).
Max pooling vs average pooling
| Max pooling | Average pooling | |
|---|---|---|
| Keeps | the strongest response in each window | the mean of the window |
| On the 4×4 map | 5 7 / 3 5 | 3 3.25 / 1.5 2.5 |
| Small, bright features | kept at full strength | diluted by the zeros around them |
| Weights to learn | none | none |
| Common use | after conv blocks (VGG16, the CIFAR-10 model) | global average pooling before the output layer |
Where you use max pooling
- After each convolution block: VGG16 halves the size five times, 224 → 112 → 56 → 28 → 14 → 7.
- In small image classifiers such as the CIFAR-10 model,
MaxPooling2D((2, 2))after each Conv2D. - To cut computation: a quarter of the values after each pool means the next convolution does a quarter of the work.
padding='same' or make the sizes even.Related
- Previous: Padding and stride
- Next: CNN architecture
- Run
pool(fmap, np.max, k=3, s=1)and check the 2×2 size against ⌊(4 − 3) / 1⌋ + 1. - Pool the board's map with
np.meanand same padding; the padded −∞ cells break the average, so pad withnp.nanand usenp.nanmeaninstead. - Move the 7 two cells to the left, into the red window, and see the max-pooled map change.
Little by little, you're building something great.