Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Convolution and filters

Convolution is an operation that slides a small grid of weights, the filter or kernel, over an image and writes the sum of the cell-by-cell products at each position into a new grid called a feature map.

Last updated: 05 Oct, 2026 · NumPy

It is the first step of every CNN in Convolutional neural networks (CNN). A filter that matches a pattern, such as a vertical edge, gives large values where the pattern is and zeros where it is not.

Sliding a 3×3 filter over a 6×6 image

The board's image is 6×6: three columns of 0 (black) and three of 255 (white). The first step divides every pixel by 255, which turns the image into three columns of 0 and three of 1. The video calls this min max scaling; dividing by 255 equals min-max scaling only when the darkest pixel is 0 and the brightest 255, so it is usually called normalising or rescaling.

A 3×3 filter is placed on the top-left corner. Each filter value is multiplied by the pixel under it, the nine products are added, and the sum goes into the first cell of the output. The filter then moves one cell to the right, which is a stride of 1, and repeats; at the end of a row it moves one cell down. A 3×3 filter fits in 4 positions across a 6-pixel row, so the output is 4×4.

The first filter on the board is 1 2 1 / 0 0 0 / −1 −2 −1, the horizontal edge filter. The board writes rows of −4 for it; on this image every window sums to 0, because all rows of the image are the same and the +1 +2 +1 row cancels the −1 −2 −1 row. A 4×4 of zeros is the right answer: the image has no horizontal edge.

Finding the vertical edge

The vertical edge filter worked out · from the Deep Learning In-depth Tutorials in 5 Hours video · 298:28 to 303:53

The new filter is 1 0 −1 / 2 0 −2 / 1 0 −1, the vertical edge filter. In the first position every pixel under it is 0, so the sum is 0. One step to the right, the right column of the filter (−1, −2, −1) sits on 1s and the rest on 0s, so the sum is −4. The third position also gives −4, and the fourth is 0 again. Every row of the image is the same, so every row of the output is 0 −4 −4 0.

The 6 by 6 image of three 0 columns and three 1 columns is convolved with the vertical edge filter; the second window sums to minus 4, each output row is 0, -4, -4, 0, and rescaled to 0-255 it becomes 255, 0, 0, 255: white sides and a black band at the edge.

To see the result as an image, the output goes back to the 0-255 range: the lowest value (−4) becomes 0 and the highest (0) becomes 255, so each row reads 255 0 0 255. 255 is white and 0 is black, so the picture is white on both sides with a black band in the middle: the vertical edge, where the image changes from 0 to 1.

The clip ends by comparing the filter with the V1 to V7 layers, each extracting one kind of information, and by calling this a correlation operation. That is the precise name: the sum above, with the filter not flipped, is cross-correlation, and it is what Keras Conv2D computes. Convolution in mathematics flips the filter first; for filters that are learned the difference does not matter.

The edge comes out dark because the image goes from dark (0) on the left to light (1) on the right, which makes the sums negative there. On an image that goes from light to dark, the same filter gives +4 and the edge shows white.

Learning the filter values with backpropagation

Filters learned by backpropagation, ReLU and stride · from the Deep Learning In-depth Tutorials in 5 Hours video · 311:29 to 315:10

A real CNN does not hard-code filters. As with the weights of an ANN, the filter values start random and are updated by backpropagation (Backpropagation and weight update), so the network finds the filters that suit its images. A layer holds many filters, for horizontal edges, vertical edges, round shapes and more, and each filter gives its own output.

After the convolution, a ReLU activation, max(0, x), is applied to every value of the output. Its derivative is easy to find, which backpropagation needs to update the filters; other activations such as PReLU also work (ReLU and its variants).

The clip also changes the stride to 2 and writes the output size as (n + 2p − f + 1) / s; the correct formula is ⌊(n + 2p − f) / s⌋ + 1, worked through in Padding and stride.

In a real layer ReLU acts on the raw output 0 −4 −4 0, which gives all zeros: this filter's edge would vanish. A learned filter can take the opposite sign, −1 0 1 / −2 0 2 / −1 0 1, and give +4, which ReLU keeps. Training settles on whichever sign is useful.

Convolving an RGB image

The video shows RGB images but convolves only the grayscale one. For an image with C channels, each filter is f×f×C: one f×f slice per channel. At each position all f·f·C products are added, plus one bias, giving a single number, so one filter always gives one 2-D feature map. A layer of k filters gives k feature maps, stacked as the k channels of the output.

The CIFAR-10 model in CNN in Keras starts with 32 filters of 3×3 on 3 channels: 3·3·3·32 + 32 = 896 parameters. VGG16's first layer has 64 filters: 3·3·3·64 + 64 = 1,792, and its second layer, on 64 channels, 3·3·64·64 + 64 = 36,928.

A 32 by 32 by 3 RGB image is convolved with k filters of 3 by 3 by 3; each filter gives one 30 by 30 feature map, the output has k channels, and the layer has f times f times C times k plus k parameters: 3 times 3 times 3 times 32 plus 32 is 896.

Running the 6×6 example in NumPy

The convolution function

Two loops place the filter at every position; at each one the window is multiplied cell by cell with the filter and summed.

python
def conv2d(img, k):
    f = k.shape[0]
    n = img.shape[0] - f + 1                       # n - f + 1
    out = np.zeros((n, n))
    for i in range(n):
        for j in range(n):
            out[i, j] = (img[i:i + f, j:j + f] * k).sum()   # multiply cell by cell, add
    return out

The two filters and the rescaling

python
vertical = np.array([[1, 0, -1], [2, 0, -2], [1, 0, -1]])
horizontal = np.array([[1, 2, 1], [0, 0, 0], [-1, -2, -1]])
v = conv2d(img, vertical)
rescaled = (v - v.min()) / (v.max() - v.min()) * 255   # -4 -> 0, 0 -> 255

Filters over three channels

For RGB, each filter's three slices are applied to the three channels and the results added, which gives one map per filter.

python
maps = np.stack([sum(conv2d(rgb[:, :, c], w[:, :, c]) for c in range(3))
                 for w in filters], axis=-1)        # one map per filter, stacked

The vertical and horizontal filters on the board's image

ExampleRun with NumPy
import numpy as np
from scipy.signal import correlate2d

def conv2d(img, k):
    f = k.shape[0]
    n = img.shape[0] - f + 1                       # n - f + 1
    out = np.zeros((n, n))
    for i in range(n):
        for j in range(n):
            out[i, j] = (img[i:i + f, j:j + f] * k).sum()   # multiply cell by cell, add
    return out

img = np.array([[0, 0, 0, 255, 255, 255]] * 6) / 255     # the 6x6 image, scaled to 0-1
vertical = np.array([[1, 0, -1], [2, 0, -2], [1, 0, -1]])
horizontal = np.array([[1, 2, 1], [0, 0, 0], [-1, -2, -1]])

v = conv2d(img, vertical)
print("vertical filter:")
print(v.astype(int))
print("rescaled to 0-255:")
print(((v - v.min()) / (v.max() - v.min()) * 255).astype(int))
print("horizontal filter:")
print(conv2d(img, horizontal).astype(int))
print("ReLU of the vertical map, row 1:", np.maximum(0, v)[0])
print("ReLU with the filter's sign flipped:", np.maximum(0, conv2d(img, -vertical))[0])
print("same as scipy correlate2d:", np.allclose(v, correlate2d(img, vertical, mode="valid")))

rgb = np.random.default_rng(0).random((32, 32, 3))                 # a random 32x32 RGB image
filters = np.random.default_rng(1).standard_normal((32, 3, 3, 3))  # k = 32 filters, each 3x3x3
maps = np.stack([sum(conv2d(rgb[:, :, c], w[:, :, c]) for c in range(3)) for w in filters], axis=-1)
print("RGB output shape:", maps.shape, " parameters:", filters.size + len(filters))

What the feature maps show

  • The vertical filter gives 0 −4 −4 0 in every row, the board's values, and rescaling turns them into 255 0 0 255.
  • The horizontal filter gives a 4×4 of zeros: the image has no horizontal edge, so the board's rows of −4 do not appear.
  • ReLU on the raw map leaves only zeros, while the filter with its sign flipped gives 0 4 4 0, which ReLU keeps.
  • scipy's correlate2d agrees with the two loops, which confirms that a CNN's convolution is cross-correlation.
  • 32 filters on a 32×32×3 image give an output of shape (30, 30, 32) and 896 parameters, the numbers of the first layer in the CIFAR-10 model.

Vertical vs horizontal edge filter

Vertical edge filterHorizontal edge filter
Values1 0 −1 / 2 0 −2 / 1 0 −11 2 1 / 0 0 0 / −1 −2 −1
Responds toa change from left to righta change from top to bottom
On the 6×6 image0 −4 −4 0 in each rowall zeros
After rescaling255 0 0 255: a dark vertical bandno edge to show
Classic nameSobel x filterSobel y filter

Where you use convolution

  • The first layers of every CNN, where learned filters play the role of these hand-made edge filters.
  • Classic image processing: the two filters here are the Sobel filters, used for edge detection in tools such as OpenCV.
  • Blurring and sharpening: a 3×3 filter of equal weights 1/9 blurs an image, and a filter with a large centre value sharpens it.
Watch out. Every convolution without padding shrinks the map: a 6×6 image becomes 4×4, and three such layers on a 32×32 photo leave 26×26. The pixels at the border take part in fewer windows than those in the middle, which is the problem padding solves.
Try it yourself
  • Use the image [[255, 255, 255, 0, 0, 0]] * 6 (light to dark) and check that the vertical filter now gives 0 4 4 0.
  • Transpose the image with img.T and run the horizontal filter on it: the edge is now horizontal, so the zeros turn into −4s.
  • Change the number of filters from 32 to 64 and check that the parameter count becomes 1,792 as in VGG16's first layer.

This is what real progress feels like.