Phase 5 · Deep LearningModule 18~42 min read

Convolutional Neural Networks

How machines see: slide small learnable filters across an image to detect edges, textures, and shapes, building from pixels up to objects.

What you'll learn

How does a machine see? With a convolutional neural network (CNN) — the architecture that cracked computer vision. Its trick is to slide small learnable filters across an image, detecting edges, then textures, then whole objects. This is one of the most beautiful ideas in AI, and it animates gorgeously.

By the end of this module you'll be able to:

  • Explain why ordinary dense networks struggle with images
  • Describe the convolution operation
  • Define filters, feature maps, and channels
  • Use pooling, stride, and padding

Why dense nets fail on images

A modest 200×200 colour image has 120,000 numbers. A fully-connected layer would need a weight from each to every neuron — billions of parameters — and it would treat a cat in the top-left as unrelated to the same cat in the bottom-right. Images have local structure and repeating patterns, and CNNs exploit both.

The convolution operation

A convolution slides a tiny grid of weights — a kernel — across the image. At each position it multiplies the overlapping pixels by the kernel and sums them, writing one number into an output feature map. Watch a vertical-line detector sweep across an image:

A convolution, one step at a time
Convolution: a filter sweeping an image
input0010000100001000010000100
kernel010010010
→
feature map

input: a bright vertical bar · kernel: a vertical-line detector

1/20A convolution slides a small kernel across the image, computing one output cell per position.
The kernel slides over the input; where it matches the pattern (a vertical line), the feature map lights up.

Key idea

The same small kernel is reused at every position — weight sharing. That's why CNNs need so few parameters and why they detect a feature anywhere in the image (translation invariance).

Filters, feature maps & channels

Each kernel (or filter) detects one kind of pattern and produces one feature map. A convolutional layer has many filters — edges at every angle, colours, blobs — so it outputs a stack of feature maps, called channels. The next layer's filters combine those into richer patterns, exactly the feature hierarchy from Module 17. Crucially, the filter weights are learned by backpropagation, not hand-designed.

Stride, padding & pooling

Three knobs control the geometry:

  • Stride: how far the kernel jumps each step. A bigger stride shrinks the output.
  • Padding: a border of zeros so the kernel can cover the edges and keep the size.
  • Pooling: downsample a feature map (e.g. max pooling keeps the strongest response in each 2×2 block), reducing size and adding robustness.

A full CNN

A complete CNN stacks convolution + activation + pooling blocks — each layer seeing more of the image and detecting higher-level features — then a few dense layers at the end to classify. In code it's a short stack:

cnn.py
import torch.nn as nn

cnn = nn.Sequential(
    nn.Conv2d(3, 32, kernel_size=3, padding=1),  # 3 colour channels -> 32 feature maps
    nn.ReLU(),
    nn.MaxPool2d(2),                              # halve the resolution
    nn.Conv2d(32, 64, kernel_size=3, padding=1),
    nn.ReLU(),
    nn.MaxPool2d(2),
    nn.Flatten(),
    nn.Linear(64 * 8 * 8, 10),                   # classify into 10 classes
)

Recap & quick check

Key takeaways

  • CNNs process images by sliding small learnable kernels to build feature maps.
  • Weight sharing gives few parameters and detects features anywhere (translation invariance).
  • Each filter makes one feature map; many filters produce many channels.
  • Stride and padding control output size; pooling downsamples and adds robustness.
  • Stacked conv-activation-pool blocks build a feature hierarchy, ending in dense classification layers.

Quick check

1. What does a convolution kernel do?

2. Why do CNNs need far fewer parameters than dense nets on images?

3. What is max pooling?

4. Where do a CNN's filter weights come from?

Convolutions taught machines to see. Next, the architectures and tricks that made vision genuinely practical. Next up: Module 19 — Modern Computer Vision.