What you'll learn
How does a machine see? With a convolutional neural network (CNN) — the architecture that cracked computer vision. Its trick is to slide small learnable filters across an image, detecting edges, then textures, then whole objects. This is one of the most beautiful ideas in AI, and it animates gorgeously.
By the end of this module you'll be able to:
- Explain why ordinary dense networks struggle with images
- Describe the convolution operation
- Define filters, feature maps, and channels
- Use pooling, stride, and padding
Why dense nets fail on images
A modest 200×200 colour image has 120,000 numbers. A fully-connected layer would need a weight from each to every neuron — billions of parameters — and it would treat a cat in the top-left as unrelated to the same cat in the bottom-right. Images have local structure and repeating patterns, and CNNs exploit both.
The convolution operation
A convolution slides a tiny grid of weights — a kernel — across the image. At each position it multiplies the overlapping pixels by the kernel and sums them, writing one number into an output feature map. Watch a vertical-line detector sweep across an image:
input: a bright vertical bar · kernel: a vertical-line detector
Key idea
Filters, feature maps & channels
Each kernel (or filter) detects one kind of pattern and produces one feature map. A convolutional layer has many filters — edges at every angle, colours, blobs — so it outputs a stack of feature maps, called channels. The next layer's filters combine those into richer patterns, exactly the feature hierarchy from Module 17. Crucially, the filter weights are learned by backpropagation, not hand-designed.
Stride, padding & pooling
Three knobs control the geometry:
- Stride: how far the kernel jumps each step. A bigger stride shrinks the output.
- Padding: a border of zeros so the kernel can cover the edges and keep the size.
- Pooling: downsample a feature map (e.g. max pooling keeps the strongest response in each 2×2 block), reducing size and adding robustness.
A full CNN
A complete CNN stacks convolution + activation + pooling blocks — each layer seeing more of the image and detecting higher-level features — then a few dense layers at the end to classify. In code it's a short stack:
import torch.nn as nn
cnn = nn.Sequential(
nn.Conv2d(3, 32, kernel_size=3, padding=1), # 3 colour channels -> 32 feature maps
nn.ReLU(),
nn.MaxPool2d(2), # halve the resolution
nn.Conv2d(32, 64, kernel_size=3, padding=1),
nn.ReLU(),
nn.MaxPool2d(2),
nn.Flatten(),
nn.Linear(64 * 8 * 8, 10), # classify into 10 classes
)Recap & quick check
Key takeaways
- CNNs process images by sliding small learnable kernels to build feature maps.
- Weight sharing gives few parameters and detects features anywhere (translation invariance).
- Each filter makes one feature map; many filters produce many channels.
- Stride and padding control output size; pooling downsamples and adds robustness.
- Stacked conv-activation-pool blocks build a feature hierarchy, ending in dense classification layers.
Quick check
1. What does a convolution kernel do?
2. Why do CNNs need far fewer parameters than dense nets on images?
3. What is max pooling?
4. Where do a CNN's filter weights come from?
Convolutions taught machines to see. Next, the architectures and tricks that made vision genuinely practical. Next up: Module 19 — Modern Computer Vision.