What you'll learn
Convolutions are the idea; this module is the practice. We'll tour the landmark architectures that pushed vision forward, meet the residual connection that made truly deep networks trainable, and learn the single most useful trick in applied deep learning: transfer learning.
By the end of this module you'll be able to:
- Place the landmark CNNs (AlexNet → ResNet) in context
- Explain what a residual connection fixes
- Use transfer learning to train strong models with little data
- Name vision tasks beyond classification
Landmark architectures
Each leap in computer vision came with a new architecture and a deeper network:
| Year | Network | Contribution |
|---|---|---|
| 1998 | LeNet-5 | First successful CNN (handwritten digits) |
| 2012 | AlexNet | Won ImageNet, launched the deep learning era |
| 2014 | VGG | Very deep, uniform 3×3 convolutions |
| 2015 | ResNet | Residual connections → 100+ layers train reliably |
Going deeper: ResNet
Beyond a point, simply stacking more layers made networks worse — the vanishing gradients of Module 15 choked training. ResNet solved it with a residual (skip) connection: each block learns a small change to its input and adds it back, giving gradients a clean shortcut to flow through. Suddenly networks with hundreds of layers trained fine — an idea so useful it reappears in Transformers.
Transfer learning
You rarely need to train from scratch. A network trained on millions of images has already learned general visual features; you can reuse it and just retrain a small final layer for your own task. This is transfer learning, and it lets you build a strong classifier from a few hundred images:
Pre-trained backbone
Layers trained on millions of images (ImageNet). They already know edges, textures, shapes. Keep & freeze.
New head
A small new layer for your classes, trained on your few hundred images.
import torchvision.models as models
import torch.nn as nn
# Load a network already trained on ImageNet
model = models.resnet50(weights="IMAGENET1K_V2")
# Freeze the backbone; replace the final layer for OUR 5 classes
for p in model.parameters():
p.requires_grad = False
model.fc = nn.Linear(model.fc.in_features, 5)
# Now train only the new head on a small dataset — fast and effective.Tip
Beyond classification
Classification ("what is in this image?") is only the start. Object detection draws boxes around multiple objects (YOLO, Faster R-CNN); segmentation labels every pixel; and modern Vision Transformers (ViT) apply the attention ideas of Phase 6 to images, increasingly rivalling CNNs. Data augmentation — flipping, cropping, and colour-shifting training images — squeezes more out of limited data across all of these.
Recap & quick check
Key takeaways
- Vision progressed through deeper architectures: LeNet → AlexNet → VGG → ResNet.
- Residual (skip) connections let gradients flow, making very deep networks trainable.
- Transfer learning reuses a pre-trained backbone and retrains a small head — strong results with little data.
- Beyond classification: object detection (boxes) and segmentation (per-pixel labels).
- Data augmentation and Vision Transformers are key parts of the modern vision toolkit.
Quick check
1. What problem do residual (skip) connections address?
2. Transfer learning means…
3. When is transfer learning especially valuable?
4. Which task labels every pixel of an image?
Images have spatial structure; language has sequential structure. Next we meet the networks built for order. Next up: Module 20 — Recurrent Networks & Sequences.