Phase 5 · Deep LearningModule 19~36 min read

Modern Computer Vision

From LeNet to today: the landmark architectures, transfer learning that reuses a trained network, and the vision tasks beyond simple classification.

What you'll learn

Convolutions are the idea; this module is the practice. We'll tour the landmark architectures that pushed vision forward, meet the residual connection that made truly deep networks trainable, and learn the single most useful trick in applied deep learning: transfer learning.

By the end of this module you'll be able to:

  • Place the landmark CNNs (AlexNet → ResNet) in context
  • Explain what a residual connection fixes
  • Use transfer learning to train strong models with little data
  • Name vision tasks beyond classification

Landmark architectures

Each leap in computer vision came with a new architecture and a deeper network:

YearNetworkContribution
1998LeNet-5First successful CNN (handwritten digits)
2012AlexNetWon ImageNet, launched the deep learning era
2014VGGVery deep, uniform 3×3 convolutions
2015ResNetResidual connections → 100+ layers train reliably
The trend is relentlessly deeper networks — enabled by better training tricks.

Going deeper: ResNet

Beyond a point, simply stacking more layers made networks worse — the vanishing gradients of Module 15 choked training. ResNet solved it with a residual (skip) connection: each block learns a small change to its input and adds it back, giving gradients a clean shortcut to flow through. Suddenly networks with hundreds of layers trained fine — an idea so useful it reappears in Transformers.

Transfer learning

You rarely need to train from scratch. A network trained on millions of images has already learned general visual features; you can reuse it and just retrain a small final layer for your own task. This is transfer learning, and it lets you build a strong classifier from a few hundred images:

Transfer learning: reuse, then specialise

Pre-trained backbone

Layers trained on millions of images (ImageNet). They already know edges, textures, shapes. Keep & freeze.

New head

A small new layer for your classes, trained on your few hundred images.

Keep the pre-trained feature extractor; attach and train a small new head for your classes.
transfer.py
import torchvision.models as models
import torch.nn as nn

# Load a network already trained on ImageNet
model = models.resnet50(weights="IMAGENET1K_V2")

# Freeze the backbone; replace the final layer for OUR 5 classes
for p in model.parameters():
    p.requires_grad = False
model.fc = nn.Linear(model.fc.in_features, 5)

# Now train only the new head on a small dataset — fast and effective.

Tip

Transfer learning is the default in practice. Unless you have millions of labelled images and a big budget, start from a pre-trained model — you'll get better results with a fraction of the data and compute.

Beyond classification

Classification ("what is in this image?") is only the start. Object detection draws boxes around multiple objects (YOLO, Faster R-CNN); segmentation labels every pixel; and modern Vision Transformers (ViT) apply the attention ideas of Phase 6 to images, increasingly rivalling CNNs. Data augmentation — flipping, cropping, and colour-shifting training images — squeezes more out of limited data across all of these.

Recap & quick check

Key takeaways

  • Vision progressed through deeper architectures: LeNet → AlexNet → VGG → ResNet.
  • Residual (skip) connections let gradients flow, making very deep networks trainable.
  • Transfer learning reuses a pre-trained backbone and retrains a small head — strong results with little data.
  • Beyond classification: object detection (boxes) and segmentation (per-pixel labels).
  • Data augmentation and Vision Transformers are key parts of the modern vision toolkit.

Quick check

1. What problem do residual (skip) connections address?

2. Transfer learning means…

3. When is transfer learning especially valuable?

4. Which task labels every pixel of an image?

Images have spatial structure; language has sequential structure. Next we meet the networks built for order. Next up: Module 20 — Recurrent Networks & Sequences.