What you'll learn
A network with random weights is useless. Backpropagation is the algorithm that teaches it — the breakthrough that made deep learning possible. It sends the error backward through the network so that every single weight learns exactly how to change. It's the most important algorithm in AI.
By the end of this module you'll be able to:
- State what backpropagation computes and why
- Explain how the chain rule assigns blame to each weight
- Trace error flowing backward through a network
- Describe the vanishing gradient problem
The learning problem
We have a loss that says how wrong the network is, and we know from Module 6 that gradient descent can reduce it — if we can compute the gradient of the loss with respect to every weight. But a deep network has millions of weights buried under many layers. How do we get a gradient for each one? That is exactly what backpropagation solves.
The chain rule
The secret is a single rule from calculus: the chain rule. If the loss depends on an output, which depends on a hidden neuron, which depends on a weight, then the effect of that weight on the loss is just the product of those links. Backpropagation is the chain rule applied systematically, from the output back to every weight.
Key idea
∂Loss/∂w for every weight w — efficiently, in one backward sweep — by multiplying local gradients along each path. Then gradient descent uses those numbers to update the weights.Propagating error backward
Here is the whole algorithm in pictures. First a forward pass makes a prediction; then the error is measured at the output and flows backward, layer by layer, until every connection knows its share of the blame:
loss compares the prediction to the truth → error at the output
One full training step
Put together, one training step is: forward (predict), loss (measure), backward (gradients), update (gradient descent). Frameworks compute the backward pass automatically with autograd, so you write four lines and they handle the calculus:
import torch, torch.nn as nn
model = nn.Sequential(nn.Linear(3, 4), nn.ReLU(), nn.Linear(4, 2))
optimizer = torch.optim.SGD(model.parameters(), lr=0.1)
loss_fn = nn.CrossEntropyLoss()
# One training step:
pred = model(x) # forward pass
loss = loss_fn(pred, target) # how wrong?
optimizer.zero_grad()
loss.backward() # BACKPROP: gradients for every weight
optimizer.step() # nudge every weight downhillRepeat that loop over your data thousands of times and the once-random weights become a trained model. Everything from image classifiers to language models is trained by this exact cycle.
Vanishing & exploding gradients
Because backprop multiplies gradients along many layers, they can shrink toward zero (vanishing) or blow up (exploding) in very deep networks — so early layers barely learn, or training destabilises. The fixes are things we'll meet next: better activations (ReLU), careful initialization, normalization, and — in Phase 5 and 6 — architectures like residual connections built specifically to keep gradients flowing.
Recap & quick check
Key takeaways
- Backpropagation computes the gradient of the loss with respect to every weight.
- It works by applying the chain rule from the output backward through the network.
- One training step = forward (predict) → loss → backward (gradients) → update (gradient descent).
- Frameworks' autograd computes the backward pass automatically.
- Multiplying gradients across many layers can make them vanish or explode in deep networks.
Quick check
1. What does backpropagation compute?
2. Which mathematical rule underlies backpropagation?
3. What are the four steps of one training iteration?
4. The vanishing gradient problem is when…
You understand how a network learns. Now let's make it train well in practice — optimizers, batches, and the tricks that matter. Next up: Module 16 — Training Neural Networks.