What you'll learn
Almost every model in this course learns the same way: it has a loss that measures how wrong it is, and it nudges its numbers to make that loss smaller. Gradient descent is the algorithm that does the nudging — and it's the single most important idea in all of machine learning.
By the end of this module you'll be able to:
- Explain what it means to minimize a loss function
- Write the gradient descent update rule and run it by hand
- Say what the learning rate does and how to pick it
- Tell apart batch, stochastic, and mini-batch descent
The problem: minimize a loss
In the last module we fit a line and measured its error with a loss function — a single number that's big when the model is wrong and small when it's right. Training is the search for the parameters that make that number as small as possible.
Picture the loss as a landscape: each setting of the parameters is a location, and its height is the loss. We want the lowest point in the valley. With millions of parameters we can't just try them all — so we need a way to walk downhill without a map.
Key idea
Follow the slope downhill
Here is gradient descent on the simplest possible loss — a single-weight bowl. Press Play and watch the ball feel the slope beneath it and roll toward the minimum. The dashed line is the tangent (the gradient); notice how its steepness shrinks as we approach the bottom, so the steps naturally get smaller.
L(w) = ½(w − 3)² · minimum at w = 3
That's the entire idea. Everything else — training a neural network with a billion weights — is this same move, done in a billion dimensions at once, over and over.
The update rule
Written as a formula, one step of gradient descent is just:
w ← w − η · ∇L(w)w the parameter · η the learning rate (step size) · ∇L the gradient of the loss.
Let's run exactly that in Python. We'll minimize the same bowl the animation used, computing the gradient by hand for now (a framework will do it for us soon):
import numpy as np
# Minimize a simple loss L(w) = 0.5 * (w - 3)**2
def loss(w): return 0.5 * (w - 3) ** 2
def gradient(w): return (w - 3) # dL/dw, worked out by hand
w = 6.4 # a random starting guess
lr = 0.32 # the learning rate (step size)
for step in range(8):
g = gradient(w) # which way is uphill?
w = w - lr * g # step the opposite way
print(f"step {step}: w={w:.3f} loss={loss(w):.3f}")The loss falls toward zero and w homes in on 3 — the bottom of the bowl — matching the animation step for step. In real projects you never differentiate by hand; a framework's autograd does it for every parameter at once:
import torch
w = torch.tensor([6.4], requires_grad=True) # a parameter to learn
optimizer = torch.optim.SGD([w], lr=0.32)
for step in range(8):
loss = 0.5 * (w - 3) ** 2
optimizer.zero_grad() # clear old gradients
loss.backward() # autograd computes dL/dw for us
optimizer.step() # w <- w - lr * w.grad
# In practice you never write the gradient by hand:
# a framework computes it for millions of parameters at once.The learning rate
The learning rate η is the one knob you tune most. It sets how big each downhill step is — and getting it wrong is the most common way training fails. Watch the same descent with three different rates:
learning rate = 0.06 (too small)
learning rate = 0.32 (just right)
learning rate = 1.9 (too large)
| Learning rate | What happens | Symptom in the loss curve |
|---|---|---|
| Too small | Converges, but agonizingly slowly | Loss creeps down; training takes forever |
| Just right | Fast, stable convergence | Loss drops smoothly and levels off |
| Too large | Overshoots; may diverge | Loss bounces around or blows up to NaN |
Watch out
NaN or shoots upward, your learning rate is almost always too high. Lower it (try 10×) before changing anything else.Batch, stochastic & mini-batch
To take one step we need the gradient — but the gradient depends on the data. How much data do we use per step? That choice gives three flavours of gradient descent:
| Variant | Data per step | Trade-off |
|---|---|---|
| Batch | The entire dataset | Smooth, accurate steps — but slow and memory-hungry |
| Stochastic (SGD) | One example | Very fast, very noisy steps; the noise can help escape bad spots |
| Mini-batch | A small batch (e.g. 32–256) | The best of both — the practical default everywhere |
Note
Local minima & saddle points
Our bowl had a single lowest point. Real loss landscapes are bumpy: they have local minima(valleys that aren't the deepest) and saddle points (flat spots that stall progress). Plain gradient descent can get stuck.
Reassuringly, in the very high-dimensional spaces of deep learning, truly bad local minima are rare, and the smarter optimizers of Module 16 (momentum, Adam) glide through flat regions. For now, hold onto the core loop — it powers everything ahead.
Recap & quick check
Key takeaways
- Training = minimizing a loss function over the model's parameters.
- The gradient points uphill; gradient descent steps in the opposite direction to go down.
- The update rule is w ← w − η·∇L, applied over and over.
- The learning rate η controls step size: too small crawls, too large diverges.
- Mini-batch gradient descent (a small batch per step) is the practical default in deep learning.
Quick check
1. Why does gradient descent step in the direction opposite to the gradient?
2. What is the role of the learning rate η?
3. Your loss suddenly becomes NaN after a few steps. The most likely cause?
4. Which variant uses a small subset of the data for each update and is the usual default?
You now know how models learn. Next we'll turn this engine on classification — teaching a model to draw the boundary between classes. Next up: Module 7 — Classification & Logistic Regression.