Phase 2 · Classical Machine LearningModule 6~42 min read

Gradient Descent

The optimization engine behind almost all of machine learning: follow the slope of the loss downhill, one small step at a time, until you reach the bottom.

What you'll learn

Almost every model in this course learns the same way: it has a loss that measures how wrong it is, and it nudges its numbers to make that loss smaller. Gradient descent is the algorithm that does the nudging — and it's the single most important idea in all of machine learning.

By the end of this module you'll be able to:

  • Explain what it means to minimize a loss function
  • Write the gradient descent update rule and run it by hand
  • Say what the learning rate does and how to pick it
  • Tell apart batch, stochastic, and mini-batch descent

The problem: minimize a loss

In the last module we fit a line and measured its error with a loss function — a single number that's big when the model is wrong and small when it's right. Training is the search for the parameters that make that number as small as possible.

Picture the loss as a landscape: each setting of the parameters is a location, and its height is the loss. We want the lowest point in the valley. With millions of parameters we can't just try them all — so we need a way to walk downhill without a map.

Key idea

The whole trick: at any point we can compute the slope of the loss — the gradient. The gradient points uphill, toward higher loss. So to go down, we simply step in the opposite direction. Repeat, and we roll to the bottom.

Follow the slope downhill

Here is gradient descent on the simplest possible loss — a single-weight bowl. Press Play and watch the ball feel the slope beneath it and roll toward the minimum. The dashed line is the tangent (the gradient); notice how its steepness shrinks as we approach the bottom, so the steps naturally get smaller.

Gradient descent, one step at a time
Rolling downhill: w ← w − lr·∇L
w=6.4weight wloss L(w)

L(w) = ½(w − 3)² · minimum at w = 3

1/18We want the weight w that makes the loss smallest — the bottom of the bowl. We start at a guess, w = 6.4.
The ball always moves opposite to the slope. Flatter ground → smaller steps.

That's the entire idea. Everything else — training a neural network with a billion weights — is this same move, done in a billion dimensions at once, over and over.

The update rule

Written as a formula, one step of gradient descent is just:

The gradient descent update
w ← w − η · ∇L(w)

w the parameter · η the learning rate (step size) · ∇L the gradient of the loss.

Let's run exactly that in Python. We'll minimize the same bowl the animation used, computing the gradient by hand for now (a framework will do it for us soon):

gradient_descent.py
import numpy as np

# Minimize a simple loss L(w) = 0.5 * (w - 3)**2
def loss(w):        return 0.5 * (w - 3) ** 2
def gradient(w):    return (w - 3)          # dL/dw, worked out by hand

w = 6.4              # a random starting guess
lr = 0.32            # the learning rate (step size)

for step in range(8):
    g = gradient(w)          # which way is uphill?
    w = w - lr * g           # step the opposite way
    print(f"step {step}: w={w:.3f}  loss={loss(w):.3f}")

The loss falls toward zero and w homes in on 3 — the bottom of the bowl — matching the animation step for step. In real projects you never differentiate by hand; a framework's autograd does it for every parameter at once:

with_pytorch.py
import torch

w = torch.tensor([6.4], requires_grad=True)   # a parameter to learn
optimizer = torch.optim.SGD([w], lr=0.32)

for step in range(8):
    loss = 0.5 * (w - 3) ** 2
    optimizer.zero_grad()      # clear old gradients
    loss.backward()            # autograd computes dL/dw for us
    optimizer.step()           # w <- w - lr * w.grad

# In practice you never write the gradient by hand:
# a framework computes it for millions of parameters at once.

The learning rate

The learning rate η is the one knob you tune most. It sets how big each downhill step is — and getting it wrong is the most common way training fails. Watch the same descent with three different rates:

Too small — safe but painfully slow
Learning rate too small
w=6.2weight wloss L(w)

learning rate = 0.06 (too small)

1/10A tiny learning rate takes safe but painfully slow steps — after 0 it is barely moving.
A tiny rate barely moves; you'd wait forever to reach the bottom.
Just right — strides in and settles
Learning rate just right
w=6.2weight wloss L(w)

learning rate = 0.32 (just right)

1/8A well-chosen rate strides toward the minimum and settles quickly (step 0).
A good rate converges quickly without overshooting.
Too large — overshoots and diverges
Learning rate too large
w=6.2weight wloss L(w)

learning rate = 1.9 (too large)

1/8A huge learning rate overshoots the minimum and bounces to ever-worse losses — it diverges.
Too big a rate leaps past the minimum to ever-worse losses.
Learning rateWhat happensSymptom in the loss curve
Too smallConverges, but agonizingly slowlyLoss creeps down; training takes forever
Just rightFast, stable convergenceLoss drops smoothly and levels off
Too largeOvershoots; may divergeLoss bounces around or blows up to NaN
The learning rate is the first thing to adjust when training misbehaves.

Watch out

If your loss suddenly becomes NaN or shoots upward, your learning rate is almost always too high. Lower it (try 10×) before changing anything else.

Batch, stochastic & mini-batch

To take one step we need the gradient — but the gradient depends on the data. How much data do we use per step? That choice gives three flavours of gradient descent:

VariantData per stepTrade-off
BatchThe entire datasetSmooth, accurate steps — but slow and memory-hungry
Stochastic (SGD)One exampleVery fast, very noisy steps; the noise can help escape bad spots
Mini-batchA small batch (e.g. 32–256)The best of both — the practical default everywhere
Nearly all deep learning uses mini-batch gradient descent.

Note

"SGD" in a deep-learning framework almost always means mini-batch descent. The noise from using a subset each step is a feature, not a bug — it helps the ball rattle out of shallow dips.

Local minima & saddle points

Our bowl had a single lowest point. Real loss landscapes are bumpy: they have local minima(valleys that aren't the deepest) and saddle points (flat spots that stall progress). Plain gradient descent can get stuck.

Reassuringly, in the very high-dimensional spaces of deep learning, truly bad local minima are rare, and the smarter optimizers of Module 16 (momentum, Adam) glide through flat regions. For now, hold onto the core loop — it powers everything ahead.

Recap & quick check

Key takeaways

  • Training = minimizing a loss function over the model's parameters.
  • The gradient points uphill; gradient descent steps in the opposite direction to go down.
  • The update rule is w ← w − η·∇L, applied over and over.
  • The learning rate η controls step size: too small crawls, too large diverges.
  • Mini-batch gradient descent (a small batch per step) is the practical default in deep learning.

Quick check

1. Why does gradient descent step in the direction opposite to the gradient?

2. What is the role of the learning rate η?

3. Your loss suddenly becomes NaN after a few steps. The most likely cause?

4. Which variant uses a small subset of the data for each update and is the usual default?

You now know how models learn. Next we'll turn this engine on classification — teaching a model to draw the boundary between classes. Next up: Module 7 — Classification & Logistic Regression.