What you'll learn
Knowing backprop is one thing; getting a network to actually train well is another. This module is the practical playbook: how data is fed in batches, which optimizer to use, the tricks that stabilise training, and how to read a loss curve like a pro.
By the end of this module you'll be able to:
- Define epoch, batch, and iteration
- Choose an optimizer (and know why Adam is popular)
- Apply dropout, batch norm, and early stopping
- Diagnose training from its loss curves
Epochs, batches & iterations
We rarely feed all the data at once. We split it into mini-batches (say 32 examples), and one update on a batch is an iteration. When the network has seen the whole dataset once, that's one epoch. Training runs for many epochs, batch by batch:
import torch
for epoch in range(num_epochs): # one pass over the data = an epoch
for xb, yb in dataloader: # mini-batches
pred = model(xb)
loss = loss_fn(pred, yb)
optimizer.zero_grad()
loss.backward() # backprop
optimizer.step() # update
val_loss = evaluate(model, val_loader)
print(f"epoch {epoch}: val_loss={val_loss:.3f}")Optimizers
Plain gradient descent works, but smarter optimizers converge faster and more reliably by adapting how they step:
| Optimizer | Idea | When |
|---|---|---|
| SGD | Plain mini-batch gradient descent | Simple, strong baseline |
| SGD + momentum | Build speed in consistent directions | Smoother, faster descent |
| Adam | Per-parameter adaptive step sizes | The popular default for deep nets |
Tip
The training toolkit
A few techniques make deep networks train reliably:
- Weight initialization: start weights at the right scale so signals neither vanish nor explode.
- Batch normalization: re-centre and re-scale each layer's activations to keep training stable and fast.
- Dropout: randomly switch off neurons during training so the network can't over-rely on any one — a powerful regulariser.
- Early stopping: stop when validation loss stops improving, capturing the best-generalising model.
Reading training curves
The loss curve is your dashboard. Watch training and validation loss together — their shapes tell you whether to keep going, stop, or change something:
blue = training loss · orange = validation loss
| What you see | Diagnosis | Fix |
|---|---|---|
| Both losses high & flat | Underfitting | Bigger model, train longer, lower regularization |
| Train low, validation rising | Overfitting | Early stopping, dropout, more data |
| Loss bouncing / NaN | Learning rate too high | Lower the learning rate |
Recap & quick check
Key takeaways
- Data is fed in mini-batches; one pass over all data is an epoch, one batch update is an iteration.
- Optimizers like momentum and Adam converge faster than plain SGD; Adam is a solid default.
- Dropout, batch normalization, good initialization, and early stopping make deep training reliable.
- The loss curve is your main diagnostic: watch training vs validation together.
- Rising validation loss while training loss falls means overfitting — stop or regularize.
Quick check
1. What is one epoch?
2. Why is Adam a popular optimizer?
3. What does dropout do?
4. Validation loss is rising while training loss keeps falling. This means…
You can now build and train a neural network from scratch. Next we scale up — why depth unlocks a new level of capability. Next up: Module 17 — What Makes Learning "Deep".