Phase 4 · Neural NetworksModule 16~40 min read

Training Neural Networks

Turning the theory into a model that actually trains: epochs and batches, smarter optimizers like Adam, learning-rate schedules, and reading the loss curve.

What you'll learn

Knowing backprop is one thing; getting a network to actually train well is another. This module is the practical playbook: how data is fed in batches, which optimizer to use, the tricks that stabilise training, and how to read a loss curve like a pro.

By the end of this module you'll be able to:

  • Define epoch, batch, and iteration
  • Choose an optimizer (and know why Adam is popular)
  • Apply dropout, batch norm, and early stopping
  • Diagnose training from its loss curves

Epochs, batches & iterations

We rarely feed all the data at once. We split it into mini-batches (say 32 examples), and one update on a batch is an iteration. When the network has seen the whole dataset once, that's one epoch. Training runs for many epochs, batch by batch:

train_loop.py
import torch

for epoch in range(num_epochs):          # one pass over the data = an epoch
    for xb, yb in dataloader:            # mini-batches
        pred = model(xb)
        loss = loss_fn(pred, yb)
        optimizer.zero_grad()
        loss.backward()                  # backprop
        optimizer.step()                 # update
    val_loss = evaluate(model, val_loader)
    print(f"epoch {epoch}: val_loss={val_loss:.3f}")

Optimizers

Plain gradient descent works, but smarter optimizers converge faster and more reliably by adapting how they step:

OptimizerIdeaWhen
SGDPlain mini-batch gradient descentSimple, strong baseline
SGD + momentumBuild speed in consistent directionsSmoother, faster descent
AdamPer-parameter adaptive step sizesThe popular default for deep nets
Start with Adam; it 'just works' for most problems.

Tip

Adam combines momentum with per-parameter learning rates, so it needs little tuning. It's the sensible default — reach for plain SGD only when you have a reason.

The training toolkit

A few techniques make deep networks train reliably:

  • Weight initialization: start weights at the right scale so signals neither vanish nor explode.
  • Batch normalization: re-centre and re-scale each layer's activations to keep training stable and fast.
  • Dropout: randomly switch off neurons during training so the network can't over-rely on any one — a powerful regulariser.
  • Early stopping: stop when validation loss stops improving, capturing the best-generalising model.

Reading training curves

The loss curve is your dashboard. Watch training and validation loss together — their shapes tell you whether to keep going, stop, or change something:

Reading the loss curves
Training vs validation loss over epochs
epoch 1epochs →loss

blue = training loss · orange = validation loss

1/8Both losses fall as the network trains. The gap between them tells you how well it will generalise.
Both falling = learning. Validation turning up while training falls = overfitting — time to stop.
What you seeDiagnosisFix
Both losses high & flatUnderfittingBigger model, train longer, lower regularization
Train low, validation risingOverfittingEarly stopping, dropout, more data
Loss bouncing / NaNLearning rate too highLower the learning rate

Recap & quick check

Key takeaways

  • Data is fed in mini-batches; one pass over all data is an epoch, one batch update is an iteration.
  • Optimizers like momentum and Adam converge faster than plain SGD; Adam is a solid default.
  • Dropout, batch normalization, good initialization, and early stopping make deep training reliable.
  • The loss curve is your main diagnostic: watch training vs validation together.
  • Rising validation loss while training loss falls means overfitting — stop or regularize.

Quick check

1. What is one epoch?

2. Why is Adam a popular optimizer?

3. What does dropout do?

4. Validation loss is rising while training loss keeps falling. This means…

You can now build and train a neural network from scratch. Next we scale up — why depth unlocks a new level of capability. Next up: Module 17 — What Makes Learning "Deep".