Phase 3 · Making Models WorkModule 11~38 min read

Overfitting, Regularization & Bias-Variance

The central struggle of ML: a model that memorizes the training data but fails in the wild. Meet the bias-variance tradeoff and the tools that tame it.

What you'll learn

Here is the single most important struggle in machine learning: a model that does brilliantly on its training data and then fails on new data. It memorised instead of learning. Understanding — and defeating — this overfitting is what separates working models from demos.

By the end of this module you'll be able to:

  • Recognise underfitting and overfitting
  • Explain the bias-variance tradeoff
  • Use cross-validation to measure real performance
  • Apply regularization and other tricks to generalise better

Under-, good, and over-fit

The same data can be fit three ways. Too simple and the model misses the pattern; too complex and it chases the noise; in between, it captures the real trend. Step through all three:

Three fits to the same data
Underfitting vs overfitting
xy

underfit · a straight line is too simple → high bias

1/3Underfitting: the model is too simple to capture the pattern. It's wrong on train AND test data (high bias).
Underfit (too stiff) → good (captures the trend) → overfit (bends to every noisy point).

Key idea

The overfit curve hits every training point perfectly — and that's exactly the problem. It memorised the noise, so it will be wrong on anything new. Perfect training accuracy is a red flag, not a trophy.

The bias-variance tradeoff

Two forces cause error. Bias is error from being too simple (underfitting); varianceis error from being too sensitive to the training data (overfitting). Lowering one tends to raise the other — so we seek the balance:

Training vs validation error as complexity grows
The bias-variance tradeoff
sweet spotmodel complexity →error

blue = training error · orange = validation error

1/1As complexity grows, training error keeps falling but validation error turns back up. The sweet spot (green) balances bias and variance.
Training error always falls; validation error is U-shaped. The bottom of the U is the goal.

The tell-tale sign of overfitting: a growing gap between training and validation error. When your training loss keeps dropping but validation loss climbs, stop — you've gone past the sweet spot.

Cross-validation

A single train/test split can be lucky or unlucky. k-fold cross-validation splits the data into k parts, trains on k−1 and tests on the held-out part, rotating through all of them, then averages. It gives a far more reliable estimate of how the model will actually perform.

Regularization

Regularization discourages the model from becoming too complex by adding a penalty for large weights to the loss. The model must now justify every bit of complexity by a real reduction in error:

  • L2 (Ridge): penalises the sum of squared weights — shrinks them all toward zero.
  • L1 (Lasso): penalises absolute weights — can drive some to exactly zero, selecting features.
regularize.py
from sklearn.linear_model import Ridge, Lasso
from sklearn.model_selection import cross_val_score

# L2 (Ridge) shrinks weights; L1 (Lasso) can zero some out
ridge = Ridge(alpha=1.0)        # alpha = regularization strength
lasso = Lasso(alpha=0.1)

# 5-fold cross-validation: a more reliable score than one split
scores = cross_val_score(ridge, X, y, cv=5)
print("mean CV score:", scores.mean())

More weapons against overfitting

Beyond regularization, the reliable cures are: more data (the best fix of all), data augmentation (creating new examples by transforming existing ones), early stopping (halt training when validation error starts rising), and dropout (randomly switching off neurons during training) — the last two we'll use directly when we train neural networks in Module 16.

Recap & quick check

Key takeaways

  • Underfitting = too simple (high bias); overfitting = too complex, memorising noise (high variance).
  • A growing gap between training and validation error is the signature of overfitting.
  • The bias-variance tradeoff: seek the complexity that minimises validation error.
  • k-fold cross-validation gives a more reliable performance estimate than a single split.
  • Fight overfitting with regularization (L1/L2), more data, augmentation, early stopping, and dropout.

Quick check

1. A model scores 100% on training data but 60% on test data. This is…

2. High bias corresponds to…

3. What does L2 (Ridge) regularization do?

4. Which is generally the best cure for overfitting when available?

You can now train, measure, and tune classical models. Time for the leap that powers modern AI — the artificial neuron. Next up: Module 12 — The Artificial Neuron.