What you'll learn
Here is the single most important struggle in machine learning: a model that does brilliantly on its training data and then fails on new data. It memorised instead of learning. Understanding — and defeating — this overfitting is what separates working models from demos.
By the end of this module you'll be able to:
- Recognise underfitting and overfitting
- Explain the bias-variance tradeoff
- Use cross-validation to measure real performance
- Apply regularization and other tricks to generalise better
Under-, good, and over-fit
The same data can be fit three ways. Too simple and the model misses the pattern; too complex and it chases the noise; in between, it captures the real trend. Step through all three:
underfit · a straight line is too simple → high bias
Key idea
The bias-variance tradeoff
Two forces cause error. Bias is error from being too simple (underfitting); varianceis error from being too sensitive to the training data (overfitting). Lowering one tends to raise the other — so we seek the balance:
blue = training error · orange = validation error
The tell-tale sign of overfitting: a growing gap between training and validation error. When your training loss keeps dropping but validation loss climbs, stop — you've gone past the sweet spot.
Cross-validation
A single train/test split can be lucky or unlucky. k-fold cross-validation splits the data into k parts, trains on k−1 and tests on the held-out part, rotating through all of them, then averages. It gives a far more reliable estimate of how the model will actually perform.
Regularization
Regularization discourages the model from becoming too complex by adding a penalty for large weights to the loss. The model must now justify every bit of complexity by a real reduction in error:
- L2 (Ridge): penalises the sum of squared weights — shrinks them all toward zero.
- L1 (Lasso): penalises absolute weights — can drive some to exactly zero, selecting features.
from sklearn.linear_model import Ridge, Lasso
from sklearn.model_selection import cross_val_score
# L2 (Ridge) shrinks weights; L1 (Lasso) can zero some out
ridge = Ridge(alpha=1.0) # alpha = regularization strength
lasso = Lasso(alpha=0.1)
# 5-fold cross-validation: a more reliable score than one split
scores = cross_val_score(ridge, X, y, cv=5)
print("mean CV score:", scores.mean())More weapons against overfitting
Beyond regularization, the reliable cures are: more data (the best fix of all), data augmentation (creating new examples by transforming existing ones), early stopping (halt training when validation error starts rising), and dropout (randomly switching off neurons during training) — the last two we'll use directly when we train neural networks in Module 16.
Recap & quick check
Key takeaways
- Underfitting = too simple (high bias); overfitting = too complex, memorising noise (high variance).
- A growing gap between training and validation error is the signature of overfitting.
- The bias-variance tradeoff: seek the complexity that minimises validation error.
- k-fold cross-validation gives a more reliable performance estimate than a single split.
- Fight overfitting with regularization (L1/L2), more data, augmentation, early stopping, and dropout.
Quick check
1. A model scores 100% on training data but 60% on test data. This is…
2. High bias corresponds to…
3. What does L2 (Ridge) regularization do?
4. Which is generally the best cure for overfitting when available?
You can now train, measure, and tune classical models. Time for the leap that powers modern AI — the artificial neuron. Next up: Module 12 — The Artificial Neuron.