What you'll learn
A model is only as trustworthy as the way you measure it — and plain accuracy hides disasters. This module gives you an honest toolkit: the confusion matrix, precision and recall, the F1 score, and the ROC curve.
By the end of this module you'll be able to:
- Explain when accuracy is misleading
- Read a confusion matrix
- Compute and interpret precision, recall, and F1
- Use the ROC curve and AUC to compare classifiers
Why accuracy misleads
Imagine a disease that affects 1 in 100 people. A model that always predicts "healthy" is 99% accurate — and utterly useless, because it never catches the disease. Whenever classes are imbalanced (fraud, disease, spam), accuracy alone is a trap.
Watch out
The confusion matrix
Every prediction on a two-class problem falls into one of four boxes. Laid out as a grid, this is the confusion matrix — the source of every other metric:
True Positive
correctly caught
False Negative
missed it
False Positive
false alarm
True Negative
correctly ignored
Precision, recall & F1
Two metrics capture the two kinds of mistake, and they usually trade off against each other:
| Metric | Question it answers | Formula |
|---|---|---|
| Precision | Of the points we flagged +, how many really were? | TP / (TP + FP) |
| Recall | Of the actual +, how many did we catch? | TP / (TP + FN) |
| F1 score | A single balance of the two | 2·P·R / (P + R) |
Note
ROC & AUC
Every probability model has a threshold for calling something positive. The ROC curveplots the trade-off as you sweep that threshold, and the area under it (AUC) summarises quality in a single number — independent of any one threshold:
AUC = area under the curve · 1.0 perfect · 0.5 random (dashed)
from sklearn.metrics import (confusion_matrix, precision_score,
recall_score, f1_score, roc_auc_score)
confusion_matrix(y_test, y_pred) # [[TN, FP], [FN, TP]]
precision_score(y_test, y_pred) # of predicted +, how many were right?
recall_score(y_test, y_pred) # of actual +, how many did we catch?
f1_score(y_test, y_pred) # harmonic mean of the two
roc_auc_score(y_test, y_proba) # threshold-independent qualityRegression metrics
For predicting numbers, we don't have a confusion matrix. Instead we measure the size of the errors: MAE (mean absolute error) is the average miss in the original units; RMSE (root mean squared error) is like MAE but punishes big misses more; and R² reports the fraction of the variance the model explains (1.0 is perfect, 0 is no better than predicting the mean).
Recap & quick check
Key takeaways
- Accuracy is misleading on imbalanced data — a lazy majority guess can score high yet be useless.
- The confusion matrix (TP, FP, FN, TN) is the basis of all classification metrics.
- Precision = of flagged positives, how many were right; recall = of actual positives, how many were caught.
- F1 balances precision and recall; ROC-AUC measures quality across all thresholds.
- For regression, use MAE, RMSE, and R² to quantify error and variance explained.
Quick check
1. Why can a 99%-accurate model still be useless?
2. Recall answers which question?
3. For a cancer screening test, which metric usually matters most?
4. An ROC-AUC of 0.5 means…
Good metrics reveal the number-one enemy of every model: fitting the training data too well. Let's confront it. Next up: Module 11 — Overfitting, Regularization & Bias-Variance.