Phase 3 · Making Models WorkModule 10~36 min read

Evaluating Models

Accuracy lies. Learn to judge a model honestly with the confusion matrix, precision and recall, the F1 score, and the ROC curve.

What you'll learn

A model is only as trustworthy as the way you measure it — and plain accuracy hides disasters. This module gives you an honest toolkit: the confusion matrix, precision and recall, the F1 score, and the ROC curve.

By the end of this module you'll be able to:

  • Explain when accuracy is misleading
  • Read a confusion matrix
  • Compute and interpret precision, recall, and F1
  • Use the ROC curve and AUC to compare classifiers

Why accuracy misleads

Imagine a disease that affects 1 in 100 people. A model that always predicts "healthy" is 99% accurate — and utterly useless, because it never catches the disease. Whenever classes are imbalanced (fraud, disease, spam), accuracy alone is a trap.

Watch out

Always ask: what does the model get wrong, and does that kind of mistake matter? Missing a cancer is not the same as a false alarm — yet accuracy treats them identically.

The confusion matrix

Every prediction on a two-class problem falls into one of four boxes. Laid out as a grid, this is the confusion matrix — the source of every other metric:

The four outcomes of a binary prediction
Predicted +
Predicted −
Actual +

True Positive

correctly caught

False Negative

missed it

Actual −

False Positive

false alarm

True Negative

correctly ignored

Green = correct, orange = the two kinds of mistake. Most metrics are just ratios of these boxes.

Precision, recall & F1

Two metrics capture the two kinds of mistake, and they usually trade off against each other:

MetricQuestion it answersFormula
PrecisionOf the points we flagged +, how many really were?TP / (TP + FP)
RecallOf the actual +, how many did we catch?TP / (TP + FN)
F1 scoreA single balance of the two2·P·R / (P + R)
Raise the threshold and precision rises but recall falls — and vice versa.

Note

Choose by what the mistake costs. For a spam filter you want high precision (never bin a real email). For cancer screening you want high recall (never miss a case).

ROC & AUC

Every probability model has a threshold for calling something positive. The ROC curveplots the trade-off as you sweep that threshold, and the area under it (AUC) summarises quality in a single number — independent of any one threshold:

The ROC curve
ROC curve & AUC
false positive ratetrue positive rate

AUC = area under the curve · 1.0 perfect · 0.5 random (dashed)

1/1The ROC curve sweeps the threshold. A good model bows toward the top-left; the dashed diagonal is random guessing.
AUC of 1.0 is perfect; 0.5 (the diagonal) is random guessing.
metrics.py
from sklearn.metrics import (confusion_matrix, precision_score,
                             recall_score, f1_score, roc_auc_score)

confusion_matrix(y_test, y_pred)     # [[TN, FP], [FN, TP]]
precision_score(y_test, y_pred)      # of predicted +, how many were right?
recall_score(y_test, y_pred)         # of actual +, how many did we catch?
f1_score(y_test, y_pred)             # harmonic mean of the two
roc_auc_score(y_test, y_proba)       # threshold-independent quality

Regression metrics

For predicting numbers, we don't have a confusion matrix. Instead we measure the size of the errors: MAE (mean absolute error) is the average miss in the original units; RMSE (root mean squared error) is like MAE but punishes big misses more; and R² reports the fraction of the variance the model explains (1.0 is perfect, 0 is no better than predicting the mean).

Recap & quick check

Key takeaways

  • Accuracy is misleading on imbalanced data — a lazy majority guess can score high yet be useless.
  • The confusion matrix (TP, FP, FN, TN) is the basis of all classification metrics.
  • Precision = of flagged positives, how many were right; recall = of actual positives, how many were caught.
  • F1 balances precision and recall; ROC-AUC measures quality across all thresholds.
  • For regression, use MAE, RMSE, and R² to quantify error and variance explained.

Quick check

1. Why can a 99%-accurate model still be useless?

2. Recall answers which question?

3. For a cancer screening test, which metric usually matters most?

4. An ROC-AUC of 0.5 means…

Good metrics reveal the number-one enemy of every model: fitting the training data too well. Let's confront it. Next up: Module 11 — Overfitting, Regularization & Bias-Variance.