Phase 2 · Classical Machine LearningModule 7~38 min read

Classification & Logistic Regression

From predicting numbers to predicting classes: the sigmoid, decision boundaries, cross-entropy loss, and how a linear model draws the line between cats and dogs.

What you'll learn

Now we predict a category instead of a number: spam or not, cat or dog, fraud or fine. Despite the name, logistic regression is a classifier — it wraps a linear score in a sigmoid to output a probability, then draws a boundary between the classes.

By the end of this module you'll be able to:

  • Tell classification apart from regression
  • Use the sigmoid to turn a score into a probability
  • Read a decision boundary in feature space
  • Explain cross-entropy loss and softmax for many classes

Classification vs regression

Regression outputs a continuous number; classification outputs a class. The machinery is similar — we still compute a linear score w·x + b — but instead of using that score directly, we convert it into a probability of belonging to a class.

The sigmoid

A raw score can be any number from −∞ to +∞, but a probability must live in [0, 1]. The sigmoid function does the squashing:

The sigmoid squashes scores into probabilities
σ(z) = 1 / (1 + e⁻ᶻ)
0.5score (w·x + b)probability

σ(z) = 1 / (1 + e^−z) — squashes any score into (0, 1)

1/1The sigmoid turns any real-valued score into a probability between 0 and 1, crossing 0.5 at score 0.
Large positive score → near 1; large negative → near 0; score 0 → exactly 0.5.

If the probability is above 0.5 we predict class 1, otherwise class 0 — though that threshold is yours to move (Module 10).

Decision boundaries

Where the probability equals 0.5, the model is undecided — and that set of points forms the decision boundary. For logistic regression it's a straight line. Training rotates that line to separate the classes as well as possible:

Learning a decision boundary
Logistic regression: separating two classes
feature 1feature 2

boundary: 1·x₁ + 0.15·x₂ + -0.55 = 0

1/5Logistic regression draws a straight boundary. This first guess mis-classifies many points.
The line moves to separate the classes; the final frame shades each side's predicted class.

Watch out

A linear boundary can only separate classes that a straight line can separate. For tangled data you need curved boundaries — from trees (Module 8), kernels, or neural networks (Phase 4).

Cross-entropy loss

We can't use mean squared error for probabilities. Instead we use cross-entropy (log loss), which punishes confident wrong answers severely: predicting 0.99 for something that was actually class 0 incurs a huge penalty, while a hesitant 0.6 costs much less.

Key idea

Pair the sigmoid with cross-entropy loss and gradient descent, and logistic regression trains exactly like every other model in this course — same loop, different loss.

Softmax & multi-class

For more than two classes (digit 0–9, say), the softmax function generalises the sigmoid: it turns a vector of scores into a set of probabilities that sum to 1 — the model's confidence spread across every class. The class with the highest probability wins. We'll meet softmax again at the output of neural networks.

logistic.py
from sklearn.linear_model import LogisticRegression

model = LogisticRegression()
model.fit(X_train, y_train)

model.predict(X_test)          # class labels: 0 or 1
model.predict_proba(X_test)    # probabilities, e.g. [0.02, 0.98]

Recap & quick check

Key takeaways

  • Classification predicts a category; logistic regression is a linear classifier, not a regressor.
  • The sigmoid squashes a linear score into a probability in (0, 1).
  • The decision boundary is where the model is 50/50; for logistic regression it's a straight line.
  • Cross-entropy (log loss) is the classification loss — it heavily penalises confident wrong predictions.
  • Softmax extends the idea to many classes, giving probabilities that sum to 1.

Quick check

1. What does the sigmoid function do?

2. The decision boundary of a model is…

3. Which loss is used for classification?

4. Softmax is used to…

Straight-line boundaries only go so far. Next we meet models that carve up space in more flexible ways. Next up: Module 8 — k-NN, Decision Trees & Random Forests.