Phase 2 · Classical Machine LearningModule 5~36 min read

Linear Regression

The 'hello world' of machine learning: fit a straight line to data, measure the error, and meet the loss function you'll then learn to minimize.

What you'll learn

Linear regression is the "hello world" of machine learning: fit a straight line through a cloud of points to predict a number. It's simple, but it introduces the exact ideas — weights, a bias, and a loss — that scale all the way up to neural networks.

By the end of this module you'll be able to:

  • Use a line y = w·x + b to make predictions
  • Explain what the weights and bias mean
  • Measure a fit with mean squared error
  • See how the line is found by minimising that error

Predicting a number

Some problems ask for a number, not a category: a house's price, tomorrow's temperature, a person's expected spend. That's regression. The simplest regressor assumes the output is a straight-line function of the input.

The line: weights & bias

A line is y = w·x + b. The weight w is the slope — how much y changes when x goes up by one. The bias b is the intercept — where the line sits when x = 0. With many features it becomes y = w₁x₁ + w₂x₂ + … + b: one weight per feature, a dot product plus a bias.

Fitting the line

"Learning" here means choosing w and b so the line runs through the middle of the data. Watch gradient descent do exactly that — starting flat and tilting until the error is as small as possible:

Fitting a line by minimising error
Linear regression: learning w and b
x (input)y (output)

y = 0·x + 0.5 · MSE = 0.032

1/9Start with a flat guess. The line is far from the points, so the error (MSE) is large.
Each step lowers the mean squared error until the line best fits the cloud.

Mean squared error

To fit the line we need to score it. For each point we take the residual — the vertical gap between the prediction and the truth — square it (so over- and under-shooting both count, and big misses hurt more), and average over all points. That's the mean squared error:

Mean squared error
MSE = (1/n) · Σ (ŷᵢ − yᵢ)²

ŷ the prediction · y the truth · squared so errors never cancel out.

Key idea

Fitting the line = minimising the MSE. That's a job for gradient descent (Module 6) — or, for a plain line, a one-shot formula called the normal equation that solves it exactly.

In practice

From scratch, it's the learning loop you already know, applied to a line:

linreg_scratch.py
import numpy as np

# Predict y from x with a line: y = w*x + b
w, b, lr = 0.0, 0.0, 0.1

for _ in range(2000):
    y_hat = w * X + b               # predictions for every point
    error = y_hat - y               # residuals
    w -= lr * np.mean(error * X)    # gradient descent on MSE
    b -= lr * np.mean(error)

print(f"y = {w:.2f}*x + {b:.2f}")

And with scikit-learn, it's three lines — the library finds the best weights for you:

linreg_sklearn.py
from sklearn.linear_model import LinearRegression

model = LinearRegression()
model.fit(X_train, y_train)          # finds the best weights & bias

print("weights:", model.coef_)       # one weight per feature
print("bias:   ", model.intercept_)  # the offset

y_pred = model.predict(X_test)

Tip

Linear regression is a great baseline. Before reaching for a deep network, fit a linear model — if it already does well, you may not need anything fancier.

Recap & quick check

Key takeaways

  • Regression predicts a continuous number; linear regression uses a straight line y = w·x + b.
  • The weight is the slope (effect of a feature); the bias is the intercept.
  • Mean squared error averages the squared residuals — the standard regression loss.
  • Fitting the line means minimising the MSE, via gradient descent or the normal equation.
  • With many features it becomes a weighted sum — a dot product of weights and features, plus a bias.

Quick check

1. In y = w·x + b, what does the weight w represent?

2. Why does mean squared error square the residuals?

3. What does it mean to 'fit' a linear regression?

4. With three features, linear regression predicts using…

Predicting a number is half the story. Next we predict a category — and draw the boundary between classes. Next up: Module 7 — Classification & Logistic Regression.