What you'll learn
Linear regression is the "hello world" of machine learning: fit a straight line through a cloud of points to predict a number. It's simple, but it introduces the exact ideas — weights, a bias, and a loss — that scale all the way up to neural networks.
By the end of this module you'll be able to:
- Use a line
y = w·x + bto make predictions - Explain what the weights and bias mean
- Measure a fit with mean squared error
- See how the line is found by minimising that error
Predicting a number
Some problems ask for a number, not a category: a house's price, tomorrow's temperature, a person's expected spend. That's regression. The simplest regressor assumes the output is a straight-line function of the input.
The line: weights & bias
A line is y = w·x + b. The weight w is the slope — how much y changes when x goes up by one. The bias b is the intercept — where the line sits when x = 0. With many features it becomes y = w₁x₁ + w₂x₂ + … + b: one weight per feature, a dot product plus a bias.
Fitting the line
"Learning" here means choosing w and b so the line runs through the middle of the data. Watch gradient descent do exactly that — starting flat and tilting until the error is as small as possible:
y = 0·x + 0.5 · MSE = 0.032
Mean squared error
To fit the line we need to score it. For each point we take the residual — the vertical gap between the prediction and the truth — square it (so over- and under-shooting both count, and big misses hurt more), and average over all points. That's the mean squared error:
MSE = (1/n) · Σ (ŷᵢ − yᵢ)²ŷ the prediction · y the truth · squared so errors never cancel out.
Key idea
In practice
From scratch, it's the learning loop you already know, applied to a line:
import numpy as np
# Predict y from x with a line: y = w*x + b
w, b, lr = 0.0, 0.0, 0.1
for _ in range(2000):
y_hat = w * X + b # predictions for every point
error = y_hat - y # residuals
w -= lr * np.mean(error * X) # gradient descent on MSE
b -= lr * np.mean(error)
print(f"y = {w:.2f}*x + {b:.2f}")And with scikit-learn, it's three lines — the library finds the best weights for you:
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train) # finds the best weights & bias
print("weights:", model.coef_) # one weight per feature
print("bias: ", model.intercept_) # the offset
y_pred = model.predict(X_test)Tip
Recap & quick check
Key takeaways
- Regression predicts a continuous number; linear regression uses a straight line y = w·x + b.
- The weight is the slope (effect of a feature); the bias is the intercept.
- Mean squared error averages the squared residuals — the standard regression loss.
- Fitting the line means minimising the MSE, via gradient descent or the normal equation.
- With many features it becomes a weighted sum — a dot product of weights and features, plus a bias.
Quick check
1. In y = w·x + b, what does the weight w represent?
2. Why does mean squared error square the residuals?
3. What does it mean to 'fit' a linear regression?
4. With three features, linear regression predicts using…
Predicting a number is half the story. Next we predict a category — and draw the boundary between classes. Next up: Module 7 — Classification & Logistic Regression.