Phase 1 · FoundationsModule 2~34 min read

How Machines Learn

The single loop behind all of machine learning: make a prediction, measure how wrong it is, and adjust — plus the three great families of learning.

What you'll learn

"Machine learning" sounds mysterious, but underneath it is one simple, repeating loop. Once you see it, every model in this course — from a straight line to a giant language model — becomes a variation on the same three steps.

By the end of this module you'll be able to:

  • Describe the predict → measure → adjust learning loop
  • Say what a model's parameters are and what training does to them
  • Tell apart supervised, unsupervised, and reinforcement learning

The learning loop

Every supervised learning system, no matter how large, runs the same loop. It makes a guess, checks how far off it was, and nudges itself to be a little less wrong — then does it again, and again.

The loop behind all learning

Model

Makes a prediction from the input

Loss

Measures how wrong the prediction is

Optimizer

Adjusts the model to reduce the loss

↑ repeat thousands of times — the model gets a little better each loop ↑

Predict, measure the error, adjust to shrink it — repeat until the error is small.

Key idea

Model, loss, optimizer. Keep these three words in mind — they are the skeleton of this entire course. A "model" is the thing that predicts, a "loss" scores how wrong it is, and an "optimizer" (gradient descent, Module 6) improves the model.

What training actually changes

A model is really just a formula with some adjustable numbers, called parameters (or weights). A straight line y = w·x + b has two of them; a large language model has hundreds of billions. Training is the process of finding good values for those numbers — nothing more.

Before training the parameters are random and the model is useless. Each loop of predict-measure-adjust moves them a tiny bit in a better direction. After enough loops, they encode the pattern in the data.

Three kinds of learning

Not all learning uses labelled answers. Machine learning splits into three great families, depending on what kind of feedback the model gets:

The three families of machine learning

Supervised

Data:
Labelled examples (input → correct answer)
Goal:
Predict the label for new inputs
Used for:
Spam detection, price prediction, image labels

Unsupervised

Data:
Data with no labels at all
Goal:
Find hidden structure or groups
Used for:
Customer segments, anomaly detection

Reinforcement

Data:
An environment that gives rewards
Goal:
Learn actions that maximise reward
Used for:
Game-playing AI, robotics, RLHF for LLMs

Most of this course is supervised learning (Phases 2–6), because it's where the core ideas live. We meet unsupervised learning in Module 9 (clustering) and reinforcement learning in Module 30 — it's also how ChatGPT is polished with human feedback.

A tiny concrete example

Let's make the loop real. We'll give a model four examples of a rule it has never seen — y = 2x — and watch it learn the number 2 entirely from the data:

learn_the_rule.py
# Learn the rule behind some examples: y = 2x.
# We don't tell the model "2" — it discovers it.
xs = [1, 2, 3, 4]
ys = [2, 4, 6, 8]

w = 0.0                       # the one number the model will learn
lr = 0.01

for epoch in range(200):
    for x, y in zip(xs, ys):
        pred = w * x                 # 1. model predicts
        error = pred - y             # 2. loss: how wrong?
        w = w - lr * 2 * error * x   # 3. optimizer nudges w

print(round(w, 3))    # -> ~2.0  (it learned the rule!)

We never wrote "multiply by 2" anywhere. The model started at w = 0 and the loop walked it to w ≈ 2. That is machine learning in miniature — and the optimizer step is exactly the gradient descent we'll unpack in Module 6.

Recap & quick check

Key takeaways

  • All supervised learning runs one loop: predict → measure the loss → adjust the model.
  • A model is a formula with adjustable parameters (weights); training finds good values for them.
  • Before training, parameters are random; each loop nudges them to reduce the loss.
  • Supervised learning uses labelled answers; unsupervised finds structure; reinforcement learns from rewards.
  • The 'adjust' step is gradient descent — the engine we study in Module 6.

Quick check

1. What are the three repeating steps of the learning loop?

2. What does 'training a model' actually change?

3. Which learning type uses labelled examples (input → correct answer)?

4. A model groups customers into segments with no labels provided. This is…

The loop runs on data — so next we look at data itself: features, the feature space, and the split that stops us fooling ourselves. Next up: Module 3 — Data, Features & Representation.