What you'll learn
"Machine learning" sounds mysterious, but underneath it is one simple, repeating loop. Once you see it, every model in this course — from a straight line to a giant language model — becomes a variation on the same three steps.
By the end of this module you'll be able to:
- Describe the predict → measure → adjust learning loop
- Say what a model's parameters are and what training does to them
- Tell apart supervised, unsupervised, and reinforcement learning
The learning loop
Every supervised learning system, no matter how large, runs the same loop. It makes a guess, checks how far off it was, and nudges itself to be a little less wrong — then does it again, and again.
Model
Makes a prediction from the input
Loss
Measures how wrong the prediction is
Optimizer
Adjusts the model to reduce the loss
↑ repeat thousands of times — the model gets a little better each loop ↑
Key idea
What training actually changes
A model is really just a formula with some adjustable numbers, called parameters (or weights). A straight line y = w·x + b has two of them; a large language model has hundreds of billions. Training is the process of finding good values for those numbers — nothing more.
Before training the parameters are random and the model is useless. Each loop of predict-measure-adjust moves them a tiny bit in a better direction. After enough loops, they encode the pattern in the data.
Three kinds of learning
Not all learning uses labelled answers. Machine learning splits into three great families, depending on what kind of feedback the model gets:
Supervised
- Data:
- Labelled examples (input → correct answer)
- Goal:
- Predict the label for new inputs
- Used for:
- Spam detection, price prediction, image labels
Unsupervised
- Data:
- Data with no labels at all
- Goal:
- Find hidden structure or groups
- Used for:
- Customer segments, anomaly detection
Reinforcement
- Data:
- An environment that gives rewards
- Goal:
- Learn actions that maximise reward
- Used for:
- Game-playing AI, robotics, RLHF for LLMs
Most of this course is supervised learning (Phases 2–6), because it's where the core ideas live. We meet unsupervised learning in Module 9 (clustering) and reinforcement learning in Module 30 — it's also how ChatGPT is polished with human feedback.
A tiny concrete example
Let's make the loop real. We'll give a model four examples of a rule it has never seen — y = 2x — and watch it learn the number 2 entirely from the data:
# Learn the rule behind some examples: y = 2x.
# We don't tell the model "2" — it discovers it.
xs = [1, 2, 3, 4]
ys = [2, 4, 6, 8]
w = 0.0 # the one number the model will learn
lr = 0.01
for epoch in range(200):
for x, y in zip(xs, ys):
pred = w * x # 1. model predicts
error = pred - y # 2. loss: how wrong?
w = w - lr * 2 * error * x # 3. optimizer nudges w
print(round(w, 3)) # -> ~2.0 (it learned the rule!)We never wrote "multiply by 2" anywhere. The model started at w = 0 and the loop walked it to w ≈ 2. That is machine learning in miniature — and the optimizer step is exactly the gradient descent we'll unpack in Module 6.
Recap & quick check
Key takeaways
- All supervised learning runs one loop: predict → measure the loss → adjust the model.
- A model is a formula with adjustable parameters (weights); training finds good values for them.
- Before training, parameters are random; each loop nudges them to reduce the loss.
- Supervised learning uses labelled answers; unsupervised finds structure; reinforcement learns from rewards.
- The 'adjust' step is gradient descent — the engine we study in Module 6.
Quick check
1. What are the three repeating steps of the learning loop?
2. What does 'training a model' actually change?
3. Which learning type uses labelled examples (input → correct answer)?
4. A model groups customers into segments with no labels provided. This is…
The loop runs on data — so next we look at data itself: features, the feature space, and the split that stops us fooling ourselves. Next up: Module 3 — Data, Features & Representation.