What you'll learn
Everything so far had labels. But most data in the world is unlabelled — so how do we learn from it? This module covers the two workhorses of unsupervised learning: k-means for finding groups, and PCA for compressing many features into a few.
By the end of this module you'll be able to:
- Explain what unsupervised learning does
- Run the k-means assign–update loop in your head
- Reason about how to choose k
- Say what PCA does and why we reduce dimensions
Learning without labels
In unsupervised learning there is no "correct answer" column — just data. The goal is to discover structure: natural groups (clustering), a simpler representation (dimensionality reduction), or unusual points (anomaly detection). The model organises the data by itself.
The k-means algorithm
k-means finds k clusters by repeating two simple steps: assign every point to its nearest centroid, then move each centroid to the average of its points. Repeat until nothing moves. Watch it discover three groups with no labels at all:
start: k centroids placed at random
Note
n_init) and keep the best. It also assumes roughly round, similar-sized clusters.Choosing k
You must tell k-means how many clusters to look for — but often you don't know. A common trick is the elbow method: run k-means for several values of k, plot how tight the clusters are, and pick the "elbow" where adding more clusters stops helping much.
PCA & dimensionality
Real data has many features, and high-dimensional space is strange: points spread out and distances lose meaning — the curse of dimensionality. Principal component analysis (PCA) fights back by finding the few directions along which the data varies most, and projecting onto them.
PCA finds the direction of greatest spread (PC1). Projecting onto it keeps most of the information in fewer numbers.
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
# Group data into 3 clusters (no labels needed)
kmeans = KMeans(n_clusters=3, n_init=10).fit(X)
print(kmeans.labels_) # a cluster id for each point
# Squeeze many features down to 2 for plotting
coords = PCA(n_components=2).fit_transform(X)Where it's used
Clustering powers customer segmentation, document grouping, and image compression; PCA speeds up models, removes noise, and turns thousand-feature data into a 2-D picture you can actually look at. And a deep-learning cousin of these ideas — learning compressed representations — is exactly what powers the embeddings behind modern AI (Module 22).
Recap & quick check
Key takeaways
- Unsupervised learning finds structure in unlabelled data: clusters, compressed representations, or anomalies.
- k-means repeats assign (nearest centroid) and update (move centroid to the mean) until it converges.
- You must choose k; the elbow method helps, and multiple restarts avoid bad initialisations.
- PCA reduces dimensions by keeping the directions of greatest variance.
- These ideas underlie both practical tools and the learned embeddings of modern deep learning.
Quick check
1. What defines unsupervised learning?
2. The two repeating steps of k-means are…
3. What does PCA do?
4. Why run k-means several times with different starts?
You now have a toolbox of classical models. Before we scale up to neural networks, let's learn to judge models honestly — and beat the trap that fools every beginner. Next up: Module 10 — Evaluating Models.