Phase 2 · Classical Machine LearningModule 9~38 min read

Unsupervised Learning: k-Means & PCA

Find structure with no labels at all: group data into clusters with k-means, and squeeze many dimensions into a few with principal component analysis.

What you'll learn

Everything so far had labels. But most data in the world is unlabelled — so how do we learn from it? This module covers the two workhorses of unsupervised learning: k-means for finding groups, and PCA for compressing many features into a few.

By the end of this module you'll be able to:

  • Explain what unsupervised learning does
  • Run the k-means assign–update loop in your head
  • Reason about how to choose k
  • Say what PCA does and why we reduce dimensions

Learning without labels

In unsupervised learning there is no "correct answer" column — just data. The goal is to discover structure: natural groups (clustering), a simpler representation (dimensionality reduction), or unusual points (anomaly detection). The model organises the data by itself.

The k-means algorithm

k-means finds k clusters by repeating two simple steps: assign every point to its nearest centroid, then move each centroid to the average of its points. Repeat until nothing moves. Watch it discover three groups with no labels at all:

k-means: assign, then update
k-means clustering (k = 3)
feature 1feature 2

start: k centroids placed at random

1/9k-means finds groups with no labels. We start by dropping 3 centroids (✕) at random.
Centroids (✕) start random, then converge as points are reassigned and centroids recentred.

Note

k-means is fast and intuitive but sensitive to where the centroids start, so libraries run it several times (n_init) and keep the best. It also assumes roughly round, similar-sized clusters.

Choosing k

You must tell k-means how many clusters to look for — but often you don't know. A common trick is the elbow method: run k-means for several values of k, plot how tight the clusters are, and pick the "elbow" where adding more clusters stops helping much.

PCA & dimensionality

Real data has many features, and high-dimensional space is strange: points spread out and distances lose meaning — the curse of dimensionality. Principal component analysis (PCA) fights back by finding the few directions along which the data varies most, and projecting onto them.

PCA finds the directions of greatest variance
PC1

PCA finds the direction of greatest spread (PC1). Projecting onto it keeps most of the information in fewer numbers.

Keep the top directions and you compress many features into a few, losing little.
unsupervised.py
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA

# Group data into 3 clusters (no labels needed)
kmeans = KMeans(n_clusters=3, n_init=10).fit(X)
print(kmeans.labels_)          # a cluster id for each point

# Squeeze many features down to 2 for plotting
coords = PCA(n_components=2).fit_transform(X)

Where it's used

Clustering powers customer segmentation, document grouping, and image compression; PCA speeds up models, removes noise, and turns thousand-feature data into a 2-D picture you can actually look at. And a deep-learning cousin of these ideas — learning compressed representations — is exactly what powers the embeddings behind modern AI (Module 22).

Recap & quick check

Key takeaways

  • Unsupervised learning finds structure in unlabelled data: clusters, compressed representations, or anomalies.
  • k-means repeats assign (nearest centroid) and update (move centroid to the mean) until it converges.
  • You must choose k; the elbow method helps, and multiple restarts avoid bad initialisations.
  • PCA reduces dimensions by keeping the directions of greatest variance.
  • These ideas underlie both practical tools and the learned embeddings of modern deep learning.

Quick check

1. What defines unsupervised learning?

2. The two repeating steps of k-means are…

3. What does PCA do?

4. Why run k-means several times with different starts?

You now have a toolbox of classical models. Before we scale up to neural networks, let's learn to judge models honestly — and beat the trap that fools every beginner. Next up: Module 10 — Evaluating Models.