Phase 1 · FoundationsModule 3~36 min read

Data, Features & Representation

Models learn from data, so data is where everything begins — features, the feature space, and the train / validation / test split that keeps you honest.

What you'll learn

Models are only as good as the data they learn from. Before any algorithm, you need to understand what data looks like to a machine: rows of features, points in a feature space, and the all-important split that keeps you from fooling yourself.

By the end of this module you'll be able to:

  • Define sample, feature, and label
  • Picture data as points in a feature space
  • Handle numerical vs categorical features and know why we scale
  • Split data into train / validation / test — and say why

Samples, features & labels

Machine-learning data is usually a table. Each row is a sample (one email, one house, one photo). Each column is a feature — a measurable property. In supervised learning one special column is the label: the answer we want to predict.

Area (m²)BedroomsAge (yrs)→ Price (label)
7228$240,000
12032$410,000
95315$300,000
Three samples, three features, one label. The model learns features → label.

The feature space

Here is the key mental picture. If each sample has two features, we can plot it as a point on a 2-D graph — its position is its features. With the labels coloured in, learning becomes a geometry problem: draw the boundary between the colours. Play through it:

Data as points in feature space
From raw samples to a train/test split
feature 1 (e.g. weight)feature 2 (e.g. height)

each dot = one sample, placed by its features

1/3A dataset is just points in a feature space — here, every animal plotted by two measurements.
Samples become points; labels become colours; a held-out test set is kept hidden.

Note

Real data has hundreds or thousands of features — a space with thousands of dimensions we can't draw. But the intuition from 2-D carries over: samples are points, and models carve up the space.

Kinds of data & scaling

Features come in two main flavours, and models need them as numbers:

TypeExamplesHow we feed it to a model
NumericalAge, price, temperatureUse directly (often scaled)
CategoricalColour, city, brandEncode as numbers (e.g. one-hot)
OrdinalSmall / medium / largeMap to ordered integers

One subtlety bites beginners: features on wildly different scales. If "income" ranges to 100,000 and "age" to 100, many models let income dominate purely because its numbers are bigger. The fix is scaling — put every feature on a comparable range:

scaling.py
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)   # learn mean/std on TRAIN
X_test  = scaler.transform(X_test)        # reuse them on TEST

# Now every feature has mean 0 and standard deviation 1,
# so "age in years" and "income in dollars" carry equal weight.

Watch out

Always compute scaling statistics on the training set only, then apply them to the test set. Peeking at the test set — even just its mean — is data leakage, and it makes your results a lie.

Train / validation / test

A model that simply memorised its training data would score perfectly on it and fail in the real world. To measure real ability, we hold data back:

SplitTypical sizeJob
Training set~60–80%The model learns its parameters from this
Validation set~10–20%Tune choices (which model, which settings)
Test set~10–20%One final, honest score — used once
The golden rule: never let the model learn from the test set.
split.py
from sklearn.model_selection import train_test_split

# X = features (one row per sample), y = labels
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# Train ONLY on the training set...
model.fit(X_train, y_train)

# ...then judge it on data it has never seen.
print("test accuracy:", model.score(X_test, y_test))

Garbage in, garbage out

No algorithm can rescue bad data. Missing values, mislabelled examples, and hidden bias all flow straight into the model. In practice, data cleaning and good features often matter more than the choice of algorithm — a theme we'll return to when we discuss bias in Module 31.

Recap & quick check

Key takeaways

  • Data is a table: rows are samples, columns are features, and the label is what we predict.
  • Every sample is a point in feature space; models carve that space into regions.
  • Categorical features must be encoded as numbers; features on different scales should be scaled.
  • Split data into train / validation / test — and never let the model learn from the test set.
  • Data leakage (peeking at test data) and bad data quietly ruin models: garbage in, garbage out.

Quick check

1. In a data table, what is a single row usually called?

2. Why do we scale features like income and age?

3. What is the cardinal rule about the test set?

4. Computing the scaler's mean using the whole dataset (train + test) is an example of…

We keep mentioning "nudging" and "downhill." Time to meet the small amount of math that makes those precise. Next up: Module 4 — The Math You Actually Need.