What you'll learn
Models are only as good as the data they learn from. Before any algorithm, you need to understand what data looks like to a machine: rows of features, points in a feature space, and the all-important split that keeps you from fooling yourself.
By the end of this module you'll be able to:
- Define sample, feature, and label
- Picture data as points in a feature space
- Handle numerical vs categorical features and know why we scale
- Split data into train / validation / test — and say why
Samples, features & labels
Machine-learning data is usually a table. Each row is a sample (one email, one house, one photo). Each column is a feature — a measurable property. In supervised learning one special column is the label: the answer we want to predict.
| Area (m²) | Bedrooms | Age (yrs) | → Price (label) |
|---|---|---|---|
| 72 | 2 | 8 | $240,000 |
| 120 | 3 | 2 | $410,000 |
| 95 | 3 | 15 | $300,000 |
The feature space
Here is the key mental picture. If each sample has two features, we can plot it as a point on a 2-D graph — its position is its features. With the labels coloured in, learning becomes a geometry problem: draw the boundary between the colours. Play through it:
each dot = one sample, placed by its features
Note
Kinds of data & scaling
Features come in two main flavours, and models need them as numbers:
| Type | Examples | How we feed it to a model |
|---|---|---|
| Numerical | Age, price, temperature | Use directly (often scaled) |
| Categorical | Colour, city, brand | Encode as numbers (e.g. one-hot) |
| Ordinal | Small / medium / large | Map to ordered integers |
One subtlety bites beginners: features on wildly different scales. If "income" ranges to 100,000 and "age" to 100, many models let income dominate purely because its numbers are bigger. The fix is scaling — put every feature on a comparable range:
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train) # learn mean/std on TRAIN
X_test = scaler.transform(X_test) # reuse them on TEST
# Now every feature has mean 0 and standard deviation 1,
# so "age in years" and "income in dollars" carry equal weight.Watch out
Train / validation / test
A model that simply memorised its training data would score perfectly on it and fail in the real world. To measure real ability, we hold data back:
| Split | Typical size | Job |
|---|---|---|
| Training set | ~60–80% | The model learns its parameters from this |
| Validation set | ~10–20% | Tune choices (which model, which settings) |
| Test set | ~10–20% | One final, honest score — used once |
from sklearn.model_selection import train_test_split
# X = features (one row per sample), y = labels
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
# Train ONLY on the training set...
model.fit(X_train, y_train)
# ...then judge it on data it has never seen.
print("test accuracy:", model.score(X_test, y_test))Garbage in, garbage out
No algorithm can rescue bad data. Missing values, mislabelled examples, and hidden bias all flow straight into the model. In practice, data cleaning and good features often matter more than the choice of algorithm — a theme we'll return to when we discuss bias in Module 31.
Recap & quick check
Key takeaways
- Data is a table: rows are samples, columns are features, and the label is what we predict.
- Every sample is a point in feature space; models carve that space into regions.
- Categorical features must be encoded as numbers; features on different scales should be scaled.
- Split data into train / validation / test — and never let the model learn from the test set.
- Data leakage (peeking at test data) and bad data quietly ruin models: garbage in, garbage out.
Quick check
1. In a data table, what is a single row usually called?
2. Why do we scale features like income and age?
3. What is the cardinal rule about the test set?
4. Computing the scaler's mean using the whole dataset (train + test) is an example of…
We keep mentioning "nudging" and "downhill." Time to meet the small amount of math that makes those precise. Next up: Module 4 — The Math You Actually Need.