Phase 1 · FoundationsModule 4~38 min read

The Math You Actually Need

Just enough math — visualized, not scary: vectors and dot products, matrices as transformations, and the gradient that tells a model which way is downhill.

What you'll learn

You do not need heavy math for this course — but a handful of ideas, seen visually, make everything click. This module is that toolkit: vectors, the dot product, matrices as transformations, and the slope-of-a-curve idea that gradient descent runs on.

By the end you'll be comfortable with:

  • Reading a row of features as a vector
  • What a dot product measures and why neurons use it
  • Thinking of a matrix as a transformation of space
  • The derivative as a slope and the gradient as its multi-dimensional cousin

Vectors: data as arrows

A vector is just an ordered list of numbers — which is exactly what a row of features is. The sample [72, 2, 8] (area, bedrooms, age) is a vector, a single point (or arrow) in 3-D space. Everything in machine learning is vectors flowing through operations.

Note

We add vectors elementwise and scale them by a number. A dataset is a stack of vectors — a matrix, with one row per sample and one column per feature.

The dot product

The dot product multiplies two vectors elementwise and sums the result: a·b = a₁b₁ + a₂b₂ + …. It measures how much two vectors point the same way — the single most important operation in all of deep learning, because a neuron is exactly a dot product of inputs and weights.

What the dot product measures

Aligned

dot = large positive

Perpendicular

dot = zero

Opposed

dot = large negative

Aligned vectors → large positive; perpendicular → zero; opposed → negative.
vectors.py
import numpy as np

a = np.array([1, 2, 3])
b = np.array([4, 5, 6])

print(a + b)          # elementwise -> [5 7 9]
print(a * 2)          # scale       -> [2 4 6]
print(np.dot(a, b))   # dot product -> 1*4 + 2*5 + 3*6 = 32

# A neuron is literally a dot product of inputs and weights:
weights = np.array([0.2, -0.5, 0.1])
inputs  = np.array([3.0,  1.0, 2.0])
print(np.dot(weights, inputs))   # -> 0.3

Matrices as transformations

A matrix is a grid of numbers, but the useful way to see it is as a machine that transforms vectors — rotating, stretching, or squashing space. When you multiply a vector by a matrix, you get a new vector in a possibly new space.

That is precisely what a neural-network layer does: it multiplies its input vector by a weight matrix to reshape it into features the next layer can use. "Deep learning" is mostly stacks of matrix multiplications with a pinch of non-linearity between them.

Slopes & derivatives

The derivative of a function at a point is the slope of the curve there — how fast the output changes as you nudge the input. Walk the point along this curve and watch the tangent tilt:

The derivative is the slope of the tangent
Derivative = slope at a point
slope 0.47xf(x)

the derivative f′(x) is the slope of the tangent — here positive → rising

1/10Where the curve rises, the tangent tilts up: the slope (derivative) is positive.
Positive slope where the curve rises, zero at peaks/troughs, negative where it falls.

Notice the flat spots: at a peak or a trough the slope is zero. That is how a machine recognises a minimum — the very thing training hunts for.

The gradient

When a function has many inputs (a loss with millions of weights), its slope becomes a vector of partial slopes — one per input. That vector is the gradient, written ∇, and it points in the direction of steepest ascent.

Key idea

The whole of training in one line: the gradient points uphill on the loss, so we step in the opposite direction to go down. That's gradient descent — and now you have the exact math it stands on. We put it to work in Module 6.

Recap & quick check

Key takeaways

  • A vector is an ordered list of numbers — a row of features is a vector.
  • The dot product measures alignment; a neuron is a dot product of inputs and weights.
  • A matrix transforms vectors (rotate/stretch/squash); a network layer is a matrix multiply plus non-linearity.
  • The derivative is the slope of a curve; it is zero at minima and maxima.
  • The gradient generalises the slope to many inputs and points in the direction of steepest ascent.

Quick check

1. What does the dot product of two vectors measure?

2. A useful way to think of a matrix is as…

3. The derivative of a function at a point is…

4. Which direction does the gradient ∇ point?

That's the entire math foundation. Now we build our first real model and fit it to data. Next up: Module 5 — Linear Regression.