What you'll learn
You do not need heavy math for this course — but a handful of ideas, seen visually, make everything click. This module is that toolkit: vectors, the dot product, matrices as transformations, and the slope-of-a-curve idea that gradient descent runs on.
By the end you'll be comfortable with:
- Reading a row of features as a vector
- What a dot product measures and why neurons use it
- Thinking of a matrix as a transformation of space
- The derivative as a slope and the gradient as its multi-dimensional cousin
Vectors: data as arrows
A vector is just an ordered list of numbers — which is exactly what a row of features is. The sample [72, 2, 8] (area, bedrooms, age) is a vector, a single point (or arrow) in 3-D space. Everything in machine learning is vectors flowing through operations.
Note
The dot product
The dot product multiplies two vectors elementwise and sums the result: a·b = a₁b₁ + a₂b₂ + …. It measures how much two vectors point the same way — the single most important operation in all of deep learning, because a neuron is exactly a dot product of inputs and weights.
Aligned
dot = large positive
Perpendicular
dot = zero
Opposed
dot = large negative
import numpy as np
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
print(a + b) # elementwise -> [5 7 9]
print(a * 2) # scale -> [2 4 6]
print(np.dot(a, b)) # dot product -> 1*4 + 2*5 + 3*6 = 32
# A neuron is literally a dot product of inputs and weights:
weights = np.array([0.2, -0.5, 0.1])
inputs = np.array([3.0, 1.0, 2.0])
print(np.dot(weights, inputs)) # -> 0.3Matrices as transformations
A matrix is a grid of numbers, but the useful way to see it is as a machine that transforms vectors — rotating, stretching, or squashing space. When you multiply a vector by a matrix, you get a new vector in a possibly new space.
That is precisely what a neural-network layer does: it multiplies its input vector by a weight matrix to reshape it into features the next layer can use. "Deep learning" is mostly stacks of matrix multiplications with a pinch of non-linearity between them.
Slopes & derivatives
The derivative of a function at a point is the slope of the curve there — how fast the output changes as you nudge the input. Walk the point along this curve and watch the tangent tilt:
the derivative f′(x) is the slope of the tangent — here positive → rising
Notice the flat spots: at a peak or a trough the slope is zero. That is how a machine recognises a minimum — the very thing training hunts for.
The gradient
When a function has many inputs (a loss with millions of weights), its slope becomes a vector of partial slopes — one per input. That vector is the gradient, written ∇, and it points in the direction of steepest ascent.
Key idea
Recap & quick check
Key takeaways
- A vector is an ordered list of numbers — a row of features is a vector.
- The dot product measures alignment; a neuron is a dot product of inputs and weights.
- A matrix transforms vectors (rotate/stretch/squash); a network layer is a matrix multiply plus non-linearity.
- The derivative is the slope of a curve; it is zero at minima and maxima.
- The gradient generalises the slope to many inputs and points in the direction of steepest ascent.
Quick check
1. What does the dot product of two vectors measure?
2. A useful way to think of a matrix is as…
3. The derivative of a function at a point is…
4. Which direction does the gradient ∇ point?
That's the entire math foundation. Now we build our first real model and fit it to data. Next up: Module 5 — Linear Regression.