What you'll learn
Wire neurons into layers and something remarkable happens: the network can learn anypattern. In this module you'll see how information flows from input to prediction — forward propagation — and why hidden layers give networks their almost unlimited power.
By the end of this module you'll be able to:
- Name the input, hidden, and output layers
- Trace a forward pass through a network
- See each layer as a matrix multiply plus an activation
- Explain why non-linearity is what makes depth matter
Layers: input, hidden, output
A neural network stacks neurons into layers. The input layer holds your features. One or more hidden layers transform them into ever more useful representations. The output layer produces the prediction. Every neuron in one layer connects to every neuron in the next — "fully connected."
Forward propagation
Forward propagation is simply running the input through the network layer by layer until a prediction pops out the other end. Watch the signal flow left to right — each layer lighting up as it computes:
forward propagation — signals flow left → right
Key idea
activation(w · x + b), layer after layer. Nothing more exotic than repeated weighted sums and squishes.Weights as matrices
Because every neuron in a layer takes the same inputs, we can bundle all their weights into a single matrix. Then a whole layer is one matrix–vector multiply plus the activation — exactly the operation from Module 4. This is why GPUs, which multiply matrices blazingly fast, are the engine of deep learning:
import numpy as np
def relu(z): return np.maximum(0, z)
def forward(x, W1, b1, W2, b2):
h = relu(W1 @ x + b1) # hidden layer (matrix multiply + activation)
y = W2 @ h + b2 # output layer
return y
# W1, W2 are weight MATRICES; each layer is one matrix multiply.Why non-linearity is essential
Here is the subtle, crucial point. If you stacked layers without an activation function, the whole network would collapse into a single linear layer — a plain line, no matter how deep. The activation functionbetween layers is what lets each layer bend the space, so that stacking them builds genuinely complex shapes.
Watch out
Universal approximation
The universal approximation theorem says a network with even one hidden layer, given enough neurons, can approximate essentially any function. In practice we prefer depth (many layers) over sheer width, because deep networks build features in stages — edges to shapes to objects — and learn far more efficiently. That staged learning is the subject of Phase 5.
Recap & quick check
Key takeaways
- A network stacks neurons into input, hidden, and output layers, usually fully connected.
- Forward propagation runs the input through each layer to produce a prediction.
- Each layer is a matrix multiply (weights) plus a bias and an activation — ideal for GPUs.
- Without non-linear activations, any number of layers collapses into a single linear layer.
- With non-linearity, networks are universal approximators; depth makes them efficient.
Quick check
1. What is forward propagation?
2. A single fully-connected layer's computation is best described as…
3. What happens if you stack layers with no activation function?
4. The universal approximation theorem roughly says…
We keep saying "activation function." Which ones, and which loss should you pair them with? That's next. Next up: Module 14 — Activation & Loss Functions.