What you'll learn
Two functions shape how a network learns: the activation that bends each neuron, and the loss that scores the whole network. Pick them well and training flies; pick them badly and it stalls. This module is your field guide to both.
By the end of this module you'll be able to:
- Say why activations must be non-linear
- Choose between sigmoid, tanh, and ReLU
- Use softmax for multi-class outputs
- Pair the right loss (MSE vs cross-entropy) with your task
Why non-linear?
We saw it in Module 13: stack linear layers and they collapse into one. The activation function is the kink that stops that collapse — it lets each layer reshape the space so the next layer has richer material to work with. Every useful activation is therefore non-linear.
The activation zoo
A handful of activations do almost all the work. Step through the classics and notice their shapes — where they saturate (flatten) and where they stay responsive:
σ(z) = 1/(1+e⁻ᶻ) · range (0,1)
| Activation | Range | Use it for |
|---|---|---|
| Sigmoid | (0, 1) | Binary output as a probability |
| Tanh | (−1, 1) | Hidden layers (older nets), zero-centred |
| ReLU | [0, ∞) | The default for hidden layers |
| Leaky ReLU | (−∞, ∞) | When ReLU neurons 'die' (stuck at 0) |
Tip
Softmax for many classes
For a multi-class output (which of 10 digits?), the final layer uses softmax. It exponentiates each score and normalises so the outputs are positive and sum to 1 — a clean probability distribution over the classes. The biggest score becomes the most confident class.
Loss functions
The loss must match what you're predicting:
| Task | Output activation | Loss function |
|---|---|---|
| Regression (a number) | None (linear) | Mean squared error |
| Binary classification | Sigmoid | Binary cross-entropy |
| Multi-class classification | Softmax | Categorical cross-entropy |
Cross-entropy is the workhorse for classification: it measures the distance between the predicted probabilities and the true answer, punishing confident mistakes hardest — which produces strong gradients exactly when the model is badly wrong.
Choosing the pair
The reliable recipe: ReLU in the hidden layers, and an output activation + loss matched to the task (linear + MSE for numbers, sigmoid/softmax + cross-entropy for classes). Nearly every network in this course follows exactly that pattern.
Recap & quick check
Key takeaways
- Activations must be non-linear, or stacked layers collapse into one linear layer.
- ReLU is the fast, simple default for hidden layers; sigmoid and tanh saturate at the extremes.
- Softmax turns output scores into a probability distribution over many classes.
- Match the loss to the task: MSE for regression, cross-entropy for classification.
- Standard recipe: ReLU hidden layers + task-matched output activation and loss.
Quick check
1. Why must activation functions be non-linear?
2. Which activation is the usual default for hidden layers?
3. Softmax is used at the output to…
4. For multi-class classification, which loss do you use?
We now have all the pieces of a network. Time for the algorithm that actually makes it learn — the crown jewel of deep learning. Next up: Module 15 — Backpropagation.