Phase 4 · Neural NetworksModule 14~34 min read

Activation & Loss Functions

The two functions that shape learning: activations (ReLU, sigmoid, tanh, softmax) that bend the network, and losses that tell it exactly how wrong it is.

What you'll learn

Two functions shape how a network learns: the activation that bends each neuron, and the loss that scores the whole network. Pick them well and training flies; pick them badly and it stalls. This module is your field guide to both.

By the end of this module you'll be able to:

  • Say why activations must be non-linear
  • Choose between sigmoid, tanh, and ReLU
  • Use softmax for multi-class outputs
  • Pair the right loss (MSE vs cross-entropy) with your task

Why non-linear?

We saw it in Module 13: stack linear layers and they collapse into one. The activation function is the kink that stops that collapse — it lets each layer reshape the space so the next layer has richer material to work with. Every useful activation is therefore non-linear.

The activation zoo

A handful of activations do almost all the work. Step through the classics and notice their shapes — where they saturate (flatten) and where they stay responsive:

The common activation functions
Activation functions
input zsigmoid(z)

σ(z) = 1/(1+e⁻ᶻ) · range (0,1)

1/4Sigmoid squashes into (0,1) — great for probabilities, but saturates and slows learning at the extremes.
Sigmoid & tanh saturate at the edges; ReLU is the fast, simple default; Leaky ReLU avoids 'dead' neurons.
ActivationRangeUse it for
Sigmoid(0, 1)Binary output as a probability
Tanh(−1, 1)Hidden layers (older nets), zero-centred
ReLU[0, ∞)The default for hidden layers
Leaky ReLU(−∞, ∞)When ReLU neurons 'die' (stuck at 0)

Tip

When in doubt, use ReLU in hidden layers. It's fast, avoids the saturation that slows sigmoid and tanh, and it just works for most networks.

Softmax for many classes

For a multi-class output (which of 10 digits?), the final layer uses softmax. It exponentiates each score and normalises so the outputs are positive and sum to 1 — a clean probability distribution over the classes. The biggest score becomes the most confident class.

Loss functions

The loss must match what you're predicting:

TaskOutput activationLoss function
Regression (a number)None (linear)Mean squared error
Binary classificationSigmoidBinary cross-entropy
Multi-class classificationSoftmaxCategorical cross-entropy
Get this pairing wrong and training struggles no matter how good the network is.

Cross-entropy is the workhorse for classification: it measures the distance between the predicted probabilities and the true answer, punishing confident mistakes hardest — which produces strong gradients exactly when the model is badly wrong.

Choosing the pair

The reliable recipe: ReLU in the hidden layers, and an output activation + loss matched to the task (linear + MSE for numbers, sigmoid/softmax + cross-entropy for classes). Nearly every network in this course follows exactly that pattern.

Recap & quick check

Key takeaways

  • Activations must be non-linear, or stacked layers collapse into one linear layer.
  • ReLU is the fast, simple default for hidden layers; sigmoid and tanh saturate at the extremes.
  • Softmax turns output scores into a probability distribution over many classes.
  • Match the loss to the task: MSE for regression, cross-entropy for classification.
  • Standard recipe: ReLU hidden layers + task-matched output activation and loss.

Quick check

1. Why must activation functions be non-linear?

2. Which activation is the usual default for hidden layers?

3. Softmax is used at the output to…

4. For multi-class classification, which loss do you use?

We now have all the pieces of a network. Time for the algorithm that actually makes it learn — the crown jewel of deep learning. Next up: Module 15 — Backpropagation.