Phase 5 · Deep LearningModule 20~38 min read

Recurrent Networks & Sequences

Text, audio, and time series arrive as sequences. Recurrent networks carry a memory from step to step — and LSTMs fix their short attention span.

What you'll learn

Text, speech, and time series arrive as sequences, where order carries meaning and length varies. Recurrent neural networks (RNNs) handle them by reading one step at a time and carrying a memory forward. They ruled language AI for years — and understanding them is the perfect run-up to attention.

By the end of this module you'll be able to:

  • Explain what makes sequential data special
  • Describe the hidden state that gives an RNN memory
  • Understand backpropagation through time
  • Say why LSTMs were invented

Sequential data

In a sentence, order changes everything: "dog bites man" ≠ "man bites dog." A plain network with fixed inputs can't handle variable-length, order-dependent data well. We need a model that processes items one at a time while remembering what came before.

The recurrent neuron & hidden state

An RNN reads the sequence step by step. At each step it combines the current input with its hidden state — a running summary of everything seen so far — to produce a new hidden state. That hidden state is the network's memory. Watch it read a sentence, carrying context forward:

An RNN reading a sentence, one word at a time
Unrolling an RNN through time
The
cat
sat
on
the
mat
h0→
h1→
h2→
h3→
h4→
h5

an RNN reads one word at a time, carrying a hidden 'memory' state

1/8A recurrent network processes a sequence one step at a time, left to right.
Each step folds the new word into the hidden state (h), which flows to the next step.

Key idea

The key is the loop: the same little network is applied at every step, feeding its hidden state back into itself. That recurrence is what lets it handle sequences of any length.

Learning through time

To train an RNN we "unroll" the loop into one long chain — one copy per time step — and run ordinary backpropagation across it. This is backpropagation through time (BPTT). Because the errors have to travel back across many steps, it inherits the vanishing-gradient problem in an especially painful form.

The long-term memory problem

Plain RNNs are forgetful. By the time they reach the end of a long paragraph, the influence of the first words has faded to almost nothing — the gradients vanished on the way back. So an RNN struggles to connect "The clouds are in the …" to a word mentioned twenty sentences earlier.

LSTMs & GRUs

The fix was the Long Short-Term Memory (LSTM) unit, and its lighter cousin the GRU. They add gates — small learned controls that decide what to keep, what to forget, and what to output — plus a protected cell state that carries information across many steps without vanishing. LSTMs powered translation, speech recognition, and text generation for years.

Note

But RNNs, even LSTMs, have two stubborn limits: they process steps strictly in order (hard to parallelise) and still struggle with very long-range links. In 2017 a new architecture removed recurrence entirely — the Transformer, which is where Phase 6 begins.

Recap & quick check

Key takeaways

  • Sequential data (text, speech, time series) is order-dependent and variable-length.
  • An RNN processes one step at a time, carrying a hidden state that acts as memory.
  • Training unrolls the loop and runs backpropagation through time (BPTT).
  • Plain RNNs forget long-range context because gradients vanish across many steps.
  • LSTMs and GRUs add gates and a protected cell state to remember longer — but still process in order.

Quick check

1. What gives an RNN its 'memory'?

2. Training an RNN across its unrolled steps is called…

3. Why do plain RNNs struggle with long-range dependencies?

4. What did LSTMs add to fix this?

RNNs read strictly left to right and forget. What if a model could look at every word at once and decide what matters? That idea changed AI forever. Next up: Module 21 — Teaching Machines Language.