What you'll learn
Take attention, add a few supporting parts, stack the result many times, and you get the Transformer — the architecture behind GPT, BERT, Claude, and essentially all modern AI. This module assembles the whole thing from pieces you already know.
By the end of this module you'll be able to:
- Explain what the Transformer replaced and why it won
- Say why positional encoding is needed
- Walk through a Transformer block
- Distinguish encoder, decoder, and encoder-decoder models
Attention is all you need
In 2017 a paper with that title threw out recurrence entirely and built a model from attention alone. Because attention looks at all words simultaneously, Transformers parallelise across a sequence — training vastly faster on GPUs than RNNs ever could. That scalability is what made billion-parameter models, trained on the whole internet, practical.
Positional encoding
Attention has no built-in sense of order — it sees a set of words, not a sequence. But "dog bites man" ≠ "man bites dog." So Transformers add a positional encoding to each token embedding: a signal that marks where the word sits. Now the model knows both what each word is and where it is.
The Transformer block
The Transformer is built from one block, repeated. Each block has two main sub-layers — multi-head self-attention and a feed-forward network — each wrapped in a residual connection and normalization (read it bottom-up):
the two "Add & Normalize" steps are residual connections — the ResNet trick, keeping gradients flowing
Key idea
Stacking blocks
Stack this block 12, 96, or more times, widen the vectors, and you scale from a small model to a giant one. Early blocks capture local grammar; deeper blocks capture long-range meaning and reasoning-like structure. The architecture barely changes — you mostly just add more of the same, which is exactly why Transformers scale so gracefully.
Encoders & decoders
Three arrangements of Transformer blocks cover most models:
| Type | Reads | Best for | Example |
|---|---|---|---|
| Encoder-only | The whole input at once | Understanding: classification, search | BERT |
| Decoder-only | Left-to-right, predicting next token | Generation | GPT, Claude |
| Encoder-decoder | Encode input, then generate output | Translation, summarization | T5, original Transformer |
Recap & quick check
Key takeaways
- The Transformer replaced recurrence with pure attention, enabling massive parallel training.
- Positional encoding injects word order, which attention alone lacks.
- A block = multi-head self-attention + feed-forward, each with residual connections and normalization.
- Stacking many identical blocks (and widening them) scales the model up gracefully.
- Encoder-only (BERT) understands; decoder-only (GPT) generates; encoder-decoder (T5) transforms.
Quick check
1. Why do Transformers train so much faster than RNNs?
2. Why is positional encoding needed?
3. A Transformer block's two main sub-layers are…
4. Most modern large language models (GPT, Claude) are…
You now understand the architecture behind ChatGPT. Time to see how it becomes a system that can write, reason, and chat. Next up: Module 25 — What Is Generative AI?