Phase 6 · Language & the TransformerModule 24~44 min read

The Transformer Architecture

Assemble attention, positional encoding, and feed-forward layers into the Transformer — the architecture that powers GPT, BERT, and essentially all modern AI.

What you'll learn

Take attention, add a few supporting parts, stack the result many times, and you get the Transformer — the architecture behind GPT, BERT, Claude, and essentially all modern AI. This module assembles the whole thing from pieces you already know.

By the end of this module you'll be able to:

  • Explain what the Transformer replaced and why it won
  • Say why positional encoding is needed
  • Walk through a Transformer block
  • Distinguish encoder, decoder, and encoder-decoder models

Attention is all you need

In 2017 a paper with that title threw out recurrence entirely and built a model from attention alone. Because attention looks at all words simultaneously, Transformers parallelise across a sequence — training vastly faster on GPUs than RNNs ever could. That scalability is what made billion-parameter models, trained on the whole internet, practical.

Positional encoding

Attention has no built-in sense of order — it sees a set of words, not a sequence. But "dog bites man" ≠ "man bites dog." So Transformers add a positional encoding to each token embedding: a signal that marks where the word sits. Now the model knows both what each word is and where it is.

The Transformer block

The Transformer is built from one block, repeated. Each block has two main sub-layers — multi-head self-attention and a feed-forward network — each wrapped in a residual connection and normalization (read it bottom-up):

One Transformer block
output to next block ↑
Add & Normalize
↑
Feed-Forward Network
↑
Add & Normalize
↑
Multi-Head Self-Attention
↑
+ Positional encoding
↑token embeddings in ↑

the two "Add & Normalize" steps are residual connections — the ResNet trick, keeping gradients flowing

Self-attention mixes information between words; the feed-forward layer processes each word; residuals keep gradients healthy.

Key idea

Two ideas do the work: self-attention lets words share information, and the feed-forwardlayer then thinks about each word individually. The residual connections (Module 19's ResNet trick) let you stack dozens of these blocks without gradients vanishing.

Stacking blocks

Stack this block 12, 96, or more times, widen the vectors, and you scale from a small model to a giant one. Early blocks capture local grammar; deeper blocks capture long-range meaning and reasoning-like structure. The architecture barely changes — you mostly just add more of the same, which is exactly why Transformers scale so gracefully.

Encoders & decoders

Three arrangements of Transformer blocks cover most models:

TypeReadsBest forExample
Encoder-onlyThe whole input at onceUnderstanding: classification, searchBERT
Decoder-onlyLeft-to-right, predicting next tokenGenerationGPT, Claude
Encoder-decoderEncode input, then generate outputTranslation, summarizationT5, original Transformer
Today's large language models are mostly decoder-only Transformers.

Recap & quick check

Key takeaways

  • The Transformer replaced recurrence with pure attention, enabling massive parallel training.
  • Positional encoding injects word order, which attention alone lacks.
  • A block = multi-head self-attention + feed-forward, each with residual connections and normalization.
  • Stacking many identical blocks (and widening them) scales the model up gracefully.
  • Encoder-only (BERT) understands; decoder-only (GPT) generates; encoder-decoder (T5) transforms.

Quick check

1. Why do Transformers train so much faster than RNNs?

2. Why is positional encoding needed?

3. A Transformer block's two main sub-layers are…

4. Most modern large language models (GPT, Claude) are…

You now understand the architecture behind ChatGPT. Time to see how it becomes a system that can write, reason, and chat. Next up: Module 25 — What Is Generative AI?