What you'll learn
This is the idea behind modern AI. Attention lets a model, for every word, decide which other words matter and focus on them — reading a whole sentence at once instead of shuffling through it in order. It replaced recurrence, unlocked the Transformer, and made today's language models possible.
By the end of this module you'll be able to:
- Explain why recurrence was a bottleneck
- Describe what attention computes, intuitively
- Define queries, keys, and values
- Understand self-attention and multi-head attention
The bottleneck of recurrence
RNNs read a sentence one word at a time, squeezing everything into a single hidden state. That forces two problems: they can't be parallelised (step 10 needs step 9), and distant words easily lose influence. What if, instead, every word could look directly at every other word — all at once?
The intuition
To understand "it" in "the cat drank the milk because it was thirsty," you look back at "cat." Attention does exactly this: each word sends out a query and pulls information from the words most relevant to it. Step through which words each word attends to:
self-attention: every word looks at every word
Key idea
Queries, keys & values
The mechanism uses three learned vectors per word — a neat library analogy:
- Query (Q): what this word is looking for.
- Key (K): what each word offers — matched against the query.
- Value (V): the actual information a word contributes if attended to.
A word's query is compared (dot product) with every word's key to get attention scores; those scores weight the values, which are summed. High query·key match ⇒ that word's value dominates the blend.
Self-attention
When the queries, keys, and values all come from the same sequence — every word attending to every word in its own sentence — it's called self-attention. This is what builds contextual meaning: the representation of "bank" now depends on whether "river" or "money" is nearby. The scores are scaled and passed through a softmax so each row is a clean probability distribution — scaled dot-product attention.
Multi-head attention
One attention pattern isn't enough — a word relates to others in many ways at once (grammar, subject, topic). So Transformers run several attention operations in parallel, called heads. Each head learns to focus on a different kind of relationship, and their results are combined. More heads, more relationships captured at once.
Recap & quick check
Key takeaways
- Attention lets each word focus on the other words most relevant to it — processed all at once.
- It's a learned weighted average: a word's new representation blends all words by relevance.
- Queries, keys, and values: query·key gives attention scores that weight the values.
- Self-attention (Q, K, V from the same sequence) builds contextual meaning.
- Multi-head attention runs several attentions in parallel to capture many relationship types.
Quick check
1. What does attention let a model do?
2. In query/key/value terms, attention scores come from…
3. What is self-attention?
4. Why use multiple attention heads?
Attention is the engine. Now let's assemble it into the full architecture that runs essentially all of modern AI. Next up: Module 24 — The Transformer Architecture.