Phase 6 · Language & the TransformerModule 23~42 min read

Attention

The one idea behind modern AI: let a model decide, for every word, which other words matter most — and weight them accordingly.

What you'll learn

This is the idea behind modern AI. Attention lets a model, for every word, decide which other words matter and focus on them — reading a whole sentence at once instead of shuffling through it in order. It replaced recurrence, unlocked the Transformer, and made today's language models possible.

By the end of this module you'll be able to:

  • Explain why recurrence was a bottleneck
  • Describe what attention computes, intuitively
  • Define queries, keys, and values
  • Understand self-attention and multi-head attention

The bottleneck of recurrence

RNNs read a sentence one word at a time, squeezing everything into a single hidden state. That forces two problems: they can't be parallelised (step 10 needs step 9), and distant words easily lose influence. What if, instead, every word could look directly at every other word — all at once?

The intuition

To understand "it" in "the cat drank the milk because it was thirsty," you look back at "cat." Attention does exactly this: each word sends out a query and pulls information from the words most relevant to it. Step through which words each word attends to:

Self-attention: who looks at whom
Self-attention across a sentence
Thecatchasedthemouse
ThecatchasethemouseThecatchasethemousekeys →

self-attention: every word looks at every word

1/6Self-attention lets each word gather information from all the others. Brighter cells = stronger attention.
Each row is a word's attention over the sentence. The verb 'chased' attends to both its subject and object.

Key idea

Attention is a weighted average: each word's new representation is a blend of all the words, weighted by relevance. The weights are learned, and they depend on the sentence — so meaning becomes contextual.

Queries, keys & values

The mechanism uses three learned vectors per word — a neat library analogy:

  • Query (Q): what this word is looking for.
  • Key (K): what each word offers — matched against the query.
  • Value (V): the actual information a word contributes if attended to.

A word's query is compared (dot product) with every word's key to get attention scores; those scores weight the values, which are summed. High query·key match ⇒ that word's value dominates the blend.

Self-attention

When the queries, keys, and values all come from the same sequence — every word attending to every word in its own sentence — it's called self-attention. This is what builds contextual meaning: the representation of "bank" now depends on whether "river" or "money" is nearby. The scores are scaled and passed through a softmax so each row is a clean probability distribution — scaled dot-product attention.

Multi-head attention

One attention pattern isn't enough — a word relates to others in many ways at once (grammar, subject, topic). So Transformers run several attention operations in parallel, called heads. Each head learns to focus on a different kind of relationship, and their results are combined. More heads, more relationships captured at once.

Recap & quick check

Key takeaways

  • Attention lets each word focus on the other words most relevant to it — processed all at once.
  • It's a learned weighted average: a word's new representation blends all words by relevance.
  • Queries, keys, and values: query·key gives attention scores that weight the values.
  • Self-attention (Q, K, V from the same sequence) builds contextual meaning.
  • Multi-head attention runs several attentions in parallel to capture many relationship types.

Quick check

1. What does attention let a model do?

2. In query/key/value terms, attention scores come from…

3. What is self-attention?

4. Why use multiple attention heads?

Attention is the engine. Now let's assemble it into the full architecture that runs essentially all of modern AI. Next up: Module 24 — The Transformer Architecture.