Phase 6 · Language & the TransformerModule 22~38 min read

Word Embeddings

The idea that meaning has geometry: represent every word as a vector so that similar words sit close together — and 'king − man + woman ≈ queen' actually works.

What you'll learn

Here is one of the most beautiful ideas in AI: meaning has geometry. If we represent each word as a vector, we can arrange them so that similar words sit close together and relationships become directions you can do arithmetic with. This is the embedding, and it's the foundation of every language model.

By the end of this module you'll be able to:

  • Explain why one-hot vectors are wasteful
  • Describe a dense embedding
  • Read the vector space of meaning
  • Understand how embeddings are learned from context

One-hot vectors

The naive way to turn a token ID into a vector is one-hot: a vector as long as the vocabulary, all zeros except a single 1. With 50,000 words that's a 50,000-long vector per word — enormous, and worse, every word is equally distant from every other. "cat" and "kitten" look exactly as unrelated as "cat" and "democracy." No meaning at all.

Dense embeddings

An embedding replaces that with a short, dense vector — maybe 300 numbers — that the model learns. Now each dimension can capture some shade of meaning, and similar words get similar vectors. Instead of 50,000 wasted slots, a few hundred rich ones.

The geometry of meaning

Plotted in space, embeddings reveal structure: words cluster by meaning, and — remarkably — relationships become consistent directions. Step through the classic example:

Words as vectors: meaning becomes geometry
Word embeddings & vector arithmetic
kingmanwomanqueencatdogpuppyappleorangemangoembedding dim 1embedding dim 2

similar meanings → nearby vectors

1/3An embedding places every word at a point in space so that words with similar meanings sit close together.
Similar words cluster; the same 'royalty' direction links man→king and woman→queen.
embeddings.py
import numpy as np

# Each word is a learned dense vector (here just 4 numbers; real ones have 100s)
king  = np.array([0.9, 0.7, 0.1, 0.8])
man   = np.array([0.8, 0.2, 0.1, 0.3])
woman = np.array([0.7, 0.2, 0.9, 0.3])

result = king - man + woman
# result is closest to the vector for "queen" -> analogies emerge from geometry

Key idea

king − man + woman ≈ queen. Analogies fall out of the geometry because the model learned to place related words in consistent directions. Meaning became math.

How embeddings are learned

Where do these vectors come from? From a simple, powerful idea: a word is known by the company it keeps. Methods like Word2Vec slide a window over huge amounts of text and train a network to predict a word from its neighbours (or vice-versa). Words that appear in similar contexts get pushed to similar vectors. No labels needed — just text.

Contextual embeddings

Classic embeddings give "bank" one fixed vector, even though it means different things by a river and in finance. Modern models produce contextual embeddings: the vector for a word depends on the whole sentence around it. The mechanism that makes that possible is attention — our next, pivotal module.

Recap & quick check

Key takeaways

  • One-hot vectors are huge and treat every word as equally unrelated — no meaning captured.
  • An embedding is a short, dense, learned vector; similar words get similar vectors.
  • In embedding space, meaning is geometry: words cluster and relationships become directions.
  • Vector arithmetic works: king − man + woman ≈ queen.
  • Embeddings are learned from context (Word2Vec); modern models make them contextual via attention.

Quick check

1. What is the main problem with one-hot word vectors?

2. What does a dense word embedding capture?

3. The result 'king − man + woman ≈ queen' shows that…

4. How are Word2Vec embeddings learned?

Fixed embeddings can't tell which "bank" you mean. The idea that fixed that — and powers every modern model — is next, and it's the big one. Next up: Module 23 — Attention.