Phase 6 · Language & the TransformerModule 21~36 min read

Teaching Machines Language

Before a model can read, text must become numbers. Tokenization, vocabularies, and subwords turn language into a sequence a network can consume.

What you'll learn

Language models work with numbers, but language is made of words. This module covers the crucial first step of every NLP system: tokenization — turning raw text into a sequence of integer tokens a network can actually consume.

By the end of this module you'll be able to:

  • Explain why raw text can't go straight into a network
  • Describe tokenization and subword splitting
  • Define a vocabulary and token IDs
  • Say why bag-of-words loses meaning

Why language is hard

Language is ambiguous, endlessly creative, and context-dependent. "Bank" means different things by a river and in finance; word order flips meaning; new words appear daily. And fundamentally, a neural network only does math on numbers — so before anything else, text must become numbers.

Tokenization

Tokenization chops text into pieces (tokens) and maps each to an integer. The obvious approach — split on spaces — breaks down on rare words, typos, and languages without spaces. Watch text move from a raw string to integer IDs:

From text to token IDs
Tokenizing a sentence
Tokenizing is fun!

raw text — a single string

1/3A model can't read text directly. First we must break it into pieces called tokens.
Raw text → tokens → integer IDs. Everything downstream operates on these numbers.

Subwords & vocabulary

Modern models use subword tokenization (like Byte-Pair Encoding). Frequent words stay whole ("the"), while rare ones break into reusable pieces ("token" + "izing"). This gives a fixed vocabulary of maybe 50,000 tokens that can spell any word — even ones never seen in training — while keeping sequences short.

tokenize.py
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("gpt2")
ids = tok.encode("Tokenizing is fun!")
print(ids)                 # -> [30642, 2890, 318, 1257, 0]
print(tok.convert_ids_to_tokens(ids))
# -> ['Token', 'izing', ' is', ' fun', '!']

Note

This is why LLM usage is billed per token, and why "context length" is measured in tokens. As a rough rule, one token ≈ 4 characters ≈ ¾ of a word in English.

Bag-of-words & its limits

An early way to feed text to models was bag-of-words: count how often each vocabulary word appears and ignore order. It works for simple tasks (spam filtering), but throwing away order throws away meaning — "dog bites man" and "man bites dog" look identical. And each word is an isolated symbol: the model has no idea "cat" and "kitten" are related.

Key idea

We need representations where a word's meaning and its relationships are captured as numbers, and where order is preserved. That is exactly what embeddings (next module) and attention (Module 23) provide.

Recap & quick check

Key takeaways

  • Networks operate on numbers, so text must first be tokenized into integer tokens.
  • Subword tokenization (e.g. BPE) keeps common words whole and splits rare ones into reusable pieces.
  • A fixed vocabulary of tokens can represent any word, even unseen ones; each token maps to an ID.
  • LLM context length and pricing are measured in tokens (≈ ¾ of a word each in English).
  • Bag-of-words discards order and treats words as unrelated symbols — losing meaning.

Quick check

1. Why must text be tokenized before a network sees it?

2. What is the advantage of subword tokenization?

3. What does bag-of-words throw away?

4. Roughly how much text is one token in English?

Tokens are still just isolated IDs. Next we give them meaning — turning each into a vector where geometry encodes semantics. Next up: Module 22 — Word Embeddings.