What you'll learn
Language models work with numbers, but language is made of words. This module covers the crucial first step of every NLP system: tokenization — turning raw text into a sequence of integer tokens a network can actually consume.
By the end of this module you'll be able to:
- Explain why raw text can't go straight into a network
- Describe tokenization and subword splitting
- Define a vocabulary and token IDs
- Say why bag-of-words loses meaning
Why language is hard
Language is ambiguous, endlessly creative, and context-dependent. "Bank" means different things by a river and in finance; word order flips meaning; new words appear daily. And fundamentally, a neural network only does math on numbers — so before anything else, text must become numbers.
Tokenization
Tokenization chops text into pieces (tokens) and maps each to an integer. The obvious approach — split on spaces — breaks down on rare words, typos, and languages without spaces. Watch text move from a raw string to integer IDs:
raw text — a single string
Subwords & vocabulary
Modern models use subword tokenization (like Byte-Pair Encoding). Frequent words stay whole ("the"), while rare ones break into reusable pieces ("token" + "izing"). This gives a fixed vocabulary of maybe 50,000 tokens that can spell any word — even ones never seen in training — while keeping sequences short.
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("gpt2")
ids = tok.encode("Tokenizing is fun!")
print(ids) # -> [30642, 2890, 318, 1257, 0]
print(tok.convert_ids_to_tokens(ids))
# -> ['Token', 'izing', ' is', ' fun', '!']Note
Bag-of-words & its limits
An early way to feed text to models was bag-of-words: count how often each vocabulary word appears and ignore order. It works for simple tasks (spam filtering), but throwing away order throws away meaning — "dog bites man" and "man bites dog" look identical. And each word is an isolated symbol: the model has no idea "cat" and "kitten" are related.
Key idea
Recap & quick check
Key takeaways
- Networks operate on numbers, so text must first be tokenized into integer tokens.
- Subword tokenization (e.g. BPE) keeps common words whole and splits rare ones into reusable pieces.
- A fixed vocabulary of tokens can represent any word, even unseen ones; each token maps to an ID.
- LLM context length and pricing are measured in tokens (≈ ¾ of a word each in English).
- Bag-of-words discards order and treats words as unrelated symbols — losing meaning.
Quick check
1. Why must text be tokenized before a network sees it?
2. What is the advantage of subword tokenization?
3. What does bag-of-words throw away?
4. Roughly how much text is one token in English?
Tokens are still just isolated IDs. Next we give them meaning — turning each into a vector where geometry encodes semantics. Next up: Module 22 — Word Embeddings.