What you'll learn
Models don't read words — they read tokens. And every model can only hold so many tokens at once, inside a fixed context window. These two facts quietly shape everything: how much you can paste in, how long a conversation can run, how fast a reply comes back, and how much it costs. Once you can "think in tokens," a lot of mysterious behavior makes sense.
By the end of this module you'll be able to:
- Explain what a token is and estimate how many a piece of text will use
- Describe the context window and what it must hold
- Predict what happens when a conversation overflows that window
- Connect tokens to latency and cost — and prompt within a budget
From text to tokens
Before a model sees your text, a tokenizer chops it into tokens — common words stay whole, while rarer or longer words split into pieces, and punctuation and spaces become tokens too. Each token maps to an integer ID, and those IDs are what the model actually processes. Step through it:
raw text — a single string
This is why a model can "spell" oddly or miscount letters: it never saw letters, it saw tokens. It's also why the same idea written more concisely uses fewer tokens — and leaves more room for everything else.
Counting tokens (rules of thumb)
You rarely need an exact count, but a good estimate helps you plan. For typical English text, these approximations are close enough:
| Rule of thumb | Roughly |
|---|---|
| 1 token | ~4 characters of English |
| 1 token | ~¾ of a word |
| 100 tokens | ~75 words |
| 1 page of prose | ~500 tokens |
| A short chat message | ~10–30 tokens |
Note
The context window
The context window is the maximum number of tokens a model can consider at once — its short-term memory for a single request. Everything has to fit inside it: the system prompt, the entire conversation so far, any documents you paste in, your new message, and the space reserved for the model's answer.
Window sizes vary widely by model — older ones held only a few thousand tokens, while many modern models hold hundreds of thousands or more. Bigger is helpful, but it's never infinite, and (as we'll see in Phase 5) simply stuffing the window full isn't automatically better.
When you run out of room
What happens when a conversation grows past the window? Something has to give. Depending on the tool, you'll typically see one of these:
- Truncation. The oldest messages get dropped — which is exactly why a long chat can suddenly "forget" what you said at the start.
- An error. A raw API request that exceeds the limit is simply rejected, so you must trim it first.
- Summarization / compaction. Smarter apps replace old turns with a short summary to save room — a technique we'll build on in the context-engineering module.
Watch out
Tokens, latency & cost
Tokens aren't just a capacity limit — they're the unit of both time and money. Because the model generates one token at a time, a longer reply literally takes longer to produce (higher latency). And commercial APIs bill per token, usually counting input and output separately. A bloated prompt is slower and pricier on every single call.
| Lever | Effect of using fewer tokens |
|---|---|
| Trim the prompt | Lower cost per call and less to process |
| Cap the output length | Faster replies and predictable cost |
| Send only relevant context | More room for reasoning; less distraction |
| Reuse a tight system prompt | Savings multiply across every request |
The goal isn't to be stingy — it's to spend tokens where they earn their keep. Include the context that improves the answer; cut the filler that doesn't. That habit is the essence of prompting on a budget.
Recap & quick check
Key takeaways
- Models process tokens, not words — common words stay whole, rare ones split, punctuation and spaces count.
- Rules of thumb: ~4 characters or ~¾ of a word per token; 100 tokens ≈ 75 words.
- The context window is the total tokens a request can hold: system prompt + history + input + the reply.
- Overflow leads to truncation, an error, or summarization — the cause of a model 'forgetting' early context.
- Tokens are the unit of latency and cost, so trim filler but keep the context that improves the answer.
Quick check
1. What is a token?
2. Roughly how many tokens is a 300-word answer?
3. What must fit inside the context window?
4. A long chat suddenly 'forgets' something you said at the very start. The most likely reason?
You now know what a prompt is made of at the token level. Next we zoom back out and look at the structure of a great prompt — the handful of parts that show up again and again. Next up: Module 4 — The Anatomy of a Prompt.