Phase 1 · FoundationsModule 3~32 min read

Tokens, Context Windows & Cost

Models don't see words — they see tokens. Understand tokenization, the context window that holds your whole conversation, and how both drive latency, limits, and cost.

What you'll learn

Models don't read words — they read tokens. And every model can only hold so many tokens at once, inside a fixed context window. These two facts quietly shape everything: how much you can paste in, how long a conversation can run, how fast a reply comes back, and how much it costs. Once you can "think in tokens," a lot of mysterious behavior makes sense.

By the end of this module you'll be able to:

  • Explain what a token is and estimate how many a piece of text will use
  • Describe the context window and what it must hold
  • Predict what happens when a conversation overflows that window
  • Connect tokens to latency and cost — and prompt within a budget

From text to tokens

Before a model sees your text, a tokenizer chops it into tokens — common words stay whole, while rarer or longer words split into pieces, and punctuation and spaces become tokens too. Each token maps to an integer ID, and those IDs are what the model actually processes. Step through it:

Turning text into tokens
Tokenizing is fun!

raw text — a single string

1/3A model can't read text directly. First we must break it into pieces called tokens.

This is why a model can "spell" oddly or miscount letters: it never saw letters, it saw tokens. It's also why the same idea written more concisely uses fewer tokens — and leaves more room for everything else.

Counting tokens (rules of thumb)

You rarely need an exact count, but a good estimate helps you plan. For typical English text, these approximations are close enough:

Rule of thumbRoughly
1 token~4 characters of English
1 token~¾ of a word
100 tokens~75 words
1 page of prose~500 tokens
A short chat message~10–30 tokens
Approximate — code, other languages, and unusual formatting tokenize differently. When it matters, use a real tokenizer.

Note

Both your prompt and the model's reply cost tokens. When you estimate a request, budget for the input tokens you send and the output tokens you expect back.

The context window

The context window is the maximum number of tokens a model can consider at once — its short-term memory for a single request. Everything has to fit inside it: the system prompt, the entire conversation so far, any documents you paste in, your new message, and the space reserved for the model's answer.

Everything shares one budget
12%
34%
30%
24%
System prompt
Conversation history
Your input / documents
Room left for the reply
The window is finite. Every token spent on history or pasted text is a token not available for the reply.

Window sizes vary widely by model — older ones held only a few thousand tokens, while many modern models hold hundreds of thousands or more. Bigger is helpful, but it's never infinite, and (as we'll see in Phase 5) simply stuffing the window full isn't automatically better.

When you run out of room

What happens when a conversation grows past the window? Something has to give. Depending on the tool, you'll typically see one of these:

  • Truncation. The oldest messages get dropped — which is exactly why a long chat can suddenly "forget" what you said at the start.
  • An error. A raw API request that exceeds the limit is simply rejected, so you must trim it first.
  • Summarization / compaction. Smarter apps replace old turns with a short summary to save room — a technique we'll build on in the context-engineering module.

Watch out

If a model seems to "forget" earlier instructions in a long session, you've probably overflowed the window. Re-state the key context, or start fresh — the information literally fell out of its memory.

Tokens, latency & cost

Tokens aren't just a capacity limit — they're the unit of both time and money. Because the model generates one token at a time, a longer reply literally takes longer to produce (higher latency). And commercial APIs bill per token, usually counting input and output separately. A bloated prompt is slower and pricier on every single call.

LeverEffect of using fewer tokens
Trim the promptLower cost per call and less to process
Cap the output lengthFaster replies and predictable cost
Send only relevant contextMore room for reasoning; less distraction
Reuse a tight system promptSavings multiply across every request
At scale, a prompt that runs a million times a day makes token efficiency a real engineering concern.

The goal isn't to be stingy — it's to spend tokens where they earn their keep. Include the context that improves the answer; cut the filler that doesn't. That habit is the essence of prompting on a budget.

Recap & quick check

Key takeaways

  • Models process tokens, not words — common words stay whole, rare ones split, punctuation and spaces count.
  • Rules of thumb: ~4 characters or ~¾ of a word per token; 100 tokens ≈ 75 words.
  • The context window is the total tokens a request can hold: system prompt + history + input + the reply.
  • Overflow leads to truncation, an error, or summarization — the cause of a model 'forgetting' early context.
  • Tokens are the unit of latency and cost, so trim filler but keep the context that improves the answer.

Quick check

1. What is a token?

2. Roughly how many tokens is a 300-word answer?

3. What must fit inside the context window?

4. A long chat suddenly 'forgets' something you said at the very start. The most likely reason?

You now know what a prompt is made of at the token level. Next we zoom back out and look at the structure of a great prompt — the handful of parts that show up again and again. Next up: Module 4 — The Anatomy of a Prompt.