What you'll learn
You don't need to build a language model to prompt one well — but a little of the machinery goes a long way. Almost every quirk you'll ever meet traces back to one simple fact: a large language model (LLM) generates text by repeatedly guessing the next token. Understand that, and the model stops feeling like magic and starts feeling like something you can steer.
By the end of this module you'll be able to:
- Explain next-token prediction in one sentence
- Say exactly what information the model does — and doesn't — have
- Understand why models sound confident even when they're wrong
- Know why the same prompt can give different answers, and what to do about it
Next-token prediction, in one picture
At its core, an LLM does one thing astonishingly well: given a stretch of text, it predicts what token comes next. A token is a chunk of text — often a word or part of a word (much more on this next module). The model assigns a probability to every possible next token, picks one, appends it, and then repeats the whole process with the slightly longer text. Press play and step through it:
prompt so far — what comes next?
That loop — predict a token, append it, feed everything back in, predict again — is called autoregressive generation. Every essay, answer, and line of code an LLM produces is built one token at a time this way. There is no plan drawn up in advance; the sentence takes shape as it goes.
Key idea
The model reads only your prompt
Because generation is just "continue this text," the text you provide is the model's entire universe for that request. It has no access to what you meant, what you did yesterday, or what's on your screen — only the tokens in front of it. This is the most important practical consequence of how LLMs work.
What the model works from
- • The exact text in your prompt
- • Patterns learned during training
- • Anything you paste in as context
- • Tools you explicitly give it (later)
What it does NOT have
- • Your intentions unless you write them
- • Live access to the internet by default
- • Events after its training cutoff
- • Memory of past chats (unless provided)
Prompt
Is it a good time to send it?
AI response
Knowledge, cutoffs & confidence
During training, the model reads an enormous amount of text and absorbs patterns from it — facts, styles, reasoning steps, code idioms. That knowledge is real, but it's frozen at a training cutoff: the model doesn't know about events after the data it was trained on, and by default it isn't browsing the web while it answers you.
There's a subtler issue too. Because the model's job is to produce likely-sounding text, it will happily generate a fluent, confident answer even when it doesn't actually know — inventing a citation, a date, or an API that doesn't exist. This is called hallucination, and confidence in the wording is not evidence of correctness. We dedicate a whole module to fighting it (Phase 5).
Watch out
Why the same prompt can vary
Ask the same question twice and you may get two different answers. That's not a bug — it's sampling. When the model picks the next token, it usually doesn't always take the single most likely option; it samples from the probability distribution, which adds variety and creativity. A setting called temperature controls how adventurous that choice is. We'll turn these dials deliberately in Phase 4; for now, just know that some randomness is expected by design.
| You'll notice… | Because… |
|---|---|
| The model continues text rather than “looking up” an answer | It is predicting likely next tokens, not querying a database |
| It sounds confident even when wrong | It optimizes for plausible-sounding text, not verified truth |
| It doesn't know recent events | Its knowledge is frozen at a training cutoff |
| The same prompt gives different answers | It samples from a probability distribution (temperature) |
| Vague prompts get generic replies | With little to go on, the most likely continuation is average |
What this means for how you prompt
Everything in this course follows from the mechanics above. Three rules to carry forward:
- Put it in the prompt. If the model needs a fact, a constraint, or context to answer well, state it — it can't infer what isn't there.
- Set up the right continuation. Give the model a beginning that naturally leads where you want: a role, an example, a format. You're steering a text-continuation engine.
- Don't trust; verify. Treat confident claims as drafts. Ground them in provided facts or ask for sources when it matters.
Recap & quick check
Key takeaways
- An LLM generates text by repeatedly predicting the next token — autoregressive generation.
- Your prompt is the starting text it continues; it's the model's entire world for that request.
- The model's knowledge is frozen at a training cutoff and it doesn't browse by default.
- Models produce likely-sounding text, so they can be confidently wrong (hallucination) — fluency ≠ truth.
- Sampling (controlled by temperature) is why the same prompt can give different answers.
Quick check
1. In one sentence, what does a large language model fundamentally do?
2. Why did the model struggle with 'Is it a good time to send it?'
3. A model gives a confident, fluent answer. What can you conclude?
4. Why might the exact same prompt produce different answers?
We keep mentioning "tokens" — so let's look at what they actually are, and the invisible budget they impose on every prompt. Next up: Module 3 — Tokens, Context Windows & Cost.