What you'll learn
A large language model is a giant Transformer trained to do one absurdly simple thing: predict the next token. Do that well enough, at enough scale, and you get something that can write essays, debug code, and hold a conversation. This module demystifies exactly how that works.
By the end of this module you'll be able to:
- Explain next-token prediction
- Describe autoregressive generation
- Say what temperature and the context window control
- Understand scaling laws and emergent abilities
Next-token prediction
At heart, an LLM is a function: given a sequence of tokens, output a probability for every possible next token. That's it. All the apparent intelligence — reasoning, translation, coding — is a side effect of getting very, very good at this one prediction. Watch it choose the next word:
prompt so far — what comes next?
Autoregressive generation
To write a whole answer, the model feeds its own output back in: predict a token, append it, then predict the next from the now-longer sequence. This loop — called autoregressive generation — is why LLMs produce text left to right, and why they generate one token at a time (which you can see as the "typing" effect in a chat interface).
Key idea
Temperature & sampling
The model outputs probabilities, but how do we pick? Always taking the single most likely token makes text repetitive. Instead we sample, and temperature controls how adventurous that sampling is: low temperature (≈0) is focused and deterministic; high temperature is more random and creative. It's the model's "creativity dial."
Context windows
An LLM can only "see" a limited number of tokens at once — its context window (from a few thousand to millions of tokens). Everything it uses to answer — your prompt, the conversation, any documents you paste — must fit inside it. Beyond that window, earlier text is simply gone. This is why long chats can "forget" the beginning.
Scale & emergence
The astonishing discovery of the last few years: as you scale up model size, data, and compute, capability improves predictably — the scaling laws. And some abilities (arithmetic, multi-step reasoning, translation) appear rather suddenly past a certain scale — emergent abilities nobody explicitly trained. Scale alone turned a next-word predictor into a general-purpose tool.
Recap & quick check
Key takeaways
- An LLM is a large Transformer trained to predict the next token given the preceding tokens.
- It generates text autoregressively: predict, append, repeat — one token at a time.
- Sampling with a temperature controls the randomness/creativity of the output.
- The context window limits how many tokens the model can attend to at once.
- Scaling laws make capability grow with scale; some abilities emerge suddenly past a threshold.
Quick check
1. What is the core task an LLM is trained to do?
2. 'Autoregressive' generation means…
3. What does a higher temperature do?
4. Why might a long conversation 'forget' its beginning?
Predicting tokens is the skill — but how does a raw model become a helpful assistant? That's the training pipeline. Next up: Module 27 — How LLMs Are Trained.