Phase 4 · Structured & Reliable OutputModule 16~32 min read

Temperature, Top-p & Sampling

The dials behind creativity and consistency. See how temperature, top-p, and top-k reshape the model's next-token choices — and when to turn them up or down.

What you'll learn

Why does the same prompt sometimes give a boring answer and sometimes a surprising one? The sampling settings — chiefly temperature and top-p — decide how adventurously the model picks each next token. Learn to set them on purpose and you control the dial between consistent and creative.

By the end of this module you'll be able to:

  • Explain where a model's randomness actually comes from
  • Describe what temperature, top-p, and top-k each do
  • Choose settings that match creative vs. deterministic tasks
  • Make outputs more reproducible when you need to

Where randomness comes from

Recall from Module 2 that at each step the model produces a probability for every possible next token. It then samples from that distribution rather than always taking the single most likely token. The sampling settings reshape that distribution before the pick — which is exactly how you tune the balance between safe and surprising.

Temperature

Temperature flattens or sharpens the probability distribution. Low temperature makes the peak steeper, so the model almost always takes the top token — consistent, focused, sometimes repetitive. High temperature flattens the curve, giving unlikely tokens a real chance — varied, creative, sometimes off the rails.

Temperature reshapes the next-token odds

Low temperature (≈0.2)

mat
86%
floor
8%
rug
3%
chair
2%
roof
1%

Peaked → picks the top token almost every time (consistent).

High temperature (≈1.0)

mat
32%
floor
26%
rug
19%
chair
14%
roof
9%

Flatter → the less-likely tokens get a real chance (varied).

Same underlying prediction, two temperatures. Low temp concentrates on 'mat'; high temp spreads the probability so other words can win.

Key idea

Low temperature = predictable. High temperature = diverse. Near 0 the model is close to deterministic (great for extraction and classification); around 0.7-1.0 it's lively (good for brainstorming and copy).

Top-p (nucleus) & top-k

Two more knobs limit which tokens are eligible before sampling:

  • Top-p (nucleus sampling) keeps only the smallest set of top tokens whose probabilities add up to p (say 0.9), then samples among those. It adapts to how confident the model is.
  • Top-k keeps only the k most likely tokens (say the top 40) and ignores the rest.

Tip

You usually don't need to touch all of them. Adjust temperature first — it's the main dial — and leave top-p near its default unless you have a specific reason. Tuning several at once makes behavior hard to reason about.

Creative vs. deterministic tasks

Match the setting to the job:

TaskTemperatureWhy
Data extraction / classification0 – 0.2You want the same, correct answer every time
Code generation0 – 0.3Precision matters; creativity introduces bugs
Q&A over documents0.2 – 0.5Mostly factual, a little natural phrasing
Brainstorming / marketing copy0.7 – 1.0Variety and surprise are the point
Poetry / fiction0.9 – 1.2Maximize novelty and voice
A rough guide, not a law — always test on your own task. When in doubt, start low and raise it if output feels flat.

Seeds & reproducibility

For testing, debugging, or anything you need to reproduce, randomness is a nuisance. Set temperature to 0 for the most deterministic behavior, and use a fixed seed if the API offers one, so the same input yields the same output.

Note

Even at temperature 0, output isn't always perfectly identical across runs — hardware and model-side details can introduce tiny variations. It's "as deterministic as it gets," not a cryptographic guarantee. This is also why self-consistency (Module 12) needs a temperature above 0 to work.

Recap & quick check

Key takeaways

  • At each step the model samples from a probability distribution; sampling settings reshape it before the pick.
  • Temperature sharpens (low → consistent) or flattens (high → creative/varied) the distribution.
  • Top-p keeps the smallest set of tokens summing to p; top-k keeps the k most likely — both limit the candidates.
  • Match the dial to the task: near 0 for extraction/code, 0.7-1.0 for brainstorming and creative writing.
  • For reproducibility use temperature 0 (and a seed if available) — but it's near-deterministic, not guaranteed identical.

Quick check

1. What does raising the temperature do?

2. You're extracting fields into JSON and want the same result every time. Best temperature?

3. What does top-p (nucleus) sampling do?

4. Which is the practical advice on tuning these knobs?

You can now shape output's structure, behavior, and variability. The final Phase 4 skill is making a prompt dependable across thousands of messy, real-world inputs. Next up: Module 17 — Consistency & Reliability at Scale.