What you'll learn
Because models sample (Module 2), one answer to a hard problem can simply be an unlucky draw. Self-consistency turns that randomness to your advantage: ask the same question several times and take the majority answer. It's one of the simplest ways to buy reliability with tokens.
By the end of this module you'll be able to:
- Explain why a single sample can be wrong even when the model "knows" the answer
- Apply self-consistency: sample multiple times and vote
- Ensemble different prompts or approaches for extra robustness
- Judge when the accuracy gain is worth the extra cost
One sample can be unlucky
On a multi-step problem, a model might reach the right answer most of the time but occasionally take a wrong turn — a dropped digit, a skipped step. Any single response could be that unlucky one. If you only ever look at one answer, you're at the mercy of that one roll of the dice.
Note
Self-consistency: sample & vote
The recipe is simple: run the same prompt several times (with sampling on), collect the answers, and take the one that appears most often. Independent reasoning paths tend to agree on the correct answer and disagree in scattered ways when they're wrong — so the majority is usually right.
Sample 1
$11.50
Sample 2
$10.50
Sample 3
$11.50
Sample 4
$11.50
Sample 5
$12.00
Consensus answer
$11.50 (3 of 5)
Key idea
Ensembling different prompts
A cousin of self-consistency is ensembling: instead of running the same prompt many times, run different prompts (or phrasings, or even models) and combine their answers. This guards against a single prompt's blind spot — if one phrasing biases the model, another may not.
| Technique | What varies | Guards against |
|---|---|---|
| Self-consistency | The random seed (same prompt, re-sampled) | Unlucky single samples |
| Prompt ensembling | The wording / approach of the prompt | A single prompt's blind spot |
| Model ensembling | Which model answers | One model's systematic weakness |
The cost / accuracy tradeoff
Nothing here is free: five samples cost roughly five times the tokens and time of one. That's an easy trade for a high-stakes decision and a poor one for a casual chat. Tune the number of samples to how much a wrong answer actually costs you.
Tip
When voting doesn't help
- Open-ended creative tasks. There's no single "correct" poem to vote on.
- Deterministic settings. At temperature 0, samples are near-identical — nothing to vote between.
- Systematic errors. If the model is consistently wrong the same way, the majority is confidently wrong too.
- Hard-to-compare outputs. Voting needs answers you can group; long free-form text is tough to tally.
Recap & quick check
Key takeaways
- Because models sample, a single answer to a hard problem can be an unlucky wrong draw.
- Self-consistency: run the same prompt several times (sampling on) and take the majority answer.
- Correct reasoning paths tend to agree; wrong ones scatter — so the majority is usually right.
- Ensembling varies the prompt or model instead of just the seed, guarding against blind spots.
- It costs N× the tokens and fails on creative, deterministic, or systematically-wrong tasks — use it selectively.
Quick check
1. What is self-consistency?
2. Why does taking the majority answer tend to work?
3. Why won't self-consistency help at temperature 0?
4. When is self-consistency a POOR choice?
Voting squeezes more reliability from a fixed approach. To go further, the model can interleave reasoning with actions, critique itself, and explore alternatives. Next up: Module 13 — Advanced Reasoning: ReAct, Reflection & Tree-of-Thought.