Phase 3 · Making Models ReasonModule 12~30 min read

Self-Consistency & Ensembling

Ask more than once. Sample several independent answers and take the consensus to squeeze more reliability out of the same model.

What you'll learn

Because models sample (Module 2), one answer to a hard problem can simply be an unlucky draw. Self-consistency turns that randomness to your advantage: ask the same question several times and take the majority answer. It's one of the simplest ways to buy reliability with tokens.

By the end of this module you'll be able to:

  • Explain why a single sample can be wrong even when the model "knows" the answer
  • Apply self-consistency: sample multiple times and vote
  • Ensemble different prompts or approaches for extra robustness
  • Judge when the accuracy gain is worth the extra cost

One sample can be unlucky

On a multi-step problem, a model might reach the right answer most of the time but occasionally take a wrong turn — a dropped digit, a skipped step. Any single response could be that unlucky one. If you only ever look at one answer, you're at the mercy of that one roll of the dice.

Note

This only applies when sampling is on (temperature above 0). At temperature 0 the model is near-deterministic, so re-asking gives essentially the same answer — and self-consistency has nothing to work with. We'll untangle temperature fully in Module 16.

Self-consistency: sample & vote

The recipe is simple: run the same prompt several times (with sampling on), collect the answers, and take the one that appears most often. Independent reasoning paths tend to agree on the correct answer and disagree in scattered ways when they're wrong — so the majority is usually right.

Sample the same prompt five times, then vote

Sample 1

$11.50

Sample 2

$10.50

Sample 3

$11.50

Sample 4

$11.50

Sample 5

$12.00

Consensus answer

$11.50 (3 of 5)

Wrong answers scatter ($10.50, $12.00); correct ones cluster. The majority wins.

Key idea

Self-consistency pairs naturally with chain-of-thought: let each run reason its own way to an answer, then vote on the final answers only. Different correct reasoning paths converge; mistakes don't.

Ensembling different prompts

A cousin of self-consistency is ensembling: instead of running the same prompt many times, run different prompts (or phrasings, or even models) and combine their answers. This guards against a single prompt's blind spot — if one phrasing biases the model, another may not.

TechniqueWhat variesGuards against
Self-consistencyThe random seed (same prompt, re-sampled)Unlucky single samples
Prompt ensemblingThe wording / approach of the promptA single prompt's blind spot
Model ensemblingWhich model answersOne model's systematic weakness
All three combine multiple independent answers into one more-reliable result.

The cost / accuracy tradeoff

Nothing here is free: five samples cost roughly five times the tokens and time of one. That's an easy trade for a high-stakes decision and a poor one for a casual chat. Tune the number of samples to how much a wrong answer actually costs you.

Tip

Use self-consistency selectively — on the few hard, high-value queries where correctness matters most, not on every request. Many teams reserve it for a "hard mode" path triggered only when a first answer looks low-confidence.

When voting doesn't help

  • Open-ended creative tasks. There's no single "correct" poem to vote on.
  • Deterministic settings. At temperature 0, samples are near-identical — nothing to vote between.
  • Systematic errors. If the model is consistently wrong the same way, the majority is confidently wrong too.
  • Hard-to-compare outputs. Voting needs answers you can group; long free-form text is tough to tally.

Recap & quick check

Key takeaways

  • Because models sample, a single answer to a hard problem can be an unlucky wrong draw.
  • Self-consistency: run the same prompt several times (sampling on) and take the majority answer.
  • Correct reasoning paths tend to agree; wrong ones scatter — so the majority is usually right.
  • Ensembling varies the prompt or model instead of just the seed, guarding against blind spots.
  • It costs N× the tokens and fails on creative, deterministic, or systematically-wrong tasks — use it selectively.

Quick check

1. What is self-consistency?

2. Why does taking the majority answer tend to work?

3. Why won't self-consistency help at temperature 0?

4. When is self-consistency a POOR choice?

Voting squeezes more reliability from a fixed approach. To go further, the model can interleave reasoning with actions, critique itself, and explore alternatives. Next up: Module 13 — Advanced Reasoning: ReAct, Reflection & Tree-of-Thought.