Home › Prompt Engineering › Self-consistency
✍️ · Ground

Self-consistency

Sample several chain-of-thought answers to the same question, then let the majority vote decide.

In one line

Instead of trusting one chain of thought, self-consistency samples many and returns the answer they most agree on.

ConceptWhat it is

Self-consistency is a decoding strategy that sits on top of chain-of-thought (CoT) prompting. Introduced by Wang and colleagues in 2022, it replaces the usual single greedy answer with several independently sampled reasoning paths, then returns the final answer that appears most often. The core intuition is marginalizing over reasoning paths: a correct answer can usually be reached by many different valid lines of thought, while wrong answers tend to scatter, so a majority vote surfaces the answer the model is most consistently drawn to.

It exists because a single CoT trace is fragile. Greedy decoding commits to one path, and if the model takes a wrong turn early, nothing corrects it. Sampling multiple paths at a nonzero temperature lets the model explore alternatives, and the vote washes out one-off mistakes without any training, fine-tuning, or change to the model itself.

How it worksThe mechanics

Take a chain-of-thought prompt and, instead of decoding once, sample N completions at a nonzero temperature so each run reasons a little differently. Parse the final answer out of every completion, normalize them into comparable buckets, and count how many paths landed on each answer. Return the answer with the most votes as the result; the size of that majority also serves as a rough confidence signal, and an even split flags the query as uncertain.

At a glanceSee it

Self-consistency diagram
Self-consistency diagram 1

Why sampling beats a single greedy chain — idiosyncratic slips scatter across different values while the correct answer recurs across independent paths, so marginalizing lets the plurality win.

Self-consistency diagram 2

A go or no-go test — self-consistency needs an extractable answer and errors that differ across paths, and it trades N-fold compute for accuracy or else silently reinforces a shared bias.

When to use itWhere it fits

  • Hard multi-step reasoning such as math word problems, logic puzzles, or structured extraction where a single pass is unreliable.
  • High-stakes decisions where a few points of extra accuracy clearly justify extra compute.
  • Tasks with a checkable, canonical final answer that votes can actually be tallied over.
  • As a cheap, prompt-only accuracy lever to try before reaching for fine-tuning.

When NOT to use itLimits & anti-patterns

  • Open-ended generative work like essays or brainstorming, where there is no single correct answer to vote on.
  • Latency- or cost-sensitive paths such as real-time chat or high-QPS endpoints.
  • Simple tasks that a single greedy pass already answers correctly.
  • Long-form outputs that are hard to normalize into comparable buckets for a clean majority.

Trade-offsAdvantages & costs

Advantages
  • Delivers real accuracy gains on reasoning benchmarks with no training and no model change.
  • Model-agnostic and prompt-only, so it wraps any existing chain-of-thought prompt.
  • The spread of the vote doubles as a built-in confidence signal.
  • Samples are independent, so they parallelize cleanly and hide much of the latency cost.
Trade-offs & costs
  • Cost and token spend scale with N, so it is roughly N times more expensive than a single call.
  • Only works when answers can be normalized and compared to find a majority.
  • Diminishing returns set in quickly; past a point, more samples barely move accuracy.
  • Voting fixes variance, not bias, so a model that is confidently wrong will simply vote for the same wrong answer every time.

ExampleIn the real world

A team runs a model to solve grade-school-style math word problems in a batch pipeline. With one greedy chain of thought, the model occasionally forgets a unit conversion and lands on the wrong number. They switch on self-consistency: for each problem they draw five reasoning paths at a moderate temperature, extract the final numeric answer from each, and keep the most common value. On a typical hard item, four paths agree on the correct figure while one drifts, and the majority vote discards the outlier. Across the batch, accuracy on the toughest multi-step problems climbs, and any problem where the five answers split evenly gets routed to a human for review.

ToolsHow to implement it

  • Model APIs that let you sample completions at a nonzero temperature: the Anthropic Messages API exposes a temperature parameter and you draw the samples with repeated calls, while the OpenAI Chat Completions API can return several samples in one call via its n parameter.
  • DSPy, whose majority utility function performs a majority vote over sampled completions.
  • LangChain or LlamaIndex to orchestrate the parallel samples and aggregate their final answers.
  • A light answer-normalization step, from a regex to a small parser, to bucket answers into comparable groups before counting votes.

Cost & effortWhat it takes

Self-consistency is among the most expensive prompting techniques per query, because you pay for N full reasoning traces instead of one, meaning roughly N times the tokens, calls, and dollar cost. The samples run independently, so if you fan them out in parallel the wall-clock latency need not grow N-fold. There is no training or infrastructure to build; the entire cost lives at inference time. In practice teams tune N, often a handful somewhere in the range of five to a few dozen, up to the point where accuracy plateaus, and reserve the technique for the hard, high-value slice of traffic rather than every request.

A living map of modern AI — kept current every morning