Home › LLMs & Foundation Models › Reasoning models
🤖 · Models

Reasoning models

Models trained to think through problems step by step before answering.

In one line

Reasoning models spend extra inference-time compute working through a problem before producing a final answer.

ConceptWhat it is

A reasoning model is explicitly trained, often via reinforcement learning on verifiable tasks like math and code, to generate an extended internal chain of reasoning before committing to a final answer, trading additional inference time for higher accuracy on hard, multi-step problems. It exists because simply scaling model size hits diminishing returns on complex reasoning, while letting the model "think longer" at inference time unlocks further gains.

This approach, sometimes called inference-time or test-time compute scaling, means the model can allocate more or less reasoning effort depending on problem difficulty, rather than always answering in one shot.

How it worksThe mechanics

During training, the model is rewarded not just for the final answer but for producing reasoning traces that lead to verifiably correct outcomes, often using reinforcement learning against automatically checkable tasks like math proofs or code execution results. At inference time the model generates a long internal reasoning sequence, sometimes hidden from the end user, before producing a concise final answer, and can be configured to reason longer for harder problems.

At a glanceSee it

Reasoning models diagram
Reasoning models diagram 1

How reinforcement learning on checkable tasks builds the skill the runtime loop uses — sample traces, auto-grade the answers, and reward the reasoning that pays off, repeated over many problems.

Reasoning models diagram 2

The extra compute need not be one fixed serial chain — it can go into a single long trace or many parallel attempts, resolved by majority vote or a reward model, each a different way to buy accuracy with tokens.

When to use itWhere it fits

  • Complex multi-step math, logic, or coding problems where accuracy matters more than speed.
  • Agentic planning tasks that require weighing multiple steps before acting.
  • Scientific or engineering analysis where a wrong intermediate step invalidates the result.
  • Tasks with verifiable correctness, like solving equations or debugging code.

When NOT to use itLimits & anti-patterns

  • Simple factual lookups or casual chat, where extended reasoning adds latency and cost with no benefit.
  • Latency-critical real-time interactions where even a few extra seconds of thinking is unacceptable.
  • High-volume, low-stakes tasks where a cheaper standard chat model is good enough.

Trade-offsAdvantages & costs

Advantages
  • Significantly higher accuracy on hard multi-step reasoning and coding benchmarks.
  • Can allocate more compute to harder problems adaptively.
  • Reduces certain classes of careless errors through self-checking.
  • Increasingly available as a configurable mode rather than a separate product.
Trade-offs & costs
  • Substantially higher latency and token cost than a standard chat response.
  • Reasoning traces can still contain confidently stated errors.
  • Overkill, and wasteful, for simple queries that do not need deep reasoning.
  • Harder to predict cost since reasoning length varies by problem.

ExampleIn the real world

OpenAI's o-series and GPT-5 reasoning modes, along with Anthropic's Claude extended thinking and DeepSeek-R1, use inference-time reasoning to solve competition math and hard coding problems.

ToolsHow to implement it

  • OpenAI o-series / GPT-5 reasoning modeAPI access to configurable reasoning effort.
  • Anthropic Claude extended thinkingexposes an adjustable reasoning budget.
  • DeepSeek-R1open-weight reasoning model trained via reinforcement learning.
  • Ragas / custom evalsused to measure reasoning accuracy gains against cost.

Cost & effortWhat it takes

Meaningfully higher per-query cost and latency due to long reasoning traces; no extra data needed to use via API; best reserved for tasks where accuracy gains justify the added spend.

A living map of modern AI — kept current every morning