A reasoning model is a language model that spends extra computation at inference time to work through problems step by step before producing a final answer. The category took shape when OpenAI o1 began trading more thinking time for higher accuracy on math, code, and logic benchmarks, and it now spans OpenAI's o-series, DeepSeek-R1, Anthropic's Claude with extended thinking, and Google DeepMind's Gemini 2.5 Pro Deep Think. The mechanism that ties them together is test-time compute (TTC): instead of generating a response in one forward pass, the model produces a stream of reasoning tokens, evaluates intermediate steps, and only then commits to an output. The result is a different cost curve, a different latency profile, and a different set of tasks where the extra compute is worth paying for.
What a Reasoning Model Actually Does
A reasoning model differs from a standard language model in one critical way: it generates a chain of intermediate steps before committing to a final answer, spending additional compute at inference time rather than producing an output in a single forward pass. IBM defines the category directly, calling a reasoning model an LLM fine-tuned to break complex problems into smaller steps, often called reasoning traces, prior to generating a final output (IBM, What Is a Reasoning Model?). The internal trace is the unit of work. Where a conventional language model maps an input to the next token immediately, OpenAI o1 first generates a private Chain-of-Thought, evaluates alternatives, and then writes the user-facing response. The mechanics of the underlying transformer do not change. What changes is the inference loop on top: more tokens are spent thinking, fewer are spent answering, and the model decides how to allocate that budget per query. For a refresher on the substrate these systems sit on, see how large language models are architected and trained, and for the wider system family, how modern generative AI systems work.
- Reasoning model: a language model trained to generate intermediate reasoning steps before its final response, rather than answering in a single pass.
- Reasoning trace: the sequence of intermediate steps the model produces while working through a problem, sometimes hidden from the user.
- Test-time compute (TTC): the additional computation spent at inference time, distinct from the compute used during training.
- Chain of thought (CoT): a step by step natural-language reasoning sequence generated before the final answer.
How Test-Time Compute Works
The core mechanism of a reasoning model is straightforward: rather than routing an input directly to a final answer, the model allocates a compute budget at inference time to generate reasoning tokens that explore the problem before the response begins. OpenAI frames this as a direct lever, reporting that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute), per OpenAI (OpenAI, Learning to Reason with LLMs). The thinking tokens are real model output. They are sampled the same way as any other tokens, count against the same context window, and consume the same per-token compute. The difference is where they live in the response: in the hidden reasoning trace rather than in the user-facing answer. OpenAI o1 can therefore burn thousands of thinking tokens on a single hard problem and emit a one-line answer, or skip the long internal trace entirely on a query it can dispatch in one pass.
The compute budget is adaptive rather than fixed. OpenAI describes its o-series as systems that can adapt their computation during inference, citing experiments on o1-preview and o1-mini that measured attack success as a function of the compute spent at inference (OpenAI, Trading Inference-Time Compute for Adversarial Robustness). RAND describes the same family of systems, including OpenAI's o1 and o3 and DeepSeek-R1, as test-time compute models that generate intermediate calculations and evaluate multiple approaches before committing (RAND, When AI Takes Time to Think). That adaptive allocation is what makes inference scaling a practical knob for production deployments, not just a research curiosity.
- Receive the prompt: the reasoning model ingests the user input into its context window like any standard language model.
- Allocate a thinking budget: the model determines how much test-time compute to spend, based on the apparent difficulty of the task.
- Generate reasoning tokens: a hidden Chain-of-Thought is produced, exploring intermediate steps and alternative approaches.
- Self-check and revise: the model evaluates the trace, discards weak branches, and refines its working answer.
- Emit the final answer: only the response is returned to the user; the reasoning tokens remain internal in most deployments.
How Reasoning Models Are Trained
Training a reasoning model requires more than standard next-token prediction: the model is taught to value step-by-step reasoning through reinforcement learning, where a reward model scores the intermediate chain of thought rather than only the final answer, an approach central to recent open reasoning models such as DeepSeek-R1. OpenAI says o1 uses a Chain-of-Thought when attempting to solve a problem and that its performance consistently improves with more reinforcement learning (RL) at train time and more time spent thinking at test time, per OpenAI. The training signal is the binding constraint. A reward model that grades only the final answer teaches the model to guess; a reward model that grades the intermediate steps teaches it to reason. That distinction is why OpenAI o1 is not simply a standard LLM with a longer prompt template; the reinforcement-learning loop, sometimes augmented with PPO and Monte Carlo Tree Search, reshapes which token sequences the model is willing to produce when it has time to think.
The two compute budgets are linked but separate. Train-time compute pays for the reinforcement learning that teaches the model to produce useful reasoning traces; test-time compute pays for the actual thinking DeepSeek-R1 or Claude with extended thinking does on each query. OpenAI distinguishes the two explicitly, which matters because it means inference scaling is not a free substitute for a better training run. A model with a weak reward signal will not magically reason better if given more thinking tokens; a model trained with a strong RL loop will. For the broader picture of how reinforcement learning from human feedback shapes model behavior, see how reinforcement learning from human feedback shapes model behavior.
- Pretraining base: the model starts as a standard pretrained language model, fluent in token prediction but not yet a step-by-step thinker.
- Reasoning-trace fine-tuning: supervised data containing worked solutions teaches the model the surface form of a Chain-of-Thought.
- Reinforcement learning with a reward model: the reward model scores intermediate steps and correct final answers, reshaping which traces the model produces, often via PPO.
- Inference-time policy: the deployed model decides per query how many thinking tokens to spend before committing to an answer.
Reasoning Model Implementations: OpenAI, Google, and Anthropic
The reasoning model category today centers on three distinct implementations: OpenAI o1 and its successors, Google DeepMind's Deep Think mode inside Gemini 2.5 Pro, and Anthropic's extended thinking capability in Claude, each taking a somewhat different approach to allocating inference-time compute. OpenAI o1 established the pattern, pairing Chain-of-Thought generation with reinforcement learning so that the model improves with both more RL and more test-time compute, per OpenAI. Google DeepMind frames Deep Think as an experimental enhanced reasoning mode for highly complex math and coding that uses new research techniques enabling Gemini 2.5 Pro to consider multiple hypotheses before responding (Google DeepMind, Gemini 2.5). DeepMind reports that Gemini 2.5 Pro Deep Think scores 84.0% on MMMU, a multimodal reasoning benchmark, and pairs that with a 1 million-token context window. Anthropic publishes interpretability research showing that Claude plans many words ahead and sometimes thinks in a conceptual space shared between languages, suggesting internal planning behavior the company has begun to trace directly (Anthropic, Tracing the Thoughts of a Language Model).
DeepSeek-R1 sits alongside these vendors as a notable open-weight system, grouped by RAND with OpenAI's o1 and o3 in the test-time compute family. Each implementation exposes a different thinking budget control surface: some surface it as a mode toggle, others as an automatic per-query decision, others as a configurable maximum. The table below maps the vendor-documented characteristics side by side.
| Implementation | Vendor positioning | Thinking budget control | Source |
|---|---|---|---|
| OpenAI o1 (and o1-mini) | Chain-of-Thought reasoning that improves with more RL and more test-time compute | Adaptive per query; model decides how much to think | OpenAI primary docs |
| Gemini 2.5 Pro Deep Think | Experimental enhanced reasoning mode for complex math and coding | Mode toggle on top of the base Gemini 2.5 Pro model | Google DeepMind blog |
| Anthropic Claude (extended thinking) | Planning many words ahead; internal conceptual space shared across languages | Capability surfaced as extended thinking, not a separate product line | Anthropic research |
| DeepSeek-R1 | Open-weight reasoning model in the o1/o3 family per RAND | Per-query adaptive thinking budget | RAND commentary |
Where Reasoning Models Outperform Standard LLMs
A reasoning model does not uniformly beat a standard language model on every task; the advantage is concentrated in problem domains that benefit from deliberate step-by-step decomposition. Competition math, formal logic, multi-step coding, and structured planning are the clearest wins, and they are exactly the workloads vendors lead with when announcing new reasoning capabilities. DeepMind reports Gemini 2.5 Pro Deep Think scoring 84.0% on MMMU and topping LiveCodeBench, a difficult benchmark for competition-level coding, per Google DeepMind. OpenAI's robustness work shows a separate axis of improvement: increasing inference-time compute, giving OpenAI o1 more time and resources to think, can improve robustness to multiple types of adversarial attacks, per OpenAI. That second finding matters because it suggests the value of test-time compute is not limited to leaderboard accuracy on MMLU, GPQA, AIME, GSM8K, or ARC-AGI; it also raises the floor against prompt-injection and jailbreak attempts.
The picture is more honest with the limits stated alongside the wins. These systems still hallucinate, still inherit the knowledge cutoff of the base language model, and still cannot retrieve external facts without tooling. Benchmark scores on MMLU, GPQA, and AIME are benchmark-specific, not proof of general intelligence. For how those scores are constructed and what they actually measure, see how AI model benchmarks are designed and what they measure.
- Competition math and proofs: tasks where the answer follows from a chain of deductions reward step by step decomposition.
- Multi-step coding: debugging and refactoring across multiple files benefit from the model's ability to plan before writing, evaluated on LiveCodeBench.
- Adversarial robustness: per OpenAI, more inference-time compute can improve robustness to several attack classes.
- Structured planning: tasks requiring the model to enumerate, evaluate, and prune options align with the inference scaling pattern, sometimes assisted by Monte Carlo Tree Search.
- Limits: hallucination, knowledge cutoff, and retrieval gaps persist; reasoning is not retrieval.
Costs and Trade-offs of Inference-Time Scaling
Scaling a reasoning model's compute budget at inference time improves accuracy on hard tasks but introduces trade-offs that matter for production deployments, including higher latency, increased token costs, and unpredictable thinking budgets for straightforward queries. Every reasoning token consumed in the hidden trace is billed and timed like any other token. A query that would have completed in 200 milliseconds with a standard language model can stretch to many seconds when DeepSeek-R1 or Claude with extended thinking decides to think hard. That latency is not a bug; it is the mechanism. OpenAI's robustness paper measured attack success as a function of the compute spent at inference, which is the same lever a developer pays for in production, per OpenAI.
Three knock-on effects are worth planning for. First, cost variance: two superficially similar prompts can land on very different thinking budgets, which makes per-query cost estimation harder than with a fixed standard LLM. Second, observability: the reasoning trace is often hidden, so a developer cannot always inspect the path the model took to its answer, which complicates debugging and auditing. Third, fit: simple classification, extraction, or single-step lookups do not benefit from inference scaling and may degrade in user experience because of the added latency. The right deployment pattern is usually a hybrid, where a standard LLM handles the high-volume simple traffic and Gemini 2.5 Pro Deep Think or OpenAI o3 is routed only to queries that clear a difficulty threshold.
- Latency: additional reasoning tokens push response times from milliseconds to seconds on hard queries.
- Token cost: hidden thinking tokens are still billed, raising the per-query cost relative to a standard LLM.
- Cost variance: adaptive thinking budgets make per-query cost harder to predict in production.
- Trace opacity: hidden reasoning traces complicate debugging, evaluation, and compliance review.
- Wrong-fit overhead: simple lookups and extractions gain nothing from inference scaling and pay the latency tax for it.
When to Use a Reasoning Model vs a Standard LLM
Choosing a reasoning model over a standard language model is a trade-off decision: the additional inference-time compute is worth the latency and cost when tasks require multi-step logic, but adds unnecessary overhead for high-volume simple queries where a standard LLM responds correctly in one pass. The decision rule that holds up in practice is task shape. If the correct answer depends on a chain of intermediate steps the model needs to work through, OpenAI o3 pays for itself in fewer wrong answers and fewer retries. If the correct answer is a single hop from the input, a standard language model is faster, cheaper, and just as accurate. The category boundary is not strict; many production systems route per query, sending hard problems to DeepSeek-R1 or OpenAI o3 and routine work to a standard LLM behind the same API surface.
The honest framing is that these systems are a tool for a specific shape of work, not a strict upgrade. The category exists because vendors found a reliable way to convert extra compute into extra accuracy on reasoning-heavy benchmarks like MMLU, GPQA, AIME, and LiveCodeBench. That conversion is genuinely useful, and it explains why the o-series, Gemini 2.5 Pro Deep Think, Claude with extended thinking, and DeepSeek-R1 have all converged on a similar pattern. It does not make them the right default for every workload.
- Route by task shape: reach for OpenAI o3 or DeepSeek-R1 on multi-step problems and a standard language model on single-step ones.
- Budget for latency: assume reasoning queries can take many seconds; design the UX around that, not around it.
- Cap the thinking budget: where the vendor exposes a maximum, set one to protect against runaway cost on pathological prompts.
- Evaluate end to end: measure accuracy and cost per task, not just per token, because the two compute budgets interact.
- Re-evaluate periodically: standard LLMs and reasoning systems both improve quickly; the routing line moves over time.
References
- OpenAI, Learning to Reason with LLMs
- OpenAI, Trading Inference-Time Compute for Adversarial Robustness
- Google DeepMind, Gemini 2.5 Pro and Deep Think
- Anthropic, Tracing the Thoughts of a Language Model
- IBM, What Is a Reasoning Model?
- RAND, When AI Takes Time to Think: Implications of Test-Time Compute
Further reading
Frequently Asked Questions
What is test-time compute in a reasoning model?
Test-time compute is the additional computation a reasoning model runs during inference, before producing its final answer. Standard language models generate a response in a single forward pass; a reasoning model uses that extra compute budget to work through intermediate steps, explore alternative approaches, and self-check before committing to an output. OpenAI describes o1 as improving consistently with more time spent thinking.
How does a reasoning model differ from a standard language model?
A reasoning model is trained to generate intermediate reasoning steps before its final response, rather than answering immediately. IBM defines a reasoning model as an LLM fine-tuned to break complex problems into smaller steps called reasoning traces. A standard LLM produces an answer in one pass; a reasoning model allocates a variable compute budget to think first.
Does a reasoning model always think longer to answer every question?
No. Reasoning models adapt their compute use to the difficulty of the task; not every query triggers extended thinking. OpenAI notes that the computation reasoning models use can adapt during inference, meaning simple questions may receive a shorter thinking budget than complex math or coding problems.
Which reasoning models are publicly available today?
OpenAI o1-preview and o1-mini are the models OpenAI used to establish the reasoning model category, per OpenAI. Google DeepMind has introduced Gemini 2.5 Pro with an experimental Deep Think mode for complex math and coding. Anthropic has published research on how Claude plans its responses, though its extended thinking is a related capability rather than a separately named product line.









