What is a reasoning model?
A reasoning model is an LLM trained to generate an extended sequence of intermediate steps, often called a chain of thought, before it produces a final answer. Instead of jumping straight to a response, it works through the problem in tokens the way a person might on scratch paper: restating what's being asked, trying an approach, checking it, sometimes backtracking. Those intermediate tokens cost inference time and memory, but on tasks that benefit from multi-step logic, math, and code, that extra work reliably improves accuracy over a model answering directly. The tradeoff is explicit: more compute and more tokens spent per answer, in exchange for a better answer on the problems where it counts.
How this differs from a model that just answers directly
A standard instruction-tuned model, given a hard multi-step problem, tends to produce a plausible-looking answer in one pass, and on anything requiring several dependent steps that answer is often wrong even when it reads fluently. A reasoning model addresses this not through a different architecture but through training: reinforcement learning on problems with checkable answers, math proofs, code that either passes tests or doesn't, rewards the model for chains of reasoning that actually arrive at correct results, and that training shapes it to spend more tokens working through a problem before answering. Same transformer, same kind of forward pass, different learned behavior about when to stop and answer.
What this costs at inference time
Every reasoning token the model generates is a token it has to decode, one at a time, the same as any other output token. A response that used to be a few hundred tokens can become several thousand once reasoning is included, and each of those tokens still costs a full decode step and adds to the KV cache the request is holding. On memory-bandwidth-bound hardware, where decode speed is set by memory bandwidth divided by active parameter bytes as our benchmarks page lays out, more output tokens per request means proportionally more decode time and more cache held per request, not a free lunch. This is also why reasoning models pair naturally with the batching and throughput questions covered in our batching in LLM inference post: the more tokens a single request spends thinking, the more it competes with other concurrent requests for decode capacity.
When it's worth the cost
Reasoning is not universally better; on simple factual lookups or short-form writing, the extra tokens mostly add latency without adding accuracy. It earns its cost on tasks with real multi-step structure: debugging, math, planning, anything where a wrong intermediate assumption compounds into a wrong final answer. Our models page notes which catalog models are tuned for reasoning versus direct response, and our evaluating a model before deploying post covers how to test whether the accuracy gain is actually worth the latency and cache cost for a specific workload rather than assuming it.