In plain terms
A standard model starts writing its answer immediately, like someone thinking aloud. A reasoning model first takes private notes: it breaks the problem down, tries a route, notices a mistake, tries another. Only then does it write the reply. You wait longer and pay for the notes, and on hard problems the answer is markedly better.
Why it matters
Reasoning models moved AI from fluent text to dependable multi-step work: debugging code, analysing a contract against a policy, planning a project with constraints. They are also slower and more expensive, and on simple tasks they add cost without benefit. Choosing when to use one is now a basic part of system design.
Example
Asked to reconcile two spreadsheets with mismatched totals, a standard model offers a plausible guess about the cause. A reasoning model checks the row counts, compares totals by category, isolates three duplicated invoice lines and states the discrepancy to the cent. It takes forty seconds, where the first took four.
Most often confused with
Reasoning Model vs. Chain of Thought (CoT)
Chain of thought is something you ask for in a prompt: “think step by step”. A reasoning model does it by design: it was trained to produce and use intermediate reasoning, often at length, without being asked. The technique came first, and the models turned it into a built-in capability.
Under the hood
Reasoning models generate a stream of reasoning tokens before the visible answer. Those tokens are billed as output and take up context, and some providers show only a summary of them. Effort is usually adjustable through a setting (a reasoning-effort level or a thinking-token budget). Training typically uses reinforcement learning on problems with checkable answers. Prompting differs from standard models: state the goal and constraints, and avoid scripting the steps. The reasoning shown is not guaranteed to reflect the computation that produced the answer, a limitation known as unfaithful reasoning.