In plain terms
Ask someone a tricky arithmetic question and demand an instant answer, and they will often get it wrong. Give them paper and they get it right. Chain of thought gives the model paper: it writes out the intermediate steps, and since each step becomes part of what it reads next, the final answer rests on worked reasoning and not on a single leap.
Why it matters
It was one of the first demonstrations that how you ask changes how well a model reasons, and it remains the cheapest way to improve accuracy on multi-step problems with standard models. It also makes answers auditable: when the steps are written down, a reviewer can see where the logic went wrong. The cost is more output tokens and a slower reply.
Example
A model is asked whether a customer qualifies for a refund under a policy with four conditions. Answering directly, it says yes. Asked to check each condition in turn and then conclude, it finds that the purchase was 34 days ago against a 30-day limit, and correctly says no.
Most often confused with
CoT vs. Reasoning Model
Chain of thought is an instruction: “think step by step, then answer”. A reasoning model has the behaviour trained in and manages its own thinking. With reasoning models, adding step-by-step instructions is usually unnecessary and can make results worse; a clear statement of the goal works better.
Origin: Introduced by Wei and colleagues at Google in 2022; the zero-shot variant, “Let's think step by step”, followed the same year.
Under the hood
Variants: zero-shot CoT (an instruction to reason first), few-shot CoT (examples that include reasoning), and structured CoT that separates thinking from the answer with tags so the reasoning can be removed before display. It works because autoregressive models condition on their own output: written steps become working memory. It helps most on arithmetic, logic and multi-condition decisions, and helps little on simple recall. One caution: the written reasoning is not always a faithful account of how the answer was reached, so it supports review without proving correctness.