In plain terms
It is the difference between asking a consultant for an answer on the spot and giving them the afternoon. With extended thinking switched on, the model is allowed to work on the problem at length before replying, and you decide how long that may be. More thinking costs more and takes longer, and helps most where the problem is hard.
Why it matters
It turns “how smart should the model be on this request” into a dial with a price on it. A product can answer routine questions instantly and spend thirty seconds of thinking on the rare complex one, using the same model. Setting that dial per task is one of the most direct ways to control both quality and cost.
Example
A financial-analysis tool uses no thinking for “what was revenue in Q2?” and a budget of 16,000 thinking tokens for “explain the three main drivers of the margin change and check them against the segment data”. The second request costs about ten times as much and is the one customers pay for.
Most often confused with
Extended Thinking vs. Reasoning Model
The terms overlap. “Reasoning model” describes the capability; “extended thinking” is the name some providers, notably Anthropic, give to the adjustable mode that switches it on and sizes it. Other providers call the same control reasoning effort or a thinking budget.
Under the hood
The developer sets a thinking budget or an effort level; newer versions let the model decide how much to think within a ceiling (adaptive thinking). Thinking tokens are billed as output tokens, count toward the context window while they are produced, and delay the first visible token, so streaming and timeouts need attention. The thinking may be returned in full, summarised or hidden, depending on the provider and model. With tool use, a model can think between tool calls (interleaved thinking). Returns diminish: beyond what the problem needs, extra budget adds cost and latency with little gain.