How models work

Extended Thinking

A mode in which a model is given a budget of extra tokens to reason before it answers, with the budget set by the user or developer.

THINKING OFFAnswerTHINKING ON · budget: for example 10,000 tokensThinking: plan, calculate, checkAnswerLow budgetsimple questions, quick repliesHigh budgethard problems, deeper checkingThinking tokens count as output and are billed; they also lengthen response time.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/extended-thinking

In plain terms

It is the difference between asking a consultant for an answer on the spot and giving them the afternoon. With extended thinking switched on, the model is allowed to work on the problem at length before replying, and you decide how long that may be. More thinking costs more and takes longer, and helps most where the problem is hard.

Why it matters

It turns “how smart should the model be on this request” into a dial with a price on it. A product can answer routine questions instantly and spend thirty seconds of thinking on the rare complex one, using the same model. Setting that dial per task is one of the most direct ways to control both quality and cost.

Example

A financial-analysis tool uses no thinking for “what was revenue in Q2?” and a budget of 16,000 thinking tokens for “explain the three main drivers of the margin change and check them against the segment data”. The second request costs about ten times as much and is the one customers pay for.

Most often confused with

Extended Thinking vs. Reasoning Model

Extended ThinkingA setting: how much the model may think on this request
Reasoning ModelA kind of model: one built to reason before answering

The terms overlap. “Reasoning model” describes the capability; “extended thinking” is the name some providers, notably Anthropic, give to the adjustable mode that switches it on and sizes it. Other providers call the same control reasoning effort or a thinking budget.

Under the hood

The developer sets a thinking budget or an effort level; newer versions let the model decide how much to think within a ceiling (adaptive thinking). Thinking tokens are billed as output tokens, count toward the context window while they are produced, and delay the first visible token, so streaming and timeouts need attention. The thinking may be returned in full, summarised or hidden, depending on the provider and model. With tool use, a model can think between tool calls (interleaved thinking). Returns diminish: beyond what the problem needs, extra budget adds cost and latency with little gain.

Written by Mehmet Erkek · Last updated: