In plain terms
At every step the model has a ranked list of candidate tokens. At temperature zero it takes the top one every time, so the same question gets nearly the same answer. Raise the temperature and it reaches further down the list more often: answers become more varied and surprising, and the risk of a poor choice rises.
Why it matters
Temperature is the simplest lever for matching a model to a task. Extraction, classification and anything checked against rules want low temperature for repeatability. Brainstorming and creative drafts benefit from more. Many reports of “the model is inconsistent” turn out to be a temperature left at its default.
Example
A team generates product descriptions at temperature 1.0 and likes the variety. The same setting on their invoice-extraction step produces a different total on one run in fifty. They set extraction to 0 and keep 1.0 for copywriting.
Most often confused with
Temperature vs. Creativity
A higher temperature makes output less predictable; it does not make the model smarter or more original in any deeper sense, and past a point it produces nonsense. A better prompt, with richer context and a clearer brief, usually does more for the quality of ideas than a higher temperature.
Under the hood
Temperature divides the logits before the softmax: values below 1 sharpen the distribution, values above 1 flatten it, and 0 is treated as greedy decoding. Typical ranges run from 0 to 1 or 0 to 2 depending on the provider. It interacts with top-p and top-k, which cut off the tail before sampling; adjusting one of them at a time is usual practice. Temperature 0 does not guarantee identical outputs, because of numerical effects and infrastructure differences; a fixed seed helps where offered. Some reasoning models fix or restrict sampling settings.