How models work

Temperature

A setting that controls how much randomness a model uses when choosing each next token: low values give consistent output, high values give varied output.

PROMPT: “The sky today is …” · illustrative probabilitiesLow temperature · 0.2consistent, predictableblue92%clear5%grey2%orange1%High temperature · 1.2varied, creative, riskierblue45%clear25%grey18%orange12%Temperature does not change what the model knows; it sets how adventurous the choice of token is.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/temperature

In plain terms

At every step the model has a ranked list of candidate tokens. At temperature zero it takes the top one every time, so the same question gets nearly the same answer. Raise the temperature and it reaches further down the list more often: answers become more varied and surprising, and the risk of a poor choice rises.

Why it matters

Temperature is the simplest lever for matching a model to a task. Extraction, classification and anything checked against rules want low temperature for repeatability. Brainstorming and creative drafts benefit from more. Many reports of “the model is inconsistent” turn out to be a temperature left at its default.

Example

A team generates product descriptions at temperature 1.0 and likes the variety. The same setting on their invoice-extraction step produces a different total on one run in fifty. They set extraction to 0 and keep 1.0 for copywriting.

Most often confused with

Temperature vs. Creativity

TemperatureAdjusts randomness in token choice
CreativityThe quality of ideas, which comes from the model and the prompt

A higher temperature makes output less predictable; it does not make the model smarter or more original in any deeper sense, and past a point it produces nonsense. A better prompt, with richer context and a clearer brief, usually does more for the quality of ideas than a higher temperature.

Under the hood

Temperature divides the logits before the softmax: values below 1 sharpen the distribution, values above 1 flatten it, and 0 is treated as greedy decoding. Typical ranges run from 0 to 1 or 0 to 2 depending on the provider. It interacts with top-p and top-k, which cut off the tail before sampling; adjusting one of them at a time is usual practice. Temperature 0 does not guarantee identical outputs, because of numerical effects and infrastructure differences; a fixed seed helps where offered. Some reasoning models fix or restrict sampling settings.

Written by Mehmet Erkek · Last updated: