How models work

Sampling

Top-p · top-k

The rules by which a model picks the next token from its probability list; top-k and top-p limit the choice to the most likely candidates.

CANDIDATE TOKENS SORTED BY PROBABILITY · illustrativeblue40%clear25%grey15%cloudy10%orange6%purple4%top-k = 3first 3top-p = 0.9sum 90%Both cut off the unlikely tail; the choice among the rest then follows the temperature.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/sampling

In plain terms

The model's list of candidate tokens has a few likely options and a very long tail of unlikely ones. Sampling settings decide how much of that tail is allowed. Top-k keeps a fixed number of top candidates. Top-p keeps as many as are needed to cover a set share of the probability. Either way, the far-fetched options are removed before the dice are rolled.

Why it matters

Most users never touch these settings, and most should not need to. They matter to teams building products: together with temperature they set the balance between reliable and varied output, and a poorly chosen combination is a common source of erratic behaviour that is then blamed on the model.

Example

A chatbot occasionally inserts an odd word that derails a sentence. The team finds top-p set to 1.0 with a high temperature, which lets the model draw from the entire tail. Lowering top-p to 0.9 removes the oddities without making answers repetitive.

Most often confused with

Sampling vs. Temperature

SamplingCuts which candidates are allowed
TemperatureReshapes the odds among the candidates

Temperature changes the probabilities of all candidates, making the distribution sharper or flatter. Top-k and top-p remove candidates outright. They are usually combined: first the distribution is reshaped, then the tail is cut, then one token is drawn from what remains.

Under the hood

Greedy decoding always takes the most probable token. Top-k sampling keeps the k most probable tokens and renormalises. Top-p, or nucleus sampling, keeps the smallest set whose cumulative probability reaches p, so the number of candidates adapts: few when the model is confident, many when it is unsure. Other controls include min-p, repetition and frequency penalties, and beam search, which is now rare for chat models. Constrained decoding goes further and masks every token that would violate a grammar or JSON schema; this is how structured outputs are guaranteed. Provider defaults differ, and some models do not expose all settings.

Written by Mehmet Erkek · Last updated: