In plain terms
The model's list of candidate tokens has a few likely options and a very long tail of unlikely ones. Sampling settings decide how much of that tail is allowed. Top-k keeps a fixed number of top candidates. Top-p keeps as many as are needed to cover a set share of the probability. Either way, the far-fetched options are removed before the dice are rolled.
Why it matters
Most users never touch these settings, and most should not need to. They matter to teams building products: together with temperature they set the balance between reliable and varied output, and a poorly chosen combination is a common source of erratic behaviour that is then blamed on the model.
Example
A chatbot occasionally inserts an odd word that derails a sentence. The team finds top-p set to 1.0 with a high temperature, which lets the model draw from the entire tail. Lowering top-p to 0.9 removes the oddities without making answers repetitive.
Most often confused with
Sampling vs. Temperature
Temperature changes the probabilities of all candidates, making the distribution sharper or flatter. Top-k and top-p remove candidates outright. They are usually combined: first the distribution is reshaped, then the tail is cut, then one token is drawn from what remains.
Under the hood
Greedy decoding always takes the most probable token. Top-k sampling keeps the k most probable tokens and renormalises. Top-p, or nucleus sampling, keeps the smallest set whose cumulative probability reaches p, so the number of candidates adapts: few when the model is confident, many when it is unsure. Other controls include min-p, repetition and frequency penalties, and beam search, which is now rare for chat models. Constrained decoding goes further and masks every token that would violate a grammar or JSON schema; this is how structured outputs are guaranteed. Provider defaults differ, and some models do not expose all settings.