How models work

Next-token Prediction

The single operation at the core of a language model: given the text so far, estimate how likely each possible next token is.

PROMPTThe capital of Türkiye is …PROBABILITIES FOR THE NEXT TOKEN · illustrativeAnkara91%Istanbul5%a2%located1%One token is picked and appended, then the prediction runs again: the answer is built token by token.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/next-token-prediction

In plain terms

Your phone's keyboard suggests the next word. A language model does the same thing with vastly more skill: it reads everything so far and assigns a probability to every token that could come next. It picks one, adds it, and repeats. An essay, a contract summary and a block of code are all produced this way, one piece at a time.

Why it matters

This one fact explains much of what surprises people about AI. The model produces what is plausible, which usually coincides with what is true and sometimes does not; that is where hallucinations come from. It also explains why wording and context matter so much: they change what is likely to come next.

Example

Given “The contract may be terminated with thirty days' written”, the model gives “notice” a probability above ninety percent and continues. Given “Our Q3 revenue was”, no token is strongly favoured unless the figure is in the context, and the model may produce a number that merely looks right.

Most often confused with

Next-token Prediction vs. Autocomplete

Next-token PredictionPredicts using an understanding built from vast training
AutocompleteSuggests from recent typing and simple statistics

The mechanism is the same in outline, which is why “fancy autocomplete” is a popular description. It undersells the result: to predict the next token of a proof, a legal argument or a program well, a model has to represent the logic behind them. The prediction is simple; what the network learned in order to predict is not.

Under the hood

The network outputs a score (logit) for every token in the vocabulary; a softmax turns the scores into probabilities; a decoding rule (greedy, temperature sampling, top-p or top-k) selects one. Generation is autoregressive, so each new token joins the input for the next step, and an early mistake can carry forward. Training minimises cross-entropy between the predicted distribution and the actual next token. Reasoning models add a phase of generated intermediate text before the final answer, but the underlying operation is still next-token prediction.

Written by Mehmet Erkek · Last updated: