In plain terms
Training is writing the textbook; inference is using it to answer a question. Each time you send a prompt, the model's fixed weights are applied to your input and an output is computed, token by token. The model is not learning anything at that moment. It is only being used.
Why it matters
Inference is the recurring cost of AI. Training is paid for once by the provider; inference is paid for by you, on every request, for as long as the product runs. Latency, capacity planning, energy use and unit economics are all inference questions, and model choice is largely a trade-off between inference cost and quality.
Example
A support assistant handles 40,000 conversations a month at an average of 6,000 tokens each. That is 240 million tokens of inference every month. Routing the simple third of those conversations to a model that costs a fifth as much changes the annual bill more than any contract negotiation would.
Most often confused with
Inference vs. Training
Training changes the weights and ends before release. Inference applies the finished weights to new input and changes nothing in the model. When people say a model “learned” something in a conversation, what happened is inference over a context that contained the information.
Under the hood
For a language model, inference is autoregressive: a forward pass produces a probability distribution, one token is sampled and appended, and the pass repeats. Cost therefore grows with input length and, more steeply, with output length. Measures: time to first token, tokens per second, throughput per GPU. Optimisations: batching, key-value caching of earlier tokens, prompt caching, quantization, distillation to smaller models, and specialised hardware. Deployment choices run from a provider's API, through cloud platforms, to self-hosted open-weight models and on-device inference for privacy or latency.