How models work

Inference

Running a trained model to produce an output: every answer, summary or tool call a model generates is an act of inference.

Trainingbuilds the modelDone once, takes monthsThe weights changeCost: one large upfront investmentInferenceuses the modelRuns again on every requestThe weights stay fixedCost: per token, on every useAs usage grows, inference is what drives the bill.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/inference

In plain terms

Training is writing the textbook; inference is using it to answer a question. Each time you send a prompt, the model's fixed weights are applied to your input and an output is computed, token by token. The model is not learning anything at that moment. It is only being used.

Why it matters

Inference is the recurring cost of AI. Training is paid for once by the provider; inference is paid for by you, on every request, for as long as the product runs. Latency, capacity planning, energy use and unit economics are all inference questions, and model choice is largely a trade-off between inference cost and quality.

Example

A support assistant handles 40,000 conversations a month at an average of 6,000 tokens each. That is 240 million tokens of inference every month. Routing the simple third of those conversations to a model that costs a fifth as much changes the annual bill more than any contract negotiation would.

Most often confused with

Inference vs. Training

InferenceUses the model; happens on every request
TrainingCreates the model; happens once

Training changes the weights and ends before release. Inference applies the finished weights to new input and changes nothing in the model. When people say a model “learned” something in a conversation, what happened is inference over a context that contained the information.

Under the hood

For a language model, inference is autoregressive: a forward pass produces a probability distribution, one token is sampled and appended, and the pass repeats. Cost therefore grows with input length and, more steeply, with output length. Measures: time to first token, tokens per second, throughput per GPU. Optimisations: batching, key-value caching of earlier tokens, prompt caching, quantization, distillation to smaller models, and specialised hardware. Deployment choices run from a provider's API, through cloud platforms, to self-hosted open-weight models and on-device inference for privacy or latency.

Written by Mehmet Erkek · Last updated: