In plain terms
Training a model is like building a power station: a huge bill, paid once, by the provider. Inference cost is the fuel: it is burned every time someone asks a question, and the bill is yours. A pilot with fifty users costs almost nothing. The same assistant rolled out to twenty thousand employees, with long documents attached to every question, produces an invoice that someone needs to have budgeted for.
Why it matters
This is the line that decides whether an AI use case pays for itself at scale. Two trends pull in opposite directions. The price of a given level of capability has fallen steeply, by roughly ten times a year on some published estimates. Consumption per task has risen just as sharply, because reasoning models think before answering and agents make dozens of model calls for one job. Many organisations therefore face a growing bill while prices fall. Cost has to be designed and monitored task by task; it will not look after itself.
Example
An insurer's claims assistant handles 50,000 cases a month. Each case uses 20,000 input tokens and 1,000 output tokens. At illustrative prices of 2 dollars per million input tokens and 10 dollars per million output tokens, a case costs 5 cents and a month 2,500 dollars. The team then turns the assistant into an agent that makes twelve model calls per case. Nothing on the price list changes, and the bill rises to about 30,000 dollars.
Most often confused with
Inference Cost vs. Token pricing
Token pricing is the provider's rate card. Inference cost is what you pay: that rate multiplied by your volume, your prompt lengths, your retries and the number of model calls each task takes. A cheaper model can produce a higher bill if it needs more attempts, and a price cut can disappear inside a design that uses more tokens. Compare models on cost per completed task.
Under the hood
For API use, cost per request is input tokens times the input price plus output tokens times the output price; output is priced several times higher, and reasoning tokens count as output. For self-hosted models it is GPU-hours divided by the tokens served, so utilisation dominates. The main levers, roughly in order of effect: use the smallest model that passes your evals and route only the hard requests to a larger one; cache repeated prompt prefixes; shorten the context through retrieval and compaction; cap output length and reasoning effort; move work that can wait to batch processing, commonly at half price; distil or fine-tune a small model for a high-volume task. Measure cost per task or per resolved case, since cost per token hides retries and multi-step agents. Controls worth having: usage attributed to teams and features, budgets and alerts, per-user limits, and a ceiling on steps and spend for every agent run. Items that are easy to miss: failed and retried calls, evaluation runs, embeddings, and tool charges such as web search.