Infrastructure & economics

Inference Cost

The recurring cost of running a trained model: it is paid on every request, in tokens or in GPU time, and grows with the number of users, the length of prompts and the length of answers.

price per token is fallingtokens per task are risingbill = price × consumptionillustrative curvestime →The price sidethe same capability gets far cheaper every yearThe consumption sidereasoning and agents use many more tokensThe bill can grow while prices fall; track cost per completed task.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/inference-cost

In plain terms

Training a model is like building a power station: a huge bill, paid once, by the provider. Inference cost is the fuel: it is burned every time someone asks a question, and the bill is yours. A pilot with fifty users costs almost nothing. The same assistant rolled out to twenty thousand employees, with long documents attached to every question, produces an invoice that someone needs to have budgeted for.

Why it matters

This is the line that decides whether an AI use case pays for itself at scale. Two trends pull in opposite directions. The price of a given level of capability has fallen steeply, by roughly ten times a year on some published estimates. Consumption per task has risen just as sharply, because reasoning models think before answering and agents make dozens of model calls for one job. Many organisations therefore face a growing bill while prices fall. Cost has to be designed and monitored task by task; it will not look after itself.

Example

An insurer's claims assistant handles 50,000 cases a month. Each case uses 20,000 input tokens and 1,000 output tokens. At illustrative prices of 2 dollars per million input tokens and 10 dollars per million output tokens, a case costs 5 cents and a month 2,500 dollars. The team then turns the assistant into an agent that makes twelve model calls per case. Nothing on the price list changes, and the bill rises to about 30,000 dollars.

Most often confused with

Inference Cost vs. Token pricing

Inference CostThe bill: price multiplied by everything you consume
Token pricingThe price list: the rate per million tokens

Token pricing is the provider's rate card. Inference cost is what you pay: that rate multiplied by your volume, your prompt lengths, your retries and the number of model calls each task takes. A cheaper model can produce a higher bill if it needs more attempts, and a price cut can disappear inside a design that uses more tokens. Compare models on cost per completed task.

Under the hood

For API use, cost per request is input tokens times the input price plus output tokens times the output price; output is priced several times higher, and reasoning tokens count as output. For self-hosted models it is GPU-hours divided by the tokens served, so utilisation dominates. The main levers, roughly in order of effect: use the smallest model that passes your evals and route only the hard requests to a larger one; cache repeated prompt prefixes; shorten the context through retrieval and compaction; cap output length and reasoning effort; move work that can wait to batch processing, commonly at half price; distil or fine-tune a small model for a high-volume task. Measure cost per task or per resolved case, since cost per token hides retries and multi-step agents. Controls worth having: usage attributed to teams and features, budgets and alerts, per-user limits, and a ceiling on steps and spend for every agent run. Items that are easy to miss: failed and retried calls, evaluation runs, embeddings, and tool charges such as web search.

Written by Mehmet Erkek · Last updated: