Infrastructure & economics

Token Pricing

Pay-per-token pricing

The standard way AI models are sold through an API: a price per million tokens, charged separately for the text sent in and the text generated, with output costing several times more.

THE BILL FOR ONE REQUEST · illustrative pricesITEMTOKENSPRICE PER MILLIONAMOUNTInput · uncached5,000$3.00$0.0150Input · read from cache25,000$0.30$0.0075Output · reasoning included1,500$15.00$0.0225Total per request$0.0450Output costs several times inputeach output token needs its own pass through the modelWithout the cache: $0.1125the same request, 2.5 times the billPrices are illustrative and change every few months; the structure is similar across providers.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/token-pricing

In plain terms

A model API is billed like a taxi with two meters. One meter counts everything you send in: instructions, documents, the conversation so far. The other counts everything the model writes back, and it runs several times faster. There is no flat fee and no price per question. A short question about a long contract can cost more than a long question with no document, because the contract passes through the first meter every time.

Why it matters

Token pricing makes cost proportional to use. That is good for starting small and awkward for budgeting, because nobody knows in advance how many tokens a new feature will consume. The price list also contains real choices. Input read from a cache is commonly discounted by fifty to ninety percent, and batch processing by around half; reasoning tokens are billed at the output rate although nobody reads them; one provider sells models at prices many times apart. Prices change every few months, so base estimates on your own measured token counts and re-check the rate card regularly.

Example

A contract-review tool sends 30,000 input tokens and receives 1,500 output tokens per review. At illustrative prices of 3 dollars per million input tokens and 15 dollars per million output tokens, one review costs 9 cents for input and about 2 cents for output: 11 cents. When 25,000 of the input tokens are a fixed review playbook, read from cache at a tenth of the price, the review drops to about 4.5 cents.

Most often confused with

Token Pricing vs. Per-seat subscription

Token PricingPay for what is consumed; the bill moves with usage
Per-seat subscriptionPay a fixed monthly fee per user

Ready-made assistants for employees are mostly sold per seat: a fixed monthly fee per person, with usage limits behind it. The APIs used to build your own applications are sold per token. A seat is predictable and can be wasteful for light users; tokens are efficient and hard to forecast. Many organisations hold both, and need to know which bill a given project lands on.

Under the hood

Prices are quoted per million tokens, with separate rates for input, cached input and output. Output usually costs several times the input rate, because every output token requires its own pass through the model. Reasoning or thinking tokens are billed at the output rate. Images, audio and video are converted to tokens and billed the same way, sometimes at their own rates. Common modifiers: a discount for reading from the cache and a surcharge for writing to it; around half price for asynchronous batch jobs; a higher rate for very long prompts on some models; dearer tiers for guaranteed or faster capacity and cheaper ones for flexible timing; a surcharge for data residency. Tool use adds the tokens of tool definitions and results, and some built-in tools carry a fee per call. At scale, committed-spend contracts and reserved capacity take the place of list prices. Token counts depend on each provider's tokenizer, and Turkish text uses more tokens than English for the same content, so compare prices on your own workload. The usage data returned with every call gives the exact counts; log them.

Written by Mehmet Erkek · Last updated: