In plain terms
A model API is billed like a taxi with two meters. One meter counts everything you send in: instructions, documents, the conversation so far. The other counts everything the model writes back, and it runs several times faster. There is no flat fee and no price per question. A short question about a long contract can cost more than a long question with no document, because the contract passes through the first meter every time.
Why it matters
Token pricing makes cost proportional to use. That is good for starting small and awkward for budgeting, because nobody knows in advance how many tokens a new feature will consume. The price list also contains real choices. Input read from a cache is commonly discounted by fifty to ninety percent, and batch processing by around half; reasoning tokens are billed at the output rate although nobody reads them; one provider sells models at prices many times apart. Prices change every few months, so base estimates on your own measured token counts and re-check the rate card regularly.
Example
A contract-review tool sends 30,000 input tokens and receives 1,500 output tokens per review. At illustrative prices of 3 dollars per million input tokens and 15 dollars per million output tokens, one review costs 9 cents for input and about 2 cents for output: 11 cents. When 25,000 of the input tokens are a fixed review playbook, read from cache at a tenth of the price, the review drops to about 4.5 cents.
Most often confused with
Token Pricing vs. Per-seat subscription
Ready-made assistants for employees are mostly sold per seat: a fixed monthly fee per person, with usage limits behind it. The APIs used to build your own applications are sold per token. A seat is predictable and can be wasteful for light users; tokens are efficient and hard to forecast. Many organisations hold both, and need to know which bill a given project lands on.
Under the hood
Prices are quoted per million tokens, with separate rates for input, cached input and output. Output usually costs several times the input rate, because every output token requires its own pass through the model. Reasoning or thinking tokens are billed at the output rate. Images, audio and video are converted to tokens and billed the same way, sometimes at their own rates. Common modifiers: a discount for reading from the cache and a surcharge for writing to it; around half price for asynchronous batch jobs; a higher rate for very long prompts on some models; dearer tiers for guaranteed or faster capacity and cheaper ones for flexible timing; a surcharge for data residency. Tool use adds the tokens of tool definitions and results, and some built-in tools carry a fee per call. At scale, committed-spend contracts and reserved capacity take the place of list prices. Token counts depend on each provider's tokenizer, and Turkish text uses more tokens than English for the same content, so compare prices on your own workload. The usage data returned with every call gives the exact counts; log them.