In plain terms
Compute is to AI what electricity is to a factory: the resource everything consumes, bought by the unit and limited at any given moment. A GPU is the machine; compute is the work it delivers over time, counted in GPU-hours. One GPU running for one hour is one GPU-hour, whether you own the card, rent it in the cloud or pay for it indirectly inside the price of a token.
Why it matters
Compute is the scarce input of AI. Access to it decides which organisations can train frontier models, which is why a few companies and governments commit sums on the scale of national infrastructure projects to data centres, and why chips have become a matter of trade policy. For most companies it appears more modestly, as a budget line: inside API prices, as cloud GPU rental, or as hardware on the balance sheet. The usual mistake is to buy capacity before knowing the workload. Rented compute is flexible and expensive per hour; owned compute is cheap per hour only when it is kept busy.
Example
A research team wants to fine-tune an open-weight model and sizes the job at 8 GPUs for 36 hours: 288 GPU-hours. At an illustrative cloud rate of 3 dollars per GPU-hour that comes to 864 dollars, with nothing to buy. Purchasing the eight cards would cost an illustrative 250,000 dollars and makes sense only if they then stay busy for most of the year.
Most often confused with
Compute vs. GPU
A GPU is an object you can count and install; compute is what you get from running it, measured in GPU-hours. A need for “more compute” can be met with more GPUs, with faster ones, by renting from a cloud or by running the existing cards around the clock. Planning in compute first and choosing hardware second keeps the buying decision open.
Under the hood
Units: GPU-hours for rented capacity; floating-point operations (FLOP) for the total size of a training run; tokens per second for serving; and, at data-centre scale, megawatts and gigawatts of power. Electricity has increasingly become the limiting factor for new capacity. Compute is consumed in three places: training (one large burst), experiments and evaluation, and inference (continuous, growing with usage and taking a rising share of the total). Ways to obtain it: bundled into an API's token price, reserved or on-demand GPU servers from cloud providers, specialist GPU clouds, or owned hardware. Cost levers: utilisation (an idle GPU costs the same as a busy one), the gap between reserved and on-demand rates, spot capacity for jobs that can be interrupted, and smaller or quantised models. Regulation also uses compute as a yardstick: the EU AI Act presumes systemic risk for general-purpose models trained with more than 10^25 floating-point operations.