Infrastructure & economics

Compute

Computing power

The processing power used to train and run AI models, supplied mostly by GPUs and measured in units such as GPU-hours; with data and algorithms, one of the three basic inputs of AI.

8 GPUs × 36 hours = 288 GPU-hourscompute: the work hardware delivers over time · illustrative exampleTHREE WAYS TO OBTAIN COMPUTEThrough an API, per tokencompute is inside the priceno fixed costRented cloud GPUspaid per GPU-hourflexible, expensive per hourYour own hardwareinvestment and operationscheap if kept busymore control, more fixed cost →Work out the workload in GPU-hours first; the hardware decision comes after that.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/compute

In plain terms

Compute is to AI what electricity is to a factory: the resource everything consumes, bought by the unit and limited at any given moment. A GPU is the machine; compute is the work it delivers over time, counted in GPU-hours. One GPU running for one hour is one GPU-hour, whether you own the card, rent it in the cloud or pay for it indirectly inside the price of a token.

Why it matters

Compute is the scarce input of AI. Access to it decides which organisations can train frontier models, which is why a few companies and governments commit sums on the scale of national infrastructure projects to data centres, and why chips have become a matter of trade policy. For most companies it appears more modestly, as a budget line: inside API prices, as cloud GPU rental, or as hardware on the balance sheet. The usual mistake is to buy capacity before knowing the workload. Rented compute is flexible and expensive per hour; owned compute is cheap per hour only when it is kept busy.

Example

A research team wants to fine-tune an open-weight model and sizes the job at 8 GPUs for 36 hours: 288 GPU-hours. At an illustrative cloud rate of 3 dollars per GPU-hour that comes to 864 dollars, with nothing to buy. Purchasing the eight cards would cost an illustrative 250,000 dollars and makes sense only if they then stay busy for most of the year.

Most often confused with

Compute vs. GPU

ComputeThe resource: processing work delivered over time
GPUThe hardware: one kind of chip that supplies it

A GPU is an object you can count and install; compute is what you get from running it, measured in GPU-hours. A need for “more compute” can be met with more GPUs, with faster ones, by renting from a cloud or by running the existing cards around the clock. Planning in compute first and choosing hardware second keeps the buying decision open.

Under the hood

Units: GPU-hours for rented capacity; floating-point operations (FLOP) for the total size of a training run; tokens per second for serving; and, at data-centre scale, megawatts and gigawatts of power. Electricity has increasingly become the limiting factor for new capacity. Compute is consumed in three places: training (one large burst), experiments and evaluation, and inference (continuous, growing with usage and taking a rising share of the total). Ways to obtain it: bundled into an API's token price, reserved or on-demand GPU servers from cloud providers, specialist GPU clouds, or owned hardware. Cost levers: utilisation (an idle GPU costs the same as a busy one), the gap between reserved and on-demand rates, spot capacity for jobs that can be interrupted, and smaller or quantised models. Regulation also uses compute as a yardstick: the EU AI Act presumes systemic risk for general-purpose models trained with more than 10^25 floating-point operations.

Written by Mehmet Erkek · Last updated: