In plain terms
A CPU, the processor in every computer, is like a small team of experts: each can do anything, and they work through tasks a few at a time. A GPU is like a hall of several thousand clerks who know only arithmetic and all work at once. A neural network is almost nothing except arithmetic: billions of multiplications and additions that do not depend on one another. The hall of clerks finishes in seconds what the experts would need hours for.
Why it matters
GPUs set the price and the pace of AI. The cost of every token traces back to GPU time, and the memory on a card decides which models it can run at all. If you use models through an API, the provider carries this. If you host models yourself, GPUs become your largest hardware line, with long delivery times and generations that are overtaken within a few years. Supply is concentrated: one company, Nvidia, sells most AI chips, most advanced chips are made by a single manufacturer in Taiwan, and exports of the most powerful ones are under government control.
Example
A manufacturer wants to run a 30-billion-parameter open-weight model in its own data centre. At 16-bit precision the weights alone take about 60 GB; the card it planned to buy has 48 GB of memory. It has two options: buy two cards and split the model across them, or run a version quantised to 4 bits, which needs about 15 GB and fits on one card with room to spare, at a small cost in quality.
Most often confused with
GPU vs. CPU
Every computer has a CPU, which runs the operating system, the database and the application code; an AI server has one too. The GPU sits beside it as a specialist for arithmetic that can be split into many identical pieces. Small models can run on a CPU alone, slowly. The difference matters when sizing servers: for AI workloads the number and memory of the GPUs is what counts, and the CPU is rarely the bottleneck.
Origin: Nvidia popularised the term in 1999 with a graphics card it marketed as the first GPU; in 2012 the success of AlexNet, an image-recognition network trained on two gaming GPUs, made the GPU the hardware of deep learning.
Under the hood
Neural networks reduce to matrix multiplication, which parallelises almost perfectly; a GPU runs it across thousands of cores with high memory bandwidth. The binding constraint is on-board memory (VRAM): the weights, plus working memory that grows with context length and with the number of simultaneous requests, must fit. As a rule of thumb, memory for the weights is parameter count times bytes per parameter: two bytes at 16-bit precision, half a byte at 4-bit. Larger models are split across several GPUs, and training clusters link thousands of them with fast networking. Nvidia dominates through its hardware and its CUDA software ecosystem; AMD sells competing GPUs, and the large cloud providers design their own accelerators, such as Google's TPU. “AI accelerator” is the wider term that covers all of these. For a buyer the practical measures are memory per card, tokens per second per card, power draw and cost per GPU-hour. Capacity is rented from cloud providers by the hour or bought; buying brings power and cooling requirements that ordinary server rooms often cannot meet.