In plain terms
A photograph saved at lower quality is still the same picture: a much smaller file, slightly less fine detail, and for most purposes nobody notices. Quantisation does that to a model. Every one of its billions of numbers is rounded to a coarser scale, as if prices were rounded from fractions of a cent to whole cents. The model keeps its shape and its knowledge, takes up a quarter of the space, and gives answers that are almost always the same.
Why it matters
Quantisation decides where a model can run. It is the difference between needing a server with several graphics cards and using one, or between the cloud and a laptop or phone, which matters for cost, for offline use and for keeping data on your own hardware. Most open-weight models are used in quantised form. The price is a loss of quality that is small at 8 bits, usually acceptable at 4 and steep below that, and it tends to show first on the hardest tasks. Measure it on your own work; published averages hide the cases that matter to you.
Example
An engineering firm wants an assistant on its field engineers' laptops, working offline. The model it chose has 7 billion parameters: 14 GB at 16 bits, too much for laptops with 8 GB of graphics memory. The 4-bit version needs about 3.5 GB and writes faster than a person reads. On the firm's 200 test questions accuracy falls from 91% to 89%. The firm accepts the two points.
Most often confused with
Quantization vs. Distillation
Quantisation is compression: no new training in its usual form, done in minutes, and undone by going back to the original file. Distillation is teaching: it needs data, a training run and a new round of testing, and produces a model with fewer parameters. Choose quantisation to fit a model you already trust onto smaller hardware; choose distillation when even the compressed model is too large or too slow.
Under the hood
Weights are normally trained as 32-bit or 16-bit floating-point numbers. Quantisation maps them to a small set of levels: 8-bit, 4-bit, sometimes lower. Memory falls in proportion: a 7-billion-parameter model takes about 14 GB at 16 bits and about 3.5 GB at 4 bits, plus working memory. Two approaches: post-training quantisation, applied to a finished model, sometimes with a small calibration dataset; and quantisation-aware training, which simulates the rounding during training and loses less. Well-known methods include GPTQ and AWQ; GGUF is the file format of the llama.cpp ecosystem, where names such as Q4_K_M indicate the bit width; the bitsandbytes library loads models in 8-bit or 4-bit form. Activations can be quantised as well as weights. Speed gains depend on hardware support for low-precision arithmetic. Pitfalls: quality that looks fine on benchmarks and fails on rare cases, and comparing models at different bit widths without saying so.