Model customization

Quantization

Quantisation

Reducing the number of bits used to store each of a model's weights, for example from 16 to 4, so that the model needs less memory and runs faster, at the price of a small loss in quality.

EVERY WEIGHT IS ROUNDED TO A COARSER SCALE16 bits: 0.73124…65,536 possible values4 bits: 0.7516 possible valuesSIZE OF A 7-BILLION-PARAMETER MODEL · illustrative16 bits14 GB4 bits3.5 GBlaptop memory: 8 GBaccuracy on 200 test questions: 91% → 89%It is the same model; its numbers are stored with fewer bits. The loss shows first on the hardest tasks.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/quantization

In plain terms

A photograph saved at lower quality is still the same picture: a much smaller file, slightly less fine detail, and for most purposes nobody notices. Quantisation does that to a model. Every one of its billions of numbers is rounded to a coarser scale, as if prices were rounded from fractions of a cent to whole cents. The model keeps its shape and its knowledge, takes up a quarter of the space, and gives answers that are almost always the same.

Why it matters

Quantisation decides where a model can run. It is the difference between needing a server with several graphics cards and using one, or between the cloud and a laptop or phone, which matters for cost, for offline use and for keeping data on your own hardware. Most open-weight models are used in quantised form. The price is a loss of quality that is small at 8 bits, usually acceptable at 4 and steep below that, and it tends to show first on the hardest tasks. Measure it on your own work; published averages hide the cases that matter to you.

Example

An engineering firm wants an assistant on its field engineers' laptops, working offline. The model it chose has 7 billion parameters: 14 GB at 16 bits, too much for laptops with 8 GB of graphics memory. The 4-bit version needs about 3.5 GB and writes faster than a person reads. On the firm's 200 test questions accuracy falls from 91% to 89%. The firm accepts the two points.

Most often confused with

Quantization vs. Distillation

QuantizationSame model, each number stored with fewer bits
DistillationA new, smaller model trained to imitate the original

Quantisation is compression: no new training in its usual form, done in minutes, and undone by going back to the original file. Distillation is teaching: it needs data, a training run and a new round of testing, and produces a model with fewer parameters. Choose quantisation to fit a model you already trust onto smaller hardware; choose distillation when even the compressed model is too large or too slow.

Under the hood

Weights are normally trained as 32-bit or 16-bit floating-point numbers. Quantisation maps them to a small set of levels: 8-bit, 4-bit, sometimes lower. Memory falls in proportion: a 7-billion-parameter model takes about 14 GB at 16 bits and about 3.5 GB at 4 bits, plus working memory. Two approaches: post-training quantisation, applied to a finished model, sometimes with a small calibration dataset; and quantisation-aware training, which simulates the rounding during training and loses less. Well-known methods include GPTQ and AWQ; GGUF is the file format of the llama.cpp ecosystem, where names such as Q4_K_M indicate the bit width; the bitsandbytes library loads models in 8-bit or 4-bit form. Activations can be quantised as well as weights. Speed gains depend on hardware support for low-precision arithmetic. Pitfalls: quality that looks fine on benchmarks and fails on rare cases, and comparing models at different bit widths without saying so.

Written by Mehmet Erkek · Last updated: