Model customization

Distillation

Knowledge distillation

Training a small model, the student, to reproduce the behaviour of a large one, the teacher, so that most of the quality is kept at a fraction of the cost and response time.

Teacherlarge model · 93% agreement300,000 decisionswith their reasonsStudentsmall model · 91% agreementlistings it is unsure about still go to the teacherCosta twentieth per listingSpeedanswers five times faster

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/distillation

In plain terms

A senior specialist cannot handle every case personally, so she trains a junior: here are thousands of cases, here is what I decided, here is my reasoning. The junior will never know everything she knows, but on that kind of case comes close, works faster and costs less. Distillation is that apprenticeship between models. The large model's answers become the small model's lessons.

Why it matters

Distillation is how a capability that is too slow or too expensive at full size becomes affordable at volume. It is also how model providers produce their small, fast tiers. For a company it offers a route from an expensive prototype to a cheap production system. The costs: the student is only good on the kind of work it was shown, it inherits the teacher's mistakes, and it must be rebuilt when the teacher or the task changes. Check the terms of use too: providers commonly forbid using their model's outputs to train a competing model.

Example

An online marketplace checks 2 million listings a month against its content rules with a large model that agrees with human reviewers 93% of the time. It has that model judge 300,000 listings, with reasons, and trains a small model on the results. The student agrees with reviewers 91% of the time, costs a twentieth as much per listing and answers five times faster. Listings it is unsure about still go to the teacher.

Most often confused with

Distillation vs. Quantization

DistillationTrains a new, smaller model to imitate a larger one
QuantizationKeeps the same model and stores its numbers more coarsely

Both make a model cheaper to run, by different routes. Distillation produces a different model with fewer parameters, and needs data and a training run. Quantisation changes no architecture and needs no training in its common form: the same weights are saved with fewer bits. Distillation can shrink a model far more; quantisation takes minutes. They are routinely combined: distil first, then quantise the student.

Origin: Introduced under this name by Geoffrey Hinton, Oriol Vinyals and Jeff Dean in the 2015 paper “Distilling the Knowledge in a Neural Network”.

Under the hood

In the original formulation the student is trained on the teacher's full output probabilities (soft targets), softened with a temperature, because the relative weight a teacher gives to wrong answers carries information that a single correct label does not. That needs access to the teacher's internals. With language models reached through an API, distillation usually means something simpler: generate a large set of teacher answers, often with reasoning, and fine-tune the student on them as synthetic training data. Variants match intermediate layers as well as outputs, or use several teachers. Evaluate the student on real held-out cases; teacher-labelled examples alone are not enough. A typical pattern in production: the student handles routine volume and passes low-confidence cases up to the teacher. Managed distillation services exist on cloud platforms such as Amazon Bedrock. The word also describes diffusion models reduced to a few generation steps.

Written by Mehmet Erkek · Last updated: