Distillation
Also: Knowledge distillation · Teacher-student training
Training a smaller model to reproduce the behaviour of a larger one, so the capability survives at a fraction of the cost to run.
A large model generates outputs, and a smaller one is trained to match them rather than to match the original training labels. What transfers is not the weights but the behaviour, including the shape of the larger model's uncertainty — which is the part that plain supervised training on the same data does not capture.
Why it matters more than scale now
Capability per parameter has been rising faster than parameter counts, and distillation is a large part of the reason. A model small enough to serve cheaply, or to run on a laptop, is worth more in most deployments than one that is marginally better and ten times the cost per token.
Where it breaks
- The student inherits the teacher's errors, confidently. A wrong answer produced fluently is exactly what the student is being trained to reproduce.
- Licence terms often prohibit training on a model's outputs. The constraint is legal rather than technical, and it is the one most frequently overlooked.
- Losses concentrate on reasoning rather than recall — the same asymmetry seen in quantisation. A distilled model that answers factual questions well may still fail multi-step problems its teacher handled.
Distillation is not compression of a file. It is retraining, and the result is a different model that behaves similarly, which means it has to be evaluated as a different model.