What Is Model Distillation?
Distillation trains a small "student" model to reproduce the behaviour of a large "teacher", so the student approaches the teacher’s quality at a fraction of the cost. The key is that the student learns from the teacher’s full output distribution — the probabilities it assigns across options — not just the single right answer. Those soft targets carry more information than a hard label, which is how a much smaller model captures a surprising amount of a larger one’s competence.
Soft targets carry dark knowledge
A hard label says "this is a cat". A teacher’s distribution says "mostly cat, a little lynx, almost never car" — and that relative structure, which Hinton called dark knowledge, tells the student how classes relate, not just which is correct. Training on it transfers more than labels alone, letting a compact student generalise better than the same model trained from scratch on hard labels.
The honest limit is that a student inherits the teacher’s errors and biases along with its skill, and cannot exceed the teacher on what the teacher cannot do. Distillation is compression of a capability that already exists, not a route to new capability — which is exactly why it matters for running capable models cheaply enough to power an agent fleet.
Related standards
Questions
Can a distilled model beat its teacher?
Generally no on the teacher’s strengths — it approximates the teacher. It can be far cheaper and faster for close to the quality.
Is distillation the same as quantization?
No — distillation trains a smaller model; quantization lowers the numerical precision of an existing one. They are often combined.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.