What Is Model Distillation?

Distillation trains a small "student" model to reproduce the behaviour of a large "teacher", so the student approaches the teacher’s quality at a fraction of the cost. The key is that the student learns from the teacher’s full output distribution — the probabilities it assigns across options — not just the single right answer. Those soft targets carry more information than a hard label, which is how a much smaller model captures a surprising amount of a larger one’s competence.

Transfers
A large teacher’s behaviour to a small student
Signal
The teacher’s full probability distribution
Introduced
Hinton et al., 2015

Soft targets carry dark knowledge

A hard label says "this is a cat". A teacher’s distribution says "mostly cat, a little lynx, almost never car" — and that relative structure, which Hinton called dark knowledge, tells the student how classes relate, not just which is correct. Training on it transfers more than labels alone, letting a compact student generalise better than the same model trained from scratch on hard labels.

The honest limit is that a student inherits the teacher’s errors and biases along with its skill, and cannot exceed the teacher on what the teacher cannot do. Distillation is compression of a capability that already exists, not a route to new capability — which is exactly why it matters for running capable models cheaply enough to power an agent fleet.

Related standards

Hinton et al., 2015 — Distilling the Knowledge in a Neural Network

Questions

Can a distilled model beat its teacher?

Generally no on the teacher’s strengths — it approximates the teacher. It can be far cheaper and faster for close to the quality.

Is distillation the same as quantization?

No — distillation trains a smaller model; quantization lowers the numerical precision of an existing one. They are often combined.

Keep reading

related
What Is Model Quantization?
related
What Is Fine-Tuning?
related
What Is a Mixture-of-Experts Model?
related
What Is a Transformer (Neural Network Architecture)?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
Retrieval-Augmented Generation (RAG) for Agents
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
Managing an Agent’s Context Window
Agent Data & Memory
What Is a Vector Database?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is an Embedding?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is the Attention Mechanism?
Agent Data & Memory
What Is Chain-of-Thought Prompting?
Agent Data & Memory
What Is RLHF (Reinforcement Learning from Human Feedback)?
Agent Data & Memory
What Is Tokenization in AI (Subword Tokens)?
Agent Data & Memory
What Is Semantic Search?
Agent Data & Memory
Transformer vs RNN: What Changed?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?
Agent Data & Memory
Supervised vs Unsupervised Learning: What’s the Difference?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.