28 Knowledge Distillation
Context
The best models are large or ensembles → expensive to deploy. Hinton, Vinyals and Dean carry the knowledge over into a small model.
The idea and the mechanism
The large teacher outputs not just a class but a full probability DISTRIBUTION, softened by a temperature T in the softmax. The small student learns to reproduce these soft targets (often alongside the ordinary hard labels). The "dark knowledge" is in the fact that a "2" looks a little like a "3" and nothing at all like a "7": that structure is present in the soft probabilities and absent from the one-hot label.
probability Temperature, soft targets and "dark knowledge"
A softmax with temperature T flattens the distribution:
At T = 1 this is the ordinary softmax (often close to one-hot); at T > 1 the distribution is "softer" and the relative odds of the wrong classes stand out — that is exactly the "dark knowledge". The student minimizes the divergence from the softened teacher (KL) plus the ordinary loss on the labels:
The T² factor compensates for the fact that gradients from a softened softmax scale as 1/T² — so that the two terms stay comparable. The student learns not "which class" but "how the teacher distributes its confidence" — and that is a far richer signal.
PyTorch The distillation loss
import torch.nn.functional as F
def distill_loss(student_logits, teacher_logits, y, T=4.0, alpha=0.7):
soft = F.kl_div(F.log_softmax(student_logits / T, -1),
F.softmax(teacher_logits / T, -1),
reduction='batchmean') * (T * T) # the teacher's soft target
hard = F.cross_entropy(student_logits, y) # hard labels
return alpha * soft + (1 - alpha) * hard
Why it matters
The basis of edge and mobile models and of LLM compression (DistilBERT and the like); the idea of "learning from the teacher's distribution" is used widely — right up to distilling huge LLMs into small ones. A striking fact from the paper: a student that never saw a single "3" in training classified threes almost perfectly — purely from the teacher's soft probabilities on the other digits.
Connections
Both are about ensembles, but in opposite directions: a forest assembles many models into one strong one; distillation compresses a large or ensemble teacher into a single small one. Distillation is often used for exactly that — packing an ensemble into one cheap network.
Modern small open models are frequently distilled from large ones: the LLaMA family and its descendants use distillation to carry the quality of the giants into models that fit on a single GPU. A 2015 idea became a key tool of the LLM era.
Distillation is simply a different target for ordinary training: the teacher's soft distribution instead of hard labels, while the mechanism (backprop through a KL loss) is the same. The novelty is in the signal, not in the way you train.
Questions worth asking
Why does a student learn better from soft targets than from the same data with hard labels?
A one-hot label carries ~log(K) bits per example and says nothing about the similarity of classes. The teacher's soft target is a dense vector over all classes: it conveys the geometry of the problem ("2 is closer to 3 than to 7"), effectively handing the student many "free" labels per example. That is a richer and less noisy signal, especially when data is limited.
Can the student beat the teacher?
Sometimes yes — in self-distillation (teacher and student sharing an architecture), or with "born-again networks" the student overtakes the teacher in places. The explanations vary: soft targets regularize, they smooth the labels, they transfer the dark knowledge. Not magic and not always, but a consistently observed effect that is still not fully explained in theory.
Distillation means running the teacher — doesn't that defeat the point if it is expensive?
The teacher is needed only during training; at inference only the cheap student runs — and inference is where the cost sits in production. On top of that, soft targets can be computed once and cached. So a one-off spend on the teacher pays for itself in a cheap deployment — which is exactly what distillation was invented for.
What to read in the original
The write-up plus the key passages is enough: temperature, soft targets and the experiment with the "missing" digit. It is worth getting a feel for the notion of dark knowledge from the original example.