Era 4 · Architectures and scale · 2015

28 Knowledge Distillation

Distilling the Knowledge in a Neural Network · Hinton, Vinyals & Dean · Google
🟧 read selectively~40 minoriginal ↗
The gist in 20 seconds. A small "student" trains on the softened probabilities of a "teacher" rather than on hard labels alone. The "dark knowledge" — the relative probabilities of the wrong classes — carries more information than a one-hot label. The student reaches the teacher's quality at a fraction of the size. The basis of model compression.

Context

The best models are large or ensembles → expensive to deploy. Hinton, Vinyals and Dean carry the knowledge over into a small model.

The idea and the mechanism

The large teacher outputs not just a class but a full probability DISTRIBUTION, softened by a temperature T in the softmax. The small student learns to reproduce these soft targets (often alongside the ordinary hard labels). The "dark knowledge" is in the fact that a "2" looks a little like a "3" and nothing at all like a "7": that structure is present in the soft probabilities and absent from the one-hot label.

probability Temperature, soft targets and "dark knowledge"

A softmax with temperature T flattens the distribution:

pi = ezi/TΣj ezj/T

At T = 1 this is the ordinary softmax (often close to one-hot); at T > 1 the distribution is "softer" and the relative odds of the wrong classes stand out — that is exactly the "dark knowledge". The student minimizes the divergence from the softened teacher (KL) plus the ordinary loss on the labels:

L = α·T²·KL(pteacherT ‖ pstudentT) + (1−α)·CE(y, pstudent)

The T² factor compensates for the fact that gradients from a softened softmax scale as 1/T² — so that the two terms stay comparable. The student learns not "which class" but "how the teacher distributes its confidence" — and that is a far richer signal.

PyTorch The distillation loss
import torch.nn.functional as F

def distill_loss(student_logits, teacher_logits, y, T=4.0, alpha=0.7):
    soft = F.kl_div(F.log_softmax(student_logits / T, -1),
                    F.softmax(teacher_logits / T, -1),
                    reduction='batchmean') * (T * T)   # the teacher's soft target
    hard = F.cross_entropy(student_logits, y)           # hard labels
    return alpha * soft + (1 - alpha) * hard
teacher (large) student (small) soft targetssoftmax(z/T) student ≈ teacher, but many times smaller
The teacher outputs a softened distribution (soft targets); the student learns to reproduce it — picking up the "dark knowledge" about how the classes resemble one another.
Analogy. An experienced doctor does not simply tell the trainee "this is flu", but adds: "looks like flu, though it slightly resembles covid and is nothing like an allergy". Those hedges about resemblance are the valuable knowledge — they teach the trainee to discriminate finely rather than memorize labels. The temperature T is how much detail the doctor goes into about their doubts.

Why it matters

The basis of edge and mobile models and of LLM compression (DistilBERT and the like); the idea of "learning from the teacher's distribution" is used widely — right up to distilling huge LLMs into small ones. A striking fact from the paper: a student that never saw a single "3" in training classified threes almost perfectly — purely from the teacher's soft probabilities on the other digits.

Connections

↔ close relative12. Random Forests

Both are about ensembles, but in opposite directions: a forest assembles many models into one strong one; distillation compresses a large or ensemble teacher into a single small one. Distillation is often used for exactly that — packing an ensemble into one cheap network.

→ applied to50. LLaMA

Modern small open models are frequently distilled from large ones: the LLaMA family and its descendants use distillation to carry the quality of the giants into models that fit on a single GPU. A 2015 idea became a key tool of the LLM era.

← builds on7. Backpropagation

Distillation is simply a different target for ordinary training: the teacher's soft distribution instead of hard labels, while the mechanism (backprop through a KL loss) is the same. The novelty is in the signal, not in the way you train.

Questions worth asking

Why does a student learn better from soft targets than from the same data with hard labels?

A one-hot label carries ~log(K) bits per example and says nothing about the similarity of classes. The teacher's soft target is a dense vector over all classes: it conveys the geometry of the problem ("2 is closer to 3 than to 7"), effectively handing the student many "free" labels per example. That is a richer and less noisy signal, especially when data is limited.

Can the student beat the teacher?

Sometimes yes — in self-distillation (teacher and student sharing an architecture), or with "born-again networks" the student overtakes the teacher in places. The explanations vary: soft targets regularize, they smooth the labels, they transfer the dark knowledge. Not magic and not always, but a consistently observed effect that is still not fully explained in theory.

Distillation means running the teacher — doesn't that defeat the point if it is expensive?

The teacher is needed only during training; at inference only the cheap student runs — and inference is where the cost sits in production. On top of that, soft targets can be computed once and cached. So a one-off spend on the teacher pays for itself in a cheap deployment — which is exactly what distillation was invented for.

What to read in the original

The write-up plus the key passages is enough: temperature, soft targets and the experiment with the "missing" digit. It is worth getting a feel for the notion of dark knowledge from the original example.