Data, Pre-Training & Post-Training · 6 min

Teaching Small Models from Big Ones: Distillation

How distillation trains a small student to copy a big teacher's output distribution — the original soft-label method, and what today's "distilled" open models actually do.

Here is the core move, plainly stated: instead of training a small model on your labeled data, you train it to imitate a big model's answers on that data. Not the label "cat" but the big model's full opinion: 86% cat, 12% dog, 2% car. That probability vector carries more information than the one-hot label, and a small model trained on it lands closer to the big model than the same small model trained on raw labels ever could.

Hinton, Vinyals, and Dean gave this its canonical form in 2015, and their MNIST result is still the cleanest demonstration. A large, well-regularized net made 67 test errors. A smaller net (two hidden layers of 800 units) trained the ordinary way on the same data made 146. The same small net, trained instead to match the large net's softened outputs, made 74. It recovered most of the gap without one change to its architecture. Same student, better teacher signal, roughly half the errors.

Why soft labels carry more

A hard label tells the student one bit: this is a 2. The teacher's distribution tells it how it's a 2, that this particular scrawl looks a little like a 7 and nothing like a 4. Hinton called these relative magnitudes over the wrong answers "dark knowledge." The teacher has learned a rich similarity structure across classes, and it leaks that structure through the tails of its output distribution.

The catch is that a confident teacher's tails are nearly zero, so the useful signal sits buried. You surface it with temperature. The softmax becomes

q_i = exp(z_i / T) / Σ_j exp(z_j / T)

with T=1 the normal softmax and higher T flattening the distribution. Watch what T does to a teacher's logits z = [4.0, 2.0, 0.5] over [cat, dog, car]:

T = 1  ->  [0.86, 0.12, 0.03]   "cat, obviously"
T = 4  ->  [0.49, 0.30, 0.21]   "cat, but dog-ish, definitely not car"

At T=1 the dog/car distinction is a rounding error. At T=4 it becomes a clean 1.5x, a signal the student can actually fit. You raise T on both teacher and student during training, then drop the student back to T=1 at inference.

The loss, and one gotcha

Distillation loss is usually a blend of two terms:

L = α · T² · KL(student_soft ‖ teacher_soft)  +  (1-α) · CE(student, hard_label)

The first term pulls the student toward the teacher's softened distribution (KL divergence at temperature T); the second keeps it honest against the true label when you have one. The T² factor is the piece people drop and then wonder why training stalls. Softening by T shrinks the gradients through the soft term by roughly 1/T², so multiplying by T² keeps that term's learning signal comparable to the hard-label term. Leave it out and your soft targets barely move the weights at high temperature.

Distillation is not quantization, and not plain SFT

Three words get used interchangeably and mean different things:

  • Quantization keeps every parameter but stores each in fewer bits, FP16 down to INT4. Same 8B weights, lower precision. (That's the next lesson.)
  • Distillation produces a model with fewer parameters. A 70B teacher, a 7B student: a genuinely smaller network, trained to behave like the big one.
  • SFT (supervised fine-tuning) trains on fixed (prompt, response) pairs, learning each token as a hard target. Classic distillation trains on the teacher's distribution over tokens, softer and richer per example.

They compose. You can distill a 7B student, quantize it to INT4, then SFT it on your domain. Orthogonal knobs.

What today's "distilled" open models actually do

Here is where the textbook and the model cards part ways, and it matters. The well-known open "distilled" reasoning models were not built with Hinton-style soft-label matching. The recipe was to have the big teacher generate roughly 800,000 worked examples, full chain-of-thought solutions, then fine-tune smaller off-the-shelf base models on that text with ordinary SFT. No logits, no temperature, no KL. The distilled 32B student still beat a strong reasoning baseline on math and code, purely from imitating generated traces.

That is sequence-level (black-box) distillation: the "labels" are the teacher's sampled outputs, and text is all you ever see. It is what you are forced into when the teacher lives behind an API. You cannot reach its logits, so you distill its behavior by training on what it writes.

WHITE-BOX (classic)              BLACK-BOX (modern open models)
teacher logits ──► KL loss        teacher text ──► SFT
  needs weights/logits              needs only samples
  richest signal, per-token         works through an API
  Hinton 2015                       "R1-Distill", most instruct clones

So when a model is labeled "distilled," ask which kind. Usually it means "SFT on a bigger model's generations," which sits closer to imitation learning than to the 2015 algorithm. Both are legitimately distillation, with very different data and access requirements.

Builder: Black-box is your default, since you rarely hold teacher logits. Spend your effort on the transfer set, not the loss function: diversity and coverage of prompts matter more than a clever objective. Keep the teacher's failures out. A wrong CoT trace teaches the student to be confidently wrong, so filter generations by verifiable correctness (unit tests, math checkers) before training on them.

What actually goes wrong

Capacity gap. Jump too far, 671B down to 1.5B, and the student cannot represent what the teacher knows. Distillation narrows the gap; it does not erase the fact that a tiny model has a lower ceiling. The small distilled reasoning students are visibly weaker than their teacher on hard problems.

You inherit the teacher's flaws, sharpened. Biases, hallucinated formats, jailbreak-susceptibility, even planted backdoors can travel through generated data. The student learns the teacher's blind spots as if they were ground truth, and you have laundered them into a fresh checkpoint with no obvious provenance.

Distribution mismatch. A student distilled on math prompts is unremarkable on dialogue it never saw. The transfer set is the curriculum; off-curriculum, the student is just its base model.

Licensing. Many API terms forbid using outputs to train competing models. Black-box distillation is technically trivial and contractually loaded, so read the ToS before you build a product on someone else's generations.

Defender: Treat unrestricted output access as a model-extraction surface. Sequence-level distillation is exactly how an attacker clones a proprietary model through its public API: query, collect, SFT. Rate limits, output watermarking, and anomaly detection on high-volume structured querying are your levers. Assume any capability you expose token-by-token can be cloned into a smaller model.
Researcher: Why soft labels beat hard ones is still unsettled. A regularization view (softer targets curb overconfidence), a privileged-information view (the teacher hands over inter-class structure the labels omit), and a gradient-variance-reduction view all have support, and none has closed the case. If you are measuring distillation, separate the two gains, richer per-example signal versus simply having more teacher-generated data, because black-box pipelines confound them.

Distillation is why the useful-model frontier keeps sliding down in size: capabilities first established at huge scale get compressed into models you can actually serve. Pair it with the next lesson on quantization. Fewer parameters and fewer bits are the two independent axes of shrinking a big model until it runs where you need it.

Sources