# Why distillation matters Large frontier models are powerful but expensive and hard to deploy. Model distillation is a technique for producing much smaller models that retain a surprising share of a larger model's behavior. The goal is practical: make high-quality models cheaper, faster, or able to run in constrained environments.
# What a model learns vs. what labels provide
# Classical distillation in plain terms Classical distillation replaces or supplements hard-label training with the teacher's output distributions. The student is trained to match those distributions as well as the original labels. Two practical elements matter:
- Temperature scaling: raising the temperature flattens the teacher's output distribution so the student can see which alternative classes the teacher finds plausible. Without it, distributions are too peaked to be informative.
- Blended loss: the student's objective mixes a term that matches the teacher's softened outputs with a term that enforces correctness against the ground-truth labels.
Hinton, Vinyals, and Dean framed this approach and showed small models could match much larger ensembles by inheriting the teacher's generalization patterns.
# Why LLMs change the problem
# Modern distillation approaches for language models Distillation for LLMs has evolved into three main families:
- Synthetic data distillation: the teacher generates large volumes of high-quality text. Students train on that synthetic dataset to learn the teacher's behaviors without directly matching enormous per-token distributions. This is now the dominant approach.
- Logit-based distillation: students attempt to match the teacher's token-level logits directly, often with adaptations to handle the massive output space.
Each approach trades off storage, compute, and the fidelity of transferred behavior. Synthetic data scales well because it converts the sequential, high-dimensional problem into a standard supervised learning dataset.
# Why distillation is contested Distilling a large, queryable model at scale can be done without the teacher owner's cooperation. That has triggered controversy because it lowers the barrier for competitors to reproduce capabilities the original provider invested to create. The tension is structural: public availability and query access enable a teacher to act as a data source that others can use to train cheaper models, which creates commercial and policy disputes.
# Practical takeaway for practitioners If you need a smaller model with teacher-like behavior, distillation is a proven path. For classification problems, classical distillation with temperature scaling is straightforward. For language models, expect to choose between generating synthetic data, matching internal features, or attempting logit matching, based on your compute budget and how closely you need to reproduce the teacher's outputs.
# Further reading The material traces the original distillation framework and its adaptation to modern LLM workflows. It also lays out why large-scale, unauthorized distillation has become a key industry flashpoint.