Home › Fine-Tuning & Alignment › Distillation
🎯 · Models

Distillation

Train a small, cheap student model to mimic a larger, expensive teacher model.

In one line

Compress a big model's behavior into a smaller one that is cheaper and faster to serve.

ConceptWhat it is

Distillation trains a smaller student model to reproduce the outputs, or the output probability distributions, of a larger teacher model, transferring much of the teacher's capability into a model that is cheaper and faster to run in production. It exists because serving a frontier-scale model for every request is often not economically or latency-wise viable, while a distilled student can approximate its behavior at a fraction of the cost.

Distillation can target either matching final answers, sometimes called behavioral cloning through generated data, or matching the teacher's soft output distributions directly during training.

How it worksThe mechanics

The large teacher model generates outputs, or soft labels over the vocabulary, for a broad set of prompts; a smaller student model is then trained, either with standard supervised loss against the teacher's generated answers or with a distillation loss that matches the teacher's output distribution, until the student approximates the teacher's behavior at a much lower parameter count.

At a glanceSee it

Distillation diagram
Distillation diagram 1

Whether you can read the teacher internals splits distillation into a white-box logit-matching path and a black-box synthetic-data path — and either way the student is capped by its teacher.

Distillation diagram 2

Dividing the teacher logits by a temperature exposes the dark knowledge of which wrong answers are plausible — teaching the student class similarity, not just the single correct label.

When to use itWhere it fits

  • Reducing serving cost and latency for high-volume production endpoints.
  • Deploying capable models on constrained hardware, like edge devices or mobile.
  • Creating a fast, cheap specialist model for one narrow task learned from a general-purpose frontier model.
  • Bootstrapping a smaller open-weight model's quality using a stronger proprietary teacher's outputs.

When NOT to use itLimits & anti-patterns

  • Tasks needing the teacher's full breadth of capability, where a distilled student will underperform on anything outside its training distribution.
  • When only a handful of production queries need the capability, where simply calling the large model directly is cheaper than building a distillation pipeline.
  • Licensing-restricted teacher models where generating training data from their outputs may violate terms of use.

Trade-offsAdvantages & costs

Advantages
  • Dramatically cuts inference cost and latency versus serving the full teacher model.
  • Enables deployment on constrained or edge hardware.
  • Can transfer a substantial fraction of teacher capability for narrow, well-defined tasks.
  • Combines well with quantization for further serving efficiency.
Trade-offs & costs
  • Student performance ceiling is bounded by and typically below the teacher's.
  • Generating enough high-quality teacher outputs for training data is itself costly.
  • Struggles to transfer capabilities outside the specific tasks it was distilled on.
  • Some teacher providers restrict using their outputs to train competing models.

ExampleIn the real world

DeepSeek released distilled versions of its reasoning model into smaller Llama and Qwen-based checkpoints, giving developers much of the reasoning behavior at a fraction of the serving cost.

ToolsHow to implement it

  • Hugging Face TRLsupports supervised distillation training loops on teacher-generated data.
  • vLLMhigh-throughput serving engine for generating large volumes of teacher outputs efficiently.
  • DistilBERT-style recipesestablished reference approach for logit-matching distillation losses.
  • Axolotlconfig-driven fine-tuning framework usable for training student models on distilled data.

Cost & effortWhat it takes

Moderate upfront cost to generate teacher outputs at scale, but large ongoing savings from cheaper student inference; meaningful engineering effort to build the data generation and training pipeline once.

A living map of modern AI — kept current every morning