What is model distillation?
Model distillation is a training technique where a smaller "student" model learns to reproduce the behavior of a larger "teacher" model, rather than learning directly from raw data alone. The goal is a compact model that keeps as much of the teacher's capability as possible while needing far less memory and compute to run, which matters most at inference time, when a model gets deployed and queried repeatedly.
How the teacher-student setup works
In its classic form, the teacher model processes training examples and produces not just a final answer but a full probability distribution over possible next tokens, which carries more information than a single correct label would. The student model is trained to match that distribution as closely as it can, in addition to or instead of matching raw ground-truth labels. Because the teacher's output distribution encodes relationships the raw data doesn't state directly, such as which wrong answers are "less wrong" than others, the student can learn a compressed version of the teacher's behavior more efficiently than training a small model from scratch on the same data alone.
A more common version in current practice
A simpler variant, sometimes called data distillation, skips matching probability distributions and instead just has the teacher generate large volumes of example outputs, question-answer pairs, reasoning traces, or completions, which then become training data for the student. This is the version most publicly discussed open-weight small models use: a large model generates the training set, and a much smaller model is trained on it directly. It's less precise than matching full output distributions but far simpler to implement, since it only requires access to the teacher's outputs rather than its internal probabilities.
What distillation trades away
A distilled student is not a compressed copy of the teacher; it's a separate model, trained to approximate the teacher's behavior on the data it saw, and its actual capability depends heavily on the amount and quality of that training signal. Distilled models tend to do well on tasks similar to what they were trained to imitate and can fall short on tasks the distillation data didn't cover well. Evaluating a specific distilled model against your actual workload matters more than trusting its parameter count or its teacher's reputation. See evaluating an LLM before deploying it for how to run that check.
Distillation vs quantization vs fine-tuning
These three are often mentioned together but do different things. Quantization keeps a model's architecture and behavior fixed and just reduces the numeric precision of its existing weights. Fine-tuning continues training an existing model on new data to specialize it. Distillation produces a new, typically smaller model designed to approximate a larger one's general behavior. They're not mutually exclusive: a distilled model can be quantized afterward for further size reduction, and can also be fine-tuned for a specific task once distilled.
Why it matters for self-hosting
A distilled model that approximates a much larger teacher's behavior at a fraction of the parameter count is often the difference between a model that fits comfortably in a single node's memory and one that doesn't fit at all. GPUwerk's dedicated Spark nodes carry 128GB of unified memory per node, which comfortably runs many distilled and mid-sized open-weight models at higher precision than a much larger model would allow on the same hardware. See what is a parameter count for how size maps to memory needs, and the hardware page for full specifications.