1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is knowledge distillation for LLMs, and when would you use it?
30-second answerSay your answer out loud first, then reveal.
Types
| Type | How | Needs |
|---|---|---|
| Logit (soft-label) distillation | Minimise KL between student and teacher token distributions | Access to teacher logits (open models) |
| Sequence-level / synthetic data | Teacher generates responses (or reasoning traces); student is fine-tuned on them | Only teacher outputs |
| On-policy distillation | Student generates; teacher scores or corrects the student's own outputs | Teacher access during training |
| Reasoning distillation | Train small models on long chain-of-thought traces from a reasoning model | Teacher traces |
Why soft labels help: the teacher's full distribution carries "dark knowledge". It says how wrong each alternative is, which gives a richer signal than one correct label.
Typical use case. Classifying 10M support tickets per day. A frontier model labels 50K examples, and a fine-tuned 3B model reproduces ~95% of its accuracy at a small fraction of the cost and latency.
Caveats
- The student is capped near the teacher's ability on that distribution and may fail outside it.
- Licence / terms: many API providers' terms restrict using outputs to train competing models. Check before distilling.
- Quality of the synthetic data matters: filter, deduplicate, verify (especially for reasoning).
Related
This is what real progress feels like.