Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q34IntermediateConcept

What is knowledge distillation for LLMs, and when would you use it?

30-second answerSay your answer out loud first, then reveal.

Types

TypeHowNeeds
Logit (soft-label) distillationMinimise KL between student and teacher token distributionsAccess to teacher logits (open models)
Sequence-level / synthetic dataTeacher generates responses (or reasoning traces); student is fine-tuned on themOnly teacher outputs
On-policy distillationStudent generates; teacher scores or corrects the student's own outputsTeacher access during training
Reasoning distillationTrain small models on long chain-of-thought traces from a reasoning modelTeacher traces

Why soft labels help: the teacher's full distribution carries "dark knowledge". It says how wrong each alternative is, which gives a richer signal than one correct label.

Typical use case. Classifying 10M support tickets per day. A frontier model labels 50K examples, and a fine-tuned 3B model reproduces ~95% of its accuracy at a small fraction of the cost and latency.

Caveats

  • The student is capped near the teacher's ability on that distribution and may fail outside it.
  • Licence / terms: many API providers' terms restrict using outputs to train competing models. Check before distilling.
  • Quality of the synthetic data matters: filter, deduplicate, verify (especially for reasoning).

This is what real progress feels like.