Knowledge Distillation
A model-compression technique in which a smaller student model is trained to reproduce informative outputs or representations from a larger teacher model.
Machine-Learning Context
Knowledge distillation trains a student model to reproduce informative outputs, probabilities, or intermediate representations from a teacher. Soft targets can encode relationships among classes that a hard label does not expose, making distillation useful for model compression and constrained deployment.
Generalization Boundary
A student is not required by definition to be smaller in every dimension, and distillation can transfer teacher errors or biases along with useful behavior. Compression benefits must be evaluated against the target task.
Related Machine-Learning Concepts
- Model Compression
- Teacher Model
- Student Model
- Quantization
Direct source: The primary paper or official specification for Knowledge Distillation is linked here for verification.