Solvable Model for Inheriting the Regularization through Knowledge Distillation

In recent years the empirical success of transfer learning with neural\nnetworks has stimulated an increasing interest in obtaining a theoretical\nunderstanding of its core properties. Knowledge distillation where a smaller\nneural network is trained using the outputs of a larger neural network is a\nparticularly interesting case of transfer learning. In the present work, we\nintroduce a statistical physics framework that allows an analytic\ncharacterization of the properties of knowledge distillation (KD) in shallow\nneural networks. Focusing the analysis on a solvable model that exhibits a\nnon-trivial generalization gap, we investigate the effectiveness of KD. We are\nable to show that, through KD, the regularization properties of the larger\nteacher model can be inherited by the smaller student and that the yielded\ngeneralization performance is closely linked to and limited by the optimality\nof the teacher. Finally, we analyze the double descent phenomenology that can\narise in the considered KD setting.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC