Knowledge distillation has been used to transfer knowledge learned by a\nsophisticated model (teacher) to a simpler model (student). This technique is\nwidely used to compress model complexity. However, in most applications the\ncompressed student model suffers from an accuracy gap with its teacher. We\npropose extracurricular learning, a novel knowledge distillation method, that\nbridges this gap by (1) modeling student and teacher output distributions; (2)\nsampling examples from an approximation to the underlying data distribution;\nand (3) matching student and teacher output distributions over this extended\nset including uncertain samples. We conduct rigorous evaluations on regression\nand classification tasks and show that compared to the standard knowledge\ndistillation, extracurricular learning reduces the gap by 46% to 68%. This\nleads to major accuracy improvements compared to the empirical risk\nminimization-based training for various recent neural network architectures:\n16% regression error reduction on the MPIIGaze dataset, +3.4% to +9.1%\nimprovement in top-1 classification accuracy on the CIFAR100 dataset, and +2.9%\ntop-1 improvement on the ImageNet dataset.\n
Paper
References (86)
Scroll for more · 38 remaining