Stochastic Precision Ensemble: Self-Knowledge Distillation for Quantized Deep Neural Networks
The quantization of deep neural networks (QDNNs) has been actively studied\nfor deployment in edge devices. Recent studies employ the knowledge\ndistillation (KD) method to improve the performance of quantized networks. In\nthis study, we propose stochastic precision ensemble training for QDNNs (SPEQ).\nSPEQ is a knowledge distillation training scheme; however, the teacher is\nformed by sharing the model parameters of the student network. We obtain the\nsoft labels of the teacher by changing the bit precision of the activation\nstochastically at each layer of the forward-pass computation. The student model\nis trained with these soft labels to reduce the activation quantization noise.\nThe cosine similarity loss is employed, instead of the KL-divergence, for KD\ntraining. As the teacher model changes continuously by random bit-precision\nassignment, it exploits the effect of stochastic ensemble KD. SPEQ outperforms\nthe existing quantization training methods in various tasks, such as image\nclassification, question-answering, and transfer learning without the need for\ncumbersome teacher networks.\n
Paper
References (53)
Scroll for more · 38 remaining