Model compression using knowledge distillation with integrated gradients

Model compression is critical for deploying deep learning models on resource-constrained devices. We introduce a novel method enhancing knowledge distillation with Integrated Gradients (IG) as a data augmentation strategy. Our approach overlays IG maps onto input images during training, providing student models with deeper insights into teacher models’ decision-making processes. Extensive evaluation on CIFAR-10 demonstrates that our IG-augmented knowledge distillation achieves 92.6% testing accuracy with a 4.1x compression factor–a significant 1.1 percentage point improvement (p < 0.001) over non-distilled models (91.5%). This compression reduces inference time from 140ms to 13ms. Our method precomputes IG maps before training, transforming substantial runtime costs into a one-time preprocessing step. We validate our approach through systematic ablation studies with attention transfer, comprehensive compression factor analysis (2.2\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\times$$\end{document}–1122\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\times$$\end{document}), Monte Carlo simulations, and cross-dataset evaluation on ImageNet subsets, demonstrating that IG-based knowledge distillation consistently outperforms conventional approaches. Our results establish this framework as a viable compression technique for real-world deployment on edge devices while maintaining competitive accuracy.

Paper

Similar papers

© 2026 NYSGPT2525 LLC