Quantisation and Pruning for Neural Network Compression and Regularisation

Deep neural networks are typically too computationally expensive to run in\nreal-time on consumer-grade hardware and low-powered devices. In this paper, we\ninvestigate reducing the computational and memory requirements of neural\nnetworks through network pruning and quantisation. We examine their efficacy on\nlarge networks like AlexNet compared to recent compact architectures:\nShuffleNet and MobileNet. Our results show that pruning and quantisation\ncompresses these networks to less than half their original size and improves\ntheir efficiency, particularly on MobileNet with a 7x speedup. We also\ndemonstrate that pruning, in addition to reducing the number of parameters in a\nnetwork, can aid in the correction of overfitting.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC