Cascade Weight Shedding in Deep Neural Networks: Benefits and Pitfalls for Network Pruning

We report, for the first time, on the cascade weight shedding phenomenon in\ndeep neural networks where in response to pruning a small percentage of a\nnetwork's weights, a large percentage of the remaining is shed over a few\nepochs during the ensuing fine-tuning phase. We show that cascade weight\nshedding, when present, can significantly improve the performance of an\notherwise sub-optimal scheme such as random pruning. This explains why some\npruning methods may perform well under certain circumstances, but poorly under\nothers, e.g., ResNet50 vs. MobileNetV3. We provide insight into why the global\nmagnitude-based pruning, i.e., GMP, despite its simplicity, provides a\ncompetitive performance for a wide range of scenarios. We also demonstrate\ncascade weight shedding's potential for improving GMP's accuracy, and reduce\nits computational complexity. In doing so, we highlight the importance of\npruning and learning-rate schedules. We shed light on weight and learning-rate\nrewinding methods of re-training, showing their possible connections to the\ncascade weight shedding and reason for their advantage over fine-tuning. We\nalso investigate cascade weight shedding's effect on the set of kept weights,\nand its implications for semi-structured pruning. Finally, we give directions\nfor future research.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC