Dense for the Price of Sparse: Improved Performance of Sparsely Initialized Networks via a Subspace Offset

That neural networks may be pruned to high sparsities and retain high\naccuracy is well established. Recent research efforts focus on pruning\nimmediately after initialization so as to allow the computational savings\nafforded by sparsity to extend to the training process. In this work, we\nintroduce a new `DCT plus Sparse' layer architecture, which maintains\ninformation propagation and trainability even with as little as 0.01% trainable\nkernel parameters remaining. We show that standard training of networks built\nwith these layers, and pruned at initialization, achieves state-of-the-art\naccuracy for extreme sparsities on a variety of benchmark network architectures\nand datasets. Moreover, these results are achieved using only simple heuristics\nto determine the locations of the trainable parameters in the network, and thus\nwithout having to initially store or compute with the full, unpruned network,\nas is required by competing prune-at-initialization algorithms. Switching from\nstandard sparse layers to DCT plus Sparse layers does not increase the storage\nfootprint of a network and incurs only a small additional computational\noverhead.\n

Paper

References (27)

Scroll for more · 15 remaining

Similar papers

© 2026 NYSGPT2525 LLC