Weight, Block or Unit? Exploring Sparsity Tradeoffs for Speech Enhancement on Tiny Neural Accelerators

We explore network sparsification strategies with the aim of compressing\nneural speech enhancement (SE) down to an optimal configuration for a new\ngeneration of low power microcontroller based neural accelerators (microNPU's).\nWe examine three unique sparsity structures: weight pruning, block pruning and\nunit pruning; and discuss their benefits and drawbacks when applied to SE. We\nfocus on the interplay between computational throughput, memory footprint and\nmodel quality. Our method supports all three structures above and jointly\nlearns integer quantized weights along with sparsity. Additionally, we\ndemonstrate offline magnitude based pruning of integer quantized models as a\nperformance baseline. Although efficient speech enhancement is an active area\nof research, our work is the first to apply block pruning to SE and the first\nto address SE model compression in the context of microNPU's. Using weight\npruning, we show that we are able to compress an already compact model's memory\nfootprint by a factor of 42x from 3.7MB to 87kB while only losing 0.1 dB SDR in\nperformance. We also show a computational speedup of 6.7x with a corresponding\nSDR drop of only 0.59 dB SDR using block pruning.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC