Existing speech separation methods utilize deep neural networks with a fixed computation graph in the training and inference phases. During inference, these networks cannot be easily re-configured to adapt to varying resource constraints. To address this challenge, we present Slim-TasNet, a slimmable neural network for speech separation, which is capable of achieving dynamic inference by utilizing different subsets of its computation graph. Specifically, network slimming is achieved by adjusting the width of the hidden layers at runtime based on an input utilization factor. Experimental evaluation shows that Slim-TasNet can be executed at different widths on the fly while maintaining similar performance to the corresponding static networks trained from scratch. We also compare the performance of Slim-TasNet with several network pruning mechanisms and demonstrate the effectiveness of the proposed method.
Paper
Full text
Slim-Tasnet: A Slimmable Neural Network for Speech Separation
Semantic Scholar · Computer Science · 2023
Abstract
Existing speech separation methods utilize deep neural networks with a fixed computation graph in the training and inference phases. During inference, these networks cannot be easily re-configured to adapt to varying resource constraints. To address this challenge, we present Slim-TasNet, a slimmable neural network for speech separation, which is capable of achieving dynamic inference by utilizing different subsets of its computation graph. Specifically, network slimming is achieved by adjusting the width of the hidden layers at runtime based on an input utilization factor. Experimental evaluation shows that Slim-TasNet can be executed at different widths on the fly while maintaining similar performance to the corresponding static networks trained from scratch. We also compare the performance of Slim-TasNet with several network pruning mechanisms and demonstrate the effectiveness of the proposed method.