Self-supervised learning (SSL) has garnered significant attention in speech processing, particularly excelling in linguistic tasks such as speech recognition. However, improving the performance of pre-trained models across various downstream tasks—each requiring distinct types of speech information—remains a significant challenge. To address this, we propose a progressive residual extraction based SSL method, named <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">PROGRE</small>. Specifically, we introduce two lightweight, specialized task modules into an encoder-style SSL backbone to enhance its ability to extract pitch variation and speaker information from speech. Furthermore, to mitigate the incompatibility between the reinforced pitch variation and speaker information and the learning of content information, we employ residual extraction, leveraging the extracted representations as references or conditioning signals to guide the subsequent modules in more effectively learning content-related information under the supervision of HuBERT-based speech masking prediction. In this manner, we can incrementally extract pitch variation, speaker, and content representations from the input speech. Finally, these multiple representations, each capturing diverse speech information, are combined using different layer weights to produce task-specific representations for various downstream tasks. Experimental results demonstrate that our <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">PROGRE</small> achieves significant performance improvements across several tasks, such as speaker identification, speech recognition, emotion recognition, speech enhancement, and voice conversion, outperforming excellent SSL methods like wav2vec2.0, HuBERT, and WavLM.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex