Progressive Residual Extraction Based Pre-Training for Speech Representation Learning

Tianrui Wang, Jin Li, Ziyang Ma, Rui Cao, Xie Chen, Longbiao Wang, Meng Ge, Xiaobao Wang, Yu Guang Wang, Jianwu Dang, Nyima Tashi · IEEE Transactions on Audio Speech and Language Processing · 2025

Self-supervised learning (SSL) has garnered significant attention in speech processing, particularly excelling in linguistic tasks such as speech recognition. However, improving the performance of pre-trained models across various downstream tasks—each requiring distinct types of speech information—remains a significant challenge. To address this, we propose a progressive residual extraction based SSL method, namedPROGRE. Specifically, we introduce two lightweight, specialized task modules into an encoder-style SSL backbone to enhance its ability to extract pitch variation and speaker information from speech. Furthermore, to mitigate the incompatibility between the reinforced pitch variation and speaker information and the learning of content information, we employ residual extraction, leveraging the extracted representations as references or conditioning signals to guide the subsequent modules in more effectively learning content-related information under the supervision of HuBERT-based speech masking prediction. In this manner, we can incrementally extract pitch variation, speaker, and content representations from the input speech. Finally, these multiple representations, each capturing diverse speech information, are combined using different layer weights to produce task-specific representations for various downstream tasks. Experimental results demonstrate that ourPROGREachieves significant performance improvements across several tasks, such as speaker identification, speech recognition, emotion recognition, speech enhancement, and voice conversion, outperforming excellent SSL methods like wav2vec2.0, HuBERT, and WavLM.

Read the paper · More papers on PaperTik