Optimizing speech representation learning for enhanced noise robustness in downstream applications
Dianwen Ng · 2025
The primary objective of this thesis is to enhance the effectiveness and efficiency of speech representations, specifically improving noise robustness for downstream applications. Current speech representation learning frameworks, despite their advanced foundational knowledge and powerful speech understanding capabilities, fall short in critical areas essential for real-world applications, such as adaptability to noise, expressiveness, and computational efficiency. For instance, their performance varies greatly across different levels of noise corruption, showing high vulnerability to distortion from external influences. Moreover, learning for domain adaptation and performing inference are computationally intensive due to the large scale and complex design of the model structures. Hence, this thesis introduces innovative solutions to bridge these gaps, offering substantial improvements over existing methods. Adaptability to Noise: We address noise robustness by integrating Barlow Twins learning, an advanced regularization technique that strategically reduces channel redundancy and effectively disentangles speech features, as presented in chapter 3. This ensures that our models capture essential, noise-free features critical for accurate speech recognition, even in adverse acoustic environments, thereby ensuring more stable and robust performance to noise distortion. Expressiveness: To enhance the expressivity of pre-trained models, we employ a parameter-efficient fine-tuning approach that incorporates the proposed deep filter tuning with Feature-wise Linear Modulation (FiLM)-inspired integration, as detailed in Chapter 4. This method refines the handling of speech nuances, allowing the models to more effectively express important information through targeted feature extraction from FiLM. This adaptation better accommodates diverse vocal attributes and acoustic variations. Efficiency: Our proposed deep filter tuning strategy enables efficient adaptation of frozen pre-trained models with the aim of minimizing the number of trainable parameters. This approach optimizes computational resources and maximizes efficiency, particularly when faced with constraints in computational memory. In addition to the work, we also innovate in data compression through multi-band quantization, optimizing bit allocation by prioritizing perceptually significant features in chapter 5. This approach leverages psychoacoustic principles to enhance the efficiency of speech processing and is particularly beneficial for speech synthesis and voice conversion tasks, ensuring high fidelity in the outputs. Our methods have been tested across tasks such as automatic speech recognition and speech reconstruction, demonstrating significant enhancements in accuracy, robustness, and processing efficiency. These improvements underscore the potential of our approaches to revolutionize speech representation learning, making it more adaptive, scalable, and context-aware. By addressing the inherent limitations of current models, this research advances the field of speech representation learning, setting a new benchmark for future developments and applications in adaptive and efficient speech technology.