Enhancing Model Robustness in Noisy Environments: Unlocking Advanced Mono-Channel Speech Enhancement With Cooperative Learning and Transformer Networks

Wei Hu, Yan Wu · IEEE Access · 2025

Enhancing mono-channel speech signals in scenarios involving complex spectral mapping poses a formidable challenge in the domain of audio signal processing. To tackle this challenge head-on, this paper introduces a novel rapid speech enhancement network that harnesses the combined strengths of Convolutional Neural Networks (ConvNets) and Transformers. The proposed network seamlessly integrates ConvNets for robust feature extraction and Transformers for comprehensive long-term sequence modeling, resulting in a substantial enhancement in speech quality. In the encoding-decoding layers of the network, we introduce the Unified Discovery Unit (UDU), employing cooperative learning techniques to augment feature extraction. Through collaborative learning across network layers, the UDU facilitates the extraction of richer feature spaces from speech data, thereby enhancing the network’s ability to capture pertinent information. Building upon recent advancements, we introduce the Temporal-Spectral Cross-Attention Component (TSCAC) in the transference layer. This innovative component, comprising temporal and spectral attention Transformers, adeptly processes both subband and entire band speech information, effectively capturing temporal and spectral dependencies to preserve crucial speech characteristics. Recognizing the significance of channel-specific features in speech signals, we integrate a channel attention segment and develop the teachable Twin-Segment Attention Integration Component (TSAIC). This component extracts semantic features from a spatial-channel perspective, further bolstering the network’s capacity to attend to relevant information. Additionally, we introduce a Gaussian-Weighted Progressive Structure (GWPS) to compensate for any lost detailed features in deeper network layers, thereby enhancing the fidelity of the enhanced speech signal and mitigating information loss. Extensive experimental evaluations conducted on datasets of varying scales and languages under diverse noise conditions underscore the effectiveness and robustness of the proposed method. Our approach showcases significant improvements in speech enhancement performance, paving the way for advancements in audio signal processing and real-world applications.

Read the paper · More papers on PaperTik