MN-Net: Speech Enhancement Network via Modeling the Noise

Ying Hu, Qin Yang, Wenbing Wei, Li Lin, Liang Ju He, Zhijian Ou, Wenzhong Yang · IEEE Transactions on Audio Speech and Language Processing · 2025

Currently, deep learning-based speech enhancement methods generally focus on target speech extraction while neglecting modeling the other sound sources in the mixture. These methods still can't distinguish the target speech from the interference well. In this paper, we present a monaural speech enhancement network via Modeling the Noise (MN-Net), which includes a shared Encoder and three separate Decoders for parallel modeling the magnitude and phase spectrogram of target speech, and the complex spectrogram of noise. Specifically, we propose a Multi-Branch Feature Extractor (MBFE) module to capture the richer contextual information in mixture, and a Spatial Reconstruction Unit (SRU) to remove the redundancy from extracted features. We compared our proposed MN-Net with 18 classical speech enhancement methods on the VoiceBank+DEMAND dataset, and with 9 ones on DNS-Challenge dataset for denoising task, and with 7 ones on the WHAMR! dataset for simultaneous denoising & de-reverberation task. Our proposed MBFE module was applied to two classical speech enhancement methods, DB-AIAT and CMGAN, replacing their DenseBlocks module. The results demonstrate that applying the MBFE module can boost their performances while keeping smaller model size. A series of visualization analysis intuitively verify that modeling the noise can enable the network to distinguish the target speech from noise and other interference more accurately.

Read the paper · More papers on PaperTik