MambaGAN: Mamba based Metric GAN for Monaural Speech Enhancement
Tianhao Luo, Feng Zhou, Zhongxin Bai · 2024
Encoder-decoder structures are widely used in deep neural network-based speech enhancement (SE), often utilizing convolutional and transformer modules as basic components. The high computational demands of transformer often limit their real-time performance. To address this issue, we propose a novel speech enhancement network called MambaGAN by combining MambaFormer and ODConv within a GAN framework. MambaFormer is a structure based on Mamba to replace transformer in SE networks. Additionally, an Omni-dimensional Dynamic Convolution (ODConv) is introduced to replace convolutional modules for capturing richer speech features more flexibly. Experimental results on the VoiceBank+DEMAND dataset show that MambaGAN achieved an impressive PESQ score of 3.56. When combined with perceptual contrast stretching, it achieved a new state-of-the-art PESQ score of 3.72, while exhibiting lower computational complexity than existing conformer-based models.