MSCTGAN: Conformer Improvements for Speech Enhancement
Zhengyu Wu, Nengheng Zheng · 2024
Convolution-augmented transformer (Conformer) has shown impressive performance in speech enhancement. However, it employs a single fixed large kernel convolutional block for local feature modeling, which limits its attention to local features. Also, modeling the entire frequency band in the same manner hinders the model from focusing on more critical sub band information, resulting in low performance. To address these two issues, we propose a GAN-based approach called Multi-Scale Convolution-augmented Transformer GAN (MSCTGAN) for speech enhancement. MSCTGAN incorporates several stacks of MSCT modules to better capture local features. Taking into account the varying importance of speech sub-bands, these features are introduced in frequency domain relative position encoding and loss function, which facilitated the fusion of both full-band and sub-band information. The objective evaluation demonstrates the competitive performance of MSCTGAN compared to models utilizing conventional Transformer and Conformer architectures.