A Two-Stage Deep Neural Network with Bounded Complex Ideal Ratio Masking for Monaural Speech Enhancement
Xiang Yan, Bing Han, Zhigang Su · 2022
Complex Ideal Ratio Mask (cIRM) estimation is an effective approach for monaural speech enhancement. However, the cIRM is extremely hard to estimate without a care designed Deep Neural Network (DNN). This work proposes a Two-stage Speech Enhancement neural Network (TSENet), where each stage is a mixed network of the U-Net and Transformer. In U-Net, we propose a Half Batch Normalization (HBN) mechanism to reserve more scale information and keep data distribution invariant as well. In addition, we deploy the stacked 2D Gated Attention Units (2D-GAUs) in U-Net to prevent the noise components from seeping into the U-Net decoder. To enrich the features information, we apply transformer blocks to capture the frequency correlations in spectrogram. Experiment results on Librispeech dataset verify that the TSENet can achieve high speech perceptual quality and intelligibility under various noise conditions. Furthermore, we compare our TSENet with other state-of-the-art baselines on VTCK dataset, and our TSENet outperforms their performance by a considerable margin.