BIUnet: Unet Network for Mask Estimation in Single-Channel Speech Enhancement

Qingge Fang, Guangcun Wei, Zhifei Pan, Jihua Fu · 2024

Deep learning-based single-channel speech enhancement techniques have demonstrated their effectiveness in enhancing speech quality. However, their performance, especially in low signal-to-noise ratio (SNR) environments, falls short of satisfaction, particularly in complex noise conditions. Therefore, we introduce a Bidirectional Long Short- Term Memory (Bi-LSTM) integrated Unet network (BIUnet), trained with Ideal Ratio Mask (IRM). By introducing a Bi-LSTM layer between the encoder and decoder, we aim to better capture contextual feature information. Additionally, we incorporate upsampling and downsampling (up- down sampling) modules at the original skip connections in the U net architecture for improved feature fusion. The experimental results demonstrate that BIU net achieves promising performance in enhancing speech on the public dataset. Under low SNR conditions, the proposed method exhibits significant superiority over several state-of-the-art (SOTA) single-channel speech enhancement techniques. Ablation experiments were conducted simultaneously to demonstrate the effectiveness of both the Bi- LSTM module and the up-down sampling module.

Read the paper · More papers on PaperTik