DCCRN-SUBNET: A DCCRN and SUBNET Fusion Model for Speech Enhancement
Xin Yuan, Qun Yang, Shaohan Liu · 2021 7th International Conference on Computer and Communications (ICCC) · 2021
Currently, most of the speech enhancement methods can’t address the performance degradation problem caused by low signal-to-noise ratios (SNR) and non-stationary noises. For better speech enhancement at the above scenarios, this paper proposes a two-stage method that fuses DCCRN and SubNet. Compared with the single stage-stage networks, two-stage networks have more powerful mapping capabilities. This paper uses complex-valued spectrogram as the training target. In the first stage, the DCCRN takes the magnitude and phase as input and estimates corresponding target of clean speech. By simulating the complex-valued operation, the DCCRN can train the complex target effectively. However, it is still difficult to handle the low SNR and non-stationary noises. This paper uses the SubNet as the second stage network for better speech enhancement. In the second stage, the SubNet further refines the magnitude of target frequency by exploiting the context frequencies. Its input is consisted of magnitude of target frequency and several context frequencies. The output is the estimation of the clean speech magnitude target for the corresponding frequency. The experimental results show that the proposed method obtains better performance than other baseline models in terms of PESQ, STOI and SI-SDR.