Two-Stage Speech Enhancement Network Based on Complex Spectrum Mapping and Voice Activity Detection

Wang Chen, Yi Xin Zhou, Liwen Tan, Yin Liu · 2024

In the multi-stage method, each stage model only focuses on one task, and the collaborative learning of multiple tasks often achieves better results than the single-stage model. Inspired by this, this paper proposes a two-stage speech enhancement model based on complex spectrum mapping and voice activity detection(VAD). In the first stage, the model directly predicts the real and imaginary parts of the clean speech. In the second stage, the model learns the frame-level information, and then applies it to the predicted complex spectrum of the first stage. In order to extract richer time-frequency information in the encoder, this paper adopts multi-scale gated convolution. At the same time, Conformer is used in the bottleneck layer to realize global and local attention calculation in the time dimension and frequency dimension. Through this design, the model proposed in this paper can effectively enhance the speech signal and accurately detect voice activity. Finally, the experimental results on the Voicebank+Demand dataset show that the model proposed in this paper is superior to other comparison models in terms of objective and subjective indicators. At the same time, in order to further verify the generalization ability of the model, the results on the Librispeech dataset show that the model proposed in this paper has good results under different noises and signal-to-noise ratios.

Read the paper · More papers on PaperTik