An End-to-End Speech Enhancement Framework Using Stacked Multi-scale Blocks
Tian Lan, Sen Li, Yilan Lyu, Chuan Peng, Qiao Liu · 2019
Speech enhancement is the task of removing additive noise from a speech signal, and the deep learning-based methods have recently seen great progress. However, it's difficult for the widely adopted short-time Fourier transform (STFT) to reconstruct target speech accurately with noisy phase. This work proposes a novel end-to-end speech framework to solve the problem. In the pre-processing stage, the time-domain signal is transformed into a two-dimensional feature representation, the speech enhancement module is used to enhance it, and the enhanced speech waveform is reconstructed by post-processing. In the speech enhancement module, we propose stacked multi-scale block (SMB) to capture more information of speech feature. To further improve the performance of the proposed model, the evaluation metrics are integrated into the loss function by using the multi-objective joint optimization training strategy. Experiments show that the proposed methods can significantly improve the ability of speech enhancement, and it also shows good performance at low SNRs.