Monaural Speech Enhancement with Deep Residual-Dense Lattice Network and Attention Mechanism in the Time Domain
Pingkuan Dou, Ting Chen Jiang, Chao Li · 2020
CNN-based encoder-decoder architecture has been one of the most popular frameworks on speech enhancement task, and it has been proven that a better performance can be obtained by adding densely connected blocks. However, it suffers the problem of parameters over-allocation in feature re-usage. For this challenging problem, we propose a novel real-time monaural model with deep residual-dense lattice (RDL) network, sub-pixel convolutional layers and attention mechanism for speech enhancement in the time domain, and alleviate the problem of feature re-usage. Besides, we introduce a time domain scale-invariant signal-to-noise ratio (SI-SNR) loss instead of mean-square error loss as the loss of the time domain, moreover, we fine-tune the training model by the perceptual evaluation of the speech quality (PESQ) loss to further improve the performance of the SE model. Experimental results show that our model has improved PESQ score by 1.26 and the short time objective intelligibility (STOI) score by 22.3% compared with the unprocessed mixture.