Speech Enhancement Based on Attention Mechanism and PG-LSTM Neural Network
Yuxi Qin, Youming Wang · 2021 International Conference on Control, Automation and Information Sciences (ICCAIS) · 2021
Speech enhancement is a challenging problem under low signal-to-noise ratio, and deep learning has been widely used in speech enhancement tasks due to its powerful learning ability for data. The long and short-term memory network(LSTM) cannot fully utilize the time-series features of speech in the training process, and it is easy to fall into the local optimal solution. At the same time, the network model has high complexity and is difficult to train. To address these problems, a speech enhancement method based on attention mechanism and particle swarm optimization-genetic algorithm (PG)-LSTM is proposed in this paper. An attention mechanism layer is added to the LSTM model before training, where the input speech is filtered using an attention scoring criterion to obtain more temporal information about the speech. In addition to this, a genetic algorithm for particle swarm optimization is used to find the optimal weight parameters of the LSTM network and to speed up the convergence of the network. Experimental results demonstrate that our proposed model can achieve considerably better performance than the other three comparison models in terms of two commonly used evaluation metrics under two noise conditions, and the enhanced speech obtains better speech quality and intelligibility.