Single-channel speech separation method based on attention mechanism
Ting Xiao, Jialing Mo, Weiping Hu, Dandan Chen, Qing Wu · Journal of Physics Conference Series · 2022
Abstract In order to improve the efficiency of speech interaction so that both humans and machines can hear it clearly, to solve the problem of single channel speech separation, a singlechannel speech separation method combined with attention mechanism is proposed. Based on the time-domain speech separation network, the method uses Encoder-separator-decoder framework and improves the Separator structure of the framework. First, the encoder transforms the input mixed speech signal into the latent space and extracts the feature of the speech signal to obtain the deep representation of the speech. Then, the deep attention features of speech were extracted by the splitter combined with the attention mechanism, and the mask was estimated for each independent sound source. Finally, the decoder inversely transforms the separated source signals back to the time domain. Experiments on WSJ0 speech separation data set show that the proposed method can effectively improve the SI-SNR of mixed speech separation, which is 0.31dB higher than the original time-domain network, and 0.94dB higher for the mixed case of boys and girls.