Monaural Speech Separation Joint With Speaker Recognition: A Modeling Approach Based On The Fusion Of Local And Global Information

Shuangqing Qian, Jing‐jing Chen, Heping Song, Qirong Mao · 2020

In recent years, the dominant speech separation methods model the sequence gradually to encompass global information with the time steps accumulated. However, due to the limited memory capacity of the model, the sequence information in the later position occupies a large proportion, making it impossible to form good interactions between sequences that are farther away. In this paper, we propose a modeling approach Dual-Path Recurrent Neural Network with Transformer (DPRNN-Transformer), which is based on the fusion of local and global information. Through this method, the interrelationships between sequences can be directly established. And the proposed model can effectively merge local and global information in the high-dimensional space by combining both sequence masking and non-masking, thus solving problems such as sequence information forgetting in speech separation area. In addition, in order for the model to extract as much speaker-related feature information as possible, we add the auxiliary task of speaker recognition after speech separation. By co-training with speaker recognition, the speech separation module will be constrained by the additional Triplet loss, thus incorporating speaker information to facilitate separation. Experimental results on WSJ0-2mix dataset show that our proposed method greatly improves the performance of speech separation.

Read the paper · More papers on PaperTik